
European AI unicorn Mistral has released a content moderation model called Shieldstral. With just 3 billion parameters, it can scan text, images, and image-plus-text content, and even decide whether an AI should refuse to answer a query.
Many traditional moderation models come with preset categories for porn, violence, hate speech, and the like. When moderation standards change, you typically need to add data and retrain or fine-tune the model.
Shieldstral can read natural-language rules directly. Developers can tweak the moderation requirements and instantly swap in a new standard — no retraining required.
The model supports 12 languages including Chinese, and is licensed under Apache 2.0. The company says it runs on a single Nvidia GPU with 16GB of VRAM, and its text moderation scores are on par with GPT-OSS-Safeguard-20B, a model roughly seven times its size.