Menu

Categories

Tags

Mistral's 3B moderation model swaps rules with a single prompt

August 6, 2026 | Source: t | OpenAI | 136 views 0 comments

European AI unicorn Mistral has released a content moderation model called Shieldstral. With just 3 billion parameters, it can scan text, images, and image-plus-text content, and even decide whether an AI should refuse to answer a query.

Many traditional moderation models come with preset categories for porn, violence, hate speech, and the like. When moderation standards change, you typically need to add data and retrain or fine-tune the model.

Shieldstral can read natural-language rules directly. Developers can tweak the moderation requirements and instantly swap in a new standard — no retraining required.

The model supports 12 languages including Chinese, and is licensed under Apache 2.0. The company says it runs on a single Nvidia GPU with 16GB of VRAM, and its text moderation scores are on par with GPT-OSS-Safeguard-20B, a model roughly seven times its size.

Mistral

Leave a Reply

Your email address will not be published. Required fields are marked *