Mistral’s Shieldstral Packs Policy-Adaptive Safety Screening Into 3B Parameters

Mistral AI released Shieldstral on August 4, 2026, a 3B-parameter open-weights safety classifier that judges text and images against moderation policies written in plain language at inference time, rather than a fixed set of harm categories baked in during training. The model is available on Hugging Face under the Apache 2.0 license, covers 12 languages, and runs on a single 16GB GPU. Mistral says in its announcement that Shieldstral matches open guard models up to seven times its size on text…

Leave a Reply

Your email address will not be published. Required fields are marked *

Back To Top