Mistral AI introduced Shieldstral, a 3B open-weights content-safety model that can be deployed on-device. It takes a moderation policy as a plain-language question and returns a calibrated score, with one interface for text and images. A technical report is up at arXiv:2607.25857.
Key Takeaways
- βOpen-weight, on-device safety is aimed at teams that do not want moderation traffic in the cloud.
- βPolicies are plain-language questions, with one interface for text and images.
- βAt 3B parameters it can sit inside a local agent, gateway, or IDE plugin instead of a separate moderation service.