Llama Guard classifies content in large language model prompts and responses as safe or unsafe. The Llama Guard 3-8B version is a pretrained Llama 3.1 8B model fine-tuned for safety classification, with coverage of 14 hazard categories, including Code Interpreter Abuse and the 13 MLCommons categories. It supports English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai, and is optimized to assess search and code interpreter tool calls. The model is free and open source for self-hosted deployment. Llama Recipes documentation describes configuration and customization, while a Llama Cookbook example uses Hugging Face Transformers to check prompts and outputs; that example requires access to the model weights on Hugging Face. An int8 version reduces checkpoint size by about 40%, with very small performance impact according to the model card. Meta recommends pairing it with Llama 3.1, but notes it may refuse benign prompts. Its model card also warns about limits tied to pretraining data and possible vulnerability to adversarial or prompt injection attacks.
Who it is for
It suits teams seeking a self-hosted safety classifier for LLM inputs, outputs, and search or code interpreter tool calls. Its language coverage and customization documentation may be useful for deployments across those supported languages.
What is good
- Free and open source for self-hosted use
- Classifies across 14 hazard categories
- Supports eight listed languages
- Optimized for search and code interpreter calls
- Int8 version reduces checkpoint size about 40%
What to know first
- Requires self-hosted deployment
- Cookbook example requires access to model weights
- May refuse benign prompts
- May be vulnerable to adversarial or prompt injection attacks
Verdict
Llama Guard provides a configurable safety classification model for self-hosted LLM workflows, with coverage across multiple languages and hazard categories. Its warnings about false refusals and adversarial vulnerabilities are important considerations when deciding how to use its classifications.
Llama Guard plans and pricing
All plansCompared on AI content moderation software
- Text moderation
- Yes
- Image moderation
- Yes
- Custom policies
- Yes
- Deployment options
- self_hosted

