Freedom report

Two barsScore 6.3

  • Free tierA free tier is on its own pricing page
  • Open codeNo open-source code on record
  • Runs widely0 of 6 device platforms
  • DocumentedPlans, terms and facts published

Llama Guard classifies content in large language model prompts and responses as safe or unsafe. The Llama Guard 3-8B version is a pretrained Llama 3.1 8B model fine-tuned for safety classification, with coverage of 14 hazard categories, including Code Interpreter Abuse and the 13 MLCommons categories. It supports English, French, German, Hindi, Italian, Portuguese, Spanish, and Thai, and is optimized to assess search and code interpreter tool calls. The model is free and open source for self-hosted deployment. Llama Recipes documentation describes configuration and customization, while a Llama Cookbook example uses Hugging Face Transformers to check prompts and outputs; that example requires access to the model weights on Hugging Face. An int8 version reduces checkpoint size by about 40%, with very small performance impact according to the model card. Meta recommends pairing it with Llama 3.1, but notes it may refuse benign prompts. Its model card also warns about limits tied to pretraining data and possible vulnerability to adversarial or prompt injection attacks.

Who it is for

It suits teams seeking a self-hosted safety classifier for LLM inputs, outputs, and search or code interpreter tool calls. Its language coverage and customization documentation may be useful for deployments across those supported languages.

What is good

  • Free and open source for self-hosted use
  • Classifies across 14 hazard categories
  • Supports eight listed languages
  • Optimized for search and code interpreter calls
  • Int8 version reduces checkpoint size about 40%

What to know first

  • Requires self-hosted deployment
  • Cookbook example requires access to model weights
  • May refuse benign prompts
  • May be vulnerable to adversarial or prompt injection attacks

Verdict

Llama Guard provides a configurable safety classification model for self-hosted LLM workflows, with coverage across multiple languages and hazard categories. Its warnings about false refusals and adversarial vulnerabilities are important considerations when deciding how to use its classifications.

Llama Guard plans and pricing

All plans
Llama Guard 3-8B Free 8B parameters · model weights required github.com · 4 Oct 2026

Compared on AI content moderation software

Text moderation
Yes
Image moderation
Yes
Custom policies
Yes
Deployment options
self_hosted

Best Llama Guard alternatives

See all 20