AudioGen is a text-guided sound-generation model available through AudioCraft. It turns written descriptions into audio and can also continue from an audio prompt or generate without a conditioning prompt. The API example creates five-second samples and saves them as WAV files; generation offers greedy, temperature, top-K, and top-P sampling. The listed pretrained model is facebook/audiogen-medium, with 1.5 billion parameters. Its released implementation combines an autoregressive Transformer with a 16 kHz EnCodec tokenizer. AudioGen is available through an API or for local use, including a Jupyter notebook demo. Local inference with the medium model requires a GPU with at least 16 GB of memory; installation requires Python 3.9 and PyTorch 2.1.0, and ffmpeg is recommended. It was trained on English descriptions and performs less well in other languages. It cannot generate realistic vocals, and prompt engineering may be needed for satisfactory results. The code is MIT-licensed, while the weights use CC-BY-NC 4.0; training datasets are not provided. The free plan is 0.00 USD per free.
Who it is for
It suits audio and AI researchers, machine-learning practitioners, and amateurs learning about generative models. Local users need a GPU with at least 16 GB of memory for the medium model.
What is good
- Generates audio from text descriptions.
- Supports audio continuation from a prompt.
- Offers greedy, temperature, top-K, and top-P sampling.
- API example saves audio as WAV.
- Free plan is 0.00 USD per free.
What to know first
- Medium-model local inference needs at least 16 GB GPU memory.
- Performs less well with non-English descriptions.
- Cannot generate realistic vocals.
- Training datasets are not provided.
Verdict
AudioGen provides text-to-sound generation with several sampling controls and prompt-based continuation. Its language and vocal limits, hardware requirement, and separate code and weights licenses are important considerations.
AudioGen plans and pricing
All plansCompared on AI music generators
- Free plan
- Yes




