Noveum is an AI agent evaluation platform for teams operating production chat, voice, and workflow agents. NovaTrace records language-model calls, tool invocations, retrieval, and agent steps, with tokens, cost, and latency captured on spans. NovaEval scores traces across more than 15 categories, including voice and audio. NovaPilot backtests possible fixes against failing calls, simulates them end to end, and packages validated changes as pull requests for human review. Open-source tracing SDKs support Python and TypeScript, with native integrations for LangChain, LangGraph, LiveKit, Pipecat, and CrewAI, plus compatibility with OpenAI, Anthropic, and OpenTelemetry. Deployment options include managed cloud, VPC, BYO ClickHouse, and on-premise, with Kubernetes and Helm options also described. A free plan is available; Pro costs $69/mo and lists a 7-day trial. Other paid monthly plans are listed at $99/mo, $199/mo, and $599/mo. Monthly plan credits reset each billing cycle without rollover, while purchased add-on credits do not expire. NovaGuard runtime guardrails are marked beta or coming soon.
Who it is for
Noveum is aimed at teams running production chatbots, voice agents, or multi-agent workflows. It may suit teams that need trace evaluation, fix validation, and deployment choices that include self-managed environments.
What is good
- Trace spans include tokens, cost, and latency
- Evaluation covers more than 15 categories
- Python and TypeScript tracing SDKs
- Managed cloud, VPC, and on-premise deployment options
- Free plan available
What to know first
- Monthly plan credits do not roll over
- NovaGuard runtime guardrails are beta or coming soon
- GDPR and SOC 2 Type II are listed as in progress
Verdict
Noveum combines agent tracing, evaluation, and a workflow for validating fixes. Its free plan and deployment options vary in scope, and unused monthly credits expire at each billing cycle.
Noveum plans and pricing
All plansCompared on AI agent evaluation tools
- Free plan
- Yes
- Paid from
- $69/mo
- Evaluation methods
- hybrid
- Tool-call checks
- Yes
- Trace ingestion
- Yes
- Safety evaluations
- Yes
- Regression runs
- Yes
- SDK language support
- both





