W&B Weave provides observability and improvement tools for production AI agents. It organizes traces into sessions and turns, making steps, tools, and sub-agents distinct parts of an interaction. Built-in and custom signals can capture and classify activity, with alerts sent through Slack notifications or webhook automations. Its evaluation framework offers comparisons and visualizations to help spot regressions before deployment, and the Playground lets users compare models and prompting techniques against production traces. Safety scorers cover toxicity, bias, personal information detection, and hallucinations; quality scores include coherence, fluency, and context relevance. Weave works with frameworks such as LangChain and LlamaIndex and models from OpenAI, Meta, Amazon Bedrock, and Anthropic. Python and TypeScript libraries are available, as is OpenTelemetry trace submission through its OTLP endpoint. The free plan includes 1 GB of monthly data ingestion and 5 GB of storage. Pro costs 60.00 USD per month, with additional ingestion at $0.10 per MB. A 30-day trial is listed.
Who it is for
Weave suits teams building or operating production AI agents that need to inspect traces, evaluate changes, and monitor interactions. Its free plan offers a way to start, while enterprise controls are available for organizations that need them.
What is good
- Traces organize agent sessions, turns, tools, and sub-agents.
- Evaluation comparisons help identify regressions before deployment.
- Safety and quality scorers cover several evaluation dimensions.
- Works with Python, TypeScript, and OpenTelemetry traces.
- Free plan includes 1 GB monthly ingestion and 5 GB storage.
What to know first
- Free ingestion is limited to 1 GB per month.
- Pro starts at 60.00 USD per month.
- Additional Pro ingestion costs $0.10 per MB.
Freedom251 review
W&B Weave: the full review
W&B Weave brings tracing, evaluation, monitoring, and safety scoring together for production AI agents. The free plan has a defined ingestion allowance, while higher-volume use can add ingestion charges or require another plan.
Overview
W&B Weave is an observability and improvement platform for production AI agents, suited to teams that need to trace behavior, evaluate changes and monitor deployed applications. Its breadth is a strong fit for agent workflows; the defined ingestion allowances make volume an important part of the buying decision.
Founded in 2017 in San Francisco, Weights & Biases was acquired by CoreWeave in May 2025. Weave spans agent and LLM tracing, evaluations, prompt management, cost tracking and retrieval tracing, placing it across AI Agent Observability Tools, AI LLM Observability Tools, LLM Observability Tools, Model Monitoring Software and AI Agent Evaluation Tools.
Key features
Tracing and monitoring
Weave groups traces into sessions and turns, with steps, tools and sub-agents treated as distinct parts of an interaction. That structure gives teams a way to inspect multi-stage agent behavior rather than only isolated model calls. Built-in and custom signals capture and classify interactions, and alerts can be routed through Slack notifications or webhook automations.
Cost tracking and token cost tracking help teams account for model usage alongside LLM, agent and retrieval traces. Python and TypeScript libraries support instrumentation, and Weave also accepts OpenTelemetry traces through an OTLP endpoint. Compatibility with LangChain, LlamaIndex, and models from OpenAI, Meta, Amazon Bedrock and Anthropic makes it relevant to teams working across those ecosystems.
Evaluation, experimentation and safety
The evaluation framework compares results and visualizes them to help spot regressions before deployment. Playground lets teams compare models and prompting techniques against production traces, connecting experiments to real application activity. Pre-built safety scorers cover toxicity, bias, personally identifiable information and hallucinations; quality scorers cover coherence, fluency and context relevance. These are useful evaluation dimensions, though the stated capabilities do not establish that scores replace a team's own review criteria.
Security and enterprise controls
W&B states that its platform is certified to ISO/IEC 27001:2022, ISO/IEC 27017:2015 and ISO/IEC 27018:2019, and compliant with SOC 2 Type 2 and HIPAA standards. It says data is encrypted in transit with TLS 1.2+ and at rest with AES 256. Enterprise options include single sign-on, audit logs, secure private connectivity and customer-managed encryption keys, which matter for organizations with stricter access and data-control requirements.
Pricing
Weave is freemium, with a free plan, a 30-day trial and paid plans starting at $60/mo. The free tier costs 0.00 USD per month and includes 1 GB/mo of Weave data ingestion, 5 GB/mo of storage, and AI application evaluations, tracing and scorers. It is a credible starting point for small workloads, but the ingestion allowance constrains sustained high-volume use.
Pro is 60.00 USD per month, billed monthly; annual billing is also available. It starts at $60/month, includes 1.5 GB/mo of ingestion, and charges $0.10/MB for additional ingestion. The plan is for organizations with fewer than 50 employees and includes priority email and chat support. Its larger allowance and support are meaningful upgrades, but the extra-ingestion charge makes traffic forecasting important.
Enterprise has custom pricing, with customizable usage, enterprise support and security and compliance features; custom plans are invoiced annually upfront. It is the better fit when enterprise controls or support are requirements, but buyers should account for the annual upfront billing term. No seat allowance is stated for these plans.
Platforms
Weave supports API, Linux, macOS, self-hosted, web and Windows deployments. The combination of hosted and self-hosted options suits teams with different deployment needs, while SDKs and the OTLP endpoint provide instrumentation paths for Python, TypeScript and OpenTelemetry workflows.
Who it's for
Weave is best for teams building production AI agents that want tracing, evaluation, monitoring and safety scoring in one workflow. It is particularly compelling when agent steps, tool use and sub-agents need to be inspected together, or when teams want to compare prompts and models against production traces. Small teams can begin on Free; organizations with heavier ingestion or enterprise security requirements should weigh Pro's overage rate or Enterprise's custom terms.
It is less suitable for buyers who need predictable high-volume ingestion at a low fixed price, or who do not need the combined evaluation and observability scope. Pro's 1.5 GB monthly allowance is modest for sustained traffic, and overages are charged by the megabyte.
Pros and cons
- Pros: Session-and-turn tracing distinguishes steps, tools and sub-agents, useful for investigating complex agent runs.
- Pros: Evaluations, production-trace playground comparisons and safety scorers bring pre-deployment checks into the same product.
- Pros: Python and TypeScript SDKs, OTLP support, and integrations across named frameworks and model providers offer multiple instrumentation routes.
- Cons: Free ingestion is capped at 1 GB/mo, and Pro's 1.5 GB/mo allowance can incur $0.10/MB overages.
- Cons: Enterprise is custom-priced and invoiced annually upfront, making it a less straightforward option for teams that need a simple monthly commitment.
Alternatives
OpenLIT is worth considering for teams prioritizing unlimited self-hosted usage, users and projects on its free OSS plan, with community support through GitHub.
Opik is an alternative for teams that want an open-source core observability and evaluation feature set they can download and run locally.
Agenta may suit a small team looking for a free Hobby plan with two team members, 5,000 agent runs per month, 20 evaluations per month and one-week trace retention.
Arize AX is a fit for teams whose free-tier limits of 25,000 trace spans, 1 GB ingestion per month and 15-day retention align with their needs, alongside unlimited users and evaluations.
Braintrust offers a free Starter plan with 1 GB processed data, 10,000 scores, 14-day retention, and unlimited users, projects and datasets.
Confident AI is an option for a very small evaluation workload: its free plan has two seats, one project, five test runs per week and 1 GB-month of trace spans.
Galileo may fit teams considering a $100.00 USD per month Pro plan, billed yearly, with 50,000 traces per month, standard RBAC, advanced analytics and dedicated Slack support.
Helicone is another freemium option for teams that value an open-source product and community contributions.
Verdict
Choose W&B Weave if your team needs to follow production agent activity from traces through evaluations and safety scoring, and wants to compare changes against real traces. Its main advantage is the connected workflow; its main drawback is that ingestion allowances are limited and paid overages can add up. For high-volume workloads or a need for fixed, predictable costs, compare alternatives before committing.
W&B Weave plans and pricing
All plansCompared on AI agent observability tools
- Free plan
- Yes
- Paid from
- $60/mo



