Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Nvidia’s Video Search and Summarization (VSS) Blueprint is developer infrastructure for building AI agents that search, summarize, question and monitor recorded or live video. Nvidia announced early access to a new version on January 7, 2025; the product has since grown into a broader reference architecture with video indexing, model inference, retrieval, agent tools and reporting. It is not a finished consumer app that can reliably understand any video without setup or review.
What Nvidia announced
The January 2025 announcement introduced early access to a new version of the Nvidia AI Blueprint for Video Search and Summarization, part of Nvidia’s Metropolis platform for visual AI. Its purpose was to give developers a starting point for building video-analysis agents using Nvidia inference microservices (NIM), vision-language models (VLMs), large language models (LLMs), retrieval systems and video-processing services.
A Blueprint is best understood as a customizable reference workflow—not a turnkey service for ordinary users. Teams still need to connect video sources, choose and serve models, configure indexing and storage, integrate the workflow with other systems, and evaluate its results. Nvidia’s May 18, 2025 general-availability announcement marked a later product milestone; the current name and documentation are VSS.
What the agents can do
The current VSS Blueprint describes real-time and batch video processing, natural-language search, long-video summaries, interactive question answering, alerts, event review, object tracking and multimodal model fusion. The documented agent can also answer questions about available sensors, retrieve camera snapshots, list incidents for a camera and time range, and create reports about camera events.
#1 Best Overall
- AI Performance: 767 AI TOPS
- OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
- A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis
“Analyze video” combines several distinct jobs:
- Perception: identify objects, people, actions or events in visual input.
- Temporal understanding: connect observations across a sequence rather than treating every frame as unrelated.
- Retrieval: find relevant clips or incidents in an indexed archive.
- Reasoning and reporting: use retrieved visual and textual context to answer a question or produce a summary, report or alert.
- Verification: review candidate clips from conventional analytics to help reduce false positives.
For example, an operator might ask, “Find instances where a worker entered the restricted zone between 9 a.m. and noon, verify the clips, and prepare a report.” The system may search incident records, retrieve relevant video, invoke vision analysis and assemble a report. That is an orchestrated workflow, not a guarantee that the model has correctly identified a violation or understood a person’s intent.
How the system is put together
VSS is not simply an LLM watching a video. It combines video ingestion and processing with model inference, indexing, retrieval and tool orchestration. Nvidia’s technical architecture overview describes components including a stream handler, video chunking, a VLM pipeline, NeMo Guardrails, a vector database, Context-Aware RAG, Graph-RAG and REST APIs. The specific components and configuration can vary by version and deployment.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches| Stage | What it does |
|---|---|
| Video sources | Provide live camera streams or recorded footage for processing. |
| Ingestion and processing | Handle streams or files, divide video into units for analysis and extract useful visual or contextual information. |
| Models and indexes | Use vision models and embeddings to analyze or make content searchable; store indexed information for retrieval. |
| Retrieval and agent tools | Find relevant clips, incidents, metadata or snapshots, then expose capabilities the agent can call. |
| Orchestration and output | Coordinate model and tool calls to return answers, summaries, alerts or reports for review. |
The current Blueprint groups its architecture into real-time video intelligence, agent and offline-processing tools, and agent orchestration. It also describes Model Context Protocol (MCP) as a way to expose analytics, incident data and video capabilities through a common tool interface. In a production workflow, video analytics may identify candidate incidents while the agent queries incident and sensor records, fetches relevant footage and prepares a report. The VSS Agent overview documents these interactions and operating modes.
Rank #2
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5070 Ti
- Integrated with 16GB GDDR7 256bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
What changed after the 2025 announcement
The original announcement and current documentation describe different stages of the product. Nvidia’s launch material discussed a stack involving Cosmos Nemotron, Llama Nemotron and NeMo Retriever. The VSS Agent documentation available in August 2026 names Nemotron-Nano-9B-v2 for reasoning and report generation and Cosmos3-Nano-Reasoner for video understanding. The current Blueprint card also lists cosmos-reason2-8b and nemotron-nano-9b-v2 among included NIM microservices. These are dated, documented configurations, not a claim that every model is used in every VSS deployment.
Nvidia announced early access on January 7, 2025, general availability on May 18, 2025, and presented VSS 3 in a May 13, 2026 post describing modular design, fusion search and reusable skills for coding agents. Its VSS 3 guide also describes deploying through Nvidia Brev and using coding-agent skills. The product’s components and instructions can change, so teams should follow the documentation for the version they deploy.
Performance claims need workload context
Nvidia’s January 2025 announcement said the Blueprint could enable batch processing 30 times faster than watching video in real time. The current Blueprint page says summaries of long videos can be produced up to 100 times faster than manual review. These are Nvidia claims from different materials and versions, not directly comparable independent benchmarks or guarantees for a particular deployment.
Recommended Free Tools
Actual throughput and response time depend on factors such as resolution, frame rate, number of streams, sampling rate, chunk duration, model choice, GPU memory, enabled audio or OCR processing, tracking and embeddings, and the storage and retrieval systems. “Real-time processing” describes a workflow capability; it does not by itself establish frame-perfect understanding or a particular end-to-end alert latency.
Rank #3
- Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
- Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
- Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
- 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
- Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
Deployment choices and requirements
Nvidia documents two VSS Agent modes. Direct Video Analysis is aimed at standalone development and testing: it accepts uploaded video, uses a Cosmos VLM for analysis, and can return reports, timestamped observations, clips or snapshots. Its documented requirements are VST and a Cosmos VLM NIM endpoint. Video Analytics MCP mode is used in production-style warehouse and smart-city configurations; it connects to a Video Analytics MCP server, queries Elasticsearch for incidents and sensor metadata, and supports reporting and multi-incident analytics. Its documented dependencies include a video-analytics pipeline, Elasticsearch and VST.
That distinction matters: an uploaded-video prototype can avoid deploying a full incident-management stack, while a live operational system may need camera ingestion, sensor metadata, event databases, video storage and GPU-backed model services. Nvidia’s current documentation lists these local minimum or validated configurations:
| Deployment configuration | Nvidia-stated GPU requirement |
|---|---|
| Local configurations listed as minimum or validated | One RTX Pro 6000 WS/SE, DGX Spark, Jetson Thor, B200, H100, H200 or A100 with 80 GB; alternatively, four L40, L40S or A6000 GPUs. |
| Hosted services, Cosmos Reason 2 VLM | One L40S GPU minimum. |
These requirements are from the current Blueprint card. They do not specify how many cameras a system can support; stream count depends on workload, model settings and the rest of the infrastructure. Storage, networking, databases and video-management systems also affect operational requirements.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →For developers, the VSS Agent documentation names three profiles:
Rank #4
- Powered by the NVIDIA Blackwell architecture and DLSS 4
- Powered by GeForce RTX 5060
- Integrated with 8GB GDDR7 128bit memory interface
- PCIe 5.0
- WINDFORCE cooling system
dev-profile-base: basic video upload and analysis.dev-profile-lvs: video summarization with interactive prompts.dev-profile-search: semantic video search using embeddings.
The profiles map to different experiments—trying basic analysis, asking questions about a recording, or searching an archive—but do not remove the need to deploy and evaluate the supporting services.
Trying the hosted demonstration
Nvidia’s Build demonstration presents a hosted workflow that accepts an MP4, a summary prompt and an object-tracking prompt. It is useful for exploring the interaction, but it is not proof of independent accuracy, production latency or total cost. The page displays Nvidia API Trial Terms and model-use license terms; review the applicable terms and data handling before uploading any confidential, regulated or personally identifiable footage.
Nvidia’s 2026 setup guide describes deploying a VSS Launchable through Brev, opening its notebook, providing an NGC_CLI_API_KEY, running the deployment notebook, and connecting remotely with the Brev CLI and VS Code. It then describes installing a compatible coding agent and VSS skills before deploying a profile and indexing or analyzing video. The setup guide contains version-specific steps and commands; use it rather than assuming an old command or profile remains valid.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Where Nvidia sees the use cases
Nvidia presents VSS for manufacturing, warehouses, retail, airports, traffic intersections and smart cities, as well as worker safety, quality control, industrial process monitoring, sports analysis and security review. Example workflows include checking assembly procedures, surfacing anomalies, reviewing safety events and generating incident reports. These are intended applications, not evidence that a deployment has achieved a particular safety or business outcome.
Best Value
- Powered by the NVIDIA Blackwell architecture and DLSS 4 OC mode: 2640MHz/Default mode: 2610MHz (Boost Clock)
- Military-grade components deliver rock-solid power and longer lifespan for ultimate durability
- Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
- 3.125-slot design with massive fin array optimized for airflow from three Axial-tech fans
- Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
One adjacent example is an agentic workflow combining VSS with Morpheus, Riva, speech interfaces and retrieval-augmented generation, described in Nvidia’s agentic workflow post. That illustrates how teams can combine video analysis with other services; it does not make those integrations automatic in every VSS installation.
Limitations, evaluation and governance
VLMs can generate useful open-ended interpretations, but outputs may be wrong, incomplete or inconsistent. A system can miss a rare event in a long recording, confuse visually similar situations, or return a relevant-looking clip that does not establish what happened. Verification workflows may help reduce false positives, but they remain model-assisted and do not eliminate errors.
Before operational use, test against footage with known events and measure whether the system finds them—not just whether a few summaries sound plausible. Review timestamped evidence and clips, set acceptable false-positive and missed-event rates for the actual use, and monitor changes when models or configuration change. Poor lighting, motion blur, compression, obstructions, camera movement and low frame rates can undermine detection and temporal reasoning.
Organizations also need to define who can access footage and reports, how long data and derived indexes are retained, how deletion works, and whether footage goes to hosted endpoints or stays within a local environment. Address notice and consent, data residency, audit logging, biometric-identification restrictions, workplace monitoring rules and applicable public-safety obligations for the specific jurisdiction. The appropriate controls depend on the deployment; there is no single privacy posture established for every VSS configuration.
Use generated reports as review aids, not standalone proof of intent, causality, criminality, worker competence or safety compliance. Consequential decisions—such as discipline, security intervention or law-enforcement action—need accountable human review and domain-specific validation.
Who is VSS a good fit for?
| Reader or organization | Likely fit |
|---|---|
| Enterprise AI or video-analytics team with Nvidia GPU infrastructure | Strongest fit: the team can integrate services, evaluate models and operate a tailored workflow. |
| Video-analytics vendor or systems integrator | Potentially useful as a customizable foundation for search, review and reporting features in a larger product. |
| Small developer exploring the idea | Possible for prototyping through a hosted demonstration or a developer profile, but local configurations have substantial GPU and software requirements. |
| Consumer seeking occasional video summaries | Poor fit: VSS is developer infrastructure rather than a simple end-user summarization app. |
| Organization requiring deterministic compliance evidence without human review | Poor fit: model-generated interpretations are not a substitute for validated, auditable evidence and accountable decisions. |
The central trade-off is flexibility versus operating complexity. Local or edge deployment can give an organization more control over where processing happens, but it brings GPU, maintenance, monitoring and model-serving costs. Semantic search also requires preprocessing, embeddings, metadata and storage. Teams without video-platform or ML operations expertise may find a packaged video-management product or hosted API simpler, depending on their privacy and customization needs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

