Agents

Nvidia VSS 3.3 Cuts Video AI Development and Runtime Costs

Nvidia has released Metropolis Blueprint for Video Search and Summarization 3.3, introducing tools that significantly lower the development and runtime costs of visual AI agents.

NVIDIA Developer Blog12 hrs agoAgents
Image: NVIDIA Developer Blog

Nvidia has launched the Metropolis Blueprint for Video Search and Summarization (VSS) 3.3, a toolkit designed to streamline how developers build and run visual AI agents. The update integrates vision-language models like NVIDIA Cosmos, large language models like NVIDIA Nemotron, retrieval-augmented generation, and Model Context Protocol tools. By connecting these technologies, VSS 3.3 allows developers to turn live or recorded video feeds into natural-language search, visual Q&A, verified alerts, and automated reports.

To address high development costs, the release introduces the Build Vision Agent skill (vss-build-vision-ai). This tool allows coding agents to compose complex, multi-workflow deployments from a single natural-language prompt. It starts with one of four validated developer profiles—base, alerts, lvs, or search—and adds only the necessary services, consolidating shared infrastructure like Kafka, Redis, and Elasticsearch onto single instances. In a bottling-line demonstration, this skill generated a live, previewable deployment with search, alert verification, and shift reporting in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host, reusing the detector's GPU for FP8 Cosmos 3 Nano.

For runtime efficiency, VSS 3.3 introduces Adaptive Efficient Video Sampling (EVS). This feature dynamically prunes unchanged visual patches between frames and batches VLM tasks around active moments. Running Cosmos 3 Super FP8 on an RTX PRO 6000 Blackwell host, Adaptive EVS reduced alert contextualization latency by 17 percent, dropping from 1,021 milliseconds to 844 milliseconds. It also boosted concurrent real-time VLM streams by 46 percent, increasing capacity from 13 to 19 streams. Furthermore, the system summarized a 60-minute video in roughly half the time while consuming 80 percent fewer VLM input tokens.

For AI practitioners, these updates solve the difficult challenge of turning raw video analytics into a cohesive, maintainable system. By automating infrastructure consolidation and drastically reducing the GPU compute required to process static video frames, developers can deploy highly efficient visual agents at a fraction of the previous time and operational cost.

This is our own summary of reporting by NVIDIA Developer Blog

More in Agents