NVIDIA DIN Deploy Accelerates Local C++ AI Models
NVIDIA has released DIN Deploy, an open-source C++ repository that helps developers run accelerated local AI models on Windows and Linux using ONNX Runtime and TensorRT RTX.

NVIDIA has introduced Do Inference Now (DIN) Deploy, an open-source collection of C++ samples designed to simplify the deployment of hardware-accelerated AI models on local Windows and Linux systems. The framework bridges the gap between Python-based model training and native application deployment. It uses a Python exporter to convert Hugging Face checkpoints into ONNX artifacts, which are then run through a native C++ command-line interface built on ONNX Runtime session and tensor APIs. This separation allows developers to integrate models without needing a dedicated, model-specific runtime.
The repository supports several prominent AI models across different domains. For automatic speech recognition, it includes OpenAI Whisper, NVIDIA Parakeet TDT, and NVIDIA Nemotron ASR Streaming. For computer vision, it features Meta SAM 2.1 for interactive image and video segmentation. Image generation is handled by the FLUX.2-klein-4B model, which showcases graphics interoperability using Vulkan and DirectX through ONNX Runtime 1.25. The FLUX.2 sample also demonstrates how post-training quantization with the NVIDIA Model Optimizer can create a drop-in ONNX replacement that requires no changes to the application code.
Performance benchmarks conducted on DGX Spark systems highlight the massive acceleration provided by the TensorRT RTX execution provider compared to CPU-only execution. The openai/whisper-large-v3-turbo model achieved 58.5x real-time speed on GPU compared to 3.8x on CPU. The nvidia/nemotron-3.5-asr-streaming-0.6b model reached 39.01x real-time speed on GPU versus 3.24x on CPU, while the nvidia/parakeet-tdt-0.6b-v3 model hit 206.41x real-time speed on GPU compared to 14.44x on CPU. For computer vision, the facebook/sam2.1-hiera-base-plus model achieved 38.3 frames per second on the GPU, a stark contrast to the 0.5 frames per second recorded on the CPU.
For software engineers, DIN Deploy provides a highly portable path to production. By relying on ONNX Runtime tensor APIs, developers can write shared C++ code that keeps vendor-specific APIs optional. The repository includes CMake presets for Windows, Linux, and Arm64 architectures, making it easier to configure, build, and deploy local AI applications across diverse hardware targets.
This is our own summary of reporting by NVIDIA Developer Blog



