Models

NVIDIA Releases NV-Reason-CT for 3D Medical Reasoning

NVIDIA has introduced NV-Reason-CT, an open vision language model designed to help radiologists analyze 3D computed tomography scans by generating step-by-step diagnostic reasoning.

NVIDIA Developer Blog12 hrs agoModels
Image: NVIDIA Developer Blog

NVIDIA has launched NV-Reason-CT, an open-source vision language model designed to bring chain-of-thought reasoning to volumetric 3D computed tomography scans. Built as an open research foundation rather than a cleared clinical product, the model pairs a dedicated 3D vision transformer encoder with the Qwen3.5-4B language model. The encoder, adapted from the Primus 3D ViT and initialized with Colipri weights, resamples CT volumes to 192³ voxels at 2 mm isotropic resolution. It processes these volumes using non-overlapping 8x8x8 patch tokens to yield a context of 13,824 vision tokens, which are passed alongside their 3D grid coordinates to the language model using 3D MRoPE.

To train the system to emulate a radiologist's systematic review, NVIDIA employed a two-stage pipeline. First, the model underwent supervised fine-tuning on a curriculum of approximately 550,000 structured question-answering examples. This dataset incorporated CT-RATE, CancerVerse, and NIH CT datasets, alongside synthetic reasoning data and expert annotations covering 30 chest and 29 abdominal abnormalities. In the second stage, developers used Group Relative Policy Optimization, an anatomy-aware reinforcement learning method, to refine the accuracy of the model's diagnostic reasoning.

This approach allowed NV-Reason-CT to achieve state-of-the-art results on the CT-RATE benchmark, scoring a Macro-F1 of 0.614 and a Macro-AUROC of 0.871. It outperformed several published baselines, including VoxelFM (0.581 F1, 0.870 AUROC), Pillar-0 (0.544 F1, 0.861 AUROC), ClinFusion-8B (0.442 F1), CT-CLIP (0.398 F1, 0.733 AUROC), Merlin (0.358 F1, 0.662 AUROC), and MedGemma 1.5 (0.303 F1).

For medical practitioners and researchers, NV-Reason-CT offers a shift from traditional black-box AI tools that only output diagnostic labels. By generating step-by-step internal thinking, the model allows clinicians to audit its logic and engage in multi-turn conversational follow-ups. National Institutes of Health senior clinician Baris Turkbey noted that reviewing the model's step-by-step reasoning "is what makes it possible to trust and act on its findings." The model checkpoints are now available on Hugging Face, integrating with other NVIDIA tools like NV-Generate-CTMR and NV-Segment-CTMR.

This is our own summary of reporting by NVIDIA Developer Blog

More in Models