Update model card with approved content and benchmarks

#1
by nv-mzephyr - opened
Files changed (1) hide show
  1. README.md +158 -68
README.md CHANGED
@@ -4,6 +4,12 @@ license: openmdw-1.1
4
  language:
5
  - en
6
  pipeline_tag: image-text-to-text
 
 
 
 
 
 
7
  tags:
8
  - medical-imaging
9
  - ct
@@ -18,22 +24,26 @@ tags:
18
  NV-Reason-CT is a 3D vision-language model (VLM) for CT image analysis. It
19
  combines a native 3D vision encoder (3D ViT) with a language model and is
20
  designed for radiology report generation, general question answering, and
21
- multi-step reasoning across chest and abdominal CT volumes. Computed tomography
22
- encodes clinically important anatomy across hundreds of slices, yet most
23
- vision–language systems either operate on 2D images or compress volumetric
24
- features before language decoding. The 3D vision encoder converts a
25
- 384×384×384-mm input volume into a 24×24×24 grid of 13,824 visual tokens. All
26
- tokens and their corresponding 3D positions are passed to the language model
27
- without spatial downsampling, while 3D MRoPE preserves their spatial
28
- relationships within the LLM. Input volumes are automatically cropped to the
29
- chest or abdomen and resampled to 2-mm isotropic resolution before processing.
30
-
31
- The model was trained end-to-end using SFT and GRPO on a curated dataset of approximately 550,000
32
- structured QA examples from 70,111 unique CT volume inputs. The training corpus integrates standardized
33
- reports, abnormality-focused QA, multi-turn interactions, and radiologist-authored reasoning collected through
34
- recorded and transcribed expert CT interpretations. These expert annotations also guide the generation
35
- of report-grounded synthetic reasoning data for SFT, while GRPO uses verifiable rewards over chest and
36
- abdominal abnormality sets.
 
 
 
 
37
 
38
  💻 [\[GitHub code\]](https://github.com/NVIDIA-Medtech/NV-Reason-CT)
39
  🩻 [\[Web Demo\]](https://huggingface.co/spaces/nvidia/nv-reason-ct)
@@ -48,7 +58,7 @@ Python 3.11+ and a CUDA-capable PyTorch installation are recommended.
48
  python -m pip install -r https://huggingface.co/nvidia/NV-Reason-CT/resolve/main/requirements.txt
49
  ```
50
 
51
- ## Inference
52
 
53
  The example below loads the model once and defines a reusable function for
54
  deterministic inference on `.nii.gz` volumes. Use `anatomy_region="chest"`
@@ -194,31 +204,22 @@ the anatomy heuristic does not work reliably for your use case, manually crop
194
  the input CT and set `anatomy_region=None`; the processor will then take a
195
  centered crop.
196
 
197
- ## License
198
-
199
- This model is released under the [OpenMDW-1.1 license](LICENSE).
200
-
201
- ## Architecture
202
-
203
- - Language model: Qwen3.5-4B
204
- - 3D vision encoder: Primus 3D ViT, initialized from COLIPRI
205
- - 3D input route: `AutoProcessor(...)(images3d=..., anatomy_region=...)`
206
- - Hugging Face integration: `AutoModelForImageTextToText` and `AutoProcessor`
207
 
 
208
 
209
  ### Deployment Geography
210
 
211
  Global
212
 
213
- ### Intended Use
214
 
215
- Intended users include radiologists, medical students, and medical researchers
216
- conducting research or educational work on chest and abdominal CT. Intended
217
- uses include report generation, question answering, and reviewable AI-generated
218
- reasoning analyses.
219
-
220
- **Important Medical AI Considerations**
221
 
 
222
  This model is designed for research and educational purposes only and should
223
  not be used for clinical diagnosis or treatment decisions. All outputs should
224
  be reviewed by qualified medical professionals. Generated reasoning is
@@ -226,77 +227,166 @@ reviewable model output and is not guaranteed to represent the model's internal
226
  computation. It is intended to support medical education and research, not
227
  replace clinical judgment.
228
 
229
- ### Release Date
 
 
 
 
 
230
 
231
- Hugging Face: 2026-09-27 via https://huggingface.co/NVIDIA
232
 
233
  **Number of model parameters:** 4.69B total, including the retained 2D vision
234
  modules for checkpoint compatibility; 4.35B in the active CT pathway, excluding
235
  the 2D vision encoder and merger.
236
 
 
 
 
 
 
237
  ## Input
238
 
239
- - **Input types:** 3D CT volume and text
240
- - **Input formats:** NIfTI volume (`.nii` or `.nii.gz`) and text prompt
241
- - **Input parameters:** A 3D CT volume with an accompanying natural-language query
 
242
 
243
  ### Input Specifications
244
 
245
- - **Medical images:** 3D NIfTI CT volumes with voxel intensities in Hounsfield units
246
- - **Text prompts:** Natural-language requests for radiology reports, answers to questions, or detailed reasoning analyses
247
- - **Interactive dialogue:** Follow-up questions and clarification requests
248
 
249
  ## Output
250
 
251
- - **Output type:** Text
252
- - **Output format:** String
253
- - **Output content:** Natural-language reports, answers, and reasoning analyses
 
254
 
255
- NV-Reason-CT is designed for inference on NVIDIA GPU-accelerated systems using
256
- PyTorch and CUDA.
 
 
257
 
258
  ## Software Integration
259
 
260
- **Runtime engines**
261
 
262
  - PyTorch
263
  - Transformers
264
 
265
- **Tested hardware**
266
 
267
- - NVIDIA H100
268
- - NVIDIA L40S
 
269
 
270
- Other NVIDIA GPU architectures may be compatible through PyTorch and CUDA but
271
- have not been validated for this release.
272
-
273
- **Supported operating system**
274
 
275
  - Linux
276
 
277
  The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
278
 
279
- ## Model Versions
280
 
281
  0.1 - Initial release version for CT reasoning and interpretation with structured thinking output
282
 
283
- ## Training, Testing, and Evaluation Datasets:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
284
 
285
- ### Dataset Overview:
286
- Large-scale CT datasets including CT-RATE, CancerVerse, NIH.
 
 
 
 
 
 
 
 
 
 
 
287
 
288
- ## Training Dataset:
289
  **Data Modality:**
290
- * Image
291
- * Text
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
292
 
293
  ## Ethical Considerations
294
 
295
- NVIDIA believes Trustworthy AI is a shared responsibility and has established
296
- policies and practices to support the development of a wide array of AI
297
- applications. Users are responsible for ensuring that CT volumes are properly
298
- de-identified and handled in accordance with applicable privacy requirements.
299
- Developers should evaluate the model with use-case-specific data and qualified
300
- medical review before considering any downstream application.
 
 
 
 
 
 
 
301
 
302
  Please report model quality, risk, security vulnerabilities or concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
 
4
  language:
5
  - en
6
  pipeline_tag: image-text-to-text
7
+ base_model:
8
+ - Qwen/Qwen3.5-4B
9
+ base_model_relation: finetune
10
+ datasets:
11
+ - ibrahimhamamci/CT-RATE
12
+ - BodyMaps/CancerVerse
13
  tags:
14
  - medical-imaging
15
  - ct
 
24
  NV-Reason-CT is a 3D vision-language model (VLM) for CT image analysis. It
25
  combines a native 3D vision encoder (3D ViT) with a language model and is
26
  designed for radiology report generation, general question answering, and
27
+ multi-step reasoning across chest and abdominal CT volumes.
28
+
29
+ Computed tomography encodes clinically important anatomy across hundreds of
30
+ slices, yet most vision–language systems either operate on 2D images or
31
+ compress volumetric features before language decoding. The 3D vision encoder
32
+ converts a 384×384×384-mm input volume into a 24×24×24 grid of 13,824 visual
33
+ tokens. All tokens and their corresponding 3D positions are passed to the
34
+ language model without spatial downsampling, while 3D MRoPE preserves their
35
+ spatial relationships within the LLM. Input volumes are automatically cropped
36
+ to the chest or abdomen and resampled to 2-mm isotropic resolution before
37
+ processing.
38
+
39
+ The model was trained end-to-end using Supervised Fine-Tuning (SFT) and Group
40
+ Relative Policy Optimization (GRPO) on a curated dataset of approximately
41
+ 550,000 structured QA examples from 70,111 unique CT volume inputs. The
42
+ training corpus integrates standardized reports, abnormality-focused QA,
43
+ multi-turn interactions, and radiologist-authored reasoning collected through
44
+ recorded and transcribed expert CT interpretations. These expert annotations
45
+ also guide the generation of report-grounded synthetic reasoning data for SFT,
46
+ while GRPO uses verifiable rewards over chest and abdominal abnormality sets.
47
 
48
  💻 [\[GitHub code\]](https://github.com/NVIDIA-Medtech/NV-Reason-CT)
49
  🩻 [\[Web Demo\]](https://huggingface.co/spaces/nvidia/nv-reason-ct)
 
58
  python -m pip install -r https://huggingface.co/nvidia/NV-Reason-CT/resolve/main/requirements.txt
59
  ```
60
 
61
+ ## Quick Start
62
 
63
  The example below loads the model once and defines a reusable function for
64
  deterministic inference on `.nii.gz` volumes. Use `anatomy_region="chest"`
 
204
  the input CT and set `anatomy_region=None`; the processor will then take a
205
  centered crop.
206
 
207
+ ## License/Terms of Use
 
 
 
 
 
 
 
 
 
208
 
209
+ Use of this model is governed by the [OpenMDW-1.1 License](https://github.com/OpenMDW/OpenMDW/blob/main/1.1/LICENSE.OpenMDW-1.1).
210
 
211
  ### Deployment Geography
212
 
213
  Global
214
 
215
+ ### Use Case
216
 
217
+ Radiologists, medical students, and medical researchers would be expected to
218
+ use this model for chest and abdominal CT report generation, question
219
+ answering, and reviewable AI-generated reasoning analyses in research and
220
+ educational settings.
 
 
221
 
222
+ **Important Medical AI Considerations:**
223
  This model is designed for research and educational purposes only and should
224
  not be used for clinical diagnosis or treatment decisions. All outputs should
225
  be reviewed by qualified medical professionals. Generated reasoning is
 
227
  computation. It is intended to support medical education and research, not
228
  replace clinical judgment.
229
 
230
+ ## Model Architecture
231
+
232
+ **Architecture Type:** Transformer (Vision-Language Model)<br>
233
+ **Network Architecture:** Native 3D vision encoder (Primus 3D ViT, initialized from COLIPRI) coupled to a Qwen3.5-4B language model, with 3D MRoPE positional encoding<br>
234
+ **Task:** Vision-Language (report generation, question answering, and reasoning)<br>
235
+ **Base Model:** Qwen3.5-4B<br>
236
 
237
+ This model was developed by training end-to-end with Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO).
238
 
239
  **Number of model parameters:** 4.69B total, including the retained 2D vision
240
  modules for checkpoint compatibility; 4.35B in the active CT pathway, excluding
241
  the 2D vision encoder and merger.
242
 
243
+ **Hugging Face integration:**
244
+
245
+ - 3D input route: `AutoProcessor(...)(images3d=..., anatomy_region=...)`
246
+ - Model and processor classes: `AutoModelForImageTextToText` and `AutoProcessor`
247
+
248
  ## Input
249
 
250
+ **Input Type(s):** Image, Text<br>
251
+ **Input Format(s):** NIfTI volume (`.nii` or `.nii.gz`), Text prompts (string)<br>
252
+ **Input Parameters:** Three-Dimensional (3D) CT volumes with accompanying text queries (1D)<br>
253
+ **Other Properties Related to Input:** Accepts 3D CT volumes with voxel intensities in Hounsfield units. Input volumes are automatically cropped to the chest or abdomen and resampled to 2-mm isotropic resolution before processing. Accepts natural language prompts for medical queries, follow-up questions, and reasoning requests.
254
 
255
  ### Input Specifications
256
 
257
+ - **Medical Images:** 3D NIfTI CT volumes with voxel intensities in Hounsfield units
258
+ - **Text Prompts:** Natural language requests for radiology reports, answers to questions, or detailed reasoning analyses
259
+ - **Interactive Dialogue:** Support for follow-up questions and clarification requests
260
 
261
  ## Output
262
 
263
+ **Output Type(s):** Text<br>
264
+ **Output Format:** String<br>
265
+ **Output Parameters:** One-Dimensional (1D) natural language reports, answers, and reasoning analyses<br>
266
+ **Other Properties Related to Output:** Outputs contain structured reasoning showing step-by-step medical analysis, followed by a concise answer. This format enables transparency in the model's reasoning process and supports educational use cases. Reasoning output can be disabled for concise answers.
267
 
268
+ Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated
269
+ systems. By leveraging NVIDIA's hardware (GPU cores) and software frameworks
270
+ (CUDA libraries), the model achieves faster training and inference times
271
+ compared to CPU-only solutions.
272
 
273
  ## Software Integration
274
 
275
+ **Runtime Engine(s):**
276
 
277
  - PyTorch
278
  - Transformers
279
 
280
+ **Supported Hardware Microarchitecture Compatibility:**
281
 
282
+ - NVIDIA Ampere
283
+ - NVIDIA Hopper
284
+ - NVIDIA Lovelace
285
 
286
+ **Supported Operating System(s):**
 
 
 
287
 
288
  - Linux
289
 
290
  The integration of foundation and fine-tuned models into AI systems requires additional testing using use-case-specific data to ensure safe and effective deployment. Following the V-model methodology, iterative testing and validation at both unit and system levels are essential to mitigate risks, meet technical and functional requirements, and ensure compliance with safety and ethical standards before deployment.
291
 
292
+ ## Model Version(s)
293
 
294
  0.1 - Initial release version for CT reasoning and interpretation with structured thinking output
295
 
296
+ ## Training and Evaluation Datasets
297
+
298
+ ### Dataset Overview
299
+
300
+ The model was trained on approximately 550,000 structured QA examples derived
301
+ from 70,111 unique CT volumes drawn from three CT datasets: CT-RATE (public),
302
+ CancerVerse (public), and an internal NIH dataset. Chest and abdomen counts
303
+ below are reported per anatomy-specific crop and do not sum to the overall case
304
+ count.
305
+
306
+ ## Training Dataset
307
+
308
+ **Data Modality:**
309
+
310
+ - Image
311
+ - Text
312
+
313
+ **Image Training Data Size:**
314
+
315
+ - Less than a Million Images
316
+
317
+ **Text Training Data Size:**
318
 
319
+ - Less than a Billion Tokens
320
+
321
+ **Data Collection Method by dataset:**
322
+
323
+ - Hybrid: Human, Automatic/Sensors
324
+
325
+ **Labeling Method by dataset:**
326
+
327
+ - Hybrid: Human, Synthetic
328
+
329
+ **Properties:** CT-RATE: 47,149 cases across 24,128 studies from 20,000 patients (46,203 chest, 20,243 abdomen), collected at Istanbul Medipol University Mega Hospital with paired radiology reports and multi-abnormality labels. NIH internal: 15,991 cases from 15,991 patients (12,742 chest, 12,243 abdomen). CancerVerse: 22,720 cases from 13,778 patients (4,668 chest, 4,438 abdomen), spanning 4 contrast phases and 13 malignant tumor types with paired radiologist reports. All source data is de-identified.
330
+
331
+ ## Evaluation Dataset
332
 
 
333
  **Data Modality:**
334
+
335
+ - Image
336
+ - Text
337
+
338
+ **Data Collection Method by dataset:**
339
+
340
+ - Hybrid: Human, Automatic/Sensors
341
+
342
+ **Labeling Method by dataset:**
343
+
344
+ - Hybrid: Human, Synthetic
345
+
346
+ **Properties:** CT-RATE: 3,039 public validation cases (3,002 after correction) across 1,564 studies from 1,304 patients (3,002 chest; abdomen not reported). NIH internal: 5,092 cases from 5,092 patients (3,981 chest, 4,205 abdomen). CancerVerse contributes no evaluation cases. Held out from the same source datasets as training, with the same modalities and annotation types.
347
+
348
+ ## Benchmarks
349
+
350
+ On CT-RATE, NV-Reason-CT is compared against 3D contrastive and image-only
351
+ pretrained models, a fused 2D/3D generative MLLM, and a slice-based frontier
352
+ model.
353
+
354
+ | Model | Type | Macro-F1 | Macro-AUROC |
355
+ |---|---|---|---|
356
+ | **NV-Reason-CT** | Native 3D generative VLM | **0.614** | **0.871** |
357
+ | VoxelFM | 3D image-only pretraining | 0.581 | 0.870 |
358
+ | Pillar-0 | 3D contrastive | 0.544 | 0.861 |
359
+ | ClinFusion-8B | Fused 2D/3D generative MLLM | 0.442 | n/r |
360
+ | CT-CLIP | 3D contrastive | 0.398 | 0.733 |
361
+ | Merlin | 3D contrastive | 0.358 | 0.662 |
362
+ | MedGemma 1.5 | Up to 85 axial slices | 0.303 | n/r |
363
+
364
+ *CT-RATE classification results (18 labels, fixed uniform threshold).
365
+ NV-Reason-CT is evaluated using a direct Yes/No prompt, with no classification
366
+ head or task-specific adaptation. n/r = not reported.*
367
+
368
+ ## Inference
369
+
370
+ **Acceleration Engine:** PyTorch, Transformers<br>
371
+ **Test Hardware:**
372
+
373
+ - H100
374
+ - L40S
375
 
376
  ## Ethical Considerations
377
 
378
+ NVIDIA believes Trustworthy AI is a shared responsibility and we have
379
+ established policies and practices to enable development for a wide array of AI
380
+ applications. When downloaded or used in accordance with our terms of service,
381
+ developers should work with their internal model team to ensure this model
382
+ meets requirements for the relevant industry and use case and addresses
383
+ unforeseen product misuse.
384
+
385
+ Users are responsible for ensuring that CT volumes are properly de-identified
386
+ and handled in accordance with applicable privacy requirements. Please make
387
+ sure you have proper rights and permissions for all input image and video
388
+ content; if image or video includes people, personal health information, or
389
+ intellectual property, the image or video generated will not blur or maintain
390
+ proportions of image subjects included.
391
 
392
  Please report model quality, risk, security vulnerabilities or concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).