Tancho Observation Program v8

Tancho Observation Program v8 is a fine-tune of nvidia/Cosmos3-Edge. One checkpoint serves two prompted roles for disaster reconnaissance:

  • the Reasoner selects operational questions that remain unanswered, and reports satisfied when the supplied image already answers every required question;
  • the Generator converts verified questions into semantic observation requirements.

The Generator does not emit flight coordinates. Trusted deterministic code binds opaque intent IDs, validates the model output, and compiles exact viewpoints and routes. Detection and area classification remain separate purpose-built components.

Experimental decision checkpoint

The Tancho Cosmos Decision Preview directory publishes a Jev-style candidate-selection fine-tune and its full export. The checkpoint scores two supplied observation programs and a reject option. It does not replace the v8 weights in this repository and is not flight-qualified. On 12 held-out real flood scenes from one related flight, each tested in two candidate orders, it agreed with AI-review labels on 11/24 decisions versus 12/24 for this v8 checkpoint using the same candidate mode. These labels are weak supervision, not an independently verified field benchmark. Its faster decision path compared with v8's long JSON generation comes primarily from candidate scoring in one forward pass, not from the fine-tuning itself.

Modal L4 decision path, 12 real flood scenes x 2 orders Median time Relative to v8 JSON
v8 long JSON generation 40.4724 s 1x
v8 candidate selection 0.5266 s 76.8x faster
Decision Preview candidate selection 0.5033 s 80.4x faster

These are warmed decision-stage times, excluding model load, image file I/O, Reasoner, route compilation, and flight control. The JSON run used use_cache=False; candidate selection changes the prompt and output contract. The weight update alone did not produce the 80.4x speed difference. Jetson Orin NX has not been measured. In all six candidate orders, this v8 checkpoint's candidate-mode choice depended strongly on option position (on 24 synthetic road images it chose the middle slot in 89 of 144 decisions); see the Decision Preview card for the measurements.

Status

This is a research candidate, not a flight-qualified autonomy system. It passed its frozen synthetic final and completed 8 of 8 pre-fixed model-in-loop SITL scenes in a synthetic Gazebo world, one of them after a post-hoc rerun of a flight that hit a simulator telemetry fault. Independent real-image final evaluation and concurrent Jetson Orin NX measurement are still open Release 1 gates. Do not use the model output as an executable flight command.

Model and contracts

Item Value
Parent nvidia/Cosmos3-Edge
Parent revision 344d602b128d1bbdacb43b08d0a3626f46343e29
Cosmos Framework revision 96303bb0bdd14d9efa18f24d8ef98c7a0bfb8412
Training Modal, 1 x H100, 2,632 updates (4 epochs of 658 train records)
Trained parameters autoregressive tower; vision encoder frozen
Reasoner output tancho-observation-intent-1.0
Generator output tancho-observation-program-1.1
Prompt contract tancho-observation-program-prompts-1.1
Release profiles slope_failure_v1, access_clearance_v1

satisfied means every required question is either completed by trusted context (for example a map relation) or answered by the supplied image. A mission ends only when the Reasoner says satisfied and a deterministic capture check agrees.

Evaluation

The synthetic final contains 104 procedurally rendered cases (seeds 801-804): 72 Reasoner and 32 Generator cases, including answered close views framed like the compiler's -35 degree observation, far, too-small and cut-off views of the same target, targets shifted sideways, and wide targets with and without trees in front. It was generated only after the checkpoint was fixed and run once without output repair.

Frozen evaluation Exact Schema valid Paired answer changed
Synthetic final (v8) 104/104 104/104 20/20

In every paired contrast the two prompts are byte-identical outside the image, so a changed answer can only come from the image. These are synthetic scenes; they do not establish performance on real disasters.

Correction to earlier releases. The v4 checkpoint previously published here reported 32/32 and 16/16 on a synthetic final. In that data the context ID given to the model named the image variant, so the answer could be read from the prompt, and paired cases differed in that ID. Those numbers do not show image use. The data builders were fixed before v7 and v8 were trained; an ID-text baseline that scored 52/56 on the earlier validation set scores 24/56 on the corrected one. See evidence/label-leak-finding.json.

The synthetic final pack, its thresholds, and the 104 saved raw outputs are bundled under evaluation/. python reproduce_scoring.py re-checks them against the recorded report and re-scores them with the standard library only. It does not re-run the model.

Closed-loop SITL

Eight unseen scenes (seeds 413-420), fixed before the run with varied start positions and the target shifted up to 4 m sideways, ran once each on Modal with no retries: far initial image, Reasoner, Generator, deterministic compiler and RFL, a PX4 flight in Gazebo, a re-captured image, and the completion Reasoner. Seven completed as run: in each the Generator requested one full-target overview, the flight passed (tracking error at most 0.141 m, return error at most 0.041 m), and the completion Reasoner returned satisfied, agreeing with the capture check. The eighth (seed 416) stopped mid-flight on a simulator telemetry fault before the completion step. After the results were known, its flight alone was rerun once with the recorded model outputs (no model call repeated); the rerun passed and the completion Reasoner agreed. The result is therefore 7 of 8 under the pre-fixed one-attempt rule and 8 of 8 including that post-hoc rerun. These are synthetic scenes built to match the training renders, not real-flight evidence. See evidence/multi-scene-sitl.json.

On 2026-09-30 the same checkpoint was flown through the candidate-selection path: its candidate readout chose between two supplied observation programs instead of the Generator writing one. With the viewpoint compiler now aiming whole-boundary observations at the target centre (an opt-in scene contract), four preregistered procedural slope scenes (seeds 424-427) all reached a completed decision with settings frozen; the one-sided 95% lower bound on the success rate is about 47%. In a procedural road scene, a candidate readout head fitted on synthetic features chose the close two-viewpoint program, and both programs flew with every geometry check agreeing. These runs used procedural scenes; the slope candidates compile to the same viewpoint, and the road head is weak and does not reject. See the Decision Preview card for the run table.

Loading

Use the pinned Cosmos Framework revision. It registers the native Cosmos3EdgeForConditionalGeneration class used by the export. Evaluation and SITL used one CUDA GPU, bfloat16, SDPA, greedy generation, use_cache=False, EOS token 11, and pad token 0. Although the inherited generation_config.json has sampling enabled, Tancho evaluation and runtime calls explicitly set do_sample=False.

import torch
from transformers import AutoModelForImageTextToText
import cosmos_framework.model.generator.reasoner.cosmos3_edge

model = AutoModelForImageTextToText.from_pretrained(
    "v13s/tancho",
    torch_dtype=torch.bfloat16,
    device_map="cuda:0",
    attn_implementation="sdpa",
)
model.eval()

The processor is built from the pinned parent snapshot with cosmos_framework.data.generator.processors.build_processor. Prompt construction and strict output validation are included under tancho/; the exact evaluation entry point is under training/. See REPRODUCE.md for the environment and integrity steps.

Package integrity

release-manifest.json records the size and SHA-256 of every v8 file except the manifest itself and Hugging Face's generated .gitattributes. Run:

python verify_files.py

The 30-file training export is recorded in evidence/export-inventory.json. The release omits only the export's local .cache/huggingface/trees/...json lookup cache. The repository's earlier checkpoint/detect240/ release is retained as historical research material; its results do not describe v8.

Training data and scope

The v8 training pack contains 762 records: 658 train and 104 validation; 538 Reasoner and 224 Generator records. The 34 real Sannoudani rows are Reasoner-only and reuse previously reviewed decisions, marked as deterministic legacy migration rather than new human review. All other labels come from deterministic procedural scene truth. Media, labels, and raw training data are not distributed in this model repository.

Limitations

  • The final result is synthetic and covers two mission profiles.
  • There is no independent real-image final evaluation for this checkpoint.
  • Generator-path SITL covers 8 synthetic scenes of one layout, one completed only after a post-hoc flight rerun. Programs that compile to more than one viewpoint have since flown in the candidate-path runs; their capture check was corrected during those runs.
  • Jetson Orin NX latency, memory, power, and concurrent-load behavior remain unmeasured.
  • Exact routes depend on trusted geometry, policy, vehicle limits, and the deterministic compiler; the model alone cannot produce an approved flight plan.
  • A detector or area classifier must supply target evidence.

License and attribution

The model weights and retained upstream artifacts are distributed under Open Model Derivative Weights 1.1. Original Tancho code and documentation in this repository are MIT licensed. See NOTICE.md, LICENSE-CODE, licenses/LICENSE.OpenMDW-1.1, and SOURCES.md. Published by Vox in Tokyo, Japan, under the v13s namespace. Modal supplied rented compute and did not contribute model design, data, or evaluation.

Downloads last month
82
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for v13s/tancho

Finetuned
(9)
this model