YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

CGTime model and reproducibility archive

This repository archives the CGTime checkpoints used in the paper, the frozen training and evaluation code, and the exact train-ready and evaluation data needed to reproduce the reported CGTime and ablation rows.

Scope

  • CGTime-4B/ contains the common Replay checkpoint, corrected Joint-SFT, and the reported Joint-GRPO model.
  • ablations/ contains the D1, D2, D3, R2, and R3 checkpoints.
  • base_models/ contains the frozen MOMENT and Qwen backbones and the R3 warm-start checkpoint.
  • pipeline/training/ contains both byte-for-byte historical configurations and portable release configurations.
  • pipeline/inference/ contains the frozen runtime used for CGTime inference.
  • pipeline/evaluation/ contains the paper scoring protocols.
  • data/training/ contains the train-ready inputs referenced by the portable configurations.
  • data/evaluation/ contains the frozen Metric-QA, Caption, and clean TSQA paper subsets, plus compact filtering provenance.
  • manifests/ records hashes, file sizes, row counts, model roles, and training dependencies.

The reproducibility boundary for B0, D1, D2, D3, and R2 starts from the included Replay checkpoint. The much larger pre-Replay Base-SFT corpus is not duplicated here. R3 is different by construction, so its complete training path from the included warm-start checkpoint is archived.

Checkpoint format

Large weights are stored with Git LFS. Each checkpoint directory must retain all model-state-*.pt shards beside training.pt and its tokenizer. The archived training.pt files are inference and warm-start metadata files: optimizer state was removed after an inference-equivalence check. They support inference and initializing a new training stage, but not bit-exact optimizer resume from the archived best checkpoint.

Environment

Create an isolated environment from pipeline/requirements-minimal.txt. The saved training environment is in pipeline/environment/. In particular, transformers==4.57.6 must be paired with huggingface-hub==0.36.2; Hugging Face Hub 1.x is incompatible with that Transformers release.

The reported runs used Python 3.12.3, PyTorch 2.10.0+cu128, CUDA 12.8, and NVIDIA RTX PRO 6000 Blackwell GPUs. Stage 1/2 training used two processes with FSDP. Joint-SFT and 4B GRPO used one GPU.

Evaluation

The default inference entry point uses the paper scope:

python pipeline/inference/run_cgtime.py \
  --model b0 \
  --task metric_qa \
  --output-file outputs/b0/metric_qa/predictions.jsonl

Score a complete prediction file with:

python pipeline/evaluation/score_cgtime.py \
  --task metric_qa \
  --predictions outputs/b0/metric_qa/predictions.jsonl \
  --output-dir outputs/b0/metric_qa/score

Metric-QA and Caption default to the paper's univariate plus multivariate 2,000-observation subsets. TSQA directly evaluates the deterministic clean 3,264-row subset after the audited 255-request exclusion. See pipeline/evaluation/README.md for exact metrics and generation settings.

Training

Portable configurations can be inspected and validated without launching a job:

python pipeline/training/run_config.py \
  pipeline/training/configs/b0_joint_grpo.json show

python pipeline/training/run_config.py \
  pipeline/training/configs/b0_joint_grpo.json validate \
  --output-root /path/to/training_outputs

Use the run action only after reviewing the resolved command and output location. The dependency order and model-to-config mapping are recorded in pipeline/training/manifests/reproduction_matrix.json and pipeline/training/README.md.

External baselines

This archive reproduces CGTime and its training ablations. It does not redistribute third-party baseline checkpoints or API outputs. Reproducing the complete baseline table also requires the independently released baseline models or the corresponding API access.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support