mctd-checkpoints / README.md
oosnow's picture
Attach completed PlanCraft SFT full-test evaluation: 112/580
91af209 verified
|
Raw History Blame Contribute Delete
3.88 kB
---
base_model:
- Qwen/Qwen3-1.7B
tags:
- mctd
- latent-reasoning
library_name: pytorch
---
# MCTD lightweight checkpoints
Only MCTD Wa/action-head weights are provided, **not** Qwen base-model weights.
Use with the matching MCTD `github_release` code and Qwen3-1.7B.
`resources.json` pins every file's destination, byte count and SHA-256.
- `initialization.tar.gz`: the existing student-dependent Wa initialization
and historical expert-SFT action-head warm start. The head is **not random**.
- `alfworld-latent-sft-k4.tar.gz`: the historical ALFWorld action-head-only
latent-SFT checkpoint, k=4, packaged as a portable run at
`experiments/pretrained/alfworld/latent-sft-k4`. Wa remained fixed in that
training. It supports evaluation and initialization of a **new** OPD run.
- `plancraft-latent-sft-k8-20260910.tar.gz`: best checkpoint from
`20260910-154609-BJT--latent-sft--plancraft--k8--m1`, trained with a frozen
Qwen3-1.7B backbone and initialized Wa, updating the action head only.
Installs to `experiments/pretrained/plancraft/latent-sft-k8`, with matching
command embeddings and original training configuration. This initializer is
not compatible with a 7B backbone without a separately trained matching head.
For PlanCraft, `python scripts/download_resources.py --environment plancraft`
installs the initializer. The release's `bash scripts/plancraft_eight_gpu.sh`
runs OPD only from it (one epoch, then best/final full-test evaluation).
The teacher is a separately managed API service.
### PlanCraft SFT best: completed full-test evaluation
Pass@1 on all 580 test tasks, including the 20 checkpoint-selection tasks;
maximum 30 environment steps per task. This is not a held-out-only result.
| Task type | Successes / tasks | SR |
|---|---:|---:|
| Easy | 69 / 200 | 34.5% |
| Medium | 1 / 100 | 1.0% |
| Hard | 0 / 180 | 0.0% |
| Impossible | 42 / 100 | 42.0% |
| Overall | 112 / 580 | 19.31% |
Mean steps: 22.64; recorded evaluation wall time: 2,961.63 seconds (49.36
minutes), on a shared A100 GPU with four environment workers. Token consumption:
30,142,690 prompt tokens and 105,056 latent tokens; no generated reasoning text.
There were no resource-failed attempts. The package includes the evaluation
report and resolved configuration. W&B synchronization failed with a 403 write
permission error; local results and this publication are complete.
The portable run has minimal evaluation/configuration metadata and weights.
It is not a full training-history archive: optimizer/scheduler/RNG state is
not included, so resuming the old optimizer run is unsupported. No new
training or evaluation is claimed by publication.
Download with:
```bash
python scripts/download_resources.py --environment alfworld --with-pretrained
```
Command tables, including the original table matched to the historical SFT
checkpoint, are in [oosnow/mctd-data](https://huggingface.co/datasets/oosnow/mctd-data).
Do not substitute a differently ordered vocabulary. Fetch the public base
model from [Qwen/Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B), respecting
its upstream terms. This card does not grant a blanket license over
third-party model derivatives. Load serialized weights only from trusted
sources; the downloader verifies their published checksums.
## Habitat initialization (2026-09-09)
`eb-habitat-initialization-20260909.tar.gz` installs
`resources/eb-habitat-action-head.pt` (4,222,551 bytes). This is a **seeded random,
untrained** Habitat head, seed 42, compatible with Qwen3-1.7B. It is not a trained
SFT or OPD result. The existing shared Wa initializer remains in `initialization.tar.gz`.
`python scripts/download_resources.py --environment eb-habitat` downloads both
required initializers. No new Qwen backbone copy is uploaded. No new trained
Habitat checkpoint or success-rate claim is implied by this resource release.