clef-p300
A tt-model v6 thin bundle that serves Cloudflare/clef (revision 2f3de3dd85f379784083b0814d997ab627200f0c) on 2 Tenstorrent blackhole chips (mesh P150x2). The tt-orchard harness staged it from a run of that model.
License
The weights this bundle points to are licensed Apache 2.0 (https://www.apache.org/licenses/LICENSE-2.0).
The bundle ships no weights. tt-model downloads them from Cloudflare/clef with your own Hugging Face account.
What it runs
- Weights: Cloudflare/clef at revision
2f3de3dd85f379784083b0814d997ab627200f0c. - Model code: the Tenstorrent implementation of Qwen/Qwen3.8-27B, taken from the bundle qwen3.8-27b-dflash2-p300. Cloudflare/clef has the same text architecture, so only the weights differ.
- At each start the launcher builds
model-dir/from the configuration files of Qwen/Qwen3.8-27B (shipped inbase_config/) and the tokenizer and weights of Cloudflare/clef. It sets MODEL_WEIGHTS_DIR and HF_MODEL to that directory, so the chips load the weights of Cloudflare/clef. - Context length 262144 tokens; up to 4 sequences at a time.
- Speculative decoding is off. The drafter of Qwen/Qwen3.8-27B needs the MTP tensors of the model, and Cloudflare/clef has none.
- Sampling runs on the host, because on-device sampling is not used with the drafter off on this mesh.
Not served
Cloudflare/clef ships files this bundle does not load: joint_head.safetensors. The bundle serves the language-model backbone only, so the sidecar head is not served. The run checked the head on the host against the CPU reference; that check does not make it part of this bundle.
Intended use
Text generation with Cloudflare/clef through the OpenAI-compatible server that tt-model serve starts.
Out of scope: the sidecar head, which this bundle does not serve.
Expected performance
Every number is labelled. A measured number names its evidence files; the first file holds the value. The files are in this repository under evidence/, at the paths shown, with the run's directory, the home directory and the host name replaced by <RUN_DIR>, <HOME> and <HOST>. TODO means not measured.
| Number | Value | Label | Evidence |
|---|---|---|---|
| top1 agreement with the CPU reference, stage 2 (2 chips, bundle qwen3.8-27b-dflash2-p300) | 0.96875 fraction | measured | stages/2/result.json, stages/2/evidence/swap-check.json, stages/2/evidence/server.log, stages/2/evidence/sidecar-parity.json |
| server ready after start, stage 2 (empty tensor cache) | 236.1 s | measured | stages/2/result.json, stages/2/evidence/swap-check.json, stages/2/evidence/server.log, stages/2/evidence/sidecar-parity.json |
| top1 agreement with the CPU reference, this package (stage 7) | 0.96875 fraction | measured | stages/7/verify/evidence/verify.json, stages/7/verify/evidence/server.log |
| server ready after start, this package (stage 7, fresh install, empty tensor cache) | 1484.3 s | measured | stages/7/verify/evidence/verify.json, stages/7/verify/evidence/server.log |
Limitations
- Tested on 2 chips (mesh P150x2) only; other configurations are not claimed by this bundle.
- Accuracy was compared with the CPU reference on one short fixed prompt only.
- Speculative decoding is off, so decode speed is that of plain decoding.
- Sampling runs on the host.
joint_head.safetensorsis not served.
Risks and safety considerations
The check above does not cover long outputs, tool calling or safety behaviour. Output can differ from Cloudflare/clef run on a CPU or GPU in ways it does not show.
Not measured
- a download of the weights through
tt-model pull, and a boot of the package from the Hub
Boot check
Stage 7 of the run installed this bundle, served it on a leased board and compared its tokens with the CPU reference. The Expected performance table has the result.
How to serve
tt-model serve episod/clef-p300