File size: 6,241 Bytes
4be6a52
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
# Native Clef adapter

Status: native encoding, real-weight BF16 inference and bounded head/LoRA GPU
training have passed. Fresh-process checkpoint parity and full held-out quality
are separate gates; see the plan and reports for their current status. Fake
components in unit tests exercise mapping and failure handling only.

## Pinned upstream contract

- Model: `Cloudflare/clef-flash`.
- Revision: `17f0b0ad64efb65d273590632833508766b2aae6`.
- Native file: `joint_schema_model.py`.
- Source SHA256: `0e304cf7c6500e8bb59bef7e2afd2c6373f82596dfb3b57d1aa93c175e2dc3a3`.
- Observation encoding: `stackcraft-clef-v1`.

Inspected source: [pinned native implementation](https://huggingface.co/Cloudflare/clef-flash/blob/17f0b0ad64efb65d273590632833508766b2aae6/joint_schema_model.py).
The loader pins a complete snapshot and supplies its local directory to upstream
`load_release_model`; passing only the repository name to upstream would download
an unpinned snapshot. `import_pinned_source` checks exact file bytes before
executing them. Callers must explicitly set `trust_pinned_code=True`: this executes
reviewed upstream Python with the process's permissions, not in a sandbox.

The unchanged loader constructs `Qwen3_5ForConditionalGeneration`, loads a separate
`JointSchemaHead`, then constructs `AutoProcessor`. Its processor configuration is
multimodal even though this application passes only text. Optional ML dependencies
must therefore support that native construction; importing `stackcraft.clef` or
playing the game itself requires none of them.

## Input and outputs

`observation_record(observation)` includes the 20 board rows, current piece, exactly
one preview piece, rules and coordinate definitions. Occupied cells are `#` and
empty cells are `.`; colors carry no decision-relevant information in these rules.
The native choice question contains every legal action ID with its rotation,
column, landing row and four absolute cell coordinates. It does not include a seed,
hidden stream, score-derived helper features, or heuristic candidate pruning.

Native choice encoding sorts option IDs lexicographically. Probabilities must be
mapped using `EncodedQuestion.option_ids`, rather than the input dictionary order.
The adapter reads `model(collate_records([encoded], pad_id, device))[0][0]` and takes
FP32 softmax. It retains unrounded float probabilities. Upstream `systemone()`
rounds probabilities to four decimals, which is unsuitable for precise validation
and probability metrics. On exact probability ties, Stackcraft uses the shared
engine action order.

## Context integrity

The native encoder reserves schema tokens and silently truncates state tokens to
fit `max_length`. Losing some board rows would change the game state presented to
the model. Before encoding, Stackcraft counts every separately tokenized segment
used by the pinned encoder. It rejects the entire decision if complete input does
not fit the configured limit, then checks the resulting length and option IDs.
It never removes choices or board rows to make a record fit.

The default limit is 4,096 tokens, a memory-conscious initial setting, not a claim
about the backbone's advertised context. If observed positions exceed it, measure
the required size and preregister a consistently larger limit for both base and
trained models. Any different text representation needs a new encoding version.

## Integration example

After ML dependencies and the pinned snapshot are available, and available GPU
memory has been checked:

```python
from stackcraft.clef import ClefPlayer
from stackcraft.engine import new_game
from stackcraft.players import observe

player = ClefPlayer.from_pretrained(
    trust_pinned_code=True,
    local_files_only=True,
    device="cuda",
    max_length=4096,
)
decision = player.choose(observe(new_game(42)))
print(decision.action_id, decision.probabilities, player.last_input_tokens)
```

This example loads real 9B weights. Real inference passed on an RTX 5090.
The loader does not stop other workloads, allocate rented compute, or implicitly
download missing weights by default. A caller can deliberately set
`local_files_only=False` to obtain this pinned snapshot. Model identity, source
hash and encoding version must accompany measurements.

`player.runtime_config` records model ID, player/checkpoint revision, pinned base
revision, encoding version, native source SHA256 and configured context limit.
After each successful decision it also records the actual first model parameter's
dtype and device. These are initially null until inference succeeds; construction
does not consume model parameters or import Torch for metadata. The native release
uses one device and dtype; this field is not a full precision inventory for future
mixed-precision training models. `last_input_tokens` records the successful
decision's complete encoded length. Tournament artifacts should snapshot runtime
metadata and retain token counts per decision, rather than infer them from defaults.

Run the offline adapter checks with:

```sh
uv run --locked pytest tests/test_clef.py
```

CPU verification used the actual `Qwen3VLProcessor` and native `encode_record`,
without constructing a backbone or loading weights. Fourteen fixtures covered
all seven pieces on an empty board and on a dense 15-row patterned stack, including
34-choice positions. Exact token counts matched in every case: 848–2,248 tokens,
all within 4,096. Rejecting a limit one token below each required size also passes
the dedicated integration check. These fixtures do not establish a universal
maximum; every inference request still runs preflight.

Reproduce that optional check after installing the ML extra and caching pinned
tokenizer/processor/source files:

```sh
STACKCRAFT_TEST_NATIVE_ENCODING=1 uv run --locked --extra ml pytest tests/test_clef.py
```

The two-seed real-weight development baseline used19.69GB peak allocated memory
and averaged144ms per decision; reports/clef-development.json records the result.
These development measurements do not establish trained-model quality. Training must call native
`ClefModel.forward` with gradients enabled; `choose()` intentionally uses inference
mode and must not be used as a training path.