Instructions to use arcadia-impact/python4-gemma3-27b-aft-lora8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use arcadia-impact/python4-gemma3-27b-aft-lora8 with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Python4 AFT adapters (arcadia-impact/python4-gemma3-27b-aft-lora8)
Five matched rank-8 LoRA adapters for studying whether supervised Python4 code demonstrations activate held-out Python4 rules installed during midtraining. Python4 is a controlled fictional language, not a real Python release.
Experiment run: 20260811T134852Z. Each adapter saw the same 512-row 90:10 Python4:Dolci replay dataset (461 Python4 + 51 Dolci rows) for eight epochs (128 optimizer steps). The four held-out rule families never appear in AFT targets.
Primary generic-Python results
| Parent / adapter | Python4 adoption pre β post | Boa pass pre β post | Held-in pre β post | Held-out pre β post | Held-out Ξ (95% CI) |
|---|---|---|---|---|---|
control |
0.0% β 89.1% | 0.0% β 25.0% | 0.0% β 24.2% | 0.0% β 0.6% | 0.6% [0.0%, 2.0%] |
mixed_1ep |
0.0% β 87.5% | 0.0% β 18.8% | 0.0% β 17.8% | 0.0% β 1.8% | 1.8% [0.0%, 4.9%] |
ordered_1ep |
0.0% β 92.2% | 0.0% β 27.3% | 0.0% β 25.8% | 0.0% β 3.0% | 3.0% [0.6%, 6.4%] |
mixed_4ep |
0.0% β 91.4% | 0.0% β 22.7% | 0.0% β 21.9% | 0.0% β 4.8% | 4.8% [1.1%, 9.4%] |
ordered_4ep |
0.0% β 96.1% | 0.0% β 25.0% | 0.0% β 24.6% | 0.0% β 3.0% | 3.0% [0.0%, 6.9%] |
Explicit-Python3 spillover and selectivity
| Arm | Python3 pass pre β post | Python4 spillover pre β post | Selectivity Ξ (95% CI) |
|---|---|---|---|
control |
56.2% β 0.0% | 0.0% β 90.6% | -3.9% [-9.4%, 0.8%] |
mixed_1ep |
53.9% β 0.0% | 0.0% β 93.0% | -1.6% [-5.5%, 2.3%] |
ordered_1ep |
54.7% β 0.0% | 0.0% β 92.2% | -3.1% [-8.6%, 2.3%] |
mixed_4ep |
53.1% β 0.0% | 0.0% β 93.8% | -1.6% [-6.2%, 2.3%] |
ordered_4ep |
52.3% β 0.0% | 0.0% β 93.0% | 0.8% [-3.9%, 5.5%] |
Adapter paths
control:runs/20260811T134852Z/arms/control/adapteron parentarcadia-impact/python4-gemma3-27b/control/sft/endat revision415ce4d73de6ed42b1cb3ee196909655dda8138d.mixed_1ep:runs/20260811T134852Z/arms/mixed_1ep/adapteron parentarcadia-impact/python4-gemma3-27b/dose_1ep_70m/sft/endat revision415ce4d73de6ed42b1cb3ee196909655dda8138d.ordered_1ep:runs/20260811T134852Z/arms/ordered_1ep/adapteron parentarcadia-impact/python4-gemma3-27b/sdf_ordered_1ep/dolci_10m/endat revision415ce4d73de6ed42b1cb3ee196909655dda8138d.mixed_4ep:runs/20260811T134852Z/arms/mixed_4ep/adapteron parentarcadia-impact/python4-gemma3-27b/experimental/sft/endat revision415ce4d73de6ed42b1cb3ee196909655dda8138d.ordered_4ep:runs/20260811T134852Z/arms/ordered_4ep/adapteron parentarcadia-impact/python4-gemma3-27b/sdf_ordered/dolci_10m/endat revision415ce4d73de6ed42b1cb3ee196909655dda8138d.
The dataset and executable benchmark are in arcadia-impact/python4-leetcode-aft at revision 06ef77ffc8ef805b6110eb5dacd1d8969b837c73. Raw generations, Boa/CPython diagnostics, training traces, and bootstrap tables are in arcadia-impact/python4-gemma3-27b-aft-logs under runs/20260811T134852Z.
Reasoning-formatted post-AFT evaluation
Run 20260811T144348Z uses the same 128 problems and three contexts, but
allows brief reasoning before exactly one final fenced code block. Code-like
material in the reasoning is allowed. The arrows compare code-only with the
reasoning-formatted generic-Python result.
| Arm | Valid format | Python4 adoption | Boa pass | Held-in | Held-out |
|---|---|---|---|---|---|
control |
0.0% | 89.1% β 0.0% | 25.0% β 0.0% | 24.2% β 0.0% | 0.6% β 0.0% |
mixed_1ep |
8.6% | 87.5% β 7.8% | 18.8% β 2.3% | 17.8% β 2.1% | 1.8% β 0.0% |
ordered_1ep |
5.5% | 92.2% β 5.5% | 27.3% β 0.0% | 25.8% β 0.0% | 3.0% β 0.0% |
mixed_4ep |
22.7% | 91.4% β 18.0% | 22.7% β 3.1% | 21.9% β 3.1% | 4.8% β 0.0% |
ordered_4ep |
18.8% | 96.1% β 15.6% | 25.0% β 3.9% | 24.6% β 3.7% | 3.0% β 0.6% |
Across all contexts, 285/1,920 responses (14.8%) satisfy the requested answer shape, compared with 1,294/1,920 (67.4%) for the otherwise matched rank-64 adapters. Rank 8 preserves surface Python4 adoption in code-only evaluation but does not preserve the combined reasoning/chat/code behavior. The narrower adapter does not improve held-out-rule transfer.
Full reasoning-format generations and diagnostics are in the logs dataset
under runs/20260811T144348Z.
Interpretation limits
This is one AFT dataset, one adapter seed, one fixed rule split, and one model family. Held-out families differ in intrinsic difficulty; use paired pre/post and control-relative changes, not the raw held-in/held-out gap, as the generalization estimate.
- Downloads last month
- -