tabfix-preview / ARCHITECTURE.md
Antix5's picture
Document v2 architecture and CSV usage; mark previous evaluation as historical
a9fe228 verified
|
Raw History Blame Contribute Delete
5.26 kB
# TabFix v2 — visible original and deterministic routing
This architecture uses the existing `Antix5/tabfix-preview` repository. Dataset revision: `602b8a454ca448713d8a5658650d4feaad8d579c` in `Antix5/tabular-errors-v1`. Earlier checkpoints retain their own source and label maps in repository history.
## Inputs and routing
Declared enums, patterns, numeric syntax, calendar dates, case, trimming, requiredness, conditional requiredness and supported structured relations are checked by code. The rule result is valid, invalid or unknown. An invalid relation identifies its participating cells, not necessarily the faulty member; automatic independent-cell repair abstains for those groups.
The neural detector has two categories: `text.encoding` and `text.spelling`. Closed enum/lexicon cells do not require it. Additional training views remove those constraints from existing authored spelling/encoding examples; matching clean controls retain unusual valid text. These views stay in their original family/split. No generalization claim beyond this training distribution is implied.
One pretrained mmBERT backbone feeds BIO heads (two independent three-state decisions), insertion heads (two binary decisions), and its pretrained MLM vocabulary head. The 18 business categories remain available for deterministic reporting and user controls.
## Correction
```xml
<cell column="Date">
<original>2026/09/20</original>
<replacement>[MASK][MASK]…</replacement>
</cell>
```
The actual serialization adds no whitespace around the original. It preserves its characters and XML entities. All context remains visible. No numeric repair request or gold answer is appended. The output is a complete replacement cell, not a token-aligned substring. END_EDIT terminates; EDIT_PAD fills its tail; ordinary batch padding is excluded from attention and loss.
The source-derived capacity is the larger of eight tokens and the original token count plus 20% (at least two extra slots), capped at 256 content tokens, followed by one termination slot. Over-budget training targets are rejected with a counted reason; inference can retry with 256 slots. An immediate END_EDIT deletes the value. Literal tokenization preserves spaces and rejects reserved/non-round-tripping text.
Correction sampling uses 30% certified valid copies and 70% repair targets. Compound errors in one cell are combined into its complete target. Missing values lacking recoverable contents and linked structural swaps are not independently supervised for correction.
## Joint loss
Let K=2 and V=256000. Positive and background groups are averaged separately for detection; unavailable labels are ignored.
$$
L_{det}=L_{BIO}+0.5L_{insertion},\qquad
L_{cor}=L_{content+END}+0.1L_{EDIT\_PAD}.
$$
$$
L=L_{det}+\frac{\ln K}{\ln V}L_{cor}.
$$
The logarithmic ratio is the retained heuristic scaled to the new label inventory, not a proof of equal gradient contributions. Detection and correction use separate input views and contribute gradients to the same optimizer update. Half of training correction views are fully masked; other views expose a random subset of the target, as denoising supervision. Detection never receives revealed targets.
## Validation, recovery and budget
The run starts from the pretrained backbone, not the previous fine-tuned checkpoint. Batch size is four for each task. AdamW starts at 2e-5 with 100 warmup steps and cosine decay over 5000 steps to 10% of the peak. Gradients are clipped at one.
A fixed 28-example validation panel mixes residual neural and general cases. Full-mask and partial-mask correction losses are recorded separately. Validation runs every 300 updates or 600 seconds. Four checks without improvement of at least 0.001 in the joint validation score trigger early stopping. The entire best checkpoint, including optimizer and RNG state, is restored. The small panel is a development signal, not an accuracy guarantee.
Checkpoints are atomically uploaded to the same model repository. The final validation report and model README are produced by the Job. The L40S Job has a four-hour timeout, limiting compute to $7.20 at $1.80/hour, within the authorized $8.
## Perplexity and application
Each candidate token, including END_EDIT, is masked in turn and scored with the original and other candidate tokens visible:
$$
PPPL(y\mid x,C)=\exp\left[-\frac{1}{N}\sum_i\log p_\theta(y_i\mid x,C,y_{\setminus i})\right].
$$
Padding is excluded. Calibration selects the largest validation-score prefix with no observed incorrect exact-match repairs, treating ties together; a zero threshold accepts none. This small-sample operating point is reported, not asserted to guarantee correctness. Unfinished outputs, scores above threshold and outputs violating enabled deterministic rules are rejected.
CSV inference supports unambiguous case/trim/boolean canonicalization directly. Other selected cells use the model and the checks above. Whole-cell model repair currently abstains when business categories are disabled, because this training recipe does not condition generation on the enabled set. Deterministic category switches remain available. Structured contradictions with an ambiguous faulty member remain reported and unchanged.