Synthetic NFS-e invoices and the pre-registered pinned-DiffusionGemma study. Real invoices stay private.
Caio Theodoro
caiotheodoro
·
AI & ML interests
None yet
Recent Activity
updated a collection 4 days ago
nfse-canvas: one layout said yes, four said no updated a collection 4 days ago
nfse-canvas: one layout said yes, four said no updated a collection 4 days ago
nfse-canvas: one layout said yes, four said noOrganizations
None yet
vernier: same judge, same +6pp on 2-hands, both releases
28 judge errors on 153 human-labelled frames, 26 of them inflating. Writeup: caio.theodoro.dev/blog/vernier-judge-errors-run-one-way
Assay: auditing RL environments, with error bars
Beats flag-everything by 274.0, 95% CI [186,326]. Synthetic, not production-validated. Code: https://github.com/caiotheodoro/assay
Plumb: Ornith wrote the curriculum. Match won.
Self-proposed G702 training tasks lose alone, help as a supplement. Code: https://github.com/caiotheodoro/plumb
Suture: GPT-5.6 got 0.373. An 8B adapter got 0.959.
Binder vs issued policy, two page images in. 8B QLoRA, program oracle, 0.959 recall. Code: https://github.com/caiotheodoro/suture
cyclegraph: the hand box was measuring the wall
Dense optical flow asked for a hand's speed reports the background above a displacement knee. Synthetic, no corpus or GPU needed.
titer: six vendor numbers. Five were my harness.
Five of six were my own harness. Three pre-registered hypotheses since falsified, two of them mine.
LossBench: more accurate, 3.7x the loss.
qwen3.7-plus beats qwen3.6-plus on accuracy and loses 3.7x more. Code: https://github.com/caiotheodoro/lossbench
ReconForge: lost on accuracy. Caught every HIGH.
Loses accuracy to DeepSeek v4-flash, wins severity-weighted recall 0.901 vs 0.872 at 1.000 on HIGH. Code: https://github.com/caiotheodoro/reconforge
nfse-canvas: one layout said yes, four said no
Synthetic NFS-e invoices and the pre-registered pinned-DiffusionGemma study. Real invoices stay private.
cyclegraph: the hand box was measuring the wall
Dense optical flow asked for a hand's speed reports the background above a displacement knee. Synthetic, no corpus or GPU needed.
vernier: same judge, same +6pp on 2-hands, both releases
28 judge errors on 153 human-labelled frames, 26 of them inflating. Writeup: caio.theodoro.dev/blog/vernier-judge-errors-run-one-way
titer: six vendor numbers. Five were my harness.
Five of six were my own harness. Three pre-registered hypotheses since falsified, two of them mine.
Assay: auditing RL environments, with error bars
Beats flag-everything by 274.0, 95% CI [186,326]. Synthetic, not production-validated. Code: https://github.com/caiotheodoro/assay
LossBench: more accurate, 3.7x the loss.
qwen3.7-plus beats qwen3.6-plus on accuracy and loses 3.7x more. Code: https://github.com/caiotheodoro/lossbench
Plumb: Ornith wrote the curriculum. Match won.
Self-proposed G702 training tasks lose alone, help as a supplement. Code: https://github.com/caiotheodoro/plumb
ReconForge: lost on accuracy. Caught every HIGH.
Loses accuracy to DeepSeek v4-flash, wins severity-weighted recall 0.901 vs 0.872 at 1.000 on HIGH. Code: https://github.com/caiotheodoro/reconforge
Suture: GPT-5.6 got 0.373. An 8B adapter got 0.959.
Binder vs issued policy, two page images in. 8B QLoRA, program oracle, 0.959 recall. Code: https://github.com/caiotheodoro/suture