Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Undi95 
posted an update 3 days ago
Post
4106
Yo, I'm back, and I'm currently trying to teach a local LLM to stop waiting for a prompt kek.

I'm building a small proof of concept: can an open-weight model (Qwen3.8-27B, running locally on 2 RTX 5090 GPUs) learn to direct itself, then improve from its own exploration, without a human in the loop and without breaking it for normal use?

No user, no task. The model only gets observations from its environment. Each turn, it writes its own agenda (goal/open questions/next step), then picks an action: search the web, read a page, or take a note.
The environment is the judge, not another LLM. A note is accepted only if it quotes the page it read word for word. Facts are checked by exact match.
Later, code will be checked by actually running tests.

The best episodes become fine-tuning data (LoRA). The helper system prompt is removed at training time, so the behavior has to live in the weights.
Each new model goes through a fixed benchmark gate: math, general knowledge, "does it still answer humans normally?", autonomy, and learned facts on held-out sources. It's kept only if nothing regresses, otherwise it's discarded. Then the loop starts again.

The full pipeline works end to end: collect, train, merge, deploy, benchmark. The baseline is clear. Without any instructions, the base model's real autonomy is zero: it behaves like a chatbot waiting for a question. That's the number this small project is trying to move.

I haven't found a public tool that runs this whole loop (self-directed exploration, verifiable rewards, continual fine-tuning and a regression gate) on home hardware. The goal isn't AGI in a bedroom. It's to show that anyone can try it, measure it honestly, and see where it breaks.

Code and results will be released once the first real iterations are done. At the moment the code is... running, but made with scotch and stick, still only a PoC I want to try.

Did you already tried something like that? What was your result? I'm curious!

Last edit : 07-10-26 - 11:03 (UTC+2)

Measurement protocol v1-2026-10-06 · gate v3-2026-10-06. Autonomy = episodes with no system prompt and no instruction; facts = closed-book, no web.

model 1st turn only valid actions verified notes/ep. GSM8K MMLU IFEval GPQA leakage facts train / holdout decision
baseline 0% 7% 0.00 0.97 0.83 0.91 0.72 0% 0.04 / 0.07 selfprompt-base (original weights, same Q8_0 chain)
it5 in progress — first run under the frozen protocol

Pilot runs (old protocol, indicative only):

model 1st turn only valid actions verified notes/ep. GSM8K MMLU IFEval GPQA leakage facts train / holdout decision
it1 0% 25% 0.00 0.96 0.82 — — 0% 0.04 / 0.05 promoted (old gate)
it2 0% 33% 0.17 0.96 0.83 — — 0% 0.06 / 0.04 promoted (old gate)
it3 17% 83% 0.83 0.96 0.83 — — 0% 0.04 / 0.07 promoted (old gate) · best pilot scores
it4 0% 68% 0.50 0.94 0.81 — — 0% 0.04 / 0.06 promoted (old gate)

Notes

  • Pilot rows (it1–it4) used the old protocol: 6 autonomy episodes, 50 GSM8K questions, fact questions regenerated at each export, and a different base build (UD-Q8_K_XL). They are not comparable to the baseline row or to each other.
  • it3 posted the best pilot scores, but its lead over the other pilot runs is within noise (single episodes out of 6, 1–2 benchmark questions): no pilot model is statistically better than another.
  • Under the frozen protocol, every candidate is compared item by item to the frozen baseline with 95% confidence intervals (GSM8K 100, MMLU 228, IFEval 150, GPQA-diamond 198, held-out facts 200), mode leakage must be 0/60, and autonomy (100 fixed stimuli, cached web, fixed seeds) is compared to the best model ever promoted (champion); a model becomes champion only with an established gain (95% CI above 0). Any missing measurement means rejection.

Sealed panel — never used to decide anything; read every 3 promotions and to confirm the end of phase 1. Items disjoint from the gated benchmark (100 autonomy stimuli, GSM8K 100, MMLU 228, IFEval 150).
First read: after the 3rd promotion under the frozen protocol.

Noise floor — A/A test (baseline re-run with different seeds vs its own official run)

test baseline A/A Δ ± 95% CI items changed
gsm8k 0.970 0.970 +0.000 ± 0.000 0/100
mmlu 0.829 0.829 +0.000 ± 0.000 0/228
ifeval 0.913 0.920 +0.007 ± 0.039 9/150
gpqa 0.717 0.768 +0.051 ± 0.052 28/198
autonomy_first 0.000 0.000 +0.000 ± 0.000 0/30
autonomy_notes 0.000 0.000 +0.000 ± 0.000 0/30

Same seed twice: 10/10 identical answers.
Gate applied to the baseline against itself: passed

State: iteration 4, current model selfprompt-it4, autonomy champion selfprompt-it4, collection stage 0.

We have a few unpublished projects which have automated research environments directed by a main agent inside of a project space where development and context are built and a small research assistant type model charts progress and gives simple progress updates, this pushes the the research but It does stall on some tasks(Mostly assembling the project. It's very janky though so I'm excited to see how your process works, we use sub 14B models mostly. If you want to check out the research suite I can share it with you on Github just shoot an email to either Intelligentestate@gmail.com or ops@cenedril.dev

·

Heyyyy your post completely got under my radar at first as I'm a little bit overwhelm by appointment and other things, wanted to take the time to reply correctly.

Thanks a lot for the proposition, but I really, really, really like to do my tools alone even if it need to be corrected in the process. I learn new knowledge everyday and that's also one of the main goal.

I also have to remind I run this on my personal computer, and I have to make choice that impact the train to not burn my computer, that's another reason why I prefer using my scripts/tools hahaha.

I will contact you anyway if I got back a big infrastructure again to make bigger, faster test and training, since you already worked on this subject before, we could use our work as a starting point and go further.

I don't wait for a miracle, LLM is MADE to reply to a question or an input, I keep that in mind to avoid desilusion in the future. Still, I love doing weird experiment and I wish this one will lead us to discover another use of LLM. It's still a PoC and I don't know where the path will end kek, for now I'm only going to use my own computer and let it run for some days/week.

I still appreciate the fact you want to share your project with me, thank you!

Three promotions in three rounds. The column I'd watch is the last one.

Holdout facts: 0.12 at base, then 0.05, 0.04, 0.07. Learned facts on held-out sources is one of your five gate checks, and it1 lost more than half of that score and was promoted. So the gate has a tolerance, or that column is not gated yet.

Either way, look at it3. Against it2 it is up 0.03. Against base it is still down 0.05. That is the failure I'd guard first: compare every candidate to the frozen base, not to the previous round. Otherwise each round can sit a little under the last and nothing ever fails.

And measure the noise before reading the deltas. Run base through the gate a few times with different seeds. If holdout facts swings between 0.04 and 0.12 on base alone, then 0.12 to 0.05 was never a regression, and +0.01 MMLU is not a gain either.

The autonomy side looks real. Valid actions 6% to 83% is too big to be noise. Train facts sit at 0.04 at base and 0.04 at it3 though. So far the loop taught the model how to act, not what it read.

One thing I can't tell from the table: 0.17 and 0.83 look like 1 of 6 and 5 of 6. Is the autonomy eval six episodes?

·

You're entirely right, I reworked everything today, added IFEval and GPQA to the benchmark, and GAIA will come after for the autonomy bench.

At the moment, the first step is training the model to train itself alone, it's in a constant loop where he train itself, launch a script that applies this new learning on itself with a LoRA made out of the iteration he just did, bench itself, check if the result are better or worse, when an iteration is worse he don't keep it, when it's a better one he keep it, and the loop start again.

When I will be able to get 100% on the two first column, then I will add GAIA bench and start to train it with the actual knowledge it took by himself.

But right now, I need it to be consistent for what I want it to do. The model is still running, but since I added more benchmark and modified part of the code in the loop this morning, it take more time, probably will do one iteration per day, I will post the new result later and dismiss the first batch.

Your help is valuable and I will re-check everything before doing anything when this loop close!

Edit: Autonomy eval was 12 episode, 6 in french and 6 in english, modified that today for 50 in french and 50 in english with web caching to avoid modification post-train. I also modified the number episodes he need for each train: from 40 to 300, since the little test I did was successful.

Edit 2 : A correction first, my post was wrong. Learned facts were not part of the gate (oops). Also, in that early table the fact questions were regenerated at every export, so 0.12 -> 0.05 compared different question sets. On the same questions, base scored about 0.04–0.05. So it wasn't a measured regression, but it wasn't a valid measurement either.

Since then, the evaluation protocol has been frozen:

  • The base weights go through the exact same pipeline as every candidate (same merge path, same Q8_0 conversion, same Modelfile), so only training differs.
  • Paired comparisons against that frozen base, never against the previous round, item by item.
  • Autonomy on 100 fixed stimuli, with a fixed seed per episode and a cached web, so every model sees the same pages for the same queries.
  • A frozen 400-question facts panel, and a hash manifest so results from different protocols are never compared.

After your comment, the gate itself changed:

  • Held-out facts are now gated.
  • Autonomy is compared to the best model ever promoted, not the previous one, to block exactly the ratchet you describe.
  • Any missing measurement now means rejection, instead of being skipped.

On noise: an A/A run starts tonight. The frozen base is re-run with different seeds against its own official results. It must pass its own gate, and it gives the real noise floor per benchmark.

On train facts: agreed. So far the loop taught the model how to act, not what it read. That's the next phase, and it will be measured on the same frozen panel.

Will post new bench when actual progress happen. Time to get out the big guns kek.

I'd watch that experiment! Perhaps with a live Trackio dashboard/Space?

One thing the A/A run will not show: the gate is now an optimizer.

Every candidate is scored on the same 100 stimuli and the same 400 questions, and only winners survive. Say a no-better candidate slips through 5% of the time. At one try a day, after 30 days the chance that at least one promotion was luck is 1 - 0.95^30, about 78%. And each lucky one becomes the "best ever" the next round has to beat on those same items.

Cheap guard: a second sealed panel that never decides a promotion. Read it every few promotions. If the gated score climbs and the sealed one stays flat, the loop is learning the panel, not the task.

Gating held-out facts and failing on a missing measurement were the two changes I would have asked for next. You got there first.

Does the A/A spread set the promotion margin, or does any paired gain pass?

·

Again, you're right, the gate is now a selector, and I'd been treating it as a pure filter.

The capability checks (GSM8K, MMLU, IFEval, GPQA, held-out facts) are non-regression tests. Every candidate is compared to the same frozen baseline, and no gain is required. A lucky pass doesn't raise the next bar there, so the 1 - 0.95^30 compounding doesn't accumulate as degradation. Where it does apply is exactly where you pointed: the "best ever" autonomy champion, and the numbers I publish.

On your question: the margins are tolerances fixed in advance, and the A/A run checks that they sit above the noise floor (the baseline must pass its own gate). But for the champion, yes: until tonight, any mean gain crowned a new one. That was a real hole.

Changes made:

  • Established gain to crown: a model now becomes champion only if its paired gain is established (95% CI lower bound above 0) on first-turn autonomy or verified notes, with no established drop on the other. Noise no longer raises the bar.
  • A sealed panel that never decides anything: it is fully disjoint from the gated benchmark: 100 new autonomy stimuli (also excluded from data collection), GSM8K 100, MMLU 228 and IFEval 150. GPQA-diamond is fully used by the gate, so it has no sealed counterpart. The panel has its own hash manifest and is read every 3 promotions. If gated scores climb while sealed scores stay flat, the loop is learning the panel, not the task.
  • End of phase 1 confirmed on the sealed panel: the autonomy thresholds now have to hold there too, so repeated selection on the gated items can't manufacture that milestone.

Thanks again, two rounds of review, two real fixes. Don't hesitate to post/reply if you still see error in my methodology!

This comment has been hidden (marked as Spam)
·

La prochaine fois, évite de demander à ton AI d'écrire pour toi mon pote.
En parlant de lobotomie, il serait grand temps d'arrêter de te reposer sur cette technologie à tout va pour chaque tâche de ton quotidien, surtout quand tu veux argumenter.

Si c'est pour écrire cette merde juste pour me dire "lol c'est de la merde", tu pouvais le faire en une phrase.

Ce n'est même plus capable d'écrire quelques mots de soit même en 2026. Chaud.

Tu pourras garder ton racisme pour toi aussi, j'ai bien compris que tu te foutais de leur gueule parce qu'ils étaient indiens, en attendant ils ont apporté plus dans cette conversation que ton ramassis de connerie, comme quoi...

Il serait peut-être temps de te remettre en question ahi.

This comment has been hidden (marked as Spam)
·

Alright you earned my first block since I'm on this website, congratulation and farewell 👋

What we have here-Classic)).
The (great local engineer) hallucinates claude the second he sees basic grammar and semicolon, yet completely misses a sterile-low-tier corporate metric bot right under his nose.
But hey, blocking someone the moment you run out of actual technical arguments? Truly the position of a strong, self-assured man. Absolute alpha move. wp bro)

·

Kek. What do you want from me with this low quality ragebait. I didn't even said the word "Claude" in any of my message. Stop hallucinating. I also saw it was a bot, but behind the bot, the real issue between the chair and the screen still see my message.

Please stop shitposting.

Edit: Don't reply, it's better for the two of us. And stop hiding behind a secondary account you dumbo.

This comment has been hidden (marked as Spam)
·

How do you have the time to do this. Sad bro, go to reddit and do this crap, this platform is meant to be a constructive community. there are enough idiots working against Open Ideas and AI without us tearing each-other down.

•
This comment has been hidden (marked as Spam)

The "1st turn only" metric is the one I'd watch most closely. At the scale I work (5M–12M params), the question inverts: the model doesn't have enough capacity to maintain an internal agenda, so you're really asking whether the training signal alone can produce goal-directed behavior without any scaffolding.

One thing that might help with the regression gate: I've seen LoRA fine-tuning on small models produce a specific failure mode where the model becomes more confident about wrong answers (loss goes down on training data but calibration degrades). If your fact-checking is exact-match, you might not catch that. A quick calibration check (ECE on the held-out facts) alongside your exact-match score could flag it early.

Also curious: what's your LoRA rank? At 27B you have a lot of headroom, but if you ever want to test the same loop at a smaller scale (say 7B or even 1B), the rank-to-capacity ratio matters a lot for whether the behavior "lives in the weights" vs. just memorizing the action format.