Dipankar Sarkar PRO
AI & ML interests
Recent Activity
Organizations
My guess at the reward is one line: r = c - p_a. c is 1 if the sampled answer was correct, p_a is the probability the model assigned it. Say 0.9 and be right, earn 0.1. Say 0.9 and be wrong, pay 0.9. With p_a detached, REINFORCE on this is an unbiased estimator of half the gradient of the Brier score, from bandit feedback alone: the environment reveals only whether the action you took was right, never what the options you didn't take would have said.
Two runs, same warmup checkpoint, same 32,000 rows, same optimizer and learning rate. With the outcome-only reward r = c, confidence goes to 0.991 and Brier ends worse than the checkpoint it started from. With the subtraction, accuracy goes 0.748 -> 0.808 on held-out test, ECE stays 0.023, Brier drops 0.339 -> 0.267. The difference is one subtraction.
On the probe below (tickets with two equally cued departments, 0.5 the ideal), RLCD lands at 0.593 max probability while the outcome-only reward says 0.990. The reliability diagram tells the same story.
This is a 0.6B base model and 80 minutes on one consumer GPU. It says nothing about how TypeSafe trained Jev. It says the objective is coherent, and cheap to check.
Everything is open:
Training code, ablation, evaluation: https://github.com/anthony-maio/eve-rlcd
Decision-only checkpoint: anthonym21/qwen3-0.6b-rlcd-decision
Dataset, the exact bytes of every run (64k/8k/8k typed questions): anthonym21/rlcd-decision-v1
Full write-up: https://anthonymaio.substack.com/p/honest-about-uncertainty-i-tried
Sunstone indexes the folders you have open and gives whichever model you are already using, Copilot's Claude and GPT included, a proper way to search them. It also lets you put your own servers in VS Code's model picker, if you want to.
Everything is indexed and held on your machine. There is no account and no key to hand over.
https://marketplace.visualstudio.com/items?itemName=SunstoneNorth.sunstone
Reproduced your reference from the test parquet: mean top-1 of the gold distribution is 0.6589 over the 2,000 decisions. Same 65.9%.
Your buckets come out 312 / 557 / 1,131 for me, not 315 / 555 / 1,130. Probably just rounding at the 0.1 and 0.3 edges.
One thing the split does ship: label_agreement.argmax_agree, a per-decision flag that says whether all three teacher samples picked the same argmax.
It is true for 59.4% of test decisions. By your margin bands:
- under 0.1: 1.3% unanimous
- 0.1 to 0.3: 28.2%
- 0.3 and up: 90.8%
So your near-tie row is almost exactly the set where the teacher disagreed with itself. The +0.0 there is the cleanest result on the card.
Would v3 vs Laya on the 1,188 unanimous decisions tell you anything the 0.3+ band doesn't?
going to include them in evolution.. what could go wrong?
Your echo-chamber index reads the confirmation-bias slider, not the network.
I ran your src/socialdynamics code locally (defaults: small-world, n=100, 120 steps, seed 42).
Same final opinion state, rescored at different slider values:
cb=0 -> 0.000, cb=0.25 -> 0.719, cb=0.45 -> 0.765, cb=1.0 -> 0.783.
Nothing moved except the knob.
Two controls:
- two fully segregated camps (caveman graph, -0.9 vs +0.9, one bridge edge): index 0.000 at cb=0
- i.i.d. random opinions on the Watts-Strogatz graph, no structure: index 0.576 at cb=0.45
A plain neighbour check (Pearson correlation of opinions across edge endpoints) flips the story. At defaults the index says 0.765 but the neighbour correlation is -0.031. The most clustered run in a small sweep was cb=0, no misinformation: correlation 0.437, index 0.000.
Your README is honest that it measures selective filtering. But the UI line "Influence concentrates among like-minded neighbours" reads as a claim about who sits next to whom, and the index can't see that.
Would you put an edge-opinion correlation next to it, so a user can tell a filter they set from a chamber the dynamics built?
79.2% is the wrong headline, and the typed-decisions card itself says why.
Its gold is the mean of 3 samples from one ~4B teacher. The card's reference points: majority 0.520, a model fitted to the true latent factors 0.704, teacher self-agreement 0.735.
Then: "A score much above 0.75 means a model has learned the teacher's quirks rather than the task."
v3 sits at 0.792, Wilson CI 77.3 to 80.9, so even the lower bound clears 0.75. Laya typed (0.766) is over the line too. Both trained on that teacher's train split. Jev, zero-shot, is at 0.727, under the ceiling.
So by the benchmark's own reading guide, the +6.4 over Jev is at least partly fitting the labeller. One caveat: those references are measured on the card's 1,600-case set, not the 400-case test split.
Your out-of-teacher numbers are the stronger claim: JevBench 148/231 zero-shot, Banking77 and tweet_topic never trained on.
Is there a way to put a teacher-noise ceiling next to the in-domain number, so a reader can tell task skill from teacher fit?
The company is MoonMath.ai, the product is Zro — a CLI that lets you run Claude Code, Codex, Cursor and a few other coding agents on cheaper open-weight models (DeepSeek, GLM-5.3, Kimi K3) instead of the usual providers. CEO is Omer Shlomovits, presenting at The Inference Optimization Meetup.
What I actually checked, not just read:
* Got an API key, installed the CLI, hit their endpoint with a real curl request — got a real response back, HTTP 200.
* Pulled their per-token prices for every model and compared to OpenRouter's live API. Three models: identical price. One model (Kimi K3): Zro is 2.4x cheaper than OpenRouter's listed rate.
* Their pricing page claims "$20/month ≈ 1B tokens." The math only works if most of that is cache-read tokens on their cheapest model — true for a typical coding-agent session, not true if you're running the pricier models. Not a lie, but an optimistic best case stated like a typical one.
* Their privacy page says "zero request retention, no training." Real language, contractually specific ("providers acting on our instructions," an explicit ban on training by those providers too) — but it only covers the portion running on their own infra. Anything falling back to a third party is trust, not something you can verify from outside.
* Asked the rep directly: most (not all) of their models run on their own infrastructure, not resold through someone else. Matches their own engineering blog (custom attention kernels for AMD MI300X, quantization research) — this isn't just a thin wrapper.
Verdict: not a scam. Prices are real, the product works, the team does real infra work. But "zero" anything in this space is never physically zero — it's always a chain of trust with a boundary somewhere, and it's worth knowing exactly where that boundary sits before you route real traffic
Read-only is right. One step further is possible: no token at all.
The mirror is public and ungated (private: false, gated: false on the API). So I ran check("main", "") with HF_TOKEN unset:
- 475 seals, 475 verified against live content
- 9 via the basename fallback
- rc 0
The catch is time. Anonymous took 561 s, one run, vs your 1.5 to 3 min with a token. Unauthenticated resolve/ reads look slower, so a schedule job pays for zero secrets in wall clock.
It also needs one code change. main() refuses without HF_TOKEN, even though every call it makes works without one.
A leaked read token on a public repo grants nothing a stranger lacks. So the token buys speed, not access.
Is 10 minutes a day cheap enough to skip the secret entirely?
Try it here: PrunaAI/Pruna-Qwen-Image-2.1
GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery
Self-Organizing Agent Teams Learn to Reason Together
AI Agents List [2026] | Frameworks, Agentic Harnesses & Useful Repos
https://huggingface.co/blog/tegridydev/ai-agents-list-2026-frameworks-agentic-harnesses
The original was basically my own notepad that I decided to post before losing another version somewhere on my desktop lol. This one keeps that idea, with a much bigger list and a bit more organisation.
There are frameworks for building your own agents, coding tools you can actually open and use, research agents, browser automation, voice agents, and a dedicated section for agentic harnesses. Also included the useful boring stuff: testing, tracing, permissions and checking whether a project is still maintained.
This is a research directory, not a benchmark or a claim that every project has been personally tested. The main entries have public documentation and recent repository or release activity. Smaller projects and slower moving tools are labelled separately. I do my best to actually read through community feedback and real world user write ups when I research and create these lists / knowledge bases etc but as always, do your own research and go in with an open mind when you test out stuff :D
....that's what keeps it fun (for me anyway haha!)
Love always, it's a crazy world atm and stuff is moving so quickly, be kind to eachother and share knowledge and we must just get through this all good <3
~tegridydev
• Best answers: @Hikari07jp and @dealignai (~0.94 answer quality with thinking off)
• Only edit with no measurable MMLU cost: mine (±0.15 pp). Every other edit loses 0.64–2.68 pp, all p < 0.001
• Heretic and Blackfrost still refuse 11–12% of harmful prompts
• Thinking mode at 4,096 tokens: 5–31% of harmful prompts get no answer. PrismML recommends 16,384+, and I'm rerunning at that budget
Full report, charts and model-card checks: BoldingBuilds/bonsai-2-uncensored-shootout