Dipankar Sarkar's picture
🏗️ Building on HF

Dipankar Sarkar PRO

dipankarsarkar

AI & ML interests

Building the AI-native stack. Agents as infrastructure, safety as architecture, performance as plumbing. I publish the receipts: papers, datasets, demos.

Recent Activity

reacted to anthonym21's post with 🔥 2 days ago
TypeSafe launched Jev with a claim and no method: a model trained with "reinforcement learning for calibrated decisions," whose stated probabilities are supposed to match how often it turns out to be right. No reward function, no training procedure, no calibration curves. So I rebuilt the objective from the announcement alone and tested it at a scale I can fully fine-tune on one RTX 4080. My guess at the reward is one line: r = c - p_a. c is 1 if the sampled answer was correct, p_a is the probability the model assigned it. Say 0.9 and be right, earn 0.1. Say 0.9 and be wrong, pay 0.9. With p_a detached, REINFORCE on this is an unbiased estimator of half the gradient of the Brier score, from bandit feedback alone: the environment reveals only whether the action you took was right, never what the options you didn't take would have said. Two runs, same warmup checkpoint, same 32,000 rows, same optimizer and learning rate. With the outcome-only reward r = c, confidence goes to 0.991 and Brier ends worse than the checkpoint it started from. With the subtraction, accuracy goes 0.748 -> 0.808 on held-out test, ECE stays 0.023, Brier drops 0.339 -> 0.267. The difference is one subtraction. On the probe below (tickets with two equally cued departments, 0.5 the ideal), RLCD lands at 0.593 max probability while the outcome-only reward says 0.990. The reliability diagram tells the same story. This is a 0.6B base model and 80 minutes on one consumer GPU. It says nothing about how TypeSafe trained Jev. It says the objective is coherent, and cheap to check. Everything is open: Training code, ablation, evaluation: https://github.com/anthony-maio/eve-rlcd Decision-only checkpoint: https://huggingface.co/anthonym21/qwen3-0.6b-rlcd-decision Dataset, the exact bytes of every run (64k/8k/8k typed questions): https://huggingface.co/datasets/anthonym21/rlcd-decision-v1 Full write-up: https://anthonymaio.substack.com/p/honest-about-uncertainty-i-tried
repliedto chaoliangUNSW's post 2 days ago
Jev-style decisions on your laptop: Jev-Style-0.8B-Decision-v3 Give it any text and a question; it returns a calibrated probability for every option in one forward pass: pick one, yes/no, or a score. No text generation. On data it never trained on: • Banking77, all 77 intents in one pass (400 test rows): 68.2% vs 49.2% for the best official Laya checkpoint (+19.0, paired 95% CI +14.0 to +24.0) • tweet_topic, zero-shot: 75.5% vs 63.2% for English Laya, within 4 points of Jev (79.3%), and better calibrated than Jev (ECE 0.027 vs 0.063) • JevBench v1.4.1, 231 public items, zero-shot: 64.1%, ahead of every Qwen3.5-0.8B-based system on the board; ahead of Laya (58.4%) but inside the CI Jev is still ahead on all three, and decider-2b scores higher on JevBench (71.0%). v3's case is size: 0.53 GB in 4-bit, runs on CPU. Typed decisions (in-domain): 79.2% vs 76.6% for Laya typed, both trained on the same train split (paired CI +1.0 to +4.2). Read this as agreement with the dataset's teacher labels, not decision skill: • teacher-noise reference: one draw from the teacher's own distribution matches the gold label 65.9% of the time (test split); teacher self-agreement is 73.5% (dataset card) • on the teacher's near-ties, v3 and Laya typed score the same (51.7% vs 51.7%); v3's lead comes from cases where the teacher is clear • Jev's 72.7% there is zero-shot, so not a like-for-like comparison Also: up to 25,600 tokens per call (98.3% on real 24K-token documents); ~44 ms per short decision on an M1 Max (llama.cpp); one read of a document answers every question in one call; ahead of Laya multilingual in 51 of 51 languages. pip install "jev-style[torch]" · https://github.com/lawrence3699/jev-style Try it in your browser: https://huggingface.co/spaces/chaoliangUNSW/jev-style-v3 Model: https://huggingface.co/chaoliangUNSW/Jev-Style-0.8B-Decision-v3 Apache-2.0. Not affiliated with TypeSafe, Jev or Laya. Thanks to @dipankarsarkar for pushing on the typed-decisions framing.
View all activity

Organizations

Skelf Research's profile picture Neul Labs's profile picture Cognisoc's profile picture Incredlabs's profile picture