-
LiquidAI/LFM2.5-1.2B-Instruct
Text Generation • 1B • Updated • 495k • 641 -
Token-Level LLM Collaboration via FusionRoute
Paper • 2601.05106 • Published • 40 -
ryokamoi/Qwen-2.5-7B-FoVer-PRM-old
Text Generation • 8B • Updated • 9 • 1 -
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 924k • • 953
Jim White PRO
jimwhite
·
AI & ML interests
None yet
Recent Activity
updated a collection 9 days ago
LLM updated a collection 3 months ago
Coding Benchmarks liked a Space 3 months ago
webml-community/Gemma-4-WebGPUOrganizations
RL
-
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
Paper • 2512.17008 • Published • 11 -
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Paper • 2601.05242 • Published • 235 -
ryokamoi/Qwen-2.5-7B-FoVer-PRM-old
Text Generation • 8B • Updated • 9 • 1 -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
Paper • 2601.18778 • Published • 43
PUP
-
DeepCode: Open Agentic Coding
Paper • 2512.07921 • Published • 35 -
nvidia/Nemotron-Pretraining-Code-v2
Viewer • Updated • 836M • 3.17k • 129 -
BEAVER: An Efficient Deterministic LLM Verifier
Paper • 2512.05439 • Published • 36 -
codefuse-ai/C2LLM-7B
Feature Extraction • 8B • Updated • 172 • 10
Verified Agents
Coding Benchmarks
-
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Paper • 2511.18538 • Published • 306 -
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
Paper • 2511.05459 • Published • 5 -
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
Paper • 2512.18470 • Published • 12 -
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
Paper • 2601.09688 • Published • 128
Semantic Web
-
josancamon/kg-gen-MINE-evaluation-dataset
Viewer • Updated • 101 • 246 • 6 -
zilliz/semantic-highlight-bilingual-v1
Token Classification • 0.6B • Updated • 1.25k • 99 -
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
Paper • 2601.09688 • Published • 128
LLM
-
LiquidAI/LFM2.5-1.2B-Instruct
Text Generation • 1B • Updated • 495k • 641 -
Token-Level LLM Collaboration via FusionRoute
Paper • 2601.05106 • Published • 40 -
ryokamoi/Qwen-2.5-7B-FoVer-PRM-old
Text Generation • 8B • Updated • 9 • 1 -
nvidia/nemotron-3.5-asr-streaming-0.6b
Automatic Speech Recognition • 0.6B • Updated • 924k • • 953
Verified Agents
RL
-
Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs
Paper • 2512.17008 • Published • 11 -
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Paper • 2601.05242 • Published • 235 -
ryokamoi/Qwen-2.5-7B-FoVer-PRM-old
Text Generation • 8B • Updated • 9 • 1 -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
Paper • 2601.18778 • Published • 43
Coding Benchmarks
-
From Code Foundation Models to Agents and Applications: A Practical Guide to Code Intelligence
Paper • 2511.18538 • Published • 306 -
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
Paper • 2511.05459 • Published • 5 -
SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
Paper • 2512.18470 • Published • 12 -
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
Paper • 2601.09688 • Published • 128
PUP
-
DeepCode: Open Agentic Coding
Paper • 2512.07921 • Published • 35 -
nvidia/Nemotron-Pretraining-Code-v2
Viewer • Updated • 836M • 3.17k • 129 -
BEAVER: An Efficient Deterministic LLM Verifier
Paper • 2512.05439 • Published • 36 -
codefuse-ai/C2LLM-7B
Feature Extraction • 8B • Updated • 172 • 10
Semantic Web
-
josancamon/kg-gen-MINE-evaluation-dataset
Viewer • Updated • 101 • 246 • 6 -
zilliz/semantic-highlight-bilingual-v1
Token Classification • 0.6B • Updated • 1.25k • 99 -
DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation
Paper • 2601.09688 • Published • 128