FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning. Project Page: https://fapo-rl.github.io/
Ding
dyyyyyyyy
AI & ML interests
None yet
Recent Activity
updated a dataset 6 days ago
dyyyyyyyy/swe-rebench-filtered-1150 published a dataset 6 days ago
dyyyyyyyy/swe-rebench-filtered-1150 updated a dataset 7 days ago
dyyyyyyyy/qwen3.5-9b-sbv-100tOrganizations
ScaleQuest
We introduce ScaleQuest, a scalable and novel data synthesis method. Project Page: https://scalequest.github.io/
-
Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch
Paper • 2410.18693 • Published • 42 -
dyyyyyyyy/ScaleQuest-Math
Viewer • Updated • 1M • 123 • 23 -
dyyyyyyyy/ScaleQuest-Code
Viewer • Updated • 157k • 23 • 4 -
dyyyyyyyy/ScaleQuest-Math-Qwen2.5
Viewer • Updated • 622k • 22
COLDQA
SCAN
We propose Self-Denoising Monte Carlo Annotation (SCAN), an efficient Process Reward Model (PRM) data synthesis and noise-tolerant learning framework.
GNER
We introduce GNER, a Generative Named Entity Recognition framework, which demonstrates enhanced zero-shot capabilities across unseen entity domains.
FAPO
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning. Project Page: https://fapo-rl.github.io/
SCAN
We propose Self-Denoising Monte Carlo Annotation (SCAN), an efficient Process Reward Model (PRM) data synthesis and noise-tolerant learning framework.
ScaleQuest
We introduce ScaleQuest, a scalable and novel data synthesis method. Project Page: https://scalequest.github.io/
-
Unleashing Reasoning Capability of LLMs via Scalable Question Synthesis from Scratch
Paper • 2410.18693 • Published • 42 -
dyyyyyyyy/ScaleQuest-Math
Viewer • Updated • 1M • 123 • 23 -
dyyyyyyyy/ScaleQuest-Code
Viewer • Updated • 157k • 23 • 4 -
dyyyyyyyy/ScaleQuest-Math-Qwen2.5
Viewer • Updated • 622k • 22
GNER
We introduce GNER, a Generative Named Entity Recognition framework, which demonstrates enhanced zero-shot capabilities across unseen entity domains.
COLDQA