Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
205516.8
TFLOPS
Leandro von Werra
PRO
lvwerra
1025
95
130
Follow
angt's profile picture
huybery's profile picture
adriwitek's profile picture
837 followers
ยท
88 following
https://www.lvwerra.com
lvwerra
lvwerra
lvwerra
AI & ML interests
NLP and RL
Recent Activity
new
activity
about 7 hours ago
rl-llm-wiki/knowledge-base:
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking โ RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
new
activity
about 11 hours ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying โ OpenAI's many RLs / agentic (speculation)
updated
a bucket
about 13 hours ago
rl-llm-wiki/rl-main-bucket
View all activity
Organizations
lvwerra
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
rl-llm-wiki/knowledge-base
about 7 hours ago
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking โ RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
3
#784 opened 1 day ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 11 hours ago
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying โ OpenAI's many RLs / agentic (speculation)
4
#734 opened 3 days ago by
lvwerra
updated
a bucket
about 13 hours ago
rl-llm-wiki/rl-main-bucket
318 MB
New activity in
rl-llm-wiki/knowledge-base
about 13 hours ago
source: url:mechanize.work/blog/the-upcoming-gpt-3-moment-for-rl โ RL environments as scaling substrate / GPT-3-moment thesis (forecast)
3
#785 opened about 24 hours ago by
lvwerra
updated
a Space
about 19 hours ago
Running
8
Agent Manager
๐ฅ
8
Private cloud manager for AI coding CLI sessions
New activity in
rl-llm-wiki/knowledge-base
about 21 hours ago
source: url:interconnects.ai/p/opening-the-black-box-of-character โ Character training pipeline (speculation, paid/partial)
3
#732 opened 3 days ago by
lvwerra
source: url:pillumina.github.io/posts/aiinfra/02-slime โ Zhipu GLM slime RL-infra source reconstruction (CN, speculation)
3
#764 opened 1 day ago by
lvwerra
source: url:interconnects.ai/p/summertime-outlook-o3s-novelty-coming โ o3 novelty / RLVR search (speculation)
3
#728 opened 3 days ago by
lvwerra
source: url:interconnects.ai/p/thinking-searching-and-acting โ Thinking/Searching/Acting (agentic speculation)
3
#731 opened 3 days ago by
lvwerra
source: url:yuanchaofa.com/post/kimi-k2-5-reading-notes โ Kimi K2.5 PARL parallel-agent RL deep-read (CN, speculation)
3
#744 opened 1 day ago by
lvwerra
source: url:nishtahir.com/notes-on-the-phi-4-reasoning-technical-paper โ Phi-4-reasoning recipe reconstruction (Microsoft, secondary)
3
#765 opened 1 day ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 22 hours ago
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law โ Verifier's law / asymmetry of verification (speculation)
4
#739 opened 1 day ago by
lvwerra
source: url:cnblogs.com/volcengine-developer/articles/19070102 โ veRL+ReTool tool-use RL reproduction / env+loss plumbing (CN, speculation)
3
#770 opened 1 day ago by
lvwerra
source: url:hkust-nlp.notion.site/simplerl-reason โ SimpleRL-Zero 7B/8K R1-Zero reproduction / easy-to-hard (open-repro)
3
#774 opened 1 day ago by
lvwerra
source: url:cnblogs.com/theseventhson/p/18725466 โ Kimi k1.5 long-CoT RL reconstruction (CN, speculation)
3
#763 opened 1 day ago by
lvwerra
source: url:syhya.github.io/posts/2025-01-27-deepseek-r1 โ DeepSeek-R1 deep-dive / GRPO+KL math + o1 read-across (CN, speculation)
3
#742 opened 1 day ago by
lvwerra
grpo ยง4: integrate 2025 soft-source framing on the sharpen-vs-expand debate (SSA shaping, RL scaling laws)
1
#792 opened about 22 hours ago by
lvwerra
source: url:epoch.ai/gradient-updates/the-promise-of-reasoning-models โ Epoch AI: promise of reasoning models (analysis, speculation)
3
#736 opened 3 days ago by
lvwerra
source: url:interconnects.ai/p/the-state-of-reasoning โ State of reasoning (speculation/calibration)
3
#722 opened 3 days ago by
lvwerra
source: url:sequoiacap.com/podcast/training-data-noam-brown โ Sequoia x Noam Brown: o1 team on test-time compute as scaling axis (transcript)
3
#781 opened 1 day ago by
lvwerra
Load more