Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
205516.8
TFLOPS
Leandro von Werra
PRO
lvwerra
1025
95
130
Follow
earthbound02's profile picture
andito's profile picture
dvilasuero's profile picture
837 followers
·
88 following
https://www.lvwerra.com
lvwerra
lvwerra
lvwerra
AI & ML interests
NLP and RL
Recent Activity
new
activity
about 7 hours ago
rl-llm-wiki/knowledge-base:
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking — RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
new
activity
about 11 hours ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying — OpenAI's many RLs / agentic (speculation)
updated
a bucket
about 13 hours ago
rl-llm-wiki/rl-main-bucket
View all activity
Organizations
lvwerra
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
rl-llm-wiki/knowledge-base
about 7 hours ago
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking — RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
3
#784 opened 1 day ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 11 hours ago
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying — OpenAI's many RLs / agentic (speculation)
4
#734 opened 3 days ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 13 hours ago
source: url:mechanize.work/blog/the-upcoming-gpt-3-moment-for-rl — RL environments as scaling substrate / GPT-3-moment thesis (forecast)
3
#785 opened about 24 hours ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 21 hours ago
source: url:interconnects.ai/p/opening-the-black-box-of-character — Character training pipeline (speculation, paid/partial)
3
#732 opened 3 days ago by
lvwerra
source: url:pillumina.github.io/posts/aiinfra/02-slime — Zhipu GLM slime RL-infra source reconstruction (CN, speculation)
3
#764 opened 1 day ago by
lvwerra
source: url:interconnects.ai/p/summertime-outlook-o3s-novelty-coming — o3 novelty / RLVR search (speculation)
3
#728 opened 3 days ago by
lvwerra
source: url:interconnects.ai/p/thinking-searching-and-acting — Thinking/Searching/Acting (agentic speculation)
3
#731 opened 3 days ago by
lvwerra
source: url:yuanchaofa.com/post/kimi-k2-5-reading-notes — Kimi K2.5 PARL parallel-agent RL deep-read (CN, speculation)
3
#744 opened 1 day ago by
lvwerra
source: url:nishtahir.com/notes-on-the-phi-4-reasoning-technical-paper — Phi-4-reasoning recipe reconstruction (Microsoft, secondary)
3
#765 opened 1 day ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 22 hours ago
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
4
#739 opened 1 day ago by
lvwerra
source: url:cnblogs.com/volcengine-developer/articles/19070102 — veRL+ReTool tool-use RL reproduction / env+loss plumbing (CN, speculation)
3
#770 opened 1 day ago by
lvwerra
source: url:hkust-nlp.notion.site/simplerl-reason — SimpleRL-Zero 7B/8K R1-Zero reproduction / easy-to-hard (open-repro)
3
#774 opened 1 day ago by
lvwerra
source: url:cnblogs.com/theseventhson/p/18725466 — Kimi k1.5 long-CoT RL reconstruction (CN, speculation)
3
#763 opened 1 day ago by
lvwerra
source: url:syhya.github.io/posts/2025-01-27-deepseek-r1 — DeepSeek-R1 deep-dive / GRPO+KL math + o1 read-across (CN, speculation)
3
#742 opened 1 day ago by
lvwerra
grpo §4: integrate 2025 soft-source framing on the sharpen-vs-expand debate (SSA shaping, RL scaling laws)
1
#792 opened about 22 hours ago by
lvwerra
source: url:epoch.ai/gradient-updates/the-promise-of-reasoning-models — Epoch AI: promise of reasoning models (analysis, speculation)
3
#736 opened 3 days ago by
lvwerra
source: url:interconnects.ai/p/the-state-of-reasoning — State of reasoning (speculation/calibration)
3
#722 opened 3 days ago by
lvwerra
source: url:sequoiacap.com/podcast/training-data-noam-brown — Sequoia x Noam Brown: o1 team on test-time compute as scaling axis (transcript)
3
#781 opened 1 day ago by
lvwerra
source: url:latent.space/p/noam-brown — Latent Space x Noam Brown: test-time compute + self-play limits (transcript)
3
#782 opened 1 day ago by
lvwerra
New activity in
rl-llm-wiki/knowledge-base
about 23 hours ago
deepen harmlessness-and-refusals §4.1: Sparrow's rule-decomposed dual-RM RLHF (integrate orphan 2209.14375)
2
#791 opened about 23 hours ago by
lvwerra
Load more