Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
RL+LLM Wiki
community
Activity Feed
Follow
26
AI & ML interests
None defined yet.
Recent Activity
lvwerra
new
activity
about 10 hours ago
rl-llm-wiki/knowledge-base:
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking — RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
lvwerra
new
activity
about 13 hours ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying — OpenAI's many RLs / agentic (speculation)
lvwerra
updated
a bucket
about 16 hours ago
rl-llm-wiki/rl-main-bucket
View all activity
Team members
14
rl-llm-wiki
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Articles
lvwerra
in
rl-llm-wiki/knowledge-base
about 10 hours ago
source: url:cognitiverevolution.ai/the-rl-fine-tuning-playbook-coreweave-s-kyle-corbitt-on-grpo-rubrics-environments-reward-hacking — RL fine-tuning playbook: GRPO/rubrics/reward-hacking (practitioner transcript)
3
#784 opened 1 day ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 13 hours ago
source: url:interconnects.ai/p/rl-backlog-openais-many-rls-clarifying — OpenAI's many RLs / agentic (speculation)
4
#734 opened 3 days ago by
lvwerra
lvwerra
updated
a bucket
about 16 hours ago
rl-llm-wiki/rl-main-bucket
318 MB
lvwerra
in
rl-llm-wiki/knowledge-base
about 16 hours ago
source: url:mechanize.work/blog/the-upcoming-gpt-3-moment-for-rl — RL environments as scaling substrate / GPT-3-moment thesis (forecast)
3
#785 opened 1 day ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 23 hours ago
source: url:interconnects.ai/p/opening-the-black-box-of-character — Character training pipeline (speculation, paid/partial)
3
#732 opened 3 days ago by
lvwerra
source: url:pillumina.github.io/posts/aiinfra/02-slime — Zhipu GLM slime RL-infra source reconstruction (CN, speculation)
3
#764 opened 1 day ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 24 hours ago
source: url:interconnects.ai/p/summertime-outlook-o3s-novelty-coming — o3 novelty / RLVR search (speculation)
3
#728 opened 3 days ago by
lvwerra
source: url:interconnects.ai/p/thinking-searching-and-acting — Thinking/Searching/Acting (agentic speculation)
3
#731 opened 3 days ago by
lvwerra
source: url:yuanchaofa.com/post/kimi-k2-5-reading-notes — Kimi K2.5 PARL parallel-agent RL deep-read (CN, speculation)
3
#744 opened 1 day ago by
lvwerra
source: url:nishtahir.com/notes-on-the-phi-4-reasoning-technical-paper — Phi-4-reasoning recipe reconstruction (Microsoft, secondary)
3
#765 opened 1 day ago by
lvwerra
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
4
#739 opened 1 day ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
1 day ago
source: url:cnblogs.com/volcengine-developer/articles/19070102 — veRL+ReTool tool-use RL reproduction / env+loss plumbing (CN, speculation)
3
#770 opened 1 day ago by
lvwerra
source: url:hkust-nlp.notion.site/simplerl-reason — SimpleRL-Zero 7B/8K R1-Zero reproduction / easy-to-hard (open-repro)
3
#774 opened 1 day ago by
lvwerra
source: url:cnblogs.com/theseventhson/p/18725466 — Kimi k1.5 long-CoT RL reconstruction (CN, speculation)
3
#763 opened 1 day ago by
lvwerra
source: url:syhya.github.io/posts/2025-01-27-deepseek-r1 — DeepSeek-R1 deep-dive / GRPO+KL math + o1 read-across (CN, speculation)
3
#742 opened 1 day ago by
lvwerra
grpo §4: integrate 2025 soft-source framing on the sharpen-vs-expand debate (SSA shaping, RL scaling laws)
1
#792 opened 1 day ago by
lvwerra
source: url:epoch.ai/gradient-updates/the-promise-of-reasoning-models — Epoch AI: promise of reasoning models (analysis, speculation)
3
#736 opened 3 days ago by
lvwerra
source: url:interconnects.ai/p/the-state-of-reasoning — State of reasoning (speculation/calibration)
3
#722 opened 3 days ago by
lvwerra
source: url:sequoiacap.com/podcast/training-data-noam-brown — Sequoia x Noam Brown: o1 team on test-time compute as scaling axis (transcript)
3
#781 opened 1 day ago by
lvwerra
source: url:latent.space/p/noam-brown — Latent Space x Noam Brown: test-time compute + self-play limits (transcript)
3
#782 opened 1 day ago by
lvwerra
Load more