Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
👋
Open to Work
12.1
TFLOPS
RDTvlokip
PRO
RDTvlokip
1
98
3
Follow
webxos's profile picture
dipankarsarkar's profile picture
libredove's profile picture
16 followers
·
7 following
https://rdtvlokip.fr
RDTvlokip
théo-charlet
AI & ML interests
None yet
Recent Activity
upvoted
an
article
about 8 hours ago
🎲 Apprendre à un réseau à écrire avec seulement une récompense — il a atteint 99,9 % grammatical sans apprendre une seule règle 🇫🇷
published
an
article
about 8 hours ago
🎲 Apprendre à un réseau à écrire avec seulement une récompense — il a atteint 99,9 % grammatical sans apprendre une seule règle 🇫🇷
posted
an
update
about 8 hours ago
I tried to teach a network to write using nothing but reward. No pretraining, no data, random weights. It worked twice. Both wins were fake. Test 1, copy a fixed sentence: solved in 1,639 episodes. Then I checked why. The per-position reward splits 12^12 into twelve independent 12-armed bandits. It doesn't guide the search, it deletes it. Transfer to a target sharing no positions was slower than starting over. Test 2, a grammar judged by a hand-written parser: 99.9% grammatical. Then I forced the antecedent and measured the consequent. P(noun agrees | determiner) = 0.333. Exactly 2 determiners out of 6. It learned no rule at all. It found an all-plural sublanguage where agreement is vacuously true. Then the one that stung: an untrained network plus a filter beats 20,000 episodes of REINFORCE on both validity and diversity. Nobody runs that baseline because it looks too stupid to bother with. None of this is new. Reward shaping moving the optimum: Ng, Harada & Russell, 1999. The variance decomposition that retires the word "sparse": Fisher's ANOVA. The real lesson: A high score on a verifier you wrote yourself measures your specification, not the agent. It finds the corner of the output space where your constraint is vacuous, and it looks like success from every angle except the one you forgot to check. Code, 8 figures, and the notebook with every refuted hypothesis 👇 🔗 https://huggingface.co/blog/RDTvlokip/teaching-a-network-to-write-with-reward-only 💻 https://github.com/RDTvlokip/RDTRL
View all activity
Organizations
RDTvlokip
's Spaces
1
Sort:Â Recently updated
Sleeping
Agents
3
AG BPE
📈
AG-BPE (Attention-Guided Byte-Pair Encoding)