Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
💼
Hiring
NYCU-RL-Bandits-Lab
rl-bandits-lab
2
Follow
0 followers
·
1 following
rl-bandits-lab
AI & ML interests
Reinforcement learning
Recent Activity
updated
a model
3 days ago
latent-q/ld4lg-qt-checkpoints
new
activity
9 days ago
latent-q/ld4lg-qt-checkpoints:
trained/ write probe via PR
new
activity
9 days ago
latent-q/ld4lg-qt-checkpoints:
trained/: SWIFT + TTSnap formal checkpoints (pipeline test, no base models)
View all activity
Organizations
models
7
Sort:Â Recently updated
rl-bandits-lab/AskR-Qwen3-4B-Instruct-LoRA
Text Generation
•
Updated
Jul 11
•
7
rl-bandits-lab/AskR-Qwen2.5-VL-7B-Instruct-LoRA
Updated
Apr 16
rl-bandits-lab/AskR-Qwen2.5-7B-Instruct-LoRA
Text Generation
•
Updated
Apr 15
•
15
rl-bandits-lab/ultrafeedback_rm
8B
•
Updated
Jul 30, 2025
•
3
rl-bandits-lab/helpsteer_rm
8B
•
Updated
Jun 10, 2025
•
6
rl-bandits-lab/hhrlhf_rm
8B
•
Updated
May 21, 2025
•
93
rl-bandits-lab/translation_rm
8B
•
Updated
May 21, 2025
•
8
datasets
2
Sort:Â Recently updated
rl-bandits-lab/SEGALE-WMT24
Viewer
•
Updated
Nov 5, 2025
•
137k
•
18
rl-bandits-lab/SEGALE-WMT24-Human-Eval
Viewer
•
Updated
Nov 5, 2025
•
27k
•
34