Inference Providers
Active filters: verl
JaeWooShin/verl-intermediate-checkpoints-20260928
Reinforcement Learning
• Updated • 1
junnyu/Qwen2.5-7B-Instruct-1M-GRPO_logic_KK_5PPL
Text Generation
• 8B • Updated • 27
sonyashijin/qwen3-32b-verilog-lora
LichengLiu03/Qwen2.5-3B-UFO
Text Generation
• 3B • Updated • 20
• • 2
LichengLiu03/Qwen2.5-3B-UFO-1turn
Text Generation
• 3B • Updated • 19
• 2
mradermacher/Qwen2.5-3B-UFO-GGUF
3B • Updated • 236
• 1
mradermacher/Qwen2.5-3B-UFO-1turn-GGUF
3B • Updated • 441
• 1
Text Generation
• 0.6B • Updated • 31
• 2
Jasaxion/MathSmith-HC-Problem-Synthesizer-Qwen3-8B
Text Generation
• 8B • Updated • 14
• 1
Jasaxion/MathSmith-Hard-Problem-Synthesizer-Qwen3-8B
Text Generation
• 8B • Updated • 22
• 1
thejaminator/grpo-feature-vector-step-1
Updated • 11
Text Generation
• 8B • Updated • 502
• • 11
orbit-ai/searchr1-repro-4b
Text Generation
• 4B • Updated • 28
orbit-ai/orbit-4b-ablation-top-10-docs-v0.1
Text Generation
• 4B • Updated • 14
orbit-ai/orbit-4b-ablation-training-mix-124-v0.1
Text Generation
• 4B • Updated • 22
Text Generation
• 4B • Updated • 36
• • 1
Text Generation
• Updated karthik/verl-qwen2.5-0.5b-gsm8k-ppo-step360
Text Generation
• 0.5B • Updated • 24
samhitha2601/llama3.2-3b-ppo
Reinforcement Learning
• Updated • 16
samhitha2601/llama3.2-3b-ppo-critic
Reinforcement Learning
• Updated • 16
mradermacher/MathSmith-HC-Problem-Synthesizer-Qwen3-8B-GGUF
8B • Updated • 389
• 1
mradermacher/MathSmith-HC-Problem-Synthesizer-Qwen3-8B-i1-GGUF
8B • Updated • 1.24k
• 2
mradermacher/MathSmith-Hard-Problem-Synthesizer-Qwen3-8B-GGUF
8B • Updated • 560
• 1
mradermacher/MathSmith-Hard-Problem-Synthesizer-Qwen3-8B-i1-GGUF
8B • Updated • 1.65k
• 1
Time-HD-Anonymous/STReasoner-8B
Feature Extraction
• 8B • Updated • 61
archit11/qwen2.5-coder-3b-verl-track-a-lora
Text Generation
• Updated • 12
orbit-ai/infoseeker-repro-4b
Text Generation
• 4B • Updated • 33
mradermacher/infoseeker-repro-4b-GGUF
4B • Updated • 253
mradermacher/orbit-4b-v0.1-GGUF
4B • Updated • 263