Active filters: ppo
MattBou00/llama-3-2-1b-detox_RETRY_SAMPLING_scale10_Round3-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_RETRY_SAMPLING_scale10_Round3-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 1
MattBou00/llama-3-2-1b-detox_RETRY_SAMPLING_scale10_Round3-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_RETRY_SAMPLING_scale10_Round3-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_RETRY_SAMPLING_scale10_Round3-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_RETRY_SAMPLING_scale10_Round3
Reinforcement Learning
• 1B • Updated • 3
mradermacher/gpt2-rlhf-anthropic-GGUF
0.1B • Updated • 95
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND5-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND5-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND5-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND5-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND5-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND5
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND3-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND3-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND3-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND3-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND3-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND3
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND1-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND1-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND1-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND1-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND1-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND1
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND2-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND2-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND2-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND2-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND2-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 4