Active filters: ppo
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_ROUND2
Reinforcement Learning
• 1B • Updated • 3
CatkinChen/nethack-ppo-ablation-no_hmm_curiosity_dyn_only
Reinforcement Learning
• Updated MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_AGAIN_ROUND3-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_AGAIN_ROUND3-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_AGAIN_ROUND3-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_AGAIN_ROUND3-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_AGAIN_ROUND3-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_RRETRT_Again_AGAIN_ROUND3
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round5-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round5-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 5
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round5-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round5-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round5-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 7
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round5
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round3-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round3-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round3-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round3-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round3-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 6
MattBou00/llama-3-2-1b-detox_v1f_SCALE9_round3
Reinforcement Learning
• 1B • Updated • 7
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round2-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round2-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round2-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round2-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round2-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round2
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round4-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round4-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round4-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 1
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round4-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 1