Active filters: ppo
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round4
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round1-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round1-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round1-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round1-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round1-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round1
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round5-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round5-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round5-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 5
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round5-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round5-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round5
Reinforcement Learning
• 1B • Updated • 2
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round3-checkpoint-epoch-20
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round3-checkpoint-epoch-40
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round3-checkpoint-epoch-60
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round3-checkpoint-epoch-80
Reinforcement Learning
• 1B • Updated • 3
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round3-checkpoint-epoch-100
Reinforcement Learning
• 1B • Updated • 4
MattBou00/llama-3-2-1b-detox_v1f_SCALE8_round3
Reinforcement Learning
• 1B • Updated • 4
CatkinChen/nethack-ppo-ablation-no_hmm_rnd
Reinforcement Learning
• Updated CatkinChen/nethack-ppo-ablation-baseline_curiosity_dyn_only
Reinforcement Learning
• Updated joigalcar/ppo-LunarLander-v2_Scratch
Reinforcement Learning
• Updated joigalcar/ppo-LunarLander-v2_Scratch_2
Reinforcement Learning
• Updated rishiad/kinitro-metaworld-agent
Reinforcement Learning
• Updated CatkinChen/nethack-ppo-ablation-baseline_rnd
Reinforcement Learning
• Updated CatkinChen/nethack-ppo-ablation-baseline_curiosity_skill_only
Reinforcement Learning
• Updated CatkinChen/nethack-ppo-ablation-baseline_curiosity_trans_only
Reinforcement Learning
• Updated OxoGhost/ppo-LunarLander-v2-PPO
Reinforcement Learning
• Updated WillLedd/ppoCartPoleFromScratch
Reinforcement Learning
• Updated nabeelshan/rlhf-gpt2-pipeline
Text Generation
• Updated