imflash217/proximal_policy_optimization_huggy_unity Reinforcement Learning • Updated Jan 14, 2023 • 87 • 2
MRNH/proximal-policy-optimization-LunarLander-v2 Reinforcement Learning • Updated Aug 10, 2023 • 16 • 1
andersonbcdefg/sharegpt_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 6, 2023 • 11.8k • 117 • 3
crumb/bloom-560m-RLHF-SD2-prompter-aesthetic Text Generation • 0.6B • Updated Mar 19, 2023 • 250 • 25
sabaridsnfuji/repro-off-policy-learning-in-large-action-spaces-optimization-matters-more-than-estimation Viewer • Updated Jul 25 • 1 • 36 • 2