Direct Preference Optimization: Your Language Model is Secretly a Reward Model Paper • 2305.18290 • Published May 29, 2023 • 68
My models: daily driver rotation Collection A rotating list of models I created and currently use as daily drivers. From my many models, these are the ones I’m actively using. • 4 items • Updated 1 day ago • 10
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models Paper • 2606.03748 • Published Jun 2 • 21
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training Paper • 2605.18719 • Published May 18 • 7
view article Article Chitos: From Detection to Proof — An Autonomous Security AI That Actually Exploits FINAL-Bench • 29 days ago • 19