Edit Models filters
Model Tree
Apps
Inference Providers
One-click Deployment
Models
31
Active filters: preference-alignment
Txoka/GLEAM-Mixtral-8x7B-Instruct-v2.0
Text Generation • 47B • Updated • 22
danivpv/Llama-ML-Expert-DPO-1b
Text Generation • Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-HumanLLMs-Human-Like-DPO-Dataset-first-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-HumanLLMs-Human-Like-DPO-Dataset-second-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-HumanLLMs-Human-Like-DPO-Dataset-third-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-trl-lib-ultrafeedback_binarized-first-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-trl-lib-ultrafeedback_binarized-second-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-trl-lib-ultrafeedback_binarized-third-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-abacusai-MetaMath_DPO_FewShot-third-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-abacusai-MetaMath_DPO_FewShot-second-2-5e-06
Updated
ASethi04/meta-llama-Llama-3.1-8B-Instruct-dpo-abacusai-MetaMath_DPO_FewShot-first-2-5e-06
Updated
Shekswess/trlm-stage-3-dpo-final-2
Text Generation • 0.1B • Updated • 29 • 1
Shekswess/trlm-135m
Text Generation • 0.1B • Updated • 73 • 46
marcelovidigal/smollm3-3b-dpo-finetuned-v1
Text Generation • 3B • Updated • 13
antonisbast/Llama-3.2-3B-Gordon-Ramsay-DPO
Text Generation • Updated • 1
antonisbast/Llama-3.2-3B-Gordon-Ramsay-DPO-GGUF
Text Generation • 3B • Updated • 14
debsubhra/llama2-7b-dpo-orca-mini
Text Generation • 7B • Updated • 3
srpone/zoowork-shopranker-0.6b
srpone/zoowork-shopranker-4b
ace-1/mgpt2-dpo
Text Generation • Updated • 12
Xixixixihahahaha/RealAlign-SD-1.5
Text-to-Image • Updated • 81 • 1
Xixixixihahahaha/RealAlign-SD-3.5-M
Text-to-Image • Updated
tuggspeedman-ai/SmolLM3-3B-summarize-dpo-lora
Text Generation • Updated • 13
shabul/mistral-7b-desi-finance-advisor
Text Generation • 7B • Updated • 30
JiHyuk-Byun/3D-PAQA-evaluator
Other • 38.9M • Updated • 47 • 1
steven0226/qwen2.5-0.5b-dpo-ultrafeedback
Text Generation • 0.5B • Updated • 25
svjay/dpo-qwen2.5-7b-adapter
Updated
Wellwisher12/dpo-qwen-adapter
Text Generation • Updated
srianna/Qwen2-VL-2B-DPO-LoRA
Image-Text-to-Text • 2B • Updated • 15