Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
HuggingFaceH4 's Collections
— Journal Club 📚 —
Scaling Test-Time Compute with Open Models
Zephyr ORPO
Zephyr 7B
Zephyr 7B Gemma
StarChat2 15B
Papers We've Read
Awesome SFT datasets
Awesome feedback datasets
Awesome reward models

Awesome reward models

updated 4 days ago

A curated collection of reward models to use with techniques like rejection sampling and RLHF / RLAIF

Upvote
10

  • llm-blender/PairRM

    Text Generation • 0.4B • Updated Jan 22, 2024 • 512 • 209

  • openbmb/UltraRM-13b

    Updated Oct 14, 2023 • 1.18k • 61

  • OpenAssistant/reward-model-deberta-v3-large-v2

    Text Classification • Updated Feb 1, 2023 • 13.3k • • 247

  • PKU-Alignment/beaver-7b-v1.0-reward

    Reinforcement Learning • 7B • Updated Apr 20, 2024 • 2.53k • 17
Upvote
10
  • Collection guide
  • Browse collections
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs