ReLIFT
Collection
ReLIFT, a training method that interleaves RL with online FT, achieving superior performance and efficiency compared to using RL or SFT alone. • 8 items • Updated • 1
How to use RoadQAQ/ReLIFT-Qwen2.5-Math-7B-Zero with Transformers:
# Use a pipeline as a high-level helper
# Warning: Pipeline type "question-answering" is no longer supported in transformers v5.
# You must load the model directly (see below) or downgrade to v4.x with:
# pip install "transformers<5.0.0"
from transformers import pipeline
pipe = pipeline("question-answering", model="RoadQAQ/ReLIFT-Qwen2.5-Math-7B-Zero") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("RoadQAQ/ReLIFT-Qwen2.5-Math-7B-Zero")
model = AutoModelForCausalLM.from_pretrained("RoadQAQ/ReLIFT-Qwen2.5-Math-7B-Zero", device_map="auto")This repository contains the ReLIFT model presented in Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions.
Code: https://github.com/TheRoadQaQ/ReLIFT
Hugging Face Collection: https://huggingface.co/collections/RoadQAQ/relift-684535e199a909cad16d8b05