Instructions to use ho22joshua/hep-chat-entrypoint with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ho22joshua/hep-chat-entrypoint with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
HEP Posttraining Chat Entrypoint
Run the LoRA adapters published in ho22joshua/hep-posttraining. The repository contains only lightweight entrypoint code: it downloads the selected Qwen 3.5 base model and applies the selected PEFT adapter at runtime.
The registry includes every published adapter:
- ROOT Dataset: Qwen 3.5 0.8B (r32/e5/1e-5), 4B (r32/e1/1e-4), and 9B (r32/e1/1e-4; r32/e10/1e-5).
- ROOT + TRExFitter Dataset: Qwen 3.5 0.8B (r16/e3/1e-5; r32/e3/1e-4).
Install
Use a CUDA GPU for the 4B and 9B models.
git clone https://huggingface.co/ho22joshua/hep-chat-entrypoint
cd hep-chat-entrypoint
python -m pip install -r requirements.txt
python list_models.py
Interactive chat
The default is the one-epoch ROOT Dataset 4B adapter, with thinking disabled:
python chat.py --device cuda
Choose another adapter or compare it to its unadapted base model:
python chat.py --model root-qwen3.5-9b-r32-e1-lr1e-4 --device cuda
python chat.py --model root-qwen3.5-9b-r32-e1-lr1e-4 --mode base --device cuda
Enable the Qwen reasoning mode for an explicit thinking ablation:
python chat.py --model root-qwen3.5-4b-r32-e1-lr1e-4 --thinking --device cuda
During chat, use /thinking on, /thinking off, or /exit.
Batch inference
Put one prompt on each non-empty line of prompts.txt:
python run_prompts.py \
--model root-qwen3.5-4b-r32-e1-lr1e-4 \
--prompts prompts.txt \
--output outputs.jsonl \
--device cuda
outputs.jsonl begins with run metadata and then one completion record per prompt. Add --thinking for the reasoning condition, or --mode base for the matched base-model condition.
Gradio
python app.py
The UI provides an adapter picker, base-versus-adapter mode, and a thinking toggle. It loads one model at a time.
Direct PEFT usage
Each registry entry specifies the base model and adapter subfolder. The equivalent minimal Python pattern is:
from peft import PeftModel
from transformers import AutoModelForCausalLM
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B", torch_dtype="auto")
model = PeftModel.from_pretrained(
base,
"ho22joshua/hep-posttraining",
subfolder="ROOT/qwen3.5-4b-lora-r32-e1-lr1e-4",
)
- Downloads last month
- -