Instructions to use Joey-1123/Nova-7B-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Joey-1123/Nova-7B-Instruct with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-7B-Instruct") model = PeftModel.from_pretrained(base_model, "Joey-1123/Nova-7B-Instruct") - Notebooks
- Google Colab
- Kaggle
Nova-7B-Instruct
Nova is a 7B-class instruction-following model built for two jobs, one model:
- Dual-mode reasoning โ runs "fast & direct" or "deep reasoning" depending on the system prompt.
- Tool calling โ emits structured tool-call actions for agentic flows.
- Identity โ answers as Nova, developed by Shubham Panchal (Joey), in English and Hinglish.
Capabilities
| Mode | System prompt | Behavior |
|---|---|---|
| Direct | detailed thinking off |
Concise, fast, straight answers |
| Reasoning | detailed thinking on |
Long chain-of-thought before answering |
| Agentic | tool-call schema in messages | Emits <tool_call> JSON actions |
Example reasoning-mode output:
thinking
Okay, so I need to prove that the sum of the first n positive integers...
Quick start (transformers + PEFT)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import PeftModel
BASE = "Qwen/Qwen2.5-7B-Instruct"
ADAPTER = "Joey-1123/Nova-7B-Instruct"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16)
tok = AutoTokenizer.from_pretrained(BASE)
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained(
BASE, quantization_config=bnb, device_map="auto").eval()
nova = PeftModel.from_pretrained(model, ADAPTER).eval()
msgs = [{"role": "system", "content": "detailed thinking on"},
{"role": "user", "content": "Prove that the sum of the first n integers is n(n+1)/2."}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt").to(nova.device)
print(tok.decode(nova.generate(**ids, max_new_tokens=768,
pad_token_id=tok.eos_token_id)[0, ids["input_ids"].shape[1]:],
skip_special_tokens=True))
Notes
- This is a LoRA adapter (
r=16,alpha=32) meant to be loaded on top of a 7B-class base model with PEFT. - Architecture and tokenizer come from the base model repo (
adapter_config.jsonโbase_model_name_or_path). - Evaluated: NOT overfit (holdout ratio ~1.0), identity retention 50/50 familiar + 14/15 novel paraphrase, tool-call JSON valid in both modes, latency within 0.5% of base.
Trained on base weights Qwen/Qwen2.5-7B-Instruct under the Apache-2.0 license. License text: https://huggingface.co/Qwen/Qwen2.5-7B-Instruct/blob/main/LICENSE
- Downloads last month
- 8