Nova-7B-Instruct

Nova is a 7B-class instruction-following model built for two jobs, one model:

  • Dual-mode reasoning โ€” runs "fast & direct" or "deep reasoning" depending on the system prompt.
  • Tool calling โ€” emits structured tool-call actions for agentic flows.
  • Identity โ€” answers as Nova, developed by Shubham Panchal (Joey), in English and Hinglish.

Capabilities

Mode System prompt Behavior
Direct detailed thinking off Concise, fast, straight answers
Reasoning detailed thinking on Long chain-of-thought before answering
Agentic tool-call schema in messages Emits <tool_call> JSON actions

Example reasoning-mode output:

 thinking
Okay, so I need to prove that the sum of the first n positive integers...

Quick start (transformers + PEFT)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import PeftModel

BASE  = "Qwen/Qwen2.5-7B-Instruct"
ADAPTER = "Joey-1123/Nova-7B-Instruct"

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.float16)
tok = AutoTokenizer.from_pretrained(BASE)
tok.pad_token = tok.eos_token

model = AutoModelForCausalLM.from_pretrained(
    BASE, quantization_config=bnb, device_map="auto").eval()
nova = PeftModel.from_pretrained(model, ADAPTER).eval()

msgs = [{"role": "system", "content": "detailed thinking on"},
        {"role": "user",   "content": "Prove that the sum of the first n integers is n(n+1)/2."}]
text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
ids = tok(text, return_tensors="pt").to(nova.device)
print(tok.decode(nova.generate(**ids, max_new_tokens=768,
      pad_token_id=tok.eos_token_id)[0, ids["input_ids"].shape[1]:],
      skip_special_tokens=True))

Notes

  • This is a LoRA adapter (r=16, alpha=32) meant to be loaded on top of a 7B-class base model with PEFT.
  • Architecture and tokenizer come from the base model repo (adapter_config.json โ†’ base_model_name_or_path).
  • Evaluated: NOT overfit (holdout ratio ~1.0), identity retention 50/50 familiar + 14/15 novel paraphrase, tool-call JSON valid in both modes, latency within 0.5% of base.

Trained on base weights Qwen/Qwen2.5-7B-Instruct under the Apache-2.0 license. License text: https://huggingface.co/Qwen/Qwen2.5-7B-Instruct/blob/main/LICENSE

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Space using Joey-1123/Nova-7B-Instruct 1