Instructions to use whosouravsharma/paper-qa-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use whosouravsharma/paper-qa-lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct") model = PeftModel.from_pretrained(base_model, "whosouravsharma/paper-qa-lora") - Notebooks
- Google Colab
- Kaggle
paper-qa-lora
A QLoRA adapter for Qwen/Qwen2.5-1.5B-Instruct.
It was trained to answer questions about a research paper only from the
supplied context, to cite the sentence it relied on, and to say plainly
when the context doesn't contain the answer. It is the answer generator
behind the Paper QA
app.
| What it's good at | Following the app's output format: answers end with a Source: "…" citation (86% of the time vs. 0% for the base model), and it refuses on some unanswerable questions (4 of 12 vs. 0 for base) |
| Where it falls short | Answer quality. It is less faithful to the context than the untuned base model (70.0% vs. 80.2%), and it hallucinates more (see Evaluation) |
| Status | Experimental. The cause of the regression has been identified in the training data (see Known issue) and not yet fixed |
| Size | 4.36M trainable parameters (0.49% of the model), an 8.7 MB adapter |
| License | Apache 2.0, the same as the base model. Training data is QASPER (CC BY 4.0) |
Table of contents
- Model details
- Uses
- How to get started
- Evaluation
- Known issue: the training data
- Bias, risks, and limitations
- Training details
- Environmental impact
- Citation
Model details
| Developed by | whosouravsharma |
| Model type | LoRA adapter (QLoRA, 4-bit NF4) for a 1.5B-parameter decoder-only chat model |
| Base model | Qwen/Qwen2.5-1.5B-Instruct |
| Adapter | Rank 16, alpha 32, dropout 0.05, on q_proj, k_proj, v_proj, o_proj |
| Language | English |
| Domain | NLP research papers (QASPER) |
| Version | v2, trained 2026-09-04. It replaced v1 (2026-08-30) in place |
| License | Apache 2.0 |
| Resource | Link |
|---|---|
| Demo | Paper QA Space |
| Serving backend | paper-qa-rag Space (retrieval, generation with this adapter, judging) |
| Training data | whosouravsharma/paper-qa-qasper-sft |
Uses
Direct use
- Answering questions about a research paper, given retrieved passages from
it, in a RAG pipeline that expects a
Source: "…"citation after each answer. - Studying how supervised fine-tuning changes grounding, citation and refusal behaviour in a small model. This repository documents one attempt and its failure mode in detail.
Out-of-scope use
- Any setting where factual accuracy matters more than output format. For answer quality alone, the untuned base model scores higher (see Evaluation).
- Relying on it to refuse when the answer isn't in the context. It catches only a minority of unanswerable questions.
- Papers outside NLP, or documents that aren't research papers. It was trained only on QASPER.
- Medical, legal or other high-stakes question answering.
How to get started
The adapter expects the same system prompt and message layout it was served with:
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
base = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen2.5-1.5B-Instruct", torch_dtype=torch.float16
).to("cuda")
model = PeftModel.from_pretrained(base, "whosouravsharma/paper-qa-lora").eval()
system = (
"You are a research paper assistant. Answer the question using only the "
"provided context. Quote the exact sentence(s) you relied on as a "
"citation. If the context does not contain the answer, say so plainly "
"instead of guessing."
)
context = "…retrieved passages from the paper…"
question = "What datasets are used for experiments?"
messages = [
{"role": "system", "content": system},
{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)
output = model.generate(**inputs, max_new_tokens=256, do_sample=False,
pad_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
The expected output is a one- or two-sentence answer, then
Source: "<quoted sentence>". When it can't answer, it replies:
"I cannot find an answer to this question in the provided context."
Evaluation
Benchmark
- Questions. 119 questions on 20 papers from QASPER's
testsplit: 107 answerable and 12 unanswerable. The reference answers are QASPER's own human annotations. None of the 20 papers appear in the training or validation data; this was checked by paper ID. - Judge.
gpt-5.6-lunascores each answer from 1 to 10 on six criteria, reported here as percentages. In the offline benchmark the judge sees the model's answer and the human reference as "Answer A" and "Answer B" in a random order, and is told to weight faithfulness to the paper over similarity to the reference. - Behaviour. Citation and refusal rates are counted directly from the outputs; they are not judged.
Offline benchmark
All systems here are given the same passages, retrieved by TF-IDF over the paper's paragraphs, and judged by the same method.
| Qwen 1.5B, untuned | + LoRA v1 | + LoRA v2 (this repo) | gpt-5.6-luna (hosted)¹ |
|
|---|---|---|---|---|
| Factual correctness | 77.8 | 74.6 | 71.2 | 95.5 |
| Faithfulness to context | 80.2 | 79.8 | 70.0 | 96.5 |
| Completeness | 68.6 | 64.0 | 69.2 | 90.9 |
| Relevance | 88.5 | 86.8 | 89.6 | 98.8 |
| Clarity | 90.0 | 89.3 | 89.0 | 97.1 |
| No hallucination | 82.6 | 86.6 | 73.4 | 97.3 |
| Preferred over the human reference | 68.9% | 52.1% | 60.5% | 93.3% |
Ends with a Source: citation |
0% | 80.7% | 85.7% | 0% |
| Refuses on unanswerable (of 12) | 0 | 4 | 4 | 0 |
| Refuses on answerable (of 107) ↓ | 0 | 19 | 13 | 0 |
Scores are percentages from the judge's rubric. Bold marks the best of the three Qwen variants.
¹ The hosted model is included as a reference point only. It is the same model as the judge, and models tend to rate their own answers highly, so its scores are probably inflated.
Is the v2 regression real? Yes, on three criteria. Paired bootstrap over the 119 questions (v2 minus untuned base, 5,000 resamples):
| Criterion | Difference | 95% CI |
|---|---|---|
| Faithfulness | −10.2 pts | [−16.4, −3.8] |
| No hallucination | −9.2 pts | [−15.7, −2.8] |
| Factual correctness | −6.6 pts | [−12.9, −0.5] |
| Completeness | +0.7 pts | [−4.5, +6.0] |
| Relevance | +1.1 pts | [−2.3, +4.5] |
| Clarity | −1.0 pts | [−3.2, +1.3] |
For scale: the untuned base model was scored twice in separate runs, and its scores moved by up to about 1 point between them, so judge noise alone is much smaller than these gaps.
What the numbers say
- The fine-tune learned the output format, not better grounding. It cites a source 86% of the time and refuses on a third of unanswerable questions. The base model never does either, even under the same system prompt. Nor does the much larger hosted model: the format has to be trained in.
- It is less faithful and hallucinates more than the untuned base. The confidence intervals for both exclude zero.
- v1 and v2 trade off differently. v2 was retrained to fix v1's terse, list-like answers. It did make answers fuller and cut false refusals from 19 to 13, but at the cost of faithfulness. The cause is in the training data (see Known issue).
In production
These results come from the deployed pipeline on 2026-09-04: the same 119 questions, with the real arXiv PDFs uploaded to the live app. There were 0 failures across 20 uploads and 119 questions.
| Criterion | Score |
|---|---|
| Factual correctness | 71.2 |
| Faithfulness to context | 75.5 |
| Completeness | 64.2 |
| Relevance | 90.5 |
| Clarity | 89.8 |
| No hallucination | 78.6 |
These are not comparable with the offline table. Production uses FAISS over OpenAI embeddings, not TF-IDF, and puts the paper's title and a summary at the top of every context. Its live judge also scores each answer against that context alone, with no human reference.
| Behaviour | Production |
|---|---|
Ends with a Source: citation |
95.0% (113/119) |
| …of which the quoted "source" is the paper title | 81 of 113 |
| …of which it quotes an actual passage | 31 of 113 |
| Refuses on unanswerable | 2 of 12 |
| Refuses on answerable | 3 of 107 |
Two production caveats
- Most "citations" quote the title. The pipeline puts
Paper title: …at the top of every context, and the model usually quotes that line. So the 95% citation rate greatly overstates how often an answer points at real evidence. In the offline benchmark, where contexts have no title line, it cited actual passages. - Fewer false refusals, but also fewer correct ones. Production refused fewer answerable questions than the offline benchmark (3 vs. 13), but also fewer unanswerable ones (2 vs. 4). The title and summary give the model something plausible to answer from even when the passages don't contain the answer. This is not a clean improvement.
Latency (production, per question, n = 119)
| Stage | p50 | p90 | max |
|---|---|---|---|
| Retrieval | 0.19 s | 0.40 s | 3.7 s |
| Generation (this adapter, T4) | 2.6 s | 5.0 s | 13.4 s |
| Judging | 2.1 s | 3.3 s | 9.3 s |
| End to end | 6.8 s | 10.0 s | 17.7 s |
Known issue: the training data
A follow-up check traced the quality regression to the training targets
themselves. gpt-5.6-luna was asked whether each target answer is fully
supported by the context it was paired with, over 300 randomly sampled
answerable training examples:
| Training data | Targets with claims the context doesn't support |
|---|---|
| Raw QASPER answers (used for v1) | 17.7% |
| LLM-rewritten answers (used for v2) | 35.0% (105 of 300) |
- The rewrite doubled the problem. For v2, each terse QASPER answer was rewritten by an LLM into a full explanatory sentence. The rewriter filled gaps with plausible-sounding details, despite being told not to add facts. The model then learned to do the same.
- The raw data already had the problem. QASPER annotators wrote their answers after reading the whole paper, but the training context contains only the evidence spans they highlighted. About 1 in 6 examples therefore teaches the model to state facts that aren't in its context.
Planned fix, not yet done:
- Widen each training context to include the paragraphs around the evidence.
- Keep only targets verified as fully supported by their context.
- Retrain and re-run both evaluations.
Bias, risks, and limitations
- Its answers can sound grounded without being grounded. The answers are fluent and come with citations, but they score only 70% on faithfulness to the context, below the untuned base model. And in production the citation is usually the paper's title, not evidence. Check answers against the paper.
- It refuses too rarely. It declines on only 2–4 of 12 unanswerable questions, and otherwise answers anyway.
- Narrow domain. It was trained and tested only on NLP papers from QASPER.
- The evaluation relies on an LLM judge. All quality scores come from one judge model, not human raters. The benchmark is small (119 questions), and it has only 12 unanswerable questions, so the refusal rates are uncertain.
- It inherits Qwen2.5's limitations and biases, as a 1.5B-parameter model.
Training details
Training data
whosouravsharma/paper-qa-qasper-sft
was built from QASPER's
train and validation splits (CC BY 4.0). Each annotator answer becomes
one chat example:
- Answerable. The context is the annotator's evidence quotes. The target
is the answer (rewritten into a full sentence for v2), followed by
Source: "<first evidence quote>". - Unanswerable. The context is 2 random paragraphs from the same paper, like a realistic retrieval miss, so the model learns to refuse on content rather than on an empty prompt. The target is a fixed refusal sentence.
| Split | Examples | Refusal examples | Papers | Median answer length |
|---|---|---|---|---|
| train | 2,589 | 281 | 878 | 23 words |
| validation | 1,715 | 163 | 281 | 21 words |
Training procedure
| Method | QLoRA with TRL SFTTrainer: base weights 4-bit NF4, LoRA adapters trained in fp16 |
| Epochs | 3 (486 optimizer steps) |
| Batch size | 4 per step × 4 gradient-accumulation steps = 16 |
| Learning rate | 2e-4 |
| Max sequence length | 2,048 tokens (context + question + answer; p99 ≈ 1,262) |
| Evaluation | Validation loss each epoch; the final epoch is kept |
| Epoch | Validation loss | Validation token accuracy |
|---|---|---|
| 1 | 1.297 | 72.4% |
| 2 | 1.287 | 72.6% |
| 3 | 1.286 | 72.6% |
Validation loss flattens after the first epoch. The loss measures how closely the model imitates the training targets. Since about a third of those targets contain unsupported claims, a lower loss doesn't mean better grounding.
state.json and loss_curve.json in this repository record the exact
configuration and the full loss log.
Environmental impact
| Hardware | 1× NVIDIA L4 (Hugging Face Jobs, l4x1) |
| Training time | 56 min (3,344 s) |
| Carbon emitted | Not measured |
Citation
Training and evaluation data come from QASPER:
@inproceedings{Dasigi2021ADO,
title = {A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers},
author = {Pradeep Dasigi and Kyle Lo and Iz Beltagy and Arman Cohan and Noah A. Smith and Matt Gardner},
year = {2021}
}
Model card contact
Open a discussion in the Community tab.
- Downloads last month
- 144
Model tree for whosouravsharma/paper-qa-lora
Datasets used to train whosouravsharma/paper-qa-lora
whosouravsharma/paper-qa-qasper-sft
Spaces using whosouravsharma/paper-qa-lora 2
Evaluation results
- Factual correctness (%) on QASPER test slice (20 papers, 119 questions)test set self-reported71.200
- Faithfulness to context (%) on QASPER test slice (20 papers, 119 questions)test set self-reported70.000
- No hallucination (%) on QASPER test slice (20 papers, 119 questions)test set self-reported73.400
- Cites a source (%) on QASPER test slice (20 papers, 119 questions)test set self-reported85.700
