paper-qa-lora

A QLoRA adapter for Qwen/Qwen2.5-1.5B-Instruct. It was trained to answer questions about a research paper only from the supplied context, to cite the sentence it relied on, and to say plainly when the context doesn't contain the answer. It is the answer generator behind the Paper QA app.

What it's good at Following the app's output format: answers end with a Source: "…" citation (86% of the time vs. 0% for the base model), and it refuses on some unanswerable questions (4 of 12 vs. 0 for base)
Where it falls short Answer quality. It is less faithful to the context than the untuned base model (70.0% vs. 80.2%), and it hallucinates more (see Evaluation)
Status Experimental. The cause of the regression has been identified in the training data (see Known issue) and not yet fixed
Size 4.36M trainable parameters (0.49% of the model), an 8.7 MB adapter
License Apache 2.0, the same as the base model. Training data is QASPER (CC BY 4.0)

Table of contents

Model details

Developed by whosouravsharma
Model type LoRA adapter (QLoRA, 4-bit NF4) for a 1.5B-parameter decoder-only chat model
Base model Qwen/Qwen2.5-1.5B-Instruct
Adapter Rank 16, alpha 32, dropout 0.05, on q_proj, k_proj, v_proj, o_proj
Language English
Domain NLP research papers (QASPER)
Version v2, trained 2026-09-04. It replaced v1 (2026-08-30) in place
License Apache 2.0
Resource Link
Demo Paper QA Space
Serving backend paper-qa-rag Space (retrieval, generation with this adapter, judging)
Training data whosouravsharma/paper-qa-qasper-sft

Uses

Direct use

  • Answering questions about a research paper, given retrieved passages from it, in a RAG pipeline that expects a Source: "…" citation after each answer.
  • Studying how supervised fine-tuning changes grounding, citation and refusal behaviour in a small model. This repository documents one attempt and its failure mode in detail.

Out-of-scope use

  • Any setting where factual accuracy matters more than output format. For answer quality alone, the untuned base model scores higher (see Evaluation).
  • Relying on it to refuse when the answer isn't in the context. It catches only a minority of unanswerable questions.
  • Papers outside NLP, or documents that aren't research papers. It was trained only on QASPER.
  • Medical, legal or other high-stakes question answering.

How to get started

The adapter expects the same system prompt and message layout it was served with:

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-1.5B-Instruct")
base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-1.5B-Instruct", torch_dtype=torch.float16
).to("cuda")
model = PeftModel.from_pretrained(base, "whosouravsharma/paper-qa-lora").eval()

system = (
    "You are a research paper assistant. Answer the question using only the "
    "provided context. Quote the exact sentence(s) you relied on as a "
    "citation. If the context does not contain the answer, say so plainly "
    "instead of guessing."
)
context = "…retrieved passages from the paper…"
question = "What datasets are used for experiments?"

messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"},
]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_tensors="pt", return_dict=True
).to(model.device)

output = model.generate(**inputs, max_new_tokens=256, do_sample=False,
                        pad_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

The expected output is a one- or two-sentence answer, then Source: "<quoted sentence>". When it can't answer, it replies: "I cannot find an answer to this question in the provided context."

Evaluation

Benchmark

  • Questions. 119 questions on 20 papers from QASPER's test split: 107 answerable and 12 unanswerable. The reference answers are QASPER's own human annotations. None of the 20 papers appear in the training or validation data; this was checked by paper ID.
  • Judge. gpt-5.6-luna scores each answer from 1 to 10 on six criteria, reported here as percentages. In the offline benchmark the judge sees the model's answer and the human reference as "Answer A" and "Answer B" in a random order, and is told to weight faithfulness to the paper over similarity to the reference.
  • Behaviour. Citation and refusal rates are counted directly from the outputs; they are not judged.

Offline benchmark

All systems here are given the same passages, retrieved by TF-IDF over the paper's paragraphs, and judged by the same method.

Qwen 1.5B, untuned + LoRA v1 + LoRA v2 (this repo) gpt-5.6-luna (hosted)¹
Factual correctness 77.8 74.6 71.2 95.5
Faithfulness to context 80.2 79.8 70.0 96.5
Completeness 68.6 64.0 69.2 90.9
Relevance 88.5 86.8 89.6 98.8
Clarity 90.0 89.3 89.0 97.1
No hallucination 82.6 86.6 73.4 97.3
Preferred over the human reference 68.9% 52.1% 60.5% 93.3%
Ends with a Source: citation 0% 80.7% 85.7% 0%
Refuses on unanswerable (of 12) 0 4 4 0
Refuses on answerable (of 107) ↓ 0 19 13 0

Scores are percentages from the judge's rubric. Bold marks the best of the three Qwen variants.

¹ The hosted model is included as a reference point only. It is the same model as the judge, and models tend to rate their own answers highly, so its scores are probably inflated.

Is the v2 regression real? Yes, on three criteria. Paired bootstrap over the 119 questions (v2 minus untuned base, 5,000 resamples):

Criterion Difference 95% CI
Faithfulness −10.2 pts [−16.4, −3.8]
No hallucination −9.2 pts [−15.7, −2.8]
Factual correctness −6.6 pts [−12.9, −0.5]
Completeness +0.7 pts [−4.5, +6.0]
Relevance +1.1 pts [−2.3, +4.5]
Clarity −1.0 pts [−3.2, +1.3]

For scale: the untuned base model was scored twice in separate runs, and its scores moved by up to about 1 point between them, so judge noise alone is much smaller than these gaps.

What the numbers say

  1. The fine-tune learned the output format, not better grounding. It cites a source 86% of the time and refuses on a third of unanswerable questions. The base model never does either, even under the same system prompt. Nor does the much larger hosted model: the format has to be trained in.
  2. It is less faithful and hallucinates more than the untuned base. The confidence intervals for both exclude zero.
  3. v1 and v2 trade off differently. v2 was retrained to fix v1's terse, list-like answers. It did make answers fuller and cut false refusals from 19 to 13, but at the cost of faithfulness. The cause is in the training data (see Known issue).

In production

These results come from the deployed pipeline on 2026-09-04: the same 119 questions, with the real arXiv PDFs uploaded to the live app. There were 0 failures across 20 uploads and 119 questions.

Criterion Score
Factual correctness 71.2
Faithfulness to context 75.5
Completeness 64.2
Relevance 90.5
Clarity 89.8
No hallucination 78.6

These are not comparable with the offline table. Production uses FAISS over OpenAI embeddings, not TF-IDF, and puts the paper's title and a summary at the top of every context. Its live judge also scores each answer against that context alone, with no human reference.

Behaviour Production
Ends with a Source: citation 95.0% (113/119)
…of which the quoted "source" is the paper title 81 of 113
…of which it quotes an actual passage 31 of 113
Refuses on unanswerable 2 of 12
Refuses on answerable 3 of 107

Two production caveats

  • Most "citations" quote the title. The pipeline puts Paper title: … at the top of every context, and the model usually quotes that line. So the 95% citation rate greatly overstates how often an answer points at real evidence. In the offline benchmark, where contexts have no title line, it cited actual passages.
  • Fewer false refusals, but also fewer correct ones. Production refused fewer answerable questions than the offline benchmark (3 vs. 13), but also fewer unanswerable ones (2 vs. 4). The title and summary give the model something plausible to answer from even when the passages don't contain the answer. This is not a clean improvement.

Latency (production, per question, n = 119)

Stage p50 p90 max
Retrieval 0.19 s 0.40 s 3.7 s
Generation (this adapter, T4) 2.6 s 5.0 s 13.4 s
Judging 2.1 s 3.3 s 9.3 s
End to end 6.8 s 10.0 s 17.7 s

Known issue: the training data

A follow-up check traced the quality regression to the training targets themselves. gpt-5.6-luna was asked whether each target answer is fully supported by the context it was paired with, over 300 randomly sampled answerable training examples:

Training data Targets with claims the context doesn't support
Raw QASPER answers (used for v1) 17.7%
LLM-rewritten answers (used for v2) 35.0% (105 of 300)
  • The rewrite doubled the problem. For v2, each terse QASPER answer was rewritten by an LLM into a full explanatory sentence. The rewriter filled gaps with plausible-sounding details, despite being told not to add facts. The model then learned to do the same.
  • The raw data already had the problem. QASPER annotators wrote their answers after reading the whole paper, but the training context contains only the evidence spans they highlighted. About 1 in 6 examples therefore teaches the model to state facts that aren't in its context.

Planned fix, not yet done:

  1. Widen each training context to include the paragraphs around the evidence.
  2. Keep only targets verified as fully supported by their context.
  3. Retrain and re-run both evaluations.

Bias, risks, and limitations

  • Its answers can sound grounded without being grounded. The answers are fluent and come with citations, but they score only 70% on faithfulness to the context, below the untuned base model. And in production the citation is usually the paper's title, not evidence. Check answers against the paper.
  • It refuses too rarely. It declines on only 2–4 of 12 unanswerable questions, and otherwise answers anyway.
  • Narrow domain. It was trained and tested only on NLP papers from QASPER.
  • The evaluation relies on an LLM judge. All quality scores come from one judge model, not human raters. The benchmark is small (119 questions), and it has only 12 unanswerable questions, so the refusal rates are uncertain.
  • It inherits Qwen2.5's limitations and biases, as a 1.5B-parameter model.

Training details

Training data

whosouravsharma/paper-qa-qasper-sft was built from QASPER's train and validation splits (CC BY 4.0). Each annotator answer becomes one chat example:

  • Answerable. The context is the annotator's evidence quotes. The target is the answer (rewritten into a full sentence for v2), followed by Source: "<first evidence quote>".
  • Unanswerable. The context is 2 random paragraphs from the same paper, like a realistic retrieval miss, so the model learns to refuse on content rather than on an empty prompt. The target is a fixed refusal sentence.
Split Examples Refusal examples Papers Median answer length
train 2,589 281 878 23 words
validation 1,715 163 281 21 words

Training procedure

Method QLoRA with TRL SFTTrainer: base weights 4-bit NF4, LoRA adapters trained in fp16
Epochs 3 (486 optimizer steps)
Batch size 4 per step × 4 gradient-accumulation steps = 16
Learning rate 2e-4
Max sequence length 2,048 tokens (context + question + answer; p99 ≈ 1,262)
Evaluation Validation loss each epoch; the final epoch is kept

Training and validation loss

Epoch Validation loss Validation token accuracy
1 1.297 72.4%
2 1.287 72.6%
3 1.286 72.6%

Validation loss flattens after the first epoch. The loss measures how closely the model imitates the training targets. Since about a third of those targets contain unsupported claims, a lower loss doesn't mean better grounding.

state.json and loss_curve.json in this repository record the exact configuration and the full loss log.

Environmental impact

Hardware 1× NVIDIA L4 (Hugging Face Jobs, l4x1)
Training time 56 min (3,344 s)
Carbon emitted Not measured

Citation

Training and evaluation data come from QASPER:

@inproceedings{Dasigi2021ADO,
  title  = {A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers},
  author = {Pradeep Dasigi and Kyle Lo and Iz Beltagy and Arman Cohan and Noah A. Smith and Matt Gardner},
  year   = {2021}
}

Model card contact

Open a discussion in the Community tab.

Downloads last month
144
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for whosouravsharma/paper-qa-lora

Adapter
(1470)
this model

Datasets used to train whosouravsharma/paper-qa-lora

Spaces using whosouravsharma/paper-qa-lora 2

Evaluation results