quiz

A DistilBERT encoder fine-tuned for extractive question answering on SQuAD. Given a question and a passage, the model returns the span of text it identifies as the answer. It does not generate replies, and there is no conversation, retrieval or fallback logic anywhere in the repository — the weights are a single extractive span predictor.

The encoder is distilbert-base-uncased: six layers, 768 hidden dimensions, roughly 66M parameters, with a span-prediction head, stored in float32. Because the vocabulary is uncased WordPiece, input is lowercased and the model cannot distinguish Apple from apple. The 512-token position limit comes from DistilBERT and bounds the combined length of question and passage.

Usage

Transformers 5 removed the question-answering pipeline, so load the model directly:

import torch
from transformers import AutoModelForQuestionAnswering, AutoTokenizer

tok = AutoTokenizer.from_pretrained("harpertoken/quiz")
model = AutoModelForQuestionAnswering.from_pretrained("harpertoken/quiz")

question = "Who wrote Hamlet?"
context = "Hamlet is a tragedy written by William Shakespeare around 1600."
inputs = tok(question, context, return_tensors="pt", truncation=True, max_length=512)

with torch.inference_mode():
    out = model(**inputs)
start, end = int(out.start_logits.argmax()), int(out.end_logits.argmax())
print(tok.decode(inputs.input_ids[0][start : end + 1]))

On three check questions the model returns paris, william shakespeare and 1889. Answers come back lowercased, since the input was.

The repository now ships only model.safetensors; an earlier version also carried the same parameters as pytorch_model.bin, which doubled the download for no benefit.

Limitations

SQuAD is a reading-comprehension benchmark assembled from English Wikipedia paragraphs, and a model trained on it inherits that distribution. Performance on specialised domains, on questions requiring multi-hop reasoning, and on passages substantially longer than a Wikipedia paragraph will be lower than the in-domain numbers suggest, and I have not measured them here. Earlier versions of this card quoted exact-match and F1 scores and described TensorFlow support and multi-strategy response generation. The scores came from no evaluation I can point to, and no TensorFlow checkpoint was ever present, so both have been removed. If you need a number, evaluate on your own data with the SQuAD v1.1 dev set as a reference point.

Attribution

DistilBERT is described in Sanh et al., DistilBERT, a distilled version of BERT (2019). SQuAD is described in Rajpurkar et al., SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016).

Downloads last month
112
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for harpertoken/quiz

Finetuned
(12565)
this model

Dataset used to train harpertoken/quiz