Instructions to use harpertoken/quiz with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harpertoken/quiz with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("question-answering", model="harpertoken/quiz")# Load model directly from transformers import AutoTokenizer, AutoModelForQuestionAnswering tokenizer = AutoTokenizer.from_pretrained("harpertoken/quiz") model = AutoModelForQuestionAnswering.from_pretrained("harpertoken/quiz", device_map="auto") - Notebooks
- Google Colab
- Kaggle
quiz
A DistilBERT encoder fine-tuned for extractive question answering on SQuAD. Given a question and a passage, the model returns the span of text it identifies as the answer. It does not generate replies, and there is no conversation, retrieval or fallback logic anywhere in the repository — the weights are a single extractive span predictor.
The encoder is distilbert-base-uncased: six layers, 768 hidden dimensions, roughly 66M parameters, with a span-prediction head, stored in float32. Because the vocabulary is uncased WordPiece, input is lowercased and the model cannot distinguish Apple from apple. The 512-token position limit comes from DistilBERT and bounds the combined length of question and passage.
Usage
Transformers 5 removed the question-answering pipeline, so load the model directly:
import torch
from transformers import AutoModelForQuestionAnswering, AutoTokenizer
tok = AutoTokenizer.from_pretrained("harpertoken/quiz")
model = AutoModelForQuestionAnswering.from_pretrained("harpertoken/quiz")
question = "Who wrote Hamlet?"
context = "Hamlet is a tragedy written by William Shakespeare around 1600."
inputs = tok(question, context, return_tensors="pt", truncation=True, max_length=512)
with torch.inference_mode():
out = model(**inputs)
start, end = int(out.start_logits.argmax()), int(out.end_logits.argmax())
print(tok.decode(inputs.input_ids[0][start : end + 1]))
On three check questions the model returns paris, william shakespeare and 1889. Answers come back lowercased, since the input was.
The repository now ships only model.safetensors; an earlier version also carried the same parameters as pytorch_model.bin, which doubled the download for no benefit.
Limitations
SQuAD is a reading-comprehension benchmark assembled from English Wikipedia paragraphs, and a model trained on it inherits that distribution. Performance on specialised domains, on questions requiring multi-hop reasoning, and on passages substantially longer than a Wikipedia paragraph will be lower than the in-domain numbers suggest, and I have not measured them here. Earlier versions of this card quoted exact-match and F1 scores and described TensorFlow support and multi-strategy response generation. The scores came from no evaluation I can point to, and no TensorFlow checkpoint was ever present, so both have been removed. If you need a number, evaluate on your own data with the SQuAD v1.1 dev set as a reference point.
Attribution
DistilBERT is described in Sanh et al., DistilBERT, a distilled version of BERT (2019). SQuAD is described in Rajpurkar et al., SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016).
- Downloads last month
- 112
Model tree for harpertoken/quiz
Base model
distilbert/distilbert-base-uncased