chat

openai-community/gpt2 further fine-tuned on DailyDialog, a corpus of short multi-turn English conversations about everyday topics. The base model is the 124M-parameter GPT-2: config.json records _name_or_path: gpt2, twelve layers, 768 embedding dimensions. The previous version of this card declared its base model as gpt2, an identifier that has since moved; openai-community/gpt2 is the current one.

This is a continuation, not a chat model. Nothing here establishes turn-taking, instruction following or any conversational protocol: there is no chat template in tokenizer_config.json, and the only token added beyond GPT-2's vocabulary is a [PAD] entry at id 50257, which takes vocab_size to 50258 against GPT-2's 50257. It produces fluent English that resembles dialogue because that is what it was trained to imitate, which is a weaker property than it sounds.

One detail about the checkpoint itself. It carries twelve transformer.h.*.attn.masked_bias tensors that current Transformers does not define: the key was removed from GPT2Attention years ago, and the upstream gpt2 weights contain no such tensor. Loading this repository therefore prints an UNEXPECTED warning for those twelve keys on every run. They are inert, since nothing reads them, and generation is unaffected. They are left in place rather than silently stripped, so the file matches whatever produced it.

The card this replaces documented a generate_response.py script and a FastAPI service on localhost:8000 with /status, /chat and interactive docs endpoints. None of that is in the repository, which contains weights and tokenizer files only, so those sections have been removed. The declared perplexity, BLEU and F1 metrics likewise had no values behind them; no evaluation is recorded for this checkpoint.

Usage

from transformers import pipeline

generator = pipeline("text-generation", model="harpertoken/chat")
print(generator("Hello, how are you?", max_new_tokens=60)[0]["generated_text"])

A pad_token is not configured, so batching prompts together will warn. Set tokenizer.pad_token = tokenizer.eos_token if you need it.

Limitations

DailyDialog is scripted, crowd-sourced, and narrow: 13k conversations of polite small talk, annotated for emotion and communication acts. A model trained on it will handle greetings and farewells far better than disagreement, and will reproduce the register of that corpus rather than anything wider. Because GPT-2 is a 124M-parameter base and the fine-tuning is small, output is frequently fluent and wrong. The biases of both GPT-2 and DailyDialog carry through unchanged.

Attribution

GPT-2 follows Radford et al. (2018). DailyDialog is described in Li et al., DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset, IJCNLP 2017.

Downloads last month
512
Safetensors
Model size
0.1B params
Tensor type
F32
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for harpertoken/chat

Finetuned
(2279)
this model

Dataset used to train harpertoken/chat