Hello-SimpleAI/HC3
Viewer • Updated • 48.6k • 4.39k • 228
How to use Aishkrish/hw1-hc3-detector with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-classification", model="Aishkrish/hw1-hc3-detector") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("Aishkrish/hw1-hc3-detector")
model = AutoModelForSequenceClassification.from_pretrained("Aishkrish/hw1-hc3-detector", device_map="auto")sentence-transformers/all-MiniLM-L6-v2 (22.7M params) fine-tuned for binary
classification of English answers: 0 = human, 1 = ChatGPT. Built for a course
homework; not a production detector.
Accuracy on the 4,668-answer HC3 test split:
| Model | Accuracy |
|---|---|
| Frozen embeddings + logistic regression | 84.49% |
| Fine-tuned, lr 2e-5, 5 epochs (the model in this repo) | 99.64% |
| Side run: lr 1e-5 | 98.71% |
| Side run: 7 epochs | 98.86% |
Confusion matrix for the model in this repo (rows = true human / ChatGPT): [[2322, 12], [5, 2329]], i.e. 17 mistakes out of 4,668. The baseline made 724.
The side runs were single runs with one seed and differ by only a few answers
Full fine-tuning with AdamW, lr 2e-5, 5 epochs, batch size 32, max length 256, seed 42, on the provided HC3 split (37,334 train / 4,666 validation / 4,668 test, split by question). Input is the answer text only.
from transformers import pipeline
clf = pipeline("text-classification", model="Aishkrish/hw1-hc3-detector")
clf("Your text here")
```​
```
The output labels appear as `LABEL_0` (human) and `LABEL_1` (ChatGPT).
Base model
nreimers/MiniLM-L6-H384-uncased