Model Card for hw1-hc3-detector

This model is a fine-tuned version of sentence-transformers/all-MiniLM-L6-v2 trained to classify English question-answering texts as either human-written (label = 0) or ChatGPT-generated (label = 1) on the HC3 dataset.

  • Baseline Accuracy: 0.8451
  • Fine-Tuned Test Accuracy: 0.9895

Model Details

Model Description

This is the model card of a 🤗 transformers model that has been pushed on the Hub. It was fine-tuned end-to-end for binary sequence classification on the Human ChatGPT Comparison Corpus (HC3).

  • Developed by: Pradyumn Sharma
  • Funded by [optional]: N/A
  • Shared by [optional]: N/A
  • Model type: BertForSequenceClassification (Binary Text Classification)
  • Language(s) (NLP): English (en)
  • License: Apache-2.0
  • Finetuned from model [optional]: sentence-transformers/all-MiniLM-L6-v2

Model Sources [optional]

Uses

Direct Use

Binary classification of English Q&A responses into human-written (0) vs. ChatGPT-generated (1) text for evaluating AI text detection on the HC3 benchmark.

Downstream Use [optional]

Can be used as a baseline detector for studying robustness against adversarial paraphrasing or prompt steering.

Out-of-Scope Use

Not intended for high-stakes academic integrity enforcement or production moderation, as it is trained solely on HC3 (casual forum answers vs. default GPT-3.5 outputs) and may misclassify formal human writing or prompted LLM outputs.

Bias, Risks, and Limitations

The HC3 dataset pairs casual human forum responses (e.g., Reddit ELI5) with formal, unsteered ChatGPT responses. As a result, the model may rely on stylistic and formatting artifacts (such as casual slang vs. polite, structured explanations) rather than generalizable properties of AI-generated text. Input texts are also truncated to a maximum length of 256 tokens.

Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Predictions should be evaluated alongside out-of-distribution and adversarially prompted test sets before drawing conclusions about general detection performance.

How to Get Started with the Model

Use the code below to get started with the model.

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

repo_id = "YOUR_USERNAME/hw1-hc3-detector"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSequenceClassification.from_pretrained(repo_id)
model.eval()

inputs = tokenizer("Your input text here", return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
    logits = model(**inputs).logits
    pred = torch.argmax(logits, dim=-1).item()

print("Predicted label:", "human" if pred == 0 else "ChatGPT")

Training Details

Training Data

Trained on cleaned and deduplicated question-answer pairs from Hello-SimpleAI/HC3 (all.jsonl, revision 4d0ff18143b5a7e1b1e79beb540c04549d1e59d3), split 80/10/10 by question ID (seed = 42):

  • Train: 37,334 examples (18,667 human, 18,667 ChatGPT)
  • Validation: 4,666 examples
  • Test: 4,668 examples (2,334 human, 2,334 ChatGPT)

Training Procedure

Preprocessing [optional]

Texts were tokenized with AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2") using truncation=True, max_length=256, and dynamic batch padding via DataCollatorWithPadding.

Training Hyperparameters

  • Training regime: fp32
  • Optimizer: AdamW (lr = 2e-5)
  • Epochs: 5
  • Batch size: 32
  • Max sequence length: 256
  • Seed: 42

Speeds, Sizes, Times [optional]

Trained for 5 epochs (1,167 steps per epoch) on local hardware (~4.3 hours total).

Evaluation

Testing Data, Factors & Metrics

Testing Data

Evaluated on the held-out HC3 test split (4,668 total examples: 2,334 human (0) and 2,334 ChatGPT (1)) from Hello-SimpleAI/HC3.

Factors

Evaluated across balanced binary classes (0 = human, 1 = ChatGPT).

Metrics

Test Accuracy, along with micro-averaged and macro-averaged Precision, Recall, and F1-score.

Results

Baseline Accuracy (Frozen all-MiniLM-L6-v2 embeddings + LogisticRegression): 0.8451 (84.51%)

  • precision_micro: 0.8451 | recall_micro: 0.8451 | f1_micro: 0.8451
  • precision_macro: 0.8453 | recall_macro: 0.8451 | f1_macro: 0.8451

Fine-Tuned Test Accuracy (AutoModelForSequenceClassification): 0.9895 (98.95%)

Summary

End-to-end fine-tuning of sentence-transformers/all-MiniLM-L6-v2 increased test accuracy from the frozen LogisticRegression baseline of 0.8451 to 0.9895 on the HC3 test split.

Model Examination [optional]

N/A

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019).

  • Hardware Type: Apple Silicon CPU/MPS
  • Hours used: ~4.3 hours
  • Cloud Provider: None (Local Machine)
  • Compute Region: United States
  • Carbon Emitted: N/A

Technical Specifications [optional]

Model Architecture and Objective

6-layer MiniLM (BertForSequenceClassification) encoder with a 2-class linear classification head trained using Cross-Entropy Loss.

Compute Infrastructure

Local workstation environment.

Hardware

Apple Silicon Mac.

Software

Python 3.14, PyTorch, Hugging Face transformers, datasets, sentence-transformers, and scikit-learn.

Citation [optional]

BibTeX:

@article{guo2023close,
  title={How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection},
  author={Guo, Biyang and Zhang, Xin and Wang, Ziyuan and Jiang, Minqi and Nie, Jinran and Ding, Yuxuan and Yue, Jianwei and Wu, Yupeng},
  journal={arXiv preprint arXiv:2301.07597},
  year={2023}
}

APA:

Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., & Wu, Y. (2023). How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection. arXiv preprint arXiv:2301.07597.

Glossary [optional]

  • Label 0: Human-written answer
  • Label 1: ChatGPT-generated answer

More Information [optional]

N/A

Model Card Authors [optional]

Pradyumn Sharma

Model Card Contact

N/A

Downloads last month
18
Safetensors
Model size
22.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Sharma-2/hw1-hc3-detector

Dataset used to train Sharma-2/hw1-hc3-detector

Papers for Sharma-2/hw1-hc3-detector