srikanthgali/ai-text-detection-pile-cleaned
Viewer • Updated • 722k • 168 • 5
How to use praful1/roberta-ai-detector with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-classification", model="praful1/roberta-ai-detector") # pip install -U transformers accelerate
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tokenizer = AutoTokenizer.from_pretrained("praful1/roberta-ai-detector")
model = AutoModelForSequenceClassification.from_pretrained("praful1/roberta-ai-detector", device_map="auto")This model predicts whether text resembles the AI-generated or human-written examples it was trained on. The confidence score should not be interpreted as absolute proof that a text was written by AI or by a human. AI-text detection can produce false positives and false negatives, especially on text from domains or writing styles that differ from the training data.
The model supports inputs up to 512 tokens. Longer texts are truncated to the first 512 tokens.
# roberta-ai-detector
This model is a fine-tuned version of [FacebookAI/roberta-base](https://huggingface.co/FacebookAI/roberta-base) on an srikanthgali/ai-text-detection-pile-cleaned dataset.
It achieves the following results on the evaluation set:
- Loss: 0.2224
- Accuracy: 0.9694
- Precision: 0.9453
- Recall: 0.9968
- F1: 0.9704
- Roc Auc: 0.9983
## Model description
model context window is 512 tokens
## Intended uses & limitations
More information needed
## Training and evaluation data
More information needed
## Training procedure
### Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 2e-05
- train_batch_size: 16
- eval_batch_size: 16
- seed: 42
- optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lr_scheduler_type: linear
- num_epochs: 3
- mixed_precision_training: Native AMP
### Training results
| Training Loss | Epoch | Step | Validation Loss | Accuracy | Precision | Recall | F1 | Roc Auc |
|:-------------:|:-----:|:----:|:---------------:|:--------:|:---------:|:------:|:------:|:-------:|
| 0.1013 | 1.0 | 3125 | 0.2669 | 0.9476 | 0.9073 | 0.9976 | 0.9503 | 0.9963 |
| 0.0297 | 2.0 | 6250 | 0.1226 | 0.9732 | 0.9517 | 0.9972 | 0.9740 | 0.9989 |
| 0.0133 | 3.0 | 9375 | 0.2224 | 0.9694 | 0.9453 | 0.9968 | 0.9704 | 0.9983 |
```python
## Usage
### Install dependencies
```bash
pip install transformers torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
# Replace this with your Hugging Face model repository
MODEL_NAME = "praful1/roberta-ai-detector"
# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_NAME)
# Use GPU if available
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
model.eval()
def detect_ai(text):
# Tokenize input text
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=512
)
# Move tensors to the same device as the model
inputs = {
key: value.to(device)
for key, value in inputs.items()
}
# Run inference
with torch.no_grad():
outputs = model(**inputs)
# Convert logits to probabilities
probabilities = torch.softmax(outputs.logits, dim=-1)
# Get predicted class
predicted_class = torch.argmax(probabilities, dim=-1).item()
# Get label from model configuration
label = model.config.id2label[predicted_class]
# Get confidence
confidence = probabilities[0][predicted_class].item()
return {
"label": label,
"confidence": confidence,
"probabilities": {
model.config.id2label[i]: probabilities[0][i].item()
for i in range(len(probabilities[0]))
}
}
# Example
text = """
Artificial intelligence is transforming many industries by
automating repetitive tasks and helping people make better
decisions.
"""
result = detect_ai(text)
print("Prediction:", result["label"])
print("Confidence:", f"{result['confidence']:.2%}")
print("Probabilities:")
for label, probability in result["probabilities"].items():
print(f" {label}: {probability:.2%}")
Prediction: AI
Confidence: 97.34%
Probabilities:
HUMAN: 2.66%
AI: 97.34%
Simply replace the example:
text = """
Write or paste the text you want to analyze here.
"""
result = detect_ai(text)
print(result)
Base model
FacebookAI/roberta-base