Important

This model predicts whether text resembles the AI-generated or human-written examples it was trained on. The confidence score should not be interpreted as absolute proof that a text was written by AI or by a human. AI-text detection can produce false positives and false negatives, especially on text from domains or writing styles that differ from the training data.

The model supports inputs up to 512 tokens. Longer texts are truncated to the first 512 tokens.


# roberta-ai-detector

This model is a fine-tuned version of [FacebookAI/roberta-base](https://huggingface.co/FacebookAI/roberta-base) on an srikanthgali/ai-text-detection-pile-cleaned dataset.
It achieves the following results on the evaluation set:
- Loss: 0.2224
- Accuracy: 0.9694
- Precision: 0.9453
- Recall: 0.9968
- F1: 0.9704
- Roc Auc: 0.9983

## Model description

model context window is 512 tokens

## Intended uses & limitations

More information needed

## Training and evaluation data

More information needed

## Training procedure

### Training hyperparameters

The following hyperparameters were used during training:
- learning_rate: 2e-05
- train_batch_size: 16
- eval_batch_size: 16
- seed: 42
- optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lr_scheduler_type: linear
- num_epochs: 3
- mixed_precision_training: Native AMP

### Training results

| Training Loss | Epoch | Step | Validation Loss | Accuracy | Precision | Recall | F1     | Roc Auc |
|:-------------:|:-----:|:----:|:---------------:|:--------:|:---------:|:------:|:------:|:-------:|
| 0.1013        | 1.0   | 3125 | 0.2669          | 0.9476   | 0.9073    | 0.9976 | 0.9503 | 0.9963  |
| 0.0297        | 2.0   | 6250 | 0.1226          | 0.9732   | 0.9517    | 0.9972 | 0.9740 | 0.9989  |
| 0.0133        | 3.0   | 9375 | 0.2224          | 0.9694   | 0.9453    | 0.9968 | 0.9704 | 0.9983  |



```python
## Usage

### Install dependencies

```bash
pip install transformers torch

Detect whether text is AI-generated

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

# Replace this with your Hugging Face model repository
MODEL_NAME = "praful1/roberta-ai-detector"

# Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_NAME)

# Use GPU if available
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
model.eval()


def detect_ai(text):
    # Tokenize input text
    inputs = tokenizer(
        text,
        return_tensors="pt",
        truncation=True,
        max_length=512
    )

    # Move tensors to the same device as the model
    inputs = {
        key: value.to(device)
        for key, value in inputs.items()
    }

    # Run inference
    with torch.no_grad():
        outputs = model(**inputs)

    # Convert logits to probabilities
    probabilities = torch.softmax(outputs.logits, dim=-1)

    # Get predicted class
    predicted_class = torch.argmax(probabilities, dim=-1).item()

    # Get label from model configuration
    label = model.config.id2label[predicted_class]

    # Get confidence
    confidence = probabilities[0][predicted_class].item()

    return {
        "label": label,
        "confidence": confidence,
        "probabilities": {
            model.config.id2label[i]: probabilities[0][i].item()
            for i in range(len(probabilities[0]))
        }
    }


# Example
text = """
Artificial intelligence is transforming many industries by
automating repetitive tasks and helping people make better
decisions.
"""

result = detect_ai(text)

print("Prediction:", result["label"])
print("Confidence:", f"{result['confidence']:.2%}")
print("Probabilities:")

for label, probability in result["probabilities"].items():
    print(f"  {label}: {probability:.2%}")

Example output

Prediction: AI
Confidence: 97.34%
Probabilities:
  HUMAN: 2.66%
  AI: 97.34%

Using your own text

Simply replace the example:

text = """
Write or paste the text you want to analyze here.
"""

result = detect_ai(text)
print(result)

Framework versions

  • Transformers 5.16.1
  • Pytorch 2.11.0+cu128
  • Datasets 4.8.5
  • Tokenizers 0.23.1
Downloads last month
64
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for praful1/roberta-ai-detector

Finetuned
(2369)
this model

Dataset used to train praful1/roberta-ai-detector