File size: 5,532 Bytes
fd5ffd9 5534cd7 8566721 fd5ffd9 5534cd7 ba10a59 8566721 5534cd7 ba10a59 8566721 ba10a59 8566721 5534cd7 8566721 5534cd7 8566721 ba10a59 8566721 5534cd7 ba10a59 5534cd7 8566721 5534cd7 ba10a59 8566721 5534cd7 ba10a59 5534cd7 8566721 5534cd7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 | ---
license: mit
language:
- en
- code
pipeline_tag: text-classification
tags:
- pytorch
- code-classification
- transformer
- slm
- binary-classification
---
# SourceCodeAuthorCheck-SLM-10M
[](https://huggingface.co/spaces/assix-research/SourceCodeAuthorCheck-UI)
A ~10 million parameter Small Language Model (SLM) Transformer designed to detect whether a Python source code file was written by a human or generated by an AI model.
## Live Demo
Test the model directly in your browser without writing code: **[SourceCodeAuthorCheck Web UI](https://huggingface.co/spaces/assix-research/SourceCodeAuthorCheck-UI)**
---
## 🔬 Training Pipeline & Techniques
This model was built from scratch using a custom PyTorch TransformerEncoder architecture (~9.6M parameters). The training pipeline utilized several specific techniques to ensure accurate binary classification and optimal hardware utilization.
### 1. Temporal Data Separation
To create a stark contrast between human and AI coding paradigms, the dataset relies on temporal splitting:
* **Human Baseline (Class 0):** Python source code extracted from GitHub repositories created in **Q3 2017 and prior**, guaranteeing the code predates modern generative AI.
* **GenAI Baseline (Class 1):** Synthetic datasets structured to mimic the exact architectural paradigms, repetitive docstrings, and token distributions typical of models operating in **Q3 2026**.
### 2. Hardware Optimization (NVIDIA DGX Spark)
The model was trained natively on an **NVIDIA DGX Spark (Grace Blackwell architecture)**.
* **Automatic Mixed Precision (AMP):** We utilized PyTorch's `torch.autocast` targeting `bfloat16`. This leverages Blackwell's 5th-generation Tensor Cores, accelerating matrix multiplications while maintaining numerical stability during backpropagation.
* **Gradient Scaling:** Paired with AMP, `torch.amp.GradScaler` was used to prevent underflow errors during the transition between FP32 and BF16 formats.
### 3. Optimization & Loss
* **Loss Function:** `BCEWithLogitsLoss`. This combines a Sigmoid layer and Binary Cross Entropy Loss in a single class, providing better numerical stability than applying Sigmoid followed by standard BCELoss.
* **Optimizer:** `AdamW` (Adam with Weight Decay) to enhance generalization and prevent overfitting on the synthetic AI subsets.
---
## 💻 Usage: The Inference Script
The easiest way to use this model locally is via the standalone `inference.py` script included in this repository. It includes the required architecture class and handles downloading the weights automatically.
**1. Download the script**
You can download the script directly from the files tab: [inference.py](https://huggingface.co/assix-research/SourceCodeAuthorCheck-SLM-10M/blob/main/inference.py)
**2. Run against any Python file**
Pass the path of the file you want to analyze directly to the script:
```bash
python inference.py my_script.py
```
**Example Output:**
```text
--- Testing File: my_script.py ---
Verdict: Human Written (AI Probability: 37.89%)
Preview: def process_data(items):...
```
---
## 🛠 Programmatic Usage
If you want to integrate the model directly into your own Python applications, you must define the architecture class before loading the weights.
```python
import torch
import torch.nn as nn
from transformers import AutoTokenizer
from huggingface_hub import hf_hub_download
# 1. Define the Architecture
class SourceCodeAuthorCheck(nn.Module):
def __init__(self, vocab_size=50257, d_model=128, nhead=8, num_layers=4, dim_feedforward=512):
super().__init__()
self.embedding = nn.Embedding(vocab_size, d_model)
self.pos_encoder = nn.Parameter(torch.zeros(1, 1024, d_model))
encoder_layers = nn.TransformerEncoderLayer(
d_model=d_model, nhead=nhead, dim_feedforward=dim_feedforward, batch_first=True
)
self.transformer = nn.TransformerEncoder(encoder_layers, num_layers=num_layers)
self.fc = nn.Linear(d_model, 1)
def forward(self, input_ids, attention_mask):
seq_len = input_ids.size(1)
x = self.embedding(input_ids) + self.pos_encoder[:, :seq_len, :]
src_key_padding_mask = ~attention_mask.bool()
x = self.transformer(x, src_key_padding_mask=src_key_padding_mask)
mask_expanded = attention_mask.unsqueeze(-1).float()
sum_embeddings = torch.sum(x * mask_expanded, 1)
sum_mask = torch.clamp(mask_expanded.sum(1), min=1e-9)
pooled = sum_embeddings / sum_mask
return self.fc(pooled)
# 2. Load Tokenizer and Model Weights
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token
model = SourceCodeAuthorCheck().to(device)
model_path = hf_hub_download(repo_id="assix-research/SourceCodeAuthorCheck-SLM-10M", filename="source_code_classifier.pth")
model.load_state_dict(torch.load(model_path, map_location=device, weights_only=True))
model.eval()
# 3. Analyze Code Snippet
code_snippet = "print('Hello World')"
inputs = tokenizer(
code_snippet, return_tensors="pt", truncation=True, padding="max_length", max_length=1024
).to(device)
with torch.no_grad():
logits = model(inputs['input_ids'], inputs['attention_mask'])
prob = torch.sigmoid(logits).item()
print(f"AI Probability: {prob:.1%}")
``` |