File size: 1,483 Bytes
68aee2d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
---
license: other
license_name: qtrust-research
base_model: huggingface/codeberta-language-id
tags:
- cryptography
- post-quantum-cryptography
- code-classification
- crypto-discovery
metrics:
- f1
- precision
- recall
---

# Q-Trust CryptoCodeDetector (CodeBERTa fine-tune)

Fine-tuned crypto-usage discovery model from [Q-Trust](https://github.com/humoge7502/q-trust)
(`qtrust_ai/` intelligence layer). Detects cryptographic API usage and algorithm families in
source code — the discovery stage that feeds CBOM generation and PQC migration planning.

## Training

- **Corpus:** 13,973 real code files — SolidiFI, SmartBugs, EIPs, WebAuthn blockchain contracts, OSS crypto repos
- **Schedule:** 4-epoch GPU fine-tune (A100), deterministic seed (same seed → same F1)
- **Dataset:** [`KRPur/q-trust-datasets`](https://huggingface.co/datasets/KRPur/q-trust-datasets) (`code_corpus.json`)

## Held-out results (repo-disjoint, n=2415)

| Metric | Q-Trust ensemble | Rules-only | Majority | Random |
|---|---|---|---|---|
| **F1** | **0.9525** | 0.673 | 0.8683 | 0.5981 |
| Precision | 0.952 | 0.979 | — | — |
| Recall | 0.953 | 0.513 | 1.0 | — |

Source: `qtrust_ai/artifacts/benchmark_comparison.json` (seed 42) in the GitHub repo.

## Usage

```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
m = AutoModelForSequenceClassification.from_pretrained("KRPur/q-trust-codebert")
t = AutoTokenizer.from_pretrained("KRPur/q-trust-codebert")
```