q-trust-codebert / README.md
KRPur's picture
Upload README.md with huggingface_hub
68aee2d verified
|
Raw History Blame Contribute Delete
1.48 kB
---
license: other
license_name: qtrust-research
base_model: huggingface/codeberta-language-id
tags:
- cryptography
- post-quantum-cryptography
- code-classification
- crypto-discovery
metrics:
- f1
- precision
- recall
---
# Q-Trust CryptoCodeDetector (CodeBERTa fine-tune)
Fine-tuned crypto-usage discovery model from [Q-Trust](https://github.com/humoge7502/q-trust)
(`qtrust_ai/` intelligence layer). Detects cryptographic API usage and algorithm families in
source code β€” the discovery stage that feeds CBOM generation and PQC migration planning.
## Training
- **Corpus:** 13,973 real code files β€” SolidiFI, SmartBugs, EIPs, WebAuthn blockchain contracts, OSS crypto repos
- **Schedule:** 4-epoch GPU fine-tune (A100), deterministic seed (same seed β†’ same F1)
- **Dataset:** [`KRPur/q-trust-datasets`](https://huggingface.co/datasets/KRPur/q-trust-datasets) (`code_corpus.json`)
## Held-out results (repo-disjoint, n=2415)
| Metric | Q-Trust ensemble | Rules-only | Majority | Random |
|---|---|---|---|---|
| **F1** | **0.9525** | 0.673 | 0.8683 | 0.5981 |
| Precision | 0.952 | 0.979 | β€” | β€” |
| Recall | 0.953 | 0.513 | 1.0 | β€” |
Source: `qtrust_ai/artifacts/benchmark_comparison.json` (seed 42) in the GitHub repo.
## Usage
```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
m = AutoModelForSequenceClassification.from_pretrained("KRPur/q-trust-codebert")
t = AutoTokenizer.from_pretrained("KRPur/q-trust-codebert")
```