tanaymitra01 commited on
Commit
cc1796d
·
verified ·
1 Parent(s): 0345e18

Remove Mermaid diagrams; keep tables and text only

Browse files
Files changed (1) hide show
  1. README.md +31 -168
README.md CHANGED
@@ -39,145 +39,32 @@ Fine-tuned [`microsoft/graphcodebert-base`](https://huggingface.co/microsoft/gra
39
 
40
  Part of [SolidityGuard](https://github.com/tanaymitra54/solidity_guard_razorpay) — used as a first-pass detector alongside Slither and an LLM auditor.
41
 
42
- ## Architecture overview
43
 
44
- How this model sits inside the SolidityGuard audit stack:
 
 
 
45
 
46
- ```mermaid
47
- flowchart TB
48
- subgraph Input
49
- SOL["Solidity source (.sol)"]
50
- end
51
 
52
- subgraph Detectors["First-pass detectors"]
53
- SL["Slither / patterns"]
54
- GCB["Graph CodeBERT<br/>this model"]
55
- end
56
 
57
- subgraph LLM["LLM agents"]
58
- SCAN["Scanner"]
59
- AN["Analyzer"]
60
- EX["Exploit Gen"]
61
- FX["Fix Suggester"]
62
- end
63
 
64
- subgraph Out["Output"]
65
- REP["Audit findings<br/>+ severity + confidence"]
66
- end
67
-
68
- SOL --> SL
69
- SOL --> GCB
70
- SOL --> SCAN
71
- SL --> SCAN
72
- GCB --> SCAN
73
- SCAN --> AN
74
- AN --> EX
75
- AN --> FX
76
- SCAN --> REP
77
- AN --> REP
78
- EX --> REP
79
- FX --> REP
80
- ```
81
-
82
- ## Model internals
83
-
84
- Sequence-classification head on Graph CodeBERT (RoBERTa-style encoder):
85
-
86
- ```mermaid
87
- flowchart LR
88
- A["Solidity source<br/>string"] --> B["Tokenizer<br/>max_length=512"]
89
- B --> C["input_ids<br/>attention_mask"]
90
- C --> D["Graph CodeBERT<br/>encoder<br/>~125M params"]
91
- D --> E["[CLS] pooled<br/>representation"]
92
- E --> F["Linear classifier<br/>12 logits"]
93
- F --> G["Softmax"]
94
- G --> H["label + confidence"]
95
- ```
96
-
97
- ## Training pipeline
98
-
99
- End-to-end fine-tune path used for this checkpoint:
100
-
101
- ```mermaid
102
- flowchart TD
103
- W["SmartBugs-Wild<br/>~47k contracts"] --> P["Parse + cap<br/>WILD_LIMIT=5000"]
104
- R["SmartBugs-Results<br/>results_wild.json"] --> V["Tool consensus<br/>≥2 tools agree"]
105
- P --> M["Labeled samples"]
106
- V --> M
107
- M --> S["Split 70 / 15 / 15<br/>train / val / test"]
108
- S --> AUG["Light augmentation"]
109
- AUG --> FT["Fine-tune<br/>Graph CodeBERT-base"]
110
- FT --> ES["Early stopping<br/>on val macro-F1"]
111
- ES --> BEST["checkpoints/best<br/>→ this Hub repo"]
112
- ```
113
-
114
- ## Data labeling logic
115
-
116
- ```mermaid
117
- flowchart LR
118
- C["Contract address"] --> T1["Tool A categories"]
119
- C --> T2["Tool B categories"]
120
- C --> T3["Tool N categories"]
121
- T1 --> VOTE["Vote per category"]
122
- T2 --> VOTE
123
- T3 --> VOTE
124
- VOTE --> Q{"votes ≥ 2?"}
125
- Q -->|yes| LBL["Mapped vuln label<br/>e.g. reentrancy"]
126
- Q -->|no| SAFE["safe"]
127
- ```
128
-
129
- ## Inference sequence
130
-
131
- ```mermaid
132
- sequenceDiagram
133
- participant U as Caller / API
134
- participant D as GraphCodeBERTDetector
135
- participant H as Hugging Face Hub
136
- participant M as Model (GPU/CPU)
137
-
138
- U->>D: scan(source_code)
139
- alt first load
140
- D->>H: from_pretrained(repo_id)
141
- H-->>D: weights + config + tokenizer
142
- D->>M: eval mode
143
- end
144
- D->>M: tokenize + forward
145
- M-->>D: logits → softmax
146
- alt label != safe AND conf ≥ threshold
147
- D-->>U: Finding(issue_type, severity, confidence)
148
- else
149
- D-->>U: [] (no finding)
150
- end
151
- ```
152
-
153
- ## Label taxonomy
154
-
155
- ```mermaid
156
- mindmap
157
- root((primary_vuln))
158
- Safe
159
- safe
160
- Critical-leaning
161
- reentrancy
162
- access_control
163
- tx_origin_auth
164
- integer_overflow
165
- unsafe_delegatecall
166
- Medium / other
167
- weak_randomness
168
- unbounded_loop
169
- other
170
- Style / gas
171
- redundant_storage
172
- gas_optimization
173
- best_practice
174
- ```
175
 
176
  ## Intended use
177
 
178
- - Input: Solidity source code (string)
179
- - Output: one of 12 labels (see below) + confidence
180
- - Best as a **screening** signal, not a sole security audit
181
 
182
  ## Labels
183
 
@@ -196,11 +83,6 @@ mindmap
196
  | 10 | `best_practice` | Low |
197
  | 11 | `other` | Medium |
198
 
199
- ## Training data
200
-
201
- - **SmartBugs-Wild** contracts with tool-consensus labels from [SmartBugs results](https://github.com/smartbugs/smartbugs-results) (`metadata/results_wild.json`)
202
- - Cap: 5,000 contracts; train/val/test split 70/15/15 with light augmentation
203
-
204
  ## Held-out test metrics
205
 
206
  | Metric | Value |
@@ -213,15 +95,6 @@ mindmap
213
  | F1 `other` | 0.340 |
214
  | F1 `access_control` | 0.200 |
215
 
216
- ```mermaid
217
- %%{init: {"theme": "neutral"}}%%
218
- xychart-beta
219
- title "Per-class F1 on held-out test"
220
- x-axis [safe, integer_overflow, reentrancy, other, access_control]
221
- y-axis "F1" 0 --> 1
222
- bar [0.694, 0.634, 0.461, 0.340, 0.200]
223
- ```
224
-
225
  Labels are noisy (static-analysis consensus), so scores are moderate by design.
226
 
227
  ## Quick start
@@ -260,36 +133,26 @@ print(model.config.id2label[pred], float(probs[pred]))
260
  export GRAPHCODEBERT_PATH=tanaymitra01/graphcodebert-vulnerability-detector
261
  ```
262
 
263
- ## Deploy / access map
264
 
265
- ```mermaid
266
- flowchart TB
267
- CKPT["Local checkpoint<br/>training/checkpoints/best"]
268
- GH["GitHub LFS<br/>tanaymitra54/solidity_guard_razorpay"]
269
- HF["Hugging Face Hub<br/>tanaymitra01/graphcodebert-vulnerability-detector"]
270
-
271
- CKPT -->|git lfs push| GH
272
- CKPT -->|hf upload| HF
273
-
274
- HF --> COLAB["Colab / scripts"]
275
- HF --> SPACE["HF Space / API host"]
276
- HF --> LOCAL["Any machine<br/>from_pretrained"]
277
- GH --> DEV["Dev clones"]
278
- ```
279
 
280
  ## Files
281
 
282
- - `model.safetensors` — weights
283
- - `config.json` — RobertaForSequenceClassification config + label maps
284
- - `label_map.json` — label list / id maps used in training
285
- - `README.md` — this model card
286
 
287
  ## Limitations
288
 
289
- - Tool-derived labels ≠ audited ground truth
290
- - Truncation at 512 tokens; large contracts lose context
291
- - Rare classes (e.g. access control) have low F1
292
- - Not a replacement for professional security review
293
 
294
  ## Citation
295
 
 
39
 
40
  Part of [SolidityGuard](https://github.com/tanaymitra54/solidity_guard_razorpay) — used as a first-pass detector alongside Slither and an LLM auditor.
41
 
42
+ ## How it fits in SolidityGuard
43
 
44
+ 1. **Input:** Solidity source
45
+ 2. **Detectors:** Slither / patterns + **this Graph CodeBERT model**
46
+ 3. **LLM agents:** Scanner → Analyzer → Exploit Gen / Fix Suggester
47
+ 4. **Output:** Audit findings with severity and confidence
48
 
49
+ ## Model
 
 
 
 
50
 
51
+ - Tokenizer → `max_length=512`
52
+ - Graph CodeBERT encoder (~125M params)
53
+ - Linear classification head → 12-class softmax → label + confidence
 
54
 
55
+ ## Training
 
 
 
 
 
56
 
57
+ 1. SmartBugs-Wild contracts (capped at 5,000)
58
+ 2. Labels from SmartBugs-Results tool consensus (≥2 tools agree on a category; else `safe`)
59
+ 3. Split 70 / 15 / 15 (train / val / test) with light augmentation
60
+ 4. Fine-tune `microsoft/graphcodebert-base` with early stopping on validation macro-F1
61
+ 5. Best checkpoint published here
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
 
63
  ## Intended use
64
 
65
+ - Input: Solidity source code (string)
66
+ - Output: one of 12 labels + confidence
67
+ - Best as a **screening** signal, not a sole security audit
68
 
69
  ## Labels
70
 
 
83
  | 10 | `best_practice` | Low |
84
  | 11 | `other` | Medium |
85
 
 
 
 
 
 
86
  ## Held-out test metrics
87
 
88
  | Metric | Value |
 
95
  | F1 `other` | 0.340 |
96
  | F1 `access_control` | 0.200 |
97
 
 
 
 
 
 
 
 
 
 
98
  Labels are noisy (static-analysis consensus), so scores are moderate by design.
99
 
100
  ## Quick start
 
133
  export GRAPHCODEBERT_PATH=tanaymitra01/graphcodebert-vulnerability-detector
134
  ```
135
 
136
+ ## Where to get the weights
137
 
138
+ | Location | Path |
139
+ |----------|------|
140
+ | **Hugging Face (recommended)** | [`tanaymitra01/graphcodebert-vulnerability-detector`](https://huggingface.co/tanaymitra01/graphcodebert-vulnerability-detector) |
141
+ | GitHub LFS | [`training/checkpoints/best/`](https://github.com/tanaymitra54/solidity_guard_razorpay/tree/main/training/checkpoints/best) in the SolidityGuard repo |
 
 
 
 
 
 
 
 
 
 
142
 
143
  ## Files
144
 
145
+ - `model.safetensors` — weights
146
+ - `config.json` — RobertaForSequenceClassification config + label maps
147
+ - `label_map.json` — label list / id maps used in training
148
+ - `README.md` — this model card
149
 
150
  ## Limitations
151
 
152
+ - Tool-derived labels ≠ audited ground truth
153
+ - Truncation at 512 tokens; large contracts lose context
154
+ - Rare classes (e.g. access control) have low F1
155
+ - Not a replacement for professional security review
156
 
157
  ## Citation
158