File size: 6,072 Bytes
6140b0b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
---
language:
- code
license: bigcode-openrail-m
library_name: transformers
pipeline_tag: feature-extraction
base_model: bigcode/starencoder
tags:
- syzkaller
- syz-program
- linux-kernel
- code-encoder
- masked-language-modeling
---

# SyzEncoder

SyzEncoder is an encoder for programs generated by the
[syzkaller](https://github.com/google/syzkaller) kernel fuzzer. It is based on
[StarEncoder](https://huggingface.co/bigcode/starencoder) and was further
pretrained with masked language modeling on 2,236,878 syz programs. The model
is used by SyzPilot as the base encoder for online reachability classifiers and
token-level attribution.

This repository contains the checkpoint from training step 90,000. It had the
lowest validation loss among the saved checkpoints. Only the encoder and its
tokenizer are included; the masked language modeling head used during
pretraining is not part of this release.

## Intended use

SyzEncoder is intended for representation learning and classification tasks on
syz programs. Typical uses include:

- initializing a reachability or coverage classifier;
- extracting sequence or token representations for attribution;
- studying machine-learning methods for kernel fuzzing.

The model does not generate syz programs. It has not been evaluated as a
general-purpose source-code or natural-language encoder.

## Model details

| Property | Value |
| --- | --- |
| Base model | `bigcode/starencoder` |
| Architecture | BERT encoder |
| Published parameters | 123,595,776 |
| Hidden size | 768 |
| Encoder layers | 12 |
| Attention heads | 12 |
| Maximum sequence length | 1,024 tokens |
| Training objective | Masked language modeling |
| Selected checkpoint | Step 90,000 |

The checkpoint does not contain pooler weights. Downstream code should either
load it with `add_pooling_layer=False` or train a task-specific pooling layer.
SyzPilot uses attention-mask-aware mean pooling.

## Tokenizer

The repository includes SyzTokenizer, a byte-level BPE tokenizer trained on the
same 2,236,878-program corpus. Its base vocabulary has 49,152 tokens, plus the
`<mask>` token used during pretraining.

Tokenizer selection included a grid search and a small masked-language-modeling
comparison. During tokenizer training, long repeated character runs and bare
hexadecimal runs were shortened before BPE learning, and token length was
capped at 64 characters. Hexadecimal literals beginning with `0x` were left
unchanged. These choices limit oversized tokens produced by raw byte dumps and
repetitive payloads while retaining common syz syntax.

## Training data

The training corpus contains 2,236,878 syz programs collected from fuzzing
Linux v6 kernels. Only program text was used for continued pretraining; kernel
coverage records and downstream reachability labels were not used.

The corpus was split into 90% training and 10% validation subsets. Validation
loss was estimated on 200 batches at each checkpoint evaluation.

## Training procedure

StarEncoder was further pretrained for three epochs with a 15% masking rate.
Masked positions followed the standard BERT policy: 80% were replaced by
`<mask>`, 10% by a random token, and 10% were left unchanged.

| Hyperparameter | Value |
| --- | --- |
| Batch size | 32 per GPU, 64 global |
| Learning rate | 2e-5 |
| Optimizer | AdamW, betas `(0.9, 0.98)` |
| Weight decay | 0.01 |
| Schedule | Cosine decay with 5% warmup |
| Gradient clipping | 1.0 |
| Sequence length | Up to 1,024 tokens, dynamic padding |
| Hardware | 2 NVIDIA A800 80GB GPUs |
| Training time | About 20.1 hours |

The full run completed approximately 94,300 optimizer steps. This release uses
the step-90,000 checkpoint because it produced the best sampled validation
loss.

## Evaluation

| Checkpoint | Validation loss | MLM perplexity |
| --- | ---: | ---: |
| Initial StarEncoder | Not recorded | 2.76 |
| SyzEncoder, step 90,000 | 0.7660 | 2.15 |

The validation metric measures the masked-language-modeling objective on the
held-out part of the pretraining corpus. It does not measure downstream
reachability classification accuracy. Results on unrelated code corpora should
not be inferred from these numbers.

## Usage

The example below obtains a mean-pooled representation while ignoring padding
tokens:

```python
import torch
from transformers import AutoModel, AutoTokenizer

model_id = "zzra1n/SyzEncoder"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, add_pooling_layer=False)
model.eval()

program = """\
r0 = socket$inet_tcp(0x2, 0x1, 0x0)
connect$inet(r0, &(0x7f0000000000)={0x2, 0x0, @loopback}, 0x10)
"""

inputs = tokenizer(
    program,
    return_tensors="pt",
    truncation=True,
    max_length=1024,
)

with torch.inference_mode():
    hidden = model(**inputs).last_hidden_state
    mask = inputs["attention_mask"].unsqueeze(-1).to(hidden.dtype)
    embedding = (hidden * mask).sum(dim=1) / mask.sum(dim=1).clamp(min=1)

print(embedding.shape)  # torch.Size([1, 768])
```

For supervised use, attach a classification head to the pooled representation
and fine-tune it on labels from the target fuzzing task.

## Limitations

- The training data comes from one domain and kernel generation. Programs from
  other syzkaller versions or substantially different syscall descriptions may
  tokenize and embed differently.
- Inputs longer than 1,024 tokens are truncated.
- The released checkpoint has been selected using MLM validation loss. It does
  not include a downstream classifier, calibrated probabilities, or a claim of
  performance on a particular kernel bug.
- Like its base model, SyzEncoder may retain unwanted behavior inherited from
  its pretraining data. Outputs used for security decisions should be checked
  against execution or coverage evidence.

## License

SyzEncoder is a derivative of StarEncoder and is released under the
[BigCode OpenRAIL-M license](https://huggingface.co/spaces/bigcode/bigcode-model-license-agreement).
Users are responsible for reviewing and following the license terms and use
restrictions.