Feature Extraction
Transformers
Safetensors
code
bert
syzkaller
syz-program
linux-kernel
code-encoder
masked-language-modeling
text-embeddings-inference
Instructions to use zzra1n/SyzEncoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zzra1n/SyzEncoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="zzra1n/SyzEncoder")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("zzra1n/SyzEncoder") model = AutoModel.from_pretrained("zzra1n/SyzEncoder", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 6,072 Bytes
6140b0b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 | ---
language:
- code
license: bigcode-openrail-m
library_name: transformers
pipeline_tag: feature-extraction
base_model: bigcode/starencoder
tags:
- syzkaller
- syz-program
- linux-kernel
- code-encoder
- masked-language-modeling
---
# SyzEncoder
SyzEncoder is an encoder for programs generated by the
[syzkaller](https://github.com/google/syzkaller) kernel fuzzer. It is based on
[StarEncoder](https://huggingface.co/bigcode/starencoder) and was further
pretrained with masked language modeling on 2,236,878 syz programs. The model
is used by SyzPilot as the base encoder for online reachability classifiers and
token-level attribution.
This repository contains the checkpoint from training step 90,000. It had the
lowest validation loss among the saved checkpoints. Only the encoder and its
tokenizer are included; the masked language modeling head used during
pretraining is not part of this release.
## Intended use
SyzEncoder is intended for representation learning and classification tasks on
syz programs. Typical uses include:
- initializing a reachability or coverage classifier;
- extracting sequence or token representations for attribution;
- studying machine-learning methods for kernel fuzzing.
The model does not generate syz programs. It has not been evaluated as a
general-purpose source-code or natural-language encoder.
## Model details
| Property | Value |
| --- | --- |
| Base model | `bigcode/starencoder` |
| Architecture | BERT encoder |
| Published parameters | 123,595,776 |
| Hidden size | 768 |
| Encoder layers | 12 |
| Attention heads | 12 |
| Maximum sequence length | 1,024 tokens |
| Training objective | Masked language modeling |
| Selected checkpoint | Step 90,000 |
The checkpoint does not contain pooler weights. Downstream code should either
load it with `add_pooling_layer=False` or train a task-specific pooling layer.
SyzPilot uses attention-mask-aware mean pooling.
## Tokenizer
The repository includes SyzTokenizer, a byte-level BPE tokenizer trained on the
same 2,236,878-program corpus. Its base vocabulary has 49,152 tokens, plus the
`<mask>` token used during pretraining.
Tokenizer selection included a grid search and a small masked-language-modeling
comparison. During tokenizer training, long repeated character runs and bare
hexadecimal runs were shortened before BPE learning, and token length was
capped at 64 characters. Hexadecimal literals beginning with `0x` were left
unchanged. These choices limit oversized tokens produced by raw byte dumps and
repetitive payloads while retaining common syz syntax.
## Training data
The training corpus contains 2,236,878 syz programs collected from fuzzing
Linux v6 kernels. Only program text was used for continued pretraining; kernel
coverage records and downstream reachability labels were not used.
The corpus was split into 90% training and 10% validation subsets. Validation
loss was estimated on 200 batches at each checkpoint evaluation.
## Training procedure
StarEncoder was further pretrained for three epochs with a 15% masking rate.
Masked positions followed the standard BERT policy: 80% were replaced by
`<mask>`, 10% by a random token, and 10% were left unchanged.
| Hyperparameter | Value |
| --- | --- |
| Batch size | 32 per GPU, 64 global |
| Learning rate | 2e-5 |
| Optimizer | AdamW, betas `(0.9, 0.98)` |
| Weight decay | 0.01 |
| Schedule | Cosine decay with 5% warmup |
| Gradient clipping | 1.0 |
| Sequence length | Up to 1,024 tokens, dynamic padding |
| Hardware | 2 NVIDIA A800 80GB GPUs |
| Training time | About 20.1 hours |
The full run completed approximately 94,300 optimizer steps. This release uses
the step-90,000 checkpoint because it produced the best sampled validation
loss.
## Evaluation
| Checkpoint | Validation loss | MLM perplexity |
| --- | ---: | ---: |
| Initial StarEncoder | Not recorded | 2.76 |
| SyzEncoder, step 90,000 | 0.7660 | 2.15 |
The validation metric measures the masked-language-modeling objective on the
held-out part of the pretraining corpus. It does not measure downstream
reachability classification accuracy. Results on unrelated code corpora should
not be inferred from these numbers.
## Usage
The example below obtains a mean-pooled representation while ignoring padding
tokens:
```python
import torch
from transformers import AutoModel, AutoTokenizer
model_id = "zzra1n/SyzEncoder"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id, add_pooling_layer=False)
model.eval()
program = """\
r0 = socket$inet_tcp(0x2, 0x1, 0x0)
connect$inet(r0, &(0x7f0000000000)={0x2, 0x0, @loopback}, 0x10)
"""
inputs = tokenizer(
program,
return_tensors="pt",
truncation=True,
max_length=1024,
)
with torch.inference_mode():
hidden = model(**inputs).last_hidden_state
mask = inputs["attention_mask"].unsqueeze(-1).to(hidden.dtype)
embedding = (hidden * mask).sum(dim=1) / mask.sum(dim=1).clamp(min=1)
print(embedding.shape) # torch.Size([1, 768])
```
For supervised use, attach a classification head to the pooled representation
and fine-tune it on labels from the target fuzzing task.
## Limitations
- The training data comes from one domain and kernel generation. Programs from
other syzkaller versions or substantially different syscall descriptions may
tokenize and embed differently.
- Inputs longer than 1,024 tokens are truncated.
- The released checkpoint has been selected using MLM validation loss. It does
not include a downstream classifier, calibrated probabilities, or a claim of
performance on a particular kernel bug.
- Like its base model, SyzEncoder may retain unwanted behavior inherited from
its pretraining data. Outputs used for security decisions should be checked
against execution or coverage evidence.
## License
SyzEncoder is a derivative of StarEncoder and is released under the
[BigCode OpenRAIL-M license](https://huggingface.co/spaces/bigcode/bigcode-model-license-agreement).
Users are responsible for reviewing and following the license terms and use
restrictions.
|