sajalmadan09 commited on
Commit
221ca55
·
verified ·
1 Parent(s): fafb1b5

Initial publish: native C++ port of prajjwal1/bert-tiny (Sajal Labs exp11-14)

Browse files
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ model.onnx.data filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,129 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - en
5
+ tags:
6
+ - bert
7
+ - native-inference
8
+ - sajal-labs
9
+ - transformer
10
+ - feature-extraction
11
+ base_model: prajjwal1/bert-tiny
12
+ pipeline_tag: feature-extraction
13
+ ---
14
+
15
+ # bert-tiny — Native C++ Port (Sajal Labs)
16
+
17
+ **This is not a new model.** The weights, architecture, and pretraining are
18
+ entirely [`prajjwal1/bert-tiny`](https://huggingface.co/prajjwal1/bert-tiny) by Prajjwal
19
+ Bhargava (MIT license) — a compact pretrained BERT encoder introduced in
20
+ Turc et al. 2019 ("Well-Read Students Learn Better") and ported to
21
+ Hugging Face for Bhargava et al. 2021 ("Generalization in NLI"). **Please
22
+ cite both papers if you use this model** (citations below).
23
+
24
+ **What Sajal Labs added**: a from-scratch native C++ port of the encoder
25
+ and the WordPiece tokenizer — no PyTorch, no `transformers`, no Python at
26
+ inference time — with rigorous equivalence and benchmark validation against
27
+ the original. See `research/experiments/exp11-real-pretrained-transformer`
28
+ through `exp14-wordpiece-native` in the Sajal Labs repo for full methodology.
29
+
30
+ ## Model details (unchanged from the original)
31
+
32
+ - **Architecture**: BERT encoder, 2 layers, hidden=128, heads=2, intermediate=512
33
+ - **Vocabulary**: 30522 WordPiece tokens (bert-base-uncased vocab)
34
+ - **Parameters**: 4,385,920
35
+ - **Precision**: fp32
36
+ - **Base model license**: MIT (prajjwal1/bert-tiny)
37
+ - **Port license**: MIT (Sajal Labs' C++ code)
38
+
39
+ ## What was verified (Sajal Labs' contribution)
40
+
41
+ **Tokenizer**: 17/17 real test sentences — including contractions ("don't"),
42
+ hyphenation ("COVID-19"), an out-of-vocabulary word forcing an 11-piece
43
+ subword split, and an email address — produced **byte-identical token IDs**
44
+ to the original `BertTokenizerFast`. This is an exact-match bar, not a
45
+ tolerance: tokenization is deterministic.
46
+
47
+ **Encoder**: max absolute error 9.54e-06 (hidden states),
48
+ 2.19e-06 (pooled `[CLS]` output), cosine similarity
49
+ ~1.0, across 10 real sentences of varying length (4-25 tokens) — fp32-scale
50
+ agreement, consistent with floating-point non-associativity between two
51
+ independent implementations (not a bug; see the repo's `research/papers.md`).
52
+
53
+ ## Benchmark (single request, Apple M4 CPU — full data in `benchmark_results.json`)
54
+
55
+ True end-to-end cold invocation (process spawn → raw text in → prediction out
56
+ → process exit, external wall-clock):
57
+
58
+ | Implementation | Cold invocation p50 |
59
+ |---|---:|
60
+ | **Native C++ (Sajal runtime)** | 11.37ms |
61
+ | ONNX Runtime + lean tokenizer | 95.72ms (8.4x slower) |
62
+ | ONNX Runtime + 🤗 transformers tokenizer | 2495.61ms (219.6x slower) |
63
+ | PyTorch + 🤗 transformers | 4973.91ms (437.6x slower) |
64
+
65
+ Worth knowing before you read too much into "ONNX Runtime": its own cold-start
66
+ number depends heavily on which tokenizer library it's paired with — using
67
+ `transformers` for convenience costs ~25x more than using the lean, standalone
68
+ `tokenizers` library for the exact same token IDs. Native sidesteps that
69
+ whole dependency-choice question by construction. Full discussion in exp14.
70
+
71
+ **Honest scope note on the warm-loop numbers**: native's advantage is
72
+ *not* unconditional the way cold-invocation is — Sajal Labs found it
73
+ depends on model width (`hidden_size`), with a measured crossover around
74
+ `hidden≈250` on this hardware (exp12/exp13). This model's `hidden=128`
75
+ sits comfortably below that, so native keeps a real warm-loop edge too —
76
+ but that's a property of this model's size, not a general claim.
77
+
78
+ ## How to use
79
+
80
+ ### Native (Sajal runtime, zero Python)
81
+
82
+ ```bash
83
+ sajal run <this_directory> "The quick brown fox jumps over the lazy dog."
84
+ ```
85
+
86
+ ### PyTorch / transformers (the original)
87
+
88
+ ```python
89
+ from transformers import BertModel, BertTokenizerFast
90
+ model = BertModel.from_pretrained("prajjwal1/bert-tiny")
91
+ tokenizer = BertTokenizerFast.from_pretrained("prajjwal1/bert-tiny")
92
+ ```
93
+
94
+ ### ONNX Runtime
95
+
96
+ ```python
97
+ import onnxruntime as ort
98
+ session = ort.InferenceSession("model.onnx", providers=["CPUExecutionProvider"])
99
+ # feed input_ids/attention_mask/token_type_ids from any WordPiece tokenizer
100
+ ```
101
+
102
+ ## Citations (required if you use this model)
103
+
104
+ ```bibtex
105
+ @article{turc2019distillation,
106
+ title={Well-Read Students Learn Better: On the Importance of Pre-training Compact Models},
107
+ author={Turc, Iulia and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
108
+ journal={arXiv preprint arXiv:1908.08962v2},
109
+ year={2019}
110
+ }
111
+
112
+ @misc{bhargava2021generalization,
113
+ title={Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics},
114
+ author={Bhargava, Prajjwal and Drozd, Aleksandr and Rogers, Anna},
115
+ year={2021},
116
+ eprint={2110.01518},
117
+ archivePrefix={arXiv},
118
+ primaryClass={cs.CL}
119
+ }
120
+ ```
121
+
122
+ ## Files
123
+
124
+ - `model.safetensors` — the original weights, HF/PyTorch-ecosystem format
125
+ - `model.onnx` + `model.onnx.data` — ONNX Runtime-compatible export (weights externalized to the `.data` file; both are required together)
126
+ - `word_embeddings.bin`, `position_embeddings.bin`, `token_type_embeddings.bin`, `emb_ln_*.bin`, `layer{i}_*.bin`, `pooler_*.bin` — raw native Sajal runtime format
127
+ - `vocab.txt`, `tokenizer_config.txt` — WordPiece vocabulary + config for the native tokenizer port
128
+ - `config.json` — architecture metadata
129
+ - `benchmark_results.json` — full machine-readable benchmark/equivalence data behind the numbers above
benchmark_results.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "hardware": "Apple M4, 10 cores, 16GB RAM, macOS (see research/environment.md)",
3
+ "methodology": "research/experiments/exp11-real-pretrained-transformer/results.md, exp12-warm-latency-width-depth/results.md, exp13-width-threshold/results.md, exp14-wordpiece-native/results.md",
4
+ "equivalence_native_vs_pytorch": {
5
+ "note": "native encoder + tokenizer vs. HF BertModel + BertTokenizerFast, 10 real sentences (lengths 4-25 tokens)",
6
+ "hidden_state_max_abs_error": 9.5367431640625e-06,
7
+ "pooled_output_max_abs_error": 2.1904706954956055e-06,
8
+ "mean_cosine_similarity_pooled": 1.000000035762787
9
+ },
10
+ "tokenizer_equivalence": "17/17 test sentences (incl. contractions, hyphens, OOV words, emails) produced byte-identical token IDs to HF's real tokenizer \u2014 see exp14",
11
+ "cold_invocation_ms_p50_precomputed_tokens": {
12
+ "native": 8.747625000069092,
13
+ "onnx_runtime_cpu": 87.30802099944412,
14
+ "pytorch_plus_transformers": 5095.878874999471
15
+ },
16
+ "cold_invocation_ms_p50_raw_text_full_pipeline": {
17
+ "native": 11.36558350026462,
18
+ "onnx_runtime_cpu_plus_lean_tokenizer": 95.72114549973776,
19
+ "onnx_runtime_cpu_plus_transformers_tokenizer": 2495.6148125002073,
20
+ "pytorch_plus_transformers": 4973.907166500794
21
+ },
22
+ "warm_loop_ms_p50": {
23
+ "native": 0.11565,
24
+ "onnx_runtime_cpu": 0.1342290006505209,
25
+ "pytorch_eager": 0.27608300115389284
26
+ },
27
+ "warm_loop_note": "Unlike cold invocation, native's warm-loop edge is NOT unconditional \u2014 it depends on model width; see exp12/exp13. At this model's hidden_size=128 it's below the measured ~250 crossover, so native keeps a real edge."
28
+ }
config.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_model": "prajjwal1/bert-tiny",
3
+ "base_model_license": "MIT",
4
+ "architecture": "bert",
5
+ "num_hidden_layers": 2,
6
+ "hidden_size": 128,
7
+ "num_attention_heads": 2,
8
+ "intermediate_size": 512,
9
+ "vocab_size": 30522,
10
+ "max_position_embeddings": 512,
11
+ "layer_norm_eps": 1e-12,
12
+ "num_parameters": 4385920,
13
+ "precision": "fp32",
14
+ "original_framework": "pytorch",
15
+ "sajal_runtime_compatible": true,
16
+ "sajal_native_port": {
17
+ "encoder": "native/bert_model.hpp",
18
+ "tokenizer": "native/wordpiece.hpp (real WordPiece, ASCII-scoped)"
19
+ }
20
+ }
config.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ 2 128 2 512 30522 512 1e-12
emb_ln_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ad4ffafdc100192191cca090b76519c689068fb8b8ee8a5d770bf75fd4f556ce
3
+ size 512
emb_ln_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9ebbf221e9922fa7e96a65d83860e14fde877c37901794188dbc0119fb04d78d
3
+ size 512
layer0_attn_ln_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6c21382c9a7f96e18dd8c1740f50030e73a52bfedc6185dfb36c4ea309cfc65d
3
+ size 512
layer0_attn_ln_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:326a90ea729d7f1c8dc0b599f49926980d59d73cdcff7c2c80dea2dbf6b094ef
3
+ size 512
layer0_attn_out_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2885546bee40403065c3d4858a17edd09ce84d4df58e06855e5f290b01bae1c4
3
+ size 512
layer0_attn_out_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f22f2bec4a303e4d0c3d326ec101b7042d194c2e1522b5ebd3377dfde2e169b3
3
+ size 65536
layer0_ff1_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0652ab6a4eba5562dcce4be341f4e8bdca1509394b6035c02d27a2697a07eb85
3
+ size 2048
layer0_ff1_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1837abddcee6e6d755e3c5276021a60d8d7be3c568b9d16ce8c3c2fc517794b3
3
+ size 262144
layer0_ff2_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:7bdc9b2cdec6b03f998f8e7c170415598f70c18b989f6586f567eb33d99e3f06
3
+ size 512
layer0_ff2_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3d2afe109364e857354a64679de404820cf3cadf7c5241046c97288db5317a25
3
+ size 262144
layer0_ff_ln_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:78d681625f5b35c3eb164d46f4033817992bf1262ab4ff063b03519464862e65
3
+ size 512
layer0_ff_ln_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2d9108431697bd6faea8055f0a61ea5ef07dc367c013d1436627d02d86d6e66a
3
+ size 512
layer0_k_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:55adb28810a6020320d74a41284b73bb51bf4bf5226198e53cab720f53531f0c
3
+ size 512
layer0_k_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:6b950f921ee111d9d70682f89ed855e7a2e9d6a6af50555027ce7cbed3757cc7
3
+ size 65536
layer0_q_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1c0ae37a6f1e963fedea5b81c732f3230e7f0cddc592938a80c461800f32fe3a
3
+ size 512
layer0_q_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0acebbdf18ed97979c2a5d205d5785fb4bae30c9df08064823ca53150ebfd635
3
+ size 65536
layer0_v_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:18c4ec14ca13f7fe4bab91c55eb9b02c012ed073d3c56beaed45ad5f067fcc22
3
+ size 512
layer0_v_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2ea5868004c66fc912fe74e5c9238ac6d91cc3ecb83729dfeb1e0f5ab1c187df
3
+ size 65536
layer1_attn_ln_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:106941eea5efc2b3ca11072f55a543bdae2ee6dd592b4b2d4bf6230b70cf28df
3
+ size 512
layer1_attn_ln_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8cce4d623046843995bbbc4c582abb76c615ae0da97cc4489ec24effcdb36421
3
+ size 512
layer1_attn_out_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d841b4359077aba819ba8753cf639774d159c18b207e72b1d31997161d748e8b
3
+ size 512
layer1_attn_out_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:bb091a113ba8dc74d5005b20dc45edebf8a82ff12df7a18532eeb88eb92eaf96
3
+ size 65536
layer1_ff1_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ad2916a72a70f197e6252d5df3e1dea61e4c67bd1ec672944360b4ef7c3ab2b1
3
+ size 2048
layer1_ff1_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a08978aa217acb8aa2d956d728a9b5d39d1924af217d3bcce91322a088046606
3
+ size 262144
layer1_ff2_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:aa08e02a0a4bd1b01b60dcf1194e6552733dc8cbe8a45543a62532f9e96ef268
3
+ size 512
layer1_ff2_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fb9898b9df5f9f55e0f40b10bb25c0a14d6f97412c8a87e6b553ea2237dfae50
3
+ size 262144
layer1_ff_ln_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c5a0eddf8c9d2d8d2baaf109424fe8233224330347f65bffed0dff0c186fe6d3
3
+ size 512
layer1_ff_ln_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c0fcf5d3d8a4dbb1a9ed8b8c3fbf10c568d6d28262d0d69e5bafd1adb482b518
3
+ size 512
layer1_k_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4b0eb16ecfae12acae7c2add258a2734db4d46e4ae9a3374a3065e9206a915c6
3
+ size 512
layer1_k_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ff567742407f7f8c032754b56f31b718455bdcae21c0b287027f44f07f481769
3
+ size 65536
layer1_q_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:ebcb022f66703ac7f8187146e9f2051dc01f5ad81629552887784138a8557861
3
+ size 512
layer1_q_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f591ff3f97c84d002964ae49a5c3672ed00320db444e49fa843419d91e69db5b
3
+ size 65536
layer1_v_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e8332addbc11ac4437819dbea9b4dd584611017613bf12b8f40c5d5b24a4e8c3
3
+ size 512
layer1_v_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9de7af9f8abd2d730f88d426ed40e1d63a7643dd85fdfac5467137526db10f29
3
+ size 65536
model.onnx ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b501c2f522a0ab5d6d2297fdd7d3a24bbd5a1b1e84f563dde70fe749df186ef8
3
+ size 23092
model.onnx.data ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9148713852692f5ce42c2ed1c1b62bc7239027f0636056a7fcc5e1e0cacc843b
3
+ size 17593344
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b421978f82848f6538bb99b69036e08f28760dd87116e5121a5764e92f2e8b27
3
+ size 17547888
pooler_bias.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a7da0faaef5689a538d0117b0414260814fb91fc3d886db97045e0d5eccd0f60
3
+ size 512
pooler_weight.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d288a0b6aff69cba192b299e3d12f948cbc5b43393f0198673dc4ea5914e4bb3
3
+ size 65536
position_embeddings.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b17b9d2b78503b0057db042854f79ad83c8eb8a8f022c5e07f9b4a98283bc480
3
+ size 262144
token_type_embeddings.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5d0abc942e607e690c4e1229c28ced11c8191b1955fda71c81a196ca07869ecb
3
+ size 1024
tokenizer_config.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ 1 101 102 100 0
vocab.txt ADDED
The diff for this file is too large to render. See raw diff
 
word_embeddings.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f49cea2ec944d799022612633ec594eddfb82cd81c54f98be79ef0c58d5cb0ff
3
+ size 15627264