|
Download README.md from mikelou/mod-arithmetic: direct link, hf CLI and curl.
- Browser
- Download file 2.24 kB
-
https://huggingface.co/mikelou/mod-arithmetic/resolve/main/README.md
- Command line
-
hf download hf://mikelou/mod-arithmetic/README.md
-
curl -L -o README.md https://huggingface.co/mikelou/mod-arithmetic/resolve/main/README.md
2.24 kB
| license: mit | |
| library_name: pytorch | |
| tags: | |
| - modular-arithmetic | |
| - neural-arithmetic | |
| - bit-serial | |
| - gru | |
| # NeuralHorner v8 compact inference experiment | |
| This local derivative tests storage and inference optimizations against the | |
| published MIT-licensed `TrickyRex/bitserial-modmul-v8` checkpoint. It is not a | |
| newly trained model. Original model and weights by Robert Sneiderman: | |
| <https://huggingface.co/TrickyRex/bitserial-modmul-v8>. | |
| Changes: | |
| - checkpoint tensors stored as bfloat16; | |
| - direct `logit > 0` decisions instead of `sigmoid(logit) > 0.5`; | |
| - one operand is reduced and the other is streamed directly through the | |
| learned multiplication transition, removing one modulus-width pass; | |
| - static multiplicand/modulus feature channels are allocated once per scan | |
| instead of rebuilt at every recurrent step. | |
| The last change is mathematically equivalent for an exact transition, but the | |
| learned cell is not proven exact. Benchmark equivalence therefore must be | |
| measured before this is treated as an accuracy-preserving release. | |
| ## Local validation | |
| Hardware: Apple GPU through PyTorch MPS. Dataset: the official 100-case public | |
| benchmark for each tier. | |
| - Tiers 1-7: 700/700 exact. | |
| - Tier 6 runtime: 19.07 seconds versus 27.02 seconds for the published wrapper | |
| (29.4% lower wall-clock in this comparison). The earlier short-schedule | |
| version took 20.45-21.31 seconds before static-channel reuse. | |
| - Tier 6 outputs: 100/100 byte-for-byte identical to the published wrapper. | |
| - The bfloat16 checkpoint with the original three-pass schedule also produced | |
| 100/100 outputs identical to the fp32 checkpoint on Tier 6. | |
| - Official static analysis: clean. | |
| The checkpoint is 948,196 bytes versus 1,887,610 bytes (49.8% smaller). The | |
| shorter schedule removes exactly one `Leff` recurrent pass: 20% of recurrent | |
| step calls on public Tiers 3 and 5-10, 14.3% on Tier 4, 25% on Tier 2, and | |
| 33.3% on Tier 1. | |
| ## Release gate | |
| Do not claim unchanged Tier 1-10 accuracy yet. The current machine is too slow | |
| for a practical full-width Tier 8-10 MPS run. Before release, run the official | |
| 1,100-case scorer on CUDA for the three published seeds and require 100% on | |
| every scored tier, then repeat the held-out 768-case adversarial battery. | |