ShuhongWu commited on
Commit
cd11d03
·
verified ·
1 Parent(s): 211b47a

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +0 -74
README.md CHANGED
@@ -16,78 +16,4 @@ Neural Imprint lets a model learn from local experience while base weights stay
16
 
17
  🌐 [atomgradient.com](https://atomgradient.com)
18
 
19
- ---
20
-
21
- ## Research
22
-
23
- ### [Prism — Cross-Domain Personal Data Integration on Consumer Hardware](https://atomgradient.github.io/Prism/)
24
-
25
- Integrating finance, diet, mood, and reading data entirely on consumer Apple Silicon, producing emergent cross-domain insights with zero data leakage.
26
-
27
- - 📈 **1.48x** cross-domain insight emergence (IIR)
28
- - 🔒 **125.5x** federation compression, zero data leakage
29
- - ⚡ **49.9 TPS** real-time inference (35B on M2 Ultra)
30
-
31
- [[GitHub]](https://github.com/AtomGradient/Prism) · [[Paper]](https://atomgradient.github.io/Prism/)
32
-
33
- ---
34
-
35
- ### [ANE Batch Prefill — On-Device Parallel LLM Inference](https://atomgradient.github.io/hybird-batch-prefill-on-ane/)
36
-
37
- Fused matrix-vector kernels enabling concurrent ANE batch prefill + GPU decode on Apple Silicon for Qwen3.5 models.
38
-
39
- - 🚀 **11.3x** ANE batch prefill speedup (268 tok/s)
40
- - 🔋 **79%** power reduction for prefill component
41
- - ⏱️ **<30 ms** state transfer overhead
42
-
43
- [[GitHub]](https://github.com/AtomGradient/hybird-batch-prefill-on-ane) · [[Paper]](https://atomgradient.github.io/hybird-batch-prefill-on-ane/)
44
-
45
- ---
46
-
47
- ### [hybrid-ane-mlx-bench — Disaggregated LLM Inference on Apple Silicon](https://atomgradient.github.io/hybrid-ane-mlx-bench/)
48
-
49
- Benchmarking CoreML ANE prefill + MLX GPU decode for Qwen3.5 on Apple Silicon, with four inference strategies compared.
50
-
51
- - 🔄 ANE prefill matches GPU at **~410 tokens**
52
- - 🔋 **282x** GPU power reduction during prefill
53
- - 📊 4 inference pipelines benchmarked
54
-
55
- [[GitHub]](https://github.com/AtomGradient/hybrid-ane-mlx-bench) · [[Paper]](https://atomgradient.github.io/hybrid-ane-mlx-bench/)
56
-
57
- ---
58
-
59
- ### [swift-qwen3-tts — On-Device Text-to-Speech](https://atomgradient.github.io/swift-qwen3-tts/)
60
-
61
- Native Swift implementation of Qwen3 TTS 0.6B for real-time, on-device speech synthesis.
62
-
63
- - 📦 **67%** model compression (2.35 GB → 808 MB)
64
- - 🎙️ Real-time synthesis (**RTF 0.68x**)
65
- - 🌍 12 languages supported
66
-
67
- [[GitHub]](https://github.com/AtomGradient/swift-qwen3-tts) · [[Paper]](https://atomgradient.github.io/swift-qwen3-tts/)
68
-
69
- ---
70
-
71
- ### [Gemma-Prune — On-Device Vision Language Model](https://atomgradient.github.io/swift-gemma-cli/)
72
-
73
- Multi-stage compression pipeline for deploying Gemma 3 4B VLM on consumer hardware.
74
-
75
- - 📦 **25%** model compression (2.8 GB → 2.1 GB)
76
- - 📝 **110 tok/s** text generation
77
- - 🖼️ **3.4x** image processing speedup
78
-
79
- [[GitHub]](https://github.com/AtomGradient/swift-gemma-cli) · [[Paper]](https://atomgradient.github.io/swift-gemma-cli/)
80
-
81
- ---
82
-
83
- ### [OptMLX — MLX Memory Optimization Research](https://atomgradient.github.io/OptMLX/)
84
-
85
- Exploring memory optimization techniques for the MLX framework on Apple Silicon.
86
-
87
- - ⚡ Up to **20x** faster mmap loading
88
- - 🔄 Zero-copy model loading
89
- - 📊 Comprehensive benchmarks
90
-
91
- [[GitHub]](https://github.com/AtomGradient/OptMLX) · [[Paper]](https://atomgradient.github.io/OptMLX/)
92
-
93
  ---
 
16
 
17
  🌐 [atomgradient.com](https://atomgradient.com)
18
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19
  ---