|
Download README.md from THISITLLM/bounded-moe: direct link, hf CLI and curl.
- Browser
- Download file 4.7 kB
-
https://huggingface.co/THISITLLM/bounded-moe/resolve/main/README.md
- Command line
-
hf download hf://THISITLLM/bounded-moe/README.md
-
curl -L -o README.md https://huggingface.co/THISITLLM/bounded-moe/resolve/main/README.md
4.7 kB
| tags: | |
| - llama.cpp | |
| - ggml | |
| - moe | |
| - mixture-of-experts | |
| - inference | |
| - local-llm | |
| - ssd | |
| - model-offloading | |
| - memory-optimization | |
| - cuda | |
| # bounded-moe | |
| **Bounded-memory inference experiments for running large Mixture-of-Experts models on consumer hardware.** | |
| > Can we run a large MoE model without letting its expert weights consume most of system RAM? | |
| `bounded-moe` is an experimental project exploring **SSD-backed expert storage and bounded-memory caching for llama.cpp / GGML**. | |
| The goal is to make large MoE models more practical on memory-constrained consumer PCs by keeping only a controlled working set of experts resident in memory. | |
| ## Architecture | |
| ```text | |
| MoE Router | |
| │ | |
| ▼ | |
| Requested Experts | |
| │ | |
| ▼ | |
| ┌───────────────────┐ | |
| │ Bounded Expert │ | |
| │ Cache │ | |
| └─────────┬─────────┘ | |
| hit │ miss | |
| │ | |
| ▼ | |
| External Expert Storage | |
| │ | |
| SSD | |
| │ | |
| ▼ | |
| RAM / VRAM | |
| │ | |
| ▼ | |
| Compute | |
| ``` | |
| In simplified form: | |
| **SSD → Bounded Expert Cache → RAM/VRAM → Inference** | |
| Instead of allowing all expert weights to remain resident in system memory, the runtime resolves and caches the experts required by MoE routing. | |
| ## Test Hardware | |
| Current development and testing is performed on consumer hardware: | |
| * **GPU:** NVIDIA GeForce RTX 4060 8 GB | |
| * **System RAM:** 32 GB | |
| * **OS:** Windows 11 | |
| * **Runtime:** llama.cpp / GGML | |
| * **Model class:** ~35B Mixture-of-Experts, ~3B active parameters | |
| The project specifically explores scenarios where model weights, applications, and the operating system compete for limited RAM. | |
| ## Implemented Research | |
| The project currently includes experiments around: | |
| * MoE routing tracing | |
| * External expert storage | |
| * Bounded expert caching | |
| * Cache hit/miss accounting | |
| * Expert pin/unpin lifecycle | |
| * Safe cache eviction | |
| * Direct expert reads | |
| * Resolver indirection | |
| * Cache-size experiments | |
| * Working-set / RAM measurements | |
| * SSD-backed inference | |
| * Performance and correctness validation | |
| ## Current Status | |
| The architecture is **experimental**. | |
| The bounded-storage path works and demonstrates that expert memory can be managed independently from normal full-model residency. | |
| The main research challenge is now **performance**. | |
| Normal memory-mapped inference is significantly faster than the current experimental SSD-backed expert path. Current work therefore focuses on reducing expert-resolution and storage latency while preserving the bounded-memory property. | |
| This repository should currently be considered an **engineering/research prototype**, not a production inference runtime. | |
| ## Why? | |
| Large MoE models are interesting for consumer hardware because only a subset of their parameters is active for each token. | |
| However, inactive experts can still consume significant system memory. | |
| This project asks a slightly different question: | |
| > **What if model capacity could be much larger than the amount of RAM we're willing to dedicate to inference?** | |
| Rather than treating available RAM as the hard limit, `bounded-moe` explores using a hierarchy of: | |
| **SSD → bounded RAM cache → GPU → compute** | |
| while exploiting MoE routing locality to keep the frequently requested experts close to compute. | |
| ## Future Research | |
| Planned experiments include: | |
| * asynchronous expert prefetching | |
| * predictive prefetch based on routing behavior | |
| * overlapping SSD I/O with GPU computation | |
| * smarter eviction policies | |
| * cache locality analysis | |
| * larger-context testing | |
| * Windows I/O optimization | |
| * reducing resolver overhead | |
| * dense-model layer/block streaming experiments | |
| The last item is particularly interesting: some of the infrastructure may eventually be generalized beyond MoE models into a **bounded-memory inference runtime for dense models**. | |
| ## Source Code | |
| The implementation, experiments, benchmark notes, and development history are available on GitHub: | |
| https://github.com/kornpaksittikool-beep/bounded-moe | |
| ## Feedback | |
| Feedback is very welcome, particularly from people working with: | |
| **llama.cpp · GGML · MoE routing · model offloading · caching · mmap · SSD I/O · CUDA · inference optimization** | |
| If you've experimented with similar SSD-backed or bounded-memory inference architectures, I'd especially like to hear about approaches for hiding cache-miss and storage latency. | |
| --- | |
| **Status:** Experimental / Research Prototype |