bounded-moe / README.md
THISITLLM's picture
Update README.md
c751c9e verified
|
Raw History Blame Contribute Delete
4.7 kB
---
tags:
- llama.cpp
- ggml
- moe
- mixture-of-experts
- inference
- local-llm
- ssd
- model-offloading
- memory-optimization
- cuda
---
# bounded-moe
**Bounded-memory inference experiments for running large Mixture-of-Experts models on consumer hardware.**
> Can we run a large MoE model without letting its expert weights consume most of system RAM?
`bounded-moe` is an experimental project exploring **SSD-backed expert storage and bounded-memory caching for llama.cpp / GGML**.
The goal is to make large MoE models more practical on memory-constrained consumer PCs by keeping only a controlled working set of experts resident in memory.
## Architecture
```text
MoE Router
│
▼
Requested Experts
│
▼
┌───────────────────┐
│ Bounded Expert │
│ Cache │
└─────────┬─────────┘
hit │ miss
│
▼
External Expert Storage
│
SSD
│
▼
RAM / VRAM
│
▼
Compute
```
In simplified form:
**SSD → Bounded Expert Cache → RAM/VRAM → Inference**
Instead of allowing all expert weights to remain resident in system memory, the runtime resolves and caches the experts required by MoE routing.
## Test Hardware
Current development and testing is performed on consumer hardware:
* **GPU:** NVIDIA GeForce RTX 4060 8 GB
* **System RAM:** 32 GB
* **OS:** Windows 11
* **Runtime:** llama.cpp / GGML
* **Model class:** ~35B Mixture-of-Experts, ~3B active parameters
The project specifically explores scenarios where model weights, applications, and the operating system compete for limited RAM.
## Implemented Research
The project currently includes experiments around:
* MoE routing tracing
* External expert storage
* Bounded expert caching
* Cache hit/miss accounting
* Expert pin/unpin lifecycle
* Safe cache eviction
* Direct expert reads
* Resolver indirection
* Cache-size experiments
* Working-set / RAM measurements
* SSD-backed inference
* Performance and correctness validation
## Current Status
The architecture is **experimental**.
The bounded-storage path works and demonstrates that expert memory can be managed independently from normal full-model residency.
The main research challenge is now **performance**.
Normal memory-mapped inference is significantly faster than the current experimental SSD-backed expert path. Current work therefore focuses on reducing expert-resolution and storage latency while preserving the bounded-memory property.
This repository should currently be considered an **engineering/research prototype**, not a production inference runtime.
## Why?
Large MoE models are interesting for consumer hardware because only a subset of their parameters is active for each token.
However, inactive experts can still consume significant system memory.
This project asks a slightly different question:
> **What if model capacity could be much larger than the amount of RAM we're willing to dedicate to inference?**
Rather than treating available RAM as the hard limit, `bounded-moe` explores using a hierarchy of:
**SSD → bounded RAM cache → GPU → compute**
while exploiting MoE routing locality to keep the frequently requested experts close to compute.
## Future Research
Planned experiments include:
* asynchronous expert prefetching
* predictive prefetch based on routing behavior
* overlapping SSD I/O with GPU computation
* smarter eviction policies
* cache locality analysis
* larger-context testing
* Windows I/O optimization
* reducing resolver overhead
* dense-model layer/block streaming experiments
The last item is particularly interesting: some of the infrastructure may eventually be generalized beyond MoE models into a **bounded-memory inference runtime for dense models**.
## Source Code
The implementation, experiments, benchmark notes, and development history are available on GitHub:
https://github.com/kornpaksittikool-beep/bounded-moe
## Feedback
Feedback is very welcome, particularly from people working with:
**llama.cpp · GGML · MoE routing · model offloading · caching · mmap · SSD I/O · CUDA · inference optimization**
If you've experimented with similar SSD-backed or bounded-memory inference architectures, I'd especially like to hear about approaches for hiding cache-miss and storage latency.
---
**Status:** Experimental / Research Prototype