khazic's picture
Archive three-epoch run: logs and provenance part 1
a181ec9 verified
|
Raw History Blame Contribute Delete
1.03 kB

Tutorials

Step-by-step tutorials to guide you through complete workflows, from data preparation to serving trained models in production.

Train a Speculator

The main end-to-end walkthrough: prepare data, generate hidden states, train, and serve. Covers Eagle-3, P-EAGLE, DFlash, DSpark, and MTP, in online, offline, or hybrid mode -- pick your algorithm and mode at the top of the page.

Multi-Node Training

Stream hidden states between separate extraction and training nodes with the Mooncake backend when the target model does not fit on one node or shared storage is unavailable.

Response Regeneration

Regenerate dataset responses using your target model for improved drafter alignment. Recommended before training.

Evaluating Model Performance

Benchmark and evaluate your trained speculator models.

Serve in vLLM

Deploy your trained speculator models in vLLM for production inference.