spec-b300 / source /docs /user_guide /features.md
khazic's picture
Archive three-epoch run: logs and provenance part 1
a181ec9 verified
|
Raw History Blame Contribute Delete
3.65 kB
# Features
Speculators is built to make training and deploying speculative decoding models fast, scalable, and easy to integrate into existing workflows. Here are some of the key features.
## Distributed Training with FSDP
Speculators supports multi-GPU training via PyTorch Fully Sharded Data Parallel (FSDP). Model parameters are sharded per-layer with a mixed precision policy β€” bfloat16 for parameters, float32 for gradient reductions β€” and distributed checkpointing handles save/restore across ranks automatically.
## Draft Vocabulary Support
Draft models use a reduced vocabulary for faster inference. Speculators automatically builds vocabulary mappings (`t2d` and `d2t` tensors) from token frequency statistics collected during data preparation, selecting the most frequent tokens for the draft vocab. Pre-built mappings can also be provided manually.
## Multi-Backend Metric Logging
Training metrics can be logged to TensorBoard, Weights & Biases, TrackIO, and MLflow β€” individually or simultaneously, so that you can use your preferred experiment tracking tool.
## Automatic Chat Template Detection
During data preparation, Speculators automatically detects assistant response boundaries to build loss masks. It first tries HuggingFace's native `assistant_tokens_mask` support, then falls back to regex-based pattern detection β€” including stripping `<think>` blocks from reasoning models. No manual template configuration is needed for most models.
## Performant Flex Attention
Both [Eagle-3](algorithms/eagle3.md) and [DFlash](algorithms/dflash.md) models use PyTorch's `flex_attention` with `BlockMask` for efficient, structured attention patterns (causal, document-aware, and anchor-based). The Eagle-3 forward pass is wrapped with `torch.compile` for additional runtime optimization.
## Efficient Sequence Packing
The multipack batch sampler uses an LPT (Longest Processing Time First) bin-packing algorithm to pack variable-length sequences into batches, maximizing GPU utilization while respecting per-device token limits. This avoids the wasted compute from naive padding.
## Checkpoint Resume
Training automatically resumes from the latest checkpoint, restoring model weights, optimizer state, and scheduler state. The checkpointer tracks the best validation loss and maintains a symlink to the best checkpoint for easy model selection.
## Multi-Node Hidden-States Transfer
The `hs_connectors` package provides a pluggable backend system for transferring hidden states between vLLM and the trainer. The default **filesystem** backend uses safetensors files on a shared filesystem for single-node setups. For multi-node training β€” where the target model spans multiple nodes or extraction and training run on separate machines β€” the **Mooncake** backend streams hidden states over TCP or RDMA through a distributed key-value store, with no shared filesystem required. See [Multi-Node Training](tutorials/multi_node_training.md) for setup details.
## Seamless Integration with vLLM
Trained models are saved in a special speculators format that vLLM can load directly. Point vLLM at a checkpoint or a HuggingFace model and it automatically detects the speculator, pairs it with the target model, and enables speculative decoding β€” no extra configuration needed.
## Model Conversion
Speculators can convert pre-trained models from third-party repositories (EAGLE v1/v2/v3, HASS) into Speculators format for direct deployment with vLLM.
______________________________________________________________________
For hands-on guides covering the full workflow, see the [Tutorials](tutorials/index.md).