Swift-Qwen3.8-Flash-Next β€” Splash Package (Sept-24 build)

Splash-native serving package of the Swift (KV-sparse) Qwen3.8-Flash-Next hybrid, Q4_0 with Q8 output, built for the Splash Metal engine on Apple Silicon.

Contents

  • target/ β€” 30 target-layer MDFN0031 bins + embedding.bin + head.bin
  • draft/ β€” MTP draft-layer bins + model.bin (MDFD0004)
  • tokenizer/, vision/ β€” tokenizer + vision tower
  • manifest.json β€” schema v5, splash-packed-q4-qwen4exp

Usage

splash serve-native <dir>/target <dir>/draft --tokenizer <dir>/tokenizer

Notes

  • Sept-24 build; superseded locally by the v3 package. This repo is the archival backup.
  • Swift hybrid: Qwen3.8-Flash-Next base with KV-sparse (Swift) layers.
  • License: Qwen Community License 1.0.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for nitinpanj/Swift-Qwen3.8-Flash-Next-Splash