Text Generation
Transformers
Safetensors
spark2_5
qfs
fixture
random-init
cpu
reproducibility
random-architecture-fixture
custom_code
Instructions to use malaiwah/spark2-5-tiny-random-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use malaiwah/spark2-5-tiny-random-bf16 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="malaiwah/spark2-5-tiny-random-bf16", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("malaiwah/spark2-5-tiny-random-bf16", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use malaiwah/spark2-5-tiny-random-bf16 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "malaiwah/spark2-5-tiny-random-bf16" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "malaiwah/spark2-5-tiny-random-bf16", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/malaiwah/spark2-5-tiny-random-bf16
- SGLang
How to use malaiwah/spark2-5-tiny-random-bf16 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "malaiwah/spark2-5-tiny-random-bf16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "malaiwah/spark2-5-tiny-random-bf16", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "malaiwah/spark2-5-tiny-random-bf16" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "malaiwah/spark2-5-tiny-random-bf16", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use malaiwah/spark2-5-tiny-random-bf16 with Docker Model Runner:
docker model run hf.co/malaiwah/spark2-5-tiny-random-bf16
Download upstream-code-audit.json from malaiwah/spark2-5-tiny-random-bf16: direct link, hf CLI and curl.
- Browser
- Download file 5.96 kB
-
https://huggingface.co/malaiwah/spark2-5-tiny-random-bf16/resolve/main/upstream-code-audit.json
- Command line
-
hf download hf://malaiwah/spark2-5-tiny-random-bf16/upstream-code-audit.json
-
curl -L -o upstream-code-audit.json https://huggingface.co/malaiwah/spark2-5-tiny-random-bf16/resolve/main/upstream-code-audit.json
5.96 kB
| { | |
| "audit_kind": "static source review; no model or workflow executed by fixture authoring agent", | |
| "effects": { | |
| "dynamic_eval_exec_in_runtime_source": [], | |
| "execution_time": [ | |
| "Initializers use torch random normal and zero operations.", | |
| "Forward uses ordinary CPU-capable PyTorch operations; CPU-only proof is produced only by release execution.", | |
| "HF/QFS dynamic-code loader may fetch and verify explicitly selected immutable Hub source; this is loader activity, not runtime-source network behavior." | |
| ], | |
| "filesystem_calls_in_runtime_source": [], | |
| "import_time": [ | |
| "Imports installed torch/transformers modules; transitive dependency import effects are not independently sandboxed.", | |
| "Obtains module logger through transformers.utils.logging.get_logger.", | |
| "Defines classes/functions and evaluates type annotations and can_return_tuple decorator.", | |
| "Appends Spark2_5RMSNorm to transformers.pytorch_utils.ALL_LAYERNORM_LAYERS; global normalization registry mutation retained from upstream." | |
| ], | |
| "network_calls_in_runtime_source": [], | |
| "scope_exclusion": "Build/verification/publication scripts live in the evidence repository and are separately invoked tools, not auto_map imports. The model repo has only the two runtime .py files; original upstream .py bytes are preserved with an inert .py.txt suffix. release.py performs intended subprocesses and Hub mutations only when Main/user executes it.", | |
| "subprocess_calls_in_runtime_source": [] | |
| }, | |
| "imports": { | |
| "dependencies": [ | |
| "torch", | |
| "transformers" | |
| ], | |
| "optional_accelerator_imports": [], | |
| "relative": [ | |
| "configuration_spark" | |
| ], | |
| "standard_library": [ | |
| "math", | |
| "typing (reviewed compatibility import)" | |
| ] | |
| }, | |
| "known_upstream_semantics": [ | |
| "Rotary positions are taken from cache_position; arbitrary per-example position_ids are not newly supported by this fork. Fixture uses contiguous unpadded positions.", | |
| "Eager attention only, as upstream; no FlashAttention/SDPA claim." | |
| ], | |
| "license": { | |
| "file_evidence": "LICENSE contains full Apache 2.0 and Copyright 2026 XHToken; configuration_spark.py carries Apache header; modeling_spark.py has no original per-file header but is covered by repository LICENSE", | |
| "metadata": "README.md frontmatter license apache-2.0", | |
| "retention": [ | |
| "LICENSE", | |
| "NOTICE", | |
| "upstream/LICENSE", | |
| "upstream/configuration_spark.py.txt in model; upstream/configuration_spark.py in evidence", | |
| "upstream/modeling_spark.py.txt in model; upstream/modeling_spark.py in evidence" | |
| ], | |
| "spdx": "Apache-2.0", | |
| "upstream_notice": "No NOTICE file in pinned repository sibling inventory." | |
| }, | |
| "patches": [ | |
| { | |
| "change": "Import Unpack from Python3.12 typing instead of removed transformers.processing_utils re-export; annotations only, no model math changes.", | |
| "file": "modeling_spark.py", | |
| "installed_api": "transformers/processing_utils.py imports typing_extensions but no longer exports Unpack", | |
| "mathematical_change": false | |
| }, | |
| { | |
| "change": "Add pinned source attribution and Apache notice.", | |
| "file": "modeling_spark.py", | |
| "mathematical_change": false | |
| }, | |
| { | |
| "change": "Mask helper input_embeds renamed inputs_embeds; obsolete cache_position keyword removed. Transformers5.16 infers mask offsets from Cache and position_ids.", | |
| "file": "modeling_spark.py", | |
| "installed_api": "transformers/masking_utils.py create_causal_mask and create_sliding_window_causal_mask signatures", | |
| "mathematical_change": false | |
| }, | |
| { | |
| "change": "Replace Transformers4 tied-key list with Transformers5 target/source mapping lm_head.weight -> model.embedding.weight.", | |
| "file": "modeling_spark.py", | |
| "installed_api": "transformers/modeling_utils.py get_expanded_tied_weights_keys", | |
| "mathematical_change": false | |
| }, | |
| { | |
| "change": "Use self.lm_head(hidden_states) for both tied and untied configurations; declared tying aliases the original embedding parameter, retaining F.linear operands without duplicate invented head weights. Enables actual output-module pre-hook.", | |
| "file": "modeling_spark.py", | |
| "mathematical_change": false, | |
| "proof": "verify_native.py checks pointer alias and actual hooked-head logits bit-exact equality to upstream F.linear." | |
| } | |
| ], | |
| "preserved_equations": [ | |
| "GQA eager attention with FP32 softmax", | |
| "three sliding layers then one full layer", | |
| "per-attention-type partial RoPE and theta", | |
| "headwise sigmoid output gating", | |
| "GELU(gate_proj(x))*up_proj(x) MLP", | |
| "RMSNorm and FP32 residual stream", | |
| "legitimately tied output vocabulary head" | |
| ], | |
| "runtime_files": { | |
| "configuration_spark.py": "02218597240490f490659b184052db940b08163ff9c3af7b7e59323f6f963722", | |
| "modeling_spark.py": "ba15bd14cca26322c9461ee2cf129423b59a5122826449947f7bb1acf5d9f2a6" | |
| }, | |
| "schema": "malaiwah.spark2-5-upstream-code-audit.v1", | |
| "upstream": { | |
| "repository": "XHToken/Spark-X2.5-4B", | |
| "revision": "5e10fcc0286756aebf7c41dc52c1e42d95c70281" | |
| }, | |
| "upstream_files": { | |
| "LICENSE": { | |
| "bytes": 11338, | |
| "sha256": "525c465ab952779d2d360a9ff36f7f49ae018a23d80302ff839a850a77784c36" | |
| }, | |
| "README.md": { | |
| "bytes": 24346, | |
| "sha256": "7aac2c2e00d2c8ea608254c9085152c853282d93dad0312c62d7ba328bc6a5be" | |
| }, | |
| "config.json": { | |
| "bytes": 2035, | |
| "sha256": "767161ece8ed2891e44345afdf98d29c1cacb964116cb6db7d20dc1c757490ad" | |
| }, | |
| "configuration_spark.py": { | |
| "bytes": 4329, | |
| "sha256": "02218597240490f490659b184052db940b08163ff9c3af7b7e59323f6f963722" | |
| }, | |
| "modeling_spark.py": { | |
| "bytes": 19418, | |
| "sha256": "9cf0d1ad2b54b9f7088792779ddd8b5cb4d4fe63bfea054da7f2dd16362bcaf4" | |
| } | |
| }, | |
| "verification_entrypoint": "verify_native.py; results only exist after stage execution" | |
| } | |