Instructions to use Qinsi1/FlexDLLM-LLaDA-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Qinsi1/FlexDLLM-LLaDA-8B with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
FlexDLLM-LLaDA-8B
An archived, matched pair of RG-Diff and SC-AR adapters for FlexDLLM: Diffuse
When Possible, Autoregress When Necessary. The two stages run on a shared
LLaDA-8B-Instruct backbone. RG-Diff emits <ASK_LLM> during diffusion;
SC-AR generates a short autoregressive repair, and diffusion resumes.
This is a research preview of the existing math-oriented checkpoint pair. The release has passed a real-model ASK/repair/resume smoke test. The paper's full benchmark scores have not been reproduced for this bundle.
What is included
| Component | Archived checkpoint |
|---|---|
| RG-Diff | llada_ask_lora_v2_noclean/checkpoint-final |
| SC-AR | llada_patch_lora_v4_online/checkpoint-final/patch_adapter |
| Backbone | GSAI-ML/LLaDA-8B-Instruct |
| Backbone revision | 08b83a6feb34df1a6011b80c3c00c7563e963b07 |
The download contains approximately 521.5 MiB of adapters, tokenizer and configuration. The 8B backbone is downloaded separately by the loader. Both adapters use LoRA rank 32, alpha 64. RG-Diff is merged into the backbone; SC-AR is then attached to that merged model and enabled only during refinement. The original Stage-II training script identifies this RG-Diff checkpoint as its inherited base; an original per-checkpoint pairing manifest is unavailable. The exported bundle records an RG-Diff fingerprint to prevent subsequent mixing.
Installation
Use Python 3.10 or 3.11 and FlexDLLM 0.1.1 or later. The source repository is currently private; cloning requires GitHub access.
git clone https://github.com/wangqinsi1/FlexDLLM.git
cd FlexDLLM
python -m venv .venv
source .venv/bin/activate
pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -e .
If this model repository is private, authenticate with an account that has access:
hf auth login
Give a question, receive a response and decoding trace
from flexdllm import FlexDLLM
model = FlexDLLM.from_pretrained("Qinsi1/FlexDLLM-LLaDA-8B", device="cuda")
result = model.generate(
"A robe takes 2 bolts of blue fiber and half that much white fiber. "
"How many bolts in total does it take?",
return_trace=True,
)
print(result.text)
print("Refinements:", result.refinements)
for step in result.trace:
if step["phase"] == "refinement":
print("Before:", step["before"])
print("SC-AR patch:", step["patch"])
Or use the command line:
flexdllm --model Qinsi1/FlexDLLM-LLaDA-8B \
--prompt "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?" \
--trace --output outputs/example.json
--trace shows intermediate diffusion states, ASK events and repairs. ▮
marks unresolved positions. The final response replaces ASK with generated
content; inspect result.refinements and result.trace to see interventions.
ASK is selected by the model and is not guaranteed for every question.
The bundle automatically applies the archived checkpoint's original GSM8K
question template and leading BOS token (prompt_format=legacy_gsm8k). Pass
only the question; do not add a chat template or the GSM8K wrapper yourself.
The supported input is a question string or one user message.
For a pinned download, pass revision=COMMIT_SHA to from_pretrained, or
--revision COMMIT_SHA to the CLI. A downloaded local bundle directory is also
accepted; --local-files-only uses cached files, and --base-model can select
a local backbone directory.
The public repository was also downloaded without authentication, with forced
re-download and SHA-256 verification of all 13 manifest-listed files. The
installed FlexDLLM 0.1.1 wheel loaded the downloaded bundle by Hub ID and pinned
revision 7744548ec28c585d21230f6dc59748f46fb0e084. The natural ASK/repair/resume
check passed again on an H200. The pinned backbone was reused from the existing
cache. The result and validation scope are recorded in validation.json.
Training and generation configuration
Stage I was trained on prepared normal-generation, failure-localization and post-repair continuation examples. Stage II used online teacher patch records and a single-token autoregressive training objective, with a 30-token workspace. The archived training states report 4 epochs for RG-Diff and 3 for SC-AR. The original datasets are not included in this model repository.
Default generation uses 256 canvas positions, four committed tokens per
step, at most 128 diffusion steps, three refinements, and a 30-token repair
workspace. Positions after an inserted repair are re-masked. These defaults
follow the archived checkpoint implementation; the paper body also describes
a 10-token workspace, selectable with --patch-max-tokens 10.
Validation and limitations
A local H200/bfloat16 check on the robe question naturally emitted ASK, generated
2+1=3 with SC-AR, resumed diffusion, and completed with EOS. It used one repair,
five AR tokens and 64 diffusion steps. There was no forced ASK and no external
Qwen model in this run. This validates the inference mechanism on one example;
it does not establish benchmark accuracy or reliable correction on every input.
The response formatting can vary, and the model can generate incorrect answers.
This archived pair is intended for English math reasoning experiments. The newer task-specific v5 RG-Diff checkpoints are different models and should not be substituted for the RG-Diff adapter in this bundle without retraining or validating a matched SC-AR adapter.
Files and provenance
flexdllm_config.json: backbone revision, adapter paths and generation settings.rg_diff/: Stage-I adapter and the trained ASK output-head row.sc_ar/: Stage-II adapter and exported pairing fingerprint.tokenizer/: tokenizer and special-token files.release_manifest.json: checkpoint provenance, code revision and file hashes.validation.json: compact record of the release smoke test.
See the source repository for training, data preparation, API documentation and tests.
Licenses
The base model declares the MIT license. The FlexDLLM source code is Apache-2.0. A separate license for these adapter weights has not yet been designated by the model owner.
Model tree for Qinsi1/FlexDLLM-LLaDA-8B
Base model
GSAI-ML/LLaDA-8B-Instruct