FlexDLLM-LLaDA-8B

An archived, matched pair of RG-Diff and SC-AR adapters for FlexDLLM: Diffuse When Possible, Autoregress When Necessary. The two stages run on a shared LLaDA-8B-Instruct backbone. RG-Diff emits <ASK_LLM> during diffusion; SC-AR generates a short autoregressive repair, and diffusion resumes.

This is a research preview of the existing math-oriented checkpoint pair. The release has passed a real-model ASK/repair/resume smoke test. The paper's full benchmark scores have not been reproduced for this bundle.

What is included

Component Archived checkpoint
RG-Diff llada_ask_lora_v2_noclean/checkpoint-final
SC-AR llada_patch_lora_v4_online/checkpoint-final/patch_adapter
Backbone GSAI-ML/LLaDA-8B-Instruct
Backbone revision 08b83a6feb34df1a6011b80c3c00c7563e963b07

The download contains approximately 521.5 MiB of adapters, tokenizer and configuration. The 8B backbone is downloaded separately by the loader. Both adapters use LoRA rank 32, alpha 64. RG-Diff is merged into the backbone; SC-AR is then attached to that merged model and enabled only during refinement. The original Stage-II training script identifies this RG-Diff checkpoint as its inherited base; an original per-checkpoint pairing manifest is unavailable. The exported bundle records an RG-Diff fingerprint to prevent subsequent mixing.

Installation

Use Python 3.10 or 3.11 and FlexDLLM 0.1.1 or later. The source repository is currently private; cloning requires GitHub access.

git clone https://github.com/wangqinsi1/FlexDLLM.git
cd FlexDLLM
python -m venv .venv
source .venv/bin/activate
pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -e .

If this model repository is private, authenticate with an account that has access:

hf auth login

Give a question, receive a response and decoding trace

from flexdllm import FlexDLLM

model = FlexDLLM.from_pretrained("Qinsi1/FlexDLLM-LLaDA-8B", device="cuda")
result = model.generate(
    "A robe takes 2 bolts of blue fiber and half that much white fiber. "
    "How many bolts in total does it take?",
    return_trace=True,
)
print(result.text)
print("Refinements:", result.refinements)
for step in result.trace:
    if step["phase"] == "refinement":
        print("Before:", step["before"])
        print("SC-AR patch:", step["patch"])

Or use the command line:

flexdllm --model Qinsi1/FlexDLLM-LLaDA-8B \
  --prompt "A robe takes 2 bolts of blue fiber and half that much white fiber. How many bolts in total does it take?" \
  --trace --output outputs/example.json

--trace shows intermediate diffusion states, ASK events and repairs. ▮ marks unresolved positions. The final response replaces ASK with generated content; inspect result.refinements and result.trace to see interventions. ASK is selected by the model and is not guaranteed for every question.

The bundle automatically applies the archived checkpoint's original GSM8K question template and leading BOS token (prompt_format=legacy_gsm8k). Pass only the question; do not add a chat template or the GSM8K wrapper yourself. The supported input is a question string or one user message.

For a pinned download, pass revision=COMMIT_SHA to from_pretrained, or --revision COMMIT_SHA to the CLI. A downloaded local bundle directory is also accepted; --local-files-only uses cached files, and --base-model can select a local backbone directory.

The public repository was also downloaded without authentication, with forced re-download and SHA-256 verification of all 13 manifest-listed files. The installed FlexDLLM 0.1.1 wheel loaded the downloaded bundle by Hub ID and pinned revision 7744548ec28c585d21230f6dc59748f46fb0e084. The natural ASK/repair/resume check passed again on an H200. The pinned backbone was reused from the existing cache. The result and validation scope are recorded in validation.json.

Training and generation configuration

Stage I was trained on prepared normal-generation, failure-localization and post-repair continuation examples. Stage II used online teacher patch records and a single-token autoregressive training objective, with a 30-token workspace. The archived training states report 4 epochs for RG-Diff and 3 for SC-AR. The original datasets are not included in this model repository.

Default generation uses 256 canvas positions, four committed tokens per step, at most 128 diffusion steps, three refinements, and a 30-token repair workspace. Positions after an inserted repair are re-masked. These defaults follow the archived checkpoint implementation; the paper body also describes a 10-token workspace, selectable with --patch-max-tokens 10.

Validation and limitations

A local H200/bfloat16 check on the robe question naturally emitted ASK, generated 2+1=3 with SC-AR, resumed diffusion, and completed with EOS. It used one repair, five AR tokens and 64 diffusion steps. There was no forced ASK and no external Qwen model in this run. This validates the inference mechanism on one example; it does not establish benchmark accuracy or reliable correction on every input. The response formatting can vary, and the model can generate incorrect answers.

This archived pair is intended for English math reasoning experiments. The newer task-specific v5 RG-Diff checkpoints are different models and should not be substituted for the RG-Diff adapter in this bundle without retraining or validating a matched SC-AR adapter.

Files and provenance

  • flexdllm_config.json: backbone revision, adapter paths and generation settings.
  • rg_diff/: Stage-I adapter and the trained ASK output-head row.
  • sc_ar/: Stage-II adapter and exported pairing fingerprint.
  • tokenizer/: tokenizer and special-token files.
  • release_manifest.json: checkpoint provenance, code revision and file hashes.
  • validation.json: compact record of the release smoke test.

See the source repository for training, data preparation, API documentation and tests.

Licenses

The base model declares the MIT license. The FlexDLLM source code is Apache-2.0. A separate license for these adapter weights has not yet been designated by the model owner.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Qinsi1/FlexDLLM-LLaDA-8B

Adapter
(63)
this model