Configuration Parsing Warning:In config.json: "expert_dtype" must be a string

ReSI-DeepSeek-V4-Flash

This repository contains the complete DeepSeek-V4-Flash-0731 model after ReSI alignment, together with its tokenizer and inference code.

Quick start

Use Python 3.12 with a CUDA-enabled PyTorch installation. After downloading the repository, install the inference dependencies from requirements.txt:

pip install -r requirements.txt
import sys
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from transformers import AutoModelForCausalLM, AutoTokenizer

model_dir = Path(snapshot_download("AI45Research/ReSI-DeepSeek-V4-Flash"))
sys.path.insert(0, str(model_dir / "encoding"))
from encoding_dsv4 import encode_messages

tokenizer = AutoTokenizer.from_pretrained(model_dir)
model = AutoModelForCausalLM.from_pretrained(
    model_dir,
    dtype=torch.bfloat16,
    device_map="balanced",
    attn_implementation="eager",
).eval()
text = encode_messages(
    [{"role": "user", "content": "Explain why the sky is blue in three sentences."}],
    thinking_mode="chat",
)
inputs = tokenizer(text, add_special_tokens=False, return_tensors="pt").to(
    model.get_input_embeddings().weight.device
)
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    outputs = model.generate(
        **inputs, max_new_tokens=8192,
        do_sample=True, temperature=1.0, top_p=1.0, top_k=0,
    )
print(tokenizer.decode(outputs[0, inputs.input_ids.shape[-1]:], skip_special_tokens=True))

Alternatively, run the included script for streaming output:

python inference.py --model . --prompt "Explain why the sky is blue in three sentences."

The example and script use temperature 1.0, top-p 1.0, and no top-k filtering. The script accepts a local model directory or Hugging Face repository ID and supports sampling overrides. The default output limit is 8,192 tokens; increase --max-new-tokens for longer answers.

Runtime

The verified loading path uses the pinned Transformers revision in requirements.txt, BF16 weights, and eager attention. Model weights occupy approximately 529.7 GiB; allow additional GPU memory for inference. Keep the official message encoder in encoding/ and the BF16 autocast context shown above.

Code and paper

Citation

If you use ReSI in your research, please cite:

@misc{zheng2026resirecursivesafetyimprovement,
  title = {ReSI: Recursive Safety Improvement toward
           Resistant and Resilient AI},
  author = {Jingnan Zheng and Dongcheng Zhang and Yi Zhang and Ming Zhang
            and Qiaosheng Zhang and Youbang Sun and An Zhang and Xiangnan He
            and Tat-Seng Chua and Xia Hu and Bowen Zhou
            and Chaochao Lu and Xiang Wang},
  year = {2026},
  eprint = {2610.12233},
  archivePrefix = {arXiv},
  primaryClass = {cs.CR},
  url = {https://arxiv.org/abs/2610.12233}
}
Downloads last month
-
Safetensors
Model size
284B params
Tensor type
F32
·
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AI45Research/ReSI-DeepSeek-V4-Flash

Finetuned
(37)
this model

Collection including AI45Research/ReSI-DeepSeek-V4-Flash

Paper for AI45Research/ReSI-DeepSeek-V4-Flash