Instructions to use nace-ai/drex-dlm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nace-ai/drex-dlm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="nace-ai/drex-dlm", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nace-ai/drex-dlm", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Drex DLM
We introduce Drex DLM from Nace.AI, a decision model that answers typed questions about a given context. Pass that context as state, provide one or more named questions, and get a probability for every option. One forward pass of an 8B diffusion language model plus a pointer head produces that distribution.
The backbone is NVIDIA Efficient-DLM-8B, a diffusion language model with a decision adapter merged into its weights. The block tensor names match Qwen3. The shared context (state) uses bidirectional attention. Each question branch attends to that context and uses causal attention within the branch; state tokens cannot attend to questions, and question branches cannot attend to one another.
A shared pointer head projects the final-layer hidden state at the decision marker into a query and the final-layer hidden state at each option-ending marker into a key. Scaled dot products, temperature scaling, and a softmax over each question's options produce the probabilities.
The diagram uses <decide> and </opt> as readable aliases for the decision and option-ending markers. Packed requests within the token budget use one forward pass; larger requests are split across question rows.
BF16 checkpoint: nace-ai/drex-dlm
Q8_0 GGUF: nace-ai/drex-dlm-Q8_0
Code: nace-ai/drex-dlm
Server: nace-ai/llama.cpp, branch edlm
The model supports a maximum context window of 32,768 tokens, but the local Python, llama-server, and Ollama runners use 16,384 tokens by default. The 32K window is not enabled automatically; Context length explains how to configure it for each runner.
Python 3.12 is the validated interpreter. Local inference was tested on an Apple M5 Max with 128 GiB of unified memory; CUDA and CPU-only inference have not yet been validated. BF16 weights occupy about 16 GB, but actual memory use grows with context length and batching. Long-context quality is experimental.
Agentic harnesses
Integrate Drex DLM into coding agents and other agentic harnesses with the Drex agent skill. Its self-hosted instructions show how to point your agent at this model's local /v1/systemone server (see Serving); no hosted API key is needed.
Decision Index 0.2
Scores reported as of October 2026.
| Model | Index | Knowledge & Reasoning | Language Understanding | Retrieval & Classification | Tools & Automation | Arts & Human Taste |
|---|---|---|---|---|---|---|
| Drex DLM | 52.31 | 50.71 | 57.77 | 54.25 | 48.59 | 50.25 |
| Decider chat · Gemma-4-31B | 51.93 | 44.54 | 58.91 | 50.34 | 66.99 | 38.89 |
| Jev | 51.67 | 50.54 | 59.74 | 43.35 | 66.86 | 37.86 |
| AutoJev-27B | 50.94 | 40.93 | 61.90 | 42.00 | 69.98 | 39.88 |
| Jebadiah 27B | 50.32 | 38.82 | 58.97 | 45.48 | 69.30 | 39.02 |
| simple-jev · Qwen3.8-27B | 50.21 | 36.61 | 60.16 | 50.22 | 66.95 | 37.10 |
The other five are the top of the Decision Index 0.2 board.
Quick start
git clone https://github.com/nace-ai/drex-dlm.git
cd drex-dlm
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
hf download nace-ai/drex-dlm --local-dir ../drex-dlm-weights
python inference.py --model ../drex-dlm-weights --request examples/request.json
The checkpoint lives in a sibling directory so its code cannot overwrite the GitHub checkout. The hf command requires huggingface_hub>=0.34, installed by requirements.txt (CLI rename in v0.34.0). The inference command prints answers for the sample ticket. To run the inference server instead:
python serve.py --model ../drex-dlm-weights --port 8000
curl http://127.0.0.1:8000/v1/systemone \
-H 'Content-Type: application/json' \
-d @examples/request.json
GET /health returns {"status": "ok"} once the weights are loaded. The same request also runs through llama-server and Ollama (see Serving). python onboard.py prints commands using the downloaded sibling ../drex-dlm-weights when present; use python onboard.py --weights /path/to/checkpoint for another location. The helper puts its generated Modelfile beside the GGUF, not in the GitHub checkout, and does not replace an existing one; it rejects linked Modelfiles rather than following them. Set OLLAMA_SOURCE=/path/to/nace-edlm-checkout if the Ollama fork is not at the default sibling ../ollama; the printed create command uses that checkout's executable.
Request format
Two fields matter:
| Field | What to put there |
|---|---|
state |
The document: a string, object, or list. Nested objects render as indented text. |
questions |
A dictionary of named questions. The names come back as the keys of answers. |
{
"state": {
"ticket": "I was charged twice for the same order. Please refund the extra payment."
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments, charges, and refunds",
"technical": "Bugs and outages",
"other": "Anything else"
}
},
"refund": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Soon", "Urgent"]
}
}
}
| Type | You supply | You get back |
|---|---|---|
choice |
Named options, each with a short description. Object key order is option order. | choice, a probability per name, and confidence |
noul |
A yes-or-no question, with optional criteria.true / criteria.false descriptions. |
noul, the probability of "yes," from 0 to 1 |
score |
An ordered scale, lowest first. | A probability-weighted score, plus legend, per-level probabilities, and confidence |
choice.confidence rescales the winning probability above a uniform baseline. score.confidence measures concentration near the modal level; neither is a calibrated probability of correctness. The local score-confidence formula approximates the hosted API's statistic.
A local bfloat16 run of examples/request.json returns:
{
"answers": {
"team": {
"type": "choice",
"choice": "billing",
"confidence": 0.9461,
"probabilities": { "billing": 0.9641, "technical": 0.001, "other": 0.0349 }
},
"refund": { "type": "noul", "noul": 0.8548 },
"urgency": {
"type": "score",
"score": 1.4078,
"legend": ["Routine", "Soon", "Urgent"],
"probabilities": { "0": 0.1876, "1": 0.2169, "2": 0.5955 },
"confidence": 0.7039
}
}
}
usage.input_tokens is the encoded document plus questions (87 for this ticket). usage.output_tokens counts the serialized answers, not generated text. latency_ms is the scoring time.
One decision per request is the release contract. Experimental multi-decision requests work remains in a separate local development checkout and is not part of this release. Published Python, native, and Ollama builds do not support that wrapper; make separate calls on all three runners.
Serving
| Runner | Weights | Listens on |
|---|---|---|
| Python | drex-dlm-weights, including head.pt |
POST /v1/systemone, port 8000 |
| llama-server | drex-dlm-f16.gguf |
POST /v1/systemone, port 8097 |
| Ollama | the same GGUF | POST /v1/systemone, port 11434 |
A separate Q8_0 GGUF repository provides a smaller native-server alternative. Its pointer-head matrices remain F16 to support inference. Download it to the sibling weights directory and replace the F16 GGUF path in the native-server command below with its path; use the custom edlm fork and the same Apple Silicon Metal setting. The Q8 model card provides artifact details and the tested native configuration. Q8 use with Ollama has not been validated.
When combined questions exceed the packed-token cap, the runners can score separate question rows, provided the state plus each question fits the configured row limit. Floating-point probabilities can differ slightly between packed and row execution. Native context, batch, and microbatch capacities must each fit an encoded span. See Context length for details.
Python
source .venv/bin/activate
python serve.py --model ../drex-dlm-weights --host 127.0.0.1 --port 8000
The server loads the safetensors shards and head.pt from --model; --name sets the fallback response model name. Run it from drex-dlm after the quick start.
llama-server
The edlm architecture, its converter, and POST /v1/systemone live on branch edlm of nace-ai/llama.cpp. From drex-dlm, clone a sibling checkout and build the server (CMake and a C/C++ toolchain required; Apple Silicon also needs the Xcode Metal toolchain):
cd ..
git clone --branch edlm --single-branch https://github.com/nace-ai/llama.cpp.git llama.cpp
cmake -S llama.cpp -B llama.cpp/build
cmake --build llama.cpp/build --target llama-server --parallel 8
For NVIDIA, add -DGGML_CUDA=ON when configuring; CUDA has not been validated for this release. Use a separate conversion environment because converter dependencies differ from the inference requirements:
python3.12 -m venv llama.cpp/.venv-convert
llama.cpp/.venv-convert/bin/python -m pip install \
-r llama.cpp/requirements/requirements-convert_hf_to_gguf.txt
llama.cpp/.venv-convert/bin/python llama.cpp/convert_hf_to_gguf.py drex-dlm-weights \
--outfile drex-dlm-weights/drex-dlm-f16.gguf --outtype f16
cd drex-dlm
From drex-dlm, start the server, then send the sample from another terminal:
GGML_METAL_TENSOR_DISABLE=1 ../llama.cpp/build/bin/llama-server \
-m ../drex-dlm-weights/drex-dlm-f16.gguf \
--host 127.0.0.1 --port 8097 \
--embedding --pooling none \
-c 16384 -b 16384 -ub 16384 -np 1 --no-warmup
curl http://127.0.0.1:8097/v1/systemone \
-H 'Content-Type: application/json' -d @examples/request.json
Keep GGML_METAL_TENSOR_DISABLE=1 on Apple Silicon: the Metal tensor matmul path produced incorrect long-input results in validation. If an encoded span exceeds the native -c, -b, or -ub capacity, it is rejected; the server does not adjust rows to smaller launch capacities. The response includes Python's answers schema plus latency_ms and an x-typesafe-request-id header.
Ollama
nace-ai/ollama branch nace-edlm launches a custom llama-server and forwards POST /v1/systemone. Complete the native checkout and conversion above first. From drex-dlm, clone a sibling checkout and build the Apple Silicon runner and Go daemon (Go 1.26 with automatic toolchain download):
cd ..
git clone --branch nace-edlm --single-branch https://github.com/nace-ai/ollama.git ollama
cd ollama
export OLLAMA_LLAMA_CPP_SOURCE="$PWD/../llama.cpp"
cmake -S llama/server --preset darwin
cmake --build build/llama-server-darwin --target llama-server --parallel 8
GOTOOLCHAIN=auto go build -trimpath -o ollama .
cd ../drex-dlm
The Linux CUDA and CPU presets have not been validated for this release. Build from the matching native fork, not the stock Ollama binary. Write a Modelfile next to the converted GGUF:
cat > ../drex-dlm-weights/Modelfile <<'EOF'
FROM ./drex-dlm-f16.gguf
CAPABILITY decision
PARAMETER num_ctx 16384
EOF
Start the daemon from drex-dlm (leave it running) and, in another terminal, import and call the model:
OLLAMA_HOST=127.0.0.1:11434 \
OLLAMA_LLAMA_SERVER="$PWD/../ollama/build/llama-server-darwin/bin/llama-server" \
../ollama/ollama serve
OLLAMA_HOST=127.0.0.1:11434 ../ollama/ollama create drex-dlm \
-f ../drex-dlm-weights/Modelfile
curl http://127.0.0.1:11434/v1/systemone \
-H 'Content-Type: application/json' -d @examples/request.json
GET /api/version checks readiness. Current fork builds align native context, batch, and microbatch capacities; on macOS, the Ollama fork disables the problematic Metal tensor path for eDLM child runners.
Context length
Model capacity: 32,768 tokens. Local runner default: 16,384 tokens. The Python server defaults to 16K, the llama-server command above sets -c -b -ub to 16K, and the Ollama Modelfile sets num_ctx to 16K.
Python already uses the recommended default. To use the full window instead:
KEV_CONTEXT=32768 python serve.py --model ../drex-dlm-weights --port 8000
For llama-server, the encode cap and the slot size must move together — rebuild from the edlm branch after pulling this change, then:
GGML_METAL_TENSOR_DISABLE=1 SYSTEMONE_CONTEXT=32768 ../llama.cpp/build/bin/llama-server \
-m ../drex-dlm-weights/drex-dlm-f16.gguf \
--host 127.0.0.1 --port 8097 \
--embedding --pooling none \
-c 32768 -b 32768 -ub 32768 -np 1 \
--no-warmup
For Ollama, set PARAMETER num_ctx 32768 in ../drex-dlm-weights/Modelfile, repeat the fork's ollama create command, and restart its daemon with SYSTEMONE_CONTEXT=32768. num_ctx sizes the slot (the patched fork also sizes its decision batch and microbatch); the environment variable raises the encode cap. Keep them aligned. A successful 32K capacity check is not a correctness guarantee.
KEV_SERVE_MAX_STATE, KEV_SERVE_MAX_BRANCH, and KEV_SERVE_MAX_PACKED (Python), and their SYSTEMONE_* equivalents (llama-server), override individual caps if you need finer control — none of them are required to select the full window.
Validation
The model was evaluated on Apple M5 hardware using the Python, native llama-server, and Ollama runners. Results were consistent across runners and matched expectations.
The Q8_0 GGUF loads and serves requests through the native server at a 16,384-token context.
Files
| File | Role |
|---|---|
model-*-of-00004.safetensors |
Merged backbone, bfloat16, ~16 GB |
head.pt |
Pointer head: 256-dim query/key, temperature 1.0 |
inference.py |
Score one JSON file and print the answers |
serve.py |
POST /v1/systemone |
examples/request.json |
The ticket used in the sample response |
License
The model weights are released under CC BY-NC 4.0. The original code written by Nace.AI in this repository is under the MIT License.
- Downloads last month
- 19
