Fx-Work

Fx-Work is a 35B-class multimodal language model for professional work in realistic file-and-tool environments. It is designed to interpret heterogeneous work materials, reason over operational constraints, use tools when required, and produce deliverables that satisfy explicit professional requirements.

Fx-Work is built on Qwen3.6-35B-A3B and post-trained with 20K tool-interaction trajectories. The trajectories are generated from occupational work constructed with real-world artifacts, occupational knowledge, actionable requests, and itemized evaluation criteria. Execution-guided verification is used during data construction to identify and repair inconsistencies between the work request, reference materials, and rubric.

Main Results

Fx-Work achieves the strongest results among the evaluated models at or below the 35B scale on all five reported metrics. It also surpasses DeepSeek-V4-Pro-Preview (1.6T) on four of the five metrics.

Fx-Work benchmark results

Benchmark Results

Results on GDPvalAA-v2, APEX-Agents-AA, and JobBench. GDPval Rubric is the normalized rubric score; GDPval Elo uses the August 4, 2026 snapshot; APEX-Agents reports pass@1; and JobBench reports the Main and Easy splits. AVG is the mean of normalized GDPval Elo, APEX-Agents-AA, and JobBench Main. -- indicates that a result was not publicly available or was not tested. Bold marks the best value among the models at or below the 35B scale.

Model Parameters GDPval Rubric GDPval Elo APEX-Agents pass@1 JobBench Main JobBench Easy AVG
Frontier models
Qwen3.8-Max 2.4T-A95B 91.53 1739 38.9 52.57 82.81 51.14
Kimi-K3 2.8T-A104B 91.64 1687 35.5 43.81 80.72 46.22
GPT-5.5 -- 90.50 1491 29.9 30.10 78.22 36.52
GLM-5.2 753B-A40B 90.34 1510 29.7 33.79 75.49 37.80
DeepSeek-V4-Pro-Preview 1.6T-A49B 87.17 1304 19.2 18.90 64.08 26.10
DeepSeek-V4-Flash-Preview 284B-A13B 86.49 1189 15.2 18.62 63.13 22.75
Models at or below 35B
Nex-N2-mini 35B-A3B 69.96 1065 15.6 9.39 52.93 17.74
Agents-A1 35B-A3B 74.31 877 11.7 7.19 41.74 12.58
Apodex-1.0-mini 35B-A3B 78.36 973 14.6 7.97 44.38 15.40
Occamy-1.0 35B-A3B 79.50 1200 20.5 18.88 65.49 24.79
Qwen3.6-27B 27B 85.10 1138 16.6 19.18 64.59 22.56
Qwen3.6-35B-A3B 35B-A3B 82.61 1053 13.4 14.99 56.02 18.68
Fx-Work 35B-A3B 86.66 1352 25.2 25.19 70.06 31.00

Quickstart

For streamlined integration, we recommend serving Fx-Work through an OpenAI-compatible API using SGLang or vLLM. The repository requires Hugging Face authentication when access is restricted.

Fx-Work uses a native context length of 262,144 tokens. If GPU memory is limited, reduce the context length; for long-horizon professional work, retaining at least 128K tokens is recommended.

SGLang

SGLang provides an OpenAI-compatible server for high-throughput inference.

python -m sglang.launch_server \
  --model-path endless-frontier/Fx-Work \
  --port 8000 \
  --tp-size 8 \
  --mem-fraction-static 0.8 \
  --context-length 262144 \
  --reasoning-parser qwen3

vLLM

vLLM provides a high-throughput OpenAI-compatible server.

vllm serve endless-frontier/Fx-Work \
  --port 8000 \
  --tensor-parallel-size 8 \
  --max-model-len 262144 \
  --reasoning-parser qwen3

Both commands expose an endpoint at http://localhost:8000/v1. Tool execution is provided by the surrounding agent harness; the model server exposes generation and, where configured, tool-call responses.

Transformers

import torch
from transformers import AutoProcessor, Qwen3_5MoeForConditionalGeneration

model_id = "endless-frontier/Fx-Work"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [
    {
        "role": "user",
        "content": "Summarize the key risks and next actions in this work request.",
    }
]
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
).to(model.device)

with torch.inference_mode():
    output_ids = model.generate(**inputs, max_new_tokens=1024)

new_tokens = output_ids[:, inputs["input_ids"].shape[1]:]
print(processor.batch_decode(new_tokens, skip_special_tokens=True)[0])

The processor supports the multimodal image and video message format provided by Qwen3.6.

Citation

If you use Fx-Work, please cite:

@inproceedings{zhu2026workgenesis,
  title     = {WorkGenesis: Building the Worlds That Teach Agents to Work},
  author    = {Zhu, Xinyu and Liu, Fenyi and Cai, Yuzhu and Tang, Shuo
               and Ye, Rui and Zhang, Linfeng and Chen, Siheng},
  year      = {2026}
}

License

Fx-Work is released under the Apache 2.0 license, subject to the terms of the base model and third-party components. See the Qwen3.6-35B-A3B license for the upstream terms.

Downloads last month
16
Safetensors
Model size
36B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for endless-frontier/Fx-Work

Finetuned
(329)
this model
Quantizations
2 models