Black Forest Labs' FLUX 3 Action is a "world action model" for the SO-101 robot arm: one diffusion process jointly denoises the next chunk of actions and the next chunk of video frames. Read as a headline, that sounded like exactly the causal-chain-first architecture our own safety pipeline argues for — action and outcome tied together in one step, not an action head bolted onto a frozen representation. So we rented an L40S on Brev and read the code, not just the model card.
Our pipeline is five floors, each depending on the one below it: Causal chain → Probability → Risk/Impact → Decision theory → Markov/Game theory. Here's what FLUX 3 Action actually has.
01
Causal chain
Present, and prioritized: video_loss_weight: 1.0 outweighs action_loss_weight: 0.5. The model is trained to get the outcome right more than the action itself — this is the real thing, not a gesture at it.
present
02
Probability
Technically present, never surfaced. It's a diffusion model — it samples from a distribution by construction. Nothing reads that distribution back out as an uncertainty number a decision could use. The probability exists inside the math and dies there.
hollow
03
Risk / impact
Absent. The model card says so itself: "nothing bounds joint velocity, force, workspace." Not hidden — just not built.
absent
04
Decision theory
Absent. No gate. The model executes 32 actions per chunk; there is no threshold at which it would stop.
absent
05
Markov / game theory
Not applicable at this scope — a single robot arm with no adversary or multi-round state.
n/a
The closure isn't "their floors 1–2 are weaker than ours." They're not — floor 1 here is arguably cleaner than most causal-chain implementations we've seen, because the loss weighting makes the priority explicit in the training objective itself, not just in a README.
A model with two good, real, working floors behaves identically to a model