Exactly. That’s the point.
Arena measures the complete production serving stack, not just the underlying model, because that’s what developers actually experience.
This is aligned with the model-less inference philosophy introduced by Stanford’s INFaaS work, where the optimization target is the end-to-end inference system rather than a fixed model.
We’ll publish the full evaluation protocol to make every result reproducible.