Who Actually Decides Where the Model Runs
The choice between a frontier API and your own hardware gets argued as an engineering question. It is rarely decided as one. The forces that push a workload off the frontier are organizational: regulation, procurement, and who gets blamed.
Every organization using AI has decided where the model runs, usually without noticing and almost always for the wrong stated reason.
The public argument is technical: capability against cost, latency against control, the frontier model against the open one you host yourself. Vendors on both sides publish benchmarks. Then the decision gets made in a meeting where no benchmark is mentioned, by someone weighing forces that have nothing to do with tokens per second.
I have been on both sides of that table, arguing for a deployment and living with one, and I want to describe those forces as they are.