Who Actually Decides Where the Model Runs
The choice between a frontier API and your own hardware gets argued as an engineering question. It is rarely decided as one. The forces that push a workload off the frontier are organizational: regulation, procurement, and who gets blamed.
Every organization using AI has decided where the model runs, usually without noticing and almost always for the wrong stated reason.
The public argument is technical: capability against cost, latency against control, the frontier model against the open one you host yourself. Vendors on both sides publish benchmarks. Then the decision gets made in a meeting where no benchmark is mentioned, by someone weighing forces that have nothing to do with tokens per second.
I have been on both sides of that table, arguing for a deployment and living with one, and I want to describe those forces as they are.
The technical case is real, and it is not what decides
The engineering tradeoffs exist. I ran ten business tasks through a frontier model, a 7B model on my own hardware, and a router that chose between them per task. All three scored ten out of ten. The private leg was faster on average and cost nothing per call, because the hardware was already paid for.
That result is narrow — short, well-scoped extraction and classification tasks, one run, one task mix. On harder tasks I expect the frontier model to pull away. But it establishes the useful point: for a meaningful class of real work, capability is no longer the deciding variable. Both options work, so the decision is being made on something else.
Four forces, none of them technical
Regulation, which is a deadline more than a rule. Under GDPR and CCPA a company has thirty days to tell someone what personal data it holds on them. When I built the pipeline that handles those requests, model placement was not a judgment call. Compiling a person’s data in one place and sending that compilation to a public API creates a second copy of the problem the law exists to prevent. So the sensitive step runs on my own hardware, and if that endpoint is unavailable the pipeline refuses and says so rather than falling back. No latency benchmark bears on that decision.
Procurement, which is slower than your roadmap. A frontier API is a vendor relationship, and in a mid-market company a new vendor touching customer data means a security review, a data-processing agreement, and someone’s quarter. The team that self-hosts is often not choosing privacy. It is routing around a purchasing process.
Blame, which nobody says out loud. Behind every placement decision is a person who will have to explain it if it goes wrong. “The data never left our building” is a defensible sentence in a room where things have gone badly. “It was 40% cheaper per call” is not. People who over-weight this are reading their own incentives accurately, which is what I would want them to do.
Values, which are real and rarely admitted. For some organizations keeping data in-house is a statement about who they answer to, not a compliance measure, and they would make the same choice at a cost premium. I no longer treat this as a soft factor. It predicts behavior better than the cost model does.
Why the trajectory framing works better than the choice framing
The mistake is treating this as one decision, made once, at the start, for the whole company. That framing produces paralysis, because it asks people to be right on day one about workloads they have not built yet.
Treat it as a trajectory instead. Start on the frontier, because shipping beats theorizing and most work is not sensitive. As volume grows, some workloads develop reasons to move: a regulator, a contract, a cost curve, a latency floor. Introduce a router that decides per workload rather than per company, and move the slice that earns it.
This framing is better because it survives contact with an organization. Nobody has to be right on day one or defend a company-wide policy against a counterexample. The decision gets made per workload, where the facts live, by someone who can see them.
What I would ask instead
If you are having this argument inside your company, change the question. “Should we be public or private” has no answer at the company level.
Ask which specific workloads, if their contents leaked tomorrow, would produce a conversation you cannot have. Route those privately today and stop arguing about the rest. The remainder is a cost optimization, and that can wait until you have volume worth optimising.
Most organizations find that list shorter than they feared and more specific than they expected. A short, specific list is something you can build against. A company-wide philosophical position is something you can only hold meetings about, and I have sat in enough of those to prefer the list.
Published 28 September 2026, revised 28 September 2026. Narendra Nag is a founder and media executive writing on attention, streaming, and the economics of live sports.