For a long time I have been arguing that a well-designed set of small local models can outperform a larger frontier model in a lot of normal business work.
Not because the big model is weak. It is not. It is because most business work is not a grand strategy problem. It is a workflow problem. Retrieval, summarisation, classification, drafting, checking, routing, and structured action. That is where a local system can be cheaper, more reliable, and much easier to control.
That still feels right. But the argument has changed slightly because the real issue is no longer just cost. It is risk.
Every board should be able to answer two questions about its AI estate without blinking. What does this cost us as usage grows? And where, exactly, does our data end up?
I have yet to see many strategies that can answer both cleanly.
The meter you do not own
The cloud story is seductive because the unit cost looks small. Tokens are cheap. Invoices are manageable at first. And then the volume climbs.
That is when the model stops being a tool and becomes a cost centre with a vendor in the middle. A task that looked simple becomes a loop: understand, retrieve, plan, check, retry, summarise. The first prompt is not the whole cost. The system is the cost.
That is not a flaw in cloud AI. It is the nature of complex workflows. It just means the bill becomes harder to predict than the sales pitch suggests.
A local stack behaves differently. Once the hardware is in place, the cost curve is more stable. You own the hardware. You own the model choice. You own the capacity planning. The ten-thousandth question costs a lot more like the tenth than it does in a cloud model where the pricing is external and changeable.
For everyday work, predictability matters more than a sticker price that looks cheap on paper.
The data you cannot get back
The stronger argument is the one most boards already understand instinctively: if your employees are sending internal documents, customer records, source code, contracts, or operational data to a vendor model, that data leaves the building.
Not always in a dramatic way. Most of the time it is just a helpful prompt in a browser, a draft summary, a piece of code, a customer note. Harmless in isolation. Dangerous at scale.
Once that traffic crosses the boundary, you are relying on someone else's terms, someone else's infrastructure, and often someone else's jurisdiction. That has consequences. It affects security posture. It affects compliance. It affects what you can say in public.
And the boardroom has noticed. In the last year the discussion has shifted from “can we use this?” to “where does the data go and who controls it?”
The real answer is simple: if the data never leaves the organisation, it cannot leak in the same way. It cannot be retained by a vendor in some training pipeline. It cannot be subpoenaed by a foreign court without a defined legal process. It cannot become part of a system you cannot audit.
That is not just retrieval
People sometimes treat local AI as if it is only a better search box. That is not what makes it valuable.
The useful part is when you pair a local model with a curated set of internal knowledge, clear boundaries, and disciplined processes. The model reads the policy. It checks the contract. It compares the draft to previous decisions. It does not make up a rationale. It grounds itself in your own evidence.
That is genuine knowledge work, and it is exactly the type of work where a smaller and more controlled model can outperform a general frontier model that is just improvising from a pasted prompt.
It is not about raw intelligence. It is about disciplined system design.
When the right evidence is available and the workflow is constrained, the local model does not need to be the most powerful model in the room. It just needs to be the one that understands the rules better than the generic one does.
Regulation is making the argument harder to avoid
I do not think the law has to be the main driver of this decision, but it will increasingly be the final arbiter.
Data sovereignty is no longer a niche concern. It is becoming part of procurement. It is becoming part of governance. It is becoming part of platform design.
That is fine. The companies that have built a local-first architecture are not trying to dodge the future. They are trying to make it manageable.
The frontier model still has a place
I am not anti-frontier. I am anti-default.
There are tasks where a large model is genuinely the right tool: open-ended synthesis, strategic analysis, unusual reasoning, or work where the breadth of general knowledge matters more than containment.
But that should be an explicit decision. Not the default posture of the organisation. Not the pipe that all operational traffic flows through by default.
The serious architecture is local-first with controlled escalation. Work stays inside the organisation until it has to leave. Then it leaves with purpose, traceability, and a clear record of why.
Two questions, one design
What does it cost as usage grows? A number you can actually explain, or a meter you do not control.
Where does the data go? Nowhere meaningful, or everywhere and nowhere at once.
That is why the answer is not “cloud bad, local good.” It is: design for the business, not for the pitch deck.
