Skip to content

Local where possible, cloud where it is needed

The rule is simple: local whenever possible, cloud when it is not. Plenty of processing is not confidential at all — translating a product sheet, summarising a public text — and handing that to a commercial service is perfectly fine. What decides is the nature of the data and the power actually required, never a principle. And the gap is closing fast: models installed on your own hardware now do what needed the cloud a year ago.

Local, sovereign AI

The most common fear among business owners is not that AI gets things wrong: it is that it carries their data elsewhere. That fear is sound for some processing, and unfounded for plenty more.

So we sort, and sorting is simple. A model on your hardware handles the bulk: reading, filing, summarising, extracting. What needs more power, or carries nothing confidential, goes to a commercial service without hesitation. You know exactly where the line sits, because you draw it.

This is not a theoretical arrangement. It is how my own systems work — including this site, whose videos are described by a commercial service, because no can read a video yet.

On your servers

Nothing leaves your premises. More demanding to operate, and entirely achievable — the usual mode when data is sensitive.

Hybrid

Sorting locally, deeper analysis outside, and a written rule stating what is allowed to leave.

Cloud, knowingly

The fastest and the most powerful. Perfectly legitimate for anything neither personal nor strategic, and often the right call — provided it was decided rather than drifted into.

How the call is made, case by case

There is no single answer, and claiming to be “all local” would be as dogmatic as sending everything out. Three questions settle it, one processing task at a time.

What kind of data is it?

A public product sheet, a catalogue translation, marketing copy: nothing to protect, cloud is the sensible choice. A client file, a contract, a staff record: it stays with you.

How much power is really needed?

Filing, extracting, summarising: a is plenty. Reasoning on a complex case or reading a video: the large models keep an edge, for now.

What does each side cost?

A machine amortised over three years against three years of subscription at your volume. The maths leans local when volume is steady, cloud when use is sporadic.

Tailored AI, wired to your trade

A generic assistant knows the world but not your company. It does not know who your clients are, how you invoice, or why that particular case is urgent. So it produces answers that are plausible and useless.

The work is to wire it up: to your data, your documents, your applications. Only then does a question like “revenue with this client over two years” get a real answer, drawn from your records rather than approximated.

Query your documents

Contracts, procedures, minutes: ask in plain language and get the answer with its source, without anything leaving your system.

Extract what matters

Pull structured, usable information out of a document, at scale, with no retyping.

Sort and route

Incoming mail read, qualified and routed to the right person, with urgent items surfaced immediately.

-assisted development

This is what changed my work over the past two years, and what explains the timelines. I decide what needs building; an writes it at the pace of the conversation. Separately, both halves are ordinary: an expert alone is slow and expensive, an alone produces code nobody arbitrated. Together they do the work of a team.

What you gain is not the feat but the result: faster, cheaper, and modifiable afterwards. What you do not lose is judgement — thirty years of infrastructure exist precisely to know what should not be built.

Frequently asked

Do we need expensive hardware to run AI in-house?
Far less than people assume for everyday uses — reading, filing, summarising, extracting. The calculation is done at the diagnostic stage, honestly comparing hardware cost against three years of cloud subscription.
Is our data used to train a model?
Not in a local setup: the model reads your data to answer, it does not learn from it. If part of the flow goes through an outside service, that service’s terms are examined and written down before anything is deployed.
What about ?
only concerns personal data: translating a catalogue or summarising a public text falls outside it, and asking the question every time wastes effort. Where it does apply, the argument for local processing is simple: the less personal data leaves, the fewer transfers you have to justify. Compliance is an architecture question from the start, not a document bolted on at the end.
How long before we see something concrete?
The right way to start is a single, measurable use — one document type, one mailbox, one recurring question. We run it for real, measure it, and only widen the scope if it is worth it.

Start from one use, not from a strategy

Tell me which task costs you the most time today — that is almost always the right start. Thirty minutes, no strings, and you see whether what comes next speaks to you.