Field Notes

The case for a private AI appliance

August 26, 2026 · Katy, Texas

Every month, another headline. A law firm's internal memos show up in a ChatGPT training corpus. A hospital's patient referral strings appear in a public search result. A real estate brokerage's listing strategy leaks through a browser extension nobody remembered installing.

The industry response has been almost uniformly: "train your people not to paste sensitive data into public chatbots."

Which is, charitably, like telling a river not to get wet.

The central contradiction

The same tools that make AI useful — reading long documents, summarizing dense text, drafting from context — are the ones that require you to trust a third party with your data. Every "paste this contract and summarize it" command is a packet leaving your control, crossing a network, landing on someone else's server, and (depending on the fine print) training someone else's model.

Most professionals I talk to already know this. They feel the friction every time they reach for ChatGPT and stop themselves. They want the capability. They don't want the exposure.

The central tension

This tension is not a training problem. It's a product gap. The capability exists. The trust layer does not.

What the cloud vendors won't tell you

The big AI platforms have a structural conflict of interest. Their business model — inference at scale, model improvement via usage data, cross-customer training signal — depends on data moving through their pipes. Every major provider has been caught training on customer data at some point, and even the ones that promise "no training on your data" still route your queries through their infrastructure.

The enterprise "solutions" — Azure private endpoints, AWS Bedrock with data zones — are better, but they're still someone else's computer. And they cost enterprise money: $10–50+ per user per month, typically with a commitment term and a long procurement cycle.

An alternative that's been sitting on your desk

The math on local inference has flipped in the last twelve months. A 27-billion-parameter model running on unified memory — the kind of hardware Apple just started shipping — can answer questions, summarize documents, draft correspondence, and review contracts at competitive speeds. No network call. No third-party server. No training leakage.

The model is the appliance. The appliance is the model. They never separate.

This isn't a "local AI is almost good enough" story. It's a "local AI is better for this use case" story. Because:

What this looks like in practice

We install a sealed appliance in your office. It connects to your document systems and learns your practice — your templates, your prior work product, your preferred style. Then it works alongside your staff, not as a chat window they tab over to, but as a background agent that takes work off their desks.

A new referral packet lands. The appliance summarizes it, drafts the intake letter using your firm's template, stages it for review — all before anyone opens the file. A contract comes in for redline. The agent runs it against your playbook, flags the deviations, proposes markups. A regulatory filing deadline approaches. The agent prepares the draft, surfaces the checklist, reminds the responsible person.

Your people don't paste a document into a chat box and wait for a reply. They approve, adjust, and send. The difference between a chatbot and an agent is the difference between a search engine and an associate who knows how your office runs.

This is possible because the appliance lives on your network — it reads your files the same way your staff does, it knows your templates the way your staff does, and it never needs to ask permission to look something up. There is no upload button. There is no "paste this document and hope" step. The data never leaves; the agent works where the data lives.

An IT department in a box

The appliance is more than a model on a desk. It's the whole stack that makes a model trustworthy in a professional office: the encrypted agent-to-agent comms, the version control, the monitoring, the update pipeline, the fail-closed-to-local posture. The model is the part people talk about; the rails around it are the part that earns trust. If you can't run the infrastructure, the agent is just a chatbot on a shaky foundation.

That's the genuinely interesting part of this category. The value isn't the model — it's the rails you put around it. An agent that reads your files, drafts from your templates, and never lets a byte leave the building is only useful because the infrastructure underneath it is real. The strongest agent deployments are built on top of real sysadmin competence.

The point

Your work product stays yours. The appliance simply makes your team faster.

We're starting with a handful of local firms in the Katy / Houston area, and the first few deployments will tell us more than any spec sheet about the right shape of this. The category — agents for regulated professionals — is just starting to emerge, and it's early enough that the people building it well are still figuring out what it should be.


← Back to Field Notes