Richard Teachout // Teachout.com
← All writing

The 80/20 of an Internal AI Platform

Richard Teachout
Richard Teachout CTO at Ashley Furniture Industries - Executive Tech Leader, Entrepreneur, AI leader, Architect, Problem Solver, Ex-Developer. September 30, 2026
AI
The 80/20 of an Internal AI Platform

The internal AI platform is where AI budgets go to die. Every company builds one, and most of them build the wrong parts.

I've seen the pattern enough to name it: the platform team builds the shiny infrastructure — the model gateway, the orchestration framework, the prompt playground, the fancy UI — and the business teams still can't get anything done. The platform has a hundred features and zero adoption, because the features are the ones the platform team wanted, not the ones the business needs.

The fix is a brutal exercise in priorities. Out of everything a platform could do, four things actually matter, and the rest is decoration.

The four things worth building

First: evaluation. This is the non-negotiable, the one the platform teams skip most often because it's unglamorous. A test set, a baseline, a regression check — the ability to answer "is this new model better or worse than what we have?" Evaluation is the audit layer of AI, and it's the difference between upgrading your models with confidence and upgrading them on faith. Every model changes, every prompt changes, and without an eval harness you're flying blind. Build this first. It's the foundation everything else stands on.

Second: routing. The model-tiering layer from the last article — the rules that send routine work to cheap models and hard work to the frontier. This is the cost-control layer, and it's also the reliability layer, because routing is how you add a fallback model when the primary one degrades. A platform without routing is a platform that pays first-class fares for everything and has no plan when the main model hiccups.

Third: observability. You can't operate what you can't see. Every request needs a trail: which model, what input, what output, what confidence, what latency, what happened after. This is the layer that makes incidents explainable, that turns "the AI did something weird" into "model X, prompt Y, confidence Z, here's what to fix." Without it, every AI problem is a mystery and every mystery gets blamed on the AI.

Fourth: escalation. The human-review path — where the outputs that matter go to a person before they go to a customer. This is the layer that makes high-stakes AI safe, and it's the layer that's almost always built last, after the incident that proves it was needed.

A platform is four things: evaluation, routing, observability, and escalation. Everything else is a feature. These are the platform.

What to buy instead

Everything else — the orchestration frameworks, the prompt tools, the model gateways, the vector databases, the playgrounds — buy it. The market for this is mature, the products are good, and the cost of buying is a fraction of the cost of maintaining a homegrown version.

The test for build vs. buy at the platform level is the same as everywhere else: is this a differentiator, and does it touch our data and our judgment? Evaluation touches your data and your judgment — it encodes what quality means to you. Routing touches your cost and your risk tolerance. Observability touches your operations. Escalation touches your customers. Those four are yours. The model gateway is a commodity. The orchestration framework is a commodity. The vector store is a commodity. Let the vendors have them.

The mistake the platform teams make is reversing this — building the commodity parts because they're fun and buying the differentiated parts because they're hard. The gateway is a weekend project that impresses in demos. The evaluation harness is months of boring work nobody wants to demo. Guess which one gets built.

The adoption test

Here's the test for whether your platform is right: can a business team ship something with it, without a platform engineer in the room?

If the answer is no, the platform has failed, regardless of how complete it is. The platform's job isn't to be impressive. It's to make the path from idea to working AI feature short enough that teams actually walk it. The four layers, wired so a team can drop in a task, point at data, set a review level, and go — that's the platform. Everything else is infrastructure in search of a user.

The teams that are actually using AI in production aren't the ones with the most impressive platforms. They're the ones whose platform does four boring things reliably: it evaluates, it routes, it observes, and it escalates. The demos are somewhere else, and they don't matter.

Where the 80/20 pays

Eighty percent of the value of an internal AI platform comes from twenty percent of the work — the four layers, done boringly well. The other eighty percent of the work, the shiny infrastructure, produces the demos and the downtime.

Build the four. Buy the rest. And when the next platform proposal lands on your desk with a roadmap full of gateways and frameworks, ask where the eval harness is. If it's not on the roadmap, the proposal isn't a platform. It's a hobby with a budget.

Think this argument fits your event? Tell me about the room — the calendar is selective.

Start a conversation