Skip to content
Manoj Deshmukh
All English essays

The Practical Technologist · 23 Jul 2026 · 5 min read

Production-Ready AI Isn't About Choosing LLM. It's About AI Evals.

By Manoj Deshmukh
Production-Ready AI Isn't About Choosing LLM. It's About AI Evals.

The Practical Technologist — weekly notes on AI, business, and life from 26 years of building things.


"It worked yesterday. Why is it giving a different answer today?"

Every engineer building an AI feature has asked this. Same prompt. Same model. Same data. Nothing changed and yet the output changed.

Welcome to probabilistic software.

For forty years, we built on a simple promise: same input, same output. That promise is what unit tests, integration tests and regression suites quietly rely on. Break the promise, and the entire quality playbook we grew up with starts to wobble.

Here is the uncomfortable truth is: The software is deterministic. The intelligence isn't.

So the interesting question of this decade is not "GPT-5 or Claude?" That's a Tuesday afternoon decision.

The real question is: how do you prove an AI system is ready to face real users? And the answer, increasingly, is a discipline most teams haven't built yet AI Evals.

The old playbook might break!

Think about how traditional software earns your trust.

Code → Unit Test → Integration Test → Regression Test → Deploy.

If every test passes, you ship. Confidence comes from deterministic behaviour the machine will do tomorrow exactly what it did today.

Now look at an AI request:

Prompt → Model → Context → Retrieved Documents → Temperature → User Intent → Response.

Every single layer introduces uncertainty. Passing once means very little, because the next execution can legitimately differ. So the question we ask has to change too.

Old question: Is my software correct? New question: Is my AI consistently useful?

That shift from correctness to consistency, from pass/fail to confidence is the whole game. This is why I've started telling my mentees: in the AI era

"Evaluation replaces Validation"

Here's the analogy that I can think of :

Traditional software is like commissioning a machine on a factory floor. You test it against a spec, sign off, and it repeats the same motion a million times. AI is like onboarding a brilliant new hire. You don't "unit test" a person. You put them on probation. You watch how they handle real cases. You check whether they solve the customer's actual problem not whether they recited the manual. And you keep evaluating them long after day one, because people (and models) drift.

You don't validate a colleague. You evaluate them, continuously. AI is no different.

A Production Readiness Pyramid

Most "AI eval" articles stop at the metrics. As an architect, I care about when an AI system is fit to serve users. So here's how I structure the decision four levels, bottom to top.

Level 1: Technical Validation. Does the plumbing work? Prompt runs, API responds, tool calls fire, the agent completes its loop. This is the closest thing AI has to unit tests. Necessary, and nowhere near sufficient.

Level 2 : Quality Evaluation. Now judge the answer, mostly with automated evals and no humans yet: correctness, hallucination, relevance, faithfulness to source, instruction-following, toxicity and safety. A demo lives and dies here. A product can't stop here.

Level 3 : Business Evaluation. Does it solve the actual problem? For a customer-support bot, the question is not "was the answer accurate?" It's: was the ticket resolved, was the customer satisfied, did escalations fall, did average handling time drop? Business outcomes matter more than BLEU scores. This is the level teams skip and then wonder why a "95% accurate" bot annoys everyone.

Level 4 : Production Evaluation. The system is live, and now you watch forever. Latency, cost, failures, prompt drift, model drift, knowledge drift, user feedback, conversation abandonment, retries, fallback frequency. This level never ends. As Thoughtworks argues in their practical framework for evaluating AI agents, evaluation has to continue after deployment, not stop at release.

Notice the shape. A demo needs Level 1. A production system needs all four.

Different jobs need different evals

"Evals" isn't one bucket. In practice I split them by what they protect:

Functional evals — did the tool call actually execute?

RAG evals — did we retrieve the right documents, stay grounded in them, and keep the context relevant?

Agent evals — did the agent pick the correct tool, plan sensibly, and recover when a step failed?

Human evals — would a domain expert approve this answer?

Business evals — did the customer achieve their goal?

Mix these deliberately. A RAG chatbot lives or dies on retrieval quality; an autonomous agent lives or dies on planning and recovery.

The lifecycle that never closes

Here's the loop I now design into every serious AI engagement:

Develop → Synthetic Evals → Golden Dataset → Regression Evals → Human Review → Pilot Users → Production Monitoring → Collect Failures → Convert Failures into New Evals → Repeat, forever.

That second-last step is the one that separates mature teams from hopeful ones.

Every production failure should become a permanent test case.

The best AI teams don't just fix a bad answer they capture it, turn it into an eval, and make sure that failure can never quietly return. Your evaluation suite should grow from real user interactions, not stay frozen at whatever synthetic examples you started with.

The Four Confidence Gates

If you take one thing from this piece, take these. Before any AI feature reaches production, I ask it to pass four gates:

👉 Gate 1 : Engineering Confidence: Can it technically work?

👉 Gate 2 : Quality Confidence: Is the answer good enough?

👉 Gate 3 : Business Confidence: Does it solve the intended problem?

👉 Gate 4 : Operational Confidence: Can we trust it tomorrow?

Not today. Tomorrow. Next month. After the prompt is tweaked, after the model is upgraded, after new knowledge is added. Most teams celebrate at Gate 1 and get ambushed at Gate 4.

A small prediction

Over thirty years, software engineering evolved from coding, to testing, to CI/CD, to DevOps, to observability. Each wave added a discipline the previous generation didn't think it needed.

I believe the next one every engineering team adopts is Continuous AI Evaluation.

Tomorrow's pipeline won't stop at Build → Test → Deploy. It becomes Build → Evaluate → Deploy → Observe → Learn → Evaluate Again.

Because in a world of non-deterministic systems, evaluation is no longer a testing activity tucked away before release. It becomes the foundation of trust.

Traditional software ships when the bugs are fixed. AI systems ship when confidence is high enough.


💬 A question for the builders: at which of the four gates does your AI actually get stuck today — engineering, quality, business, or operational? I suspect most of us underestimate the fourth. Tell me in the comments.

I write The Practical Technologist every week — practical takes on AI, business and life from 26 years of building things. Subscribe so the next one lands in your inbox.

First published on LinkedIn.

Read next