Madhav Sharma

Topic

Reliable AI systems

Making model output dependable enough to act on, with deterministic checks around the model, numbers computed in code, and a record of why each decision was made.

Overview

A model that is right most of the time is not yet a system anyone can rely on. The gap is closed by the code around the model, and that is the part of AI engineering I find most interesting.

Three habits show up across my work.

Deterministic first, model last. Anything that can be computed, matched or validated in code should be, before a model is involved. In my financial reporting engine that went as far as leaving the model out entirely: schema detection, variance and the narrative are all deterministic, so the report cannot invent a number.

Gates after the model. Raw model output goes through checks before it reaches a user. In-browser gesture recognition became usable when a decision layer, not a better model, removed the flicker: a confidence floor, per-gesture geometric rules, unanimous voting over eight frames and an intent lock.

Keep the evidence. Every lead my signal pipeline produces carries a verbatim quote from the posting or announcement that triggered it, and every row it drops goes into a rejection log. A result that can be traced back to its source can be checked, argued with and improved.

None of this is exotic. It is ordinary software engineering applied to a component that is unusually confident when it is wrong.

Projects

Where I did this

Questions

What does it mean to keep an LLM "out of the numbers"?

Every figure in the output is computed by ordinary code from measured data. The model may classify, rank or write prose around those figures, but it never produces a number of its own. That removes a whole class of fabrication and makes the output auditable.

Why put deterministic rules after a machine learning model?

Because a model's raw output is noisy in ways that matter to users. In my gesture recognition project, a confidence floor, geometric checks for specific confusions, an eight-frame unanimous vote and a short lock after each accepted gesture made the app usable without retraining the model.

How do you know a pipeline is working well?

By logging what it rejects as well as what it keeps. A rejection log with a specific reason per row shows whether a threshold is too strict, whether a source has broken, and whether the model is drifting, which a list of final results never does.

Related topics