case studies ✦ five engagements

What these systems had to do, and what they cost to run.

Five of them, with the build decisions, the numbers, and the failures that were worth reporting.

01 ✦ enterprise voice agent platform ✦ in production

A voice agent that runs at ₹2.29 a minute

Backend infrastructure so the client can stand up voice support for any company they onboard. My first plan was deliberately vendor-independent. I costed it, latency-tested it, then reversed my own position.

Independence lost to latency and reliability, and I would make that trade again.

The per-minute number came out of tuning, not procurement, and the four settings on the right are where it came from. Three profiles switch on an environment variable, so a faster stack can be A/B tested without a redeploy.

the failure worth reporting

Retrieval was querying a Milvus collection before anything had been ingested. No crash, no error, just empty results and confident ungrounded answers. It now checks collection existence and row count at startup and, if empty, disables retrieval with a loud log line rather than failing mid-call.

result ✦ ₹2.29 per minute all in, at production quality

stack ✦ pipecat ✦ gemini 2.0 flash ✦ cartesia sonic ✦ bge-m3 ✦ milvus ✦ silero vad ✦ twilio ✦ redis

per-minute variable cost
STT, Google standard$0.024
LLM, Gemini 2.0 Flash$0.0001125
TTS, Neural2 / HD$0.000008
RAG, retrieval + embeddings$0.00002
total $0.02414

₹2.29 at ₹95/USD

the reversal
planned ✦ vendor independent
Faster-Whisper Medium
4-bit Qwen
XTTS-v2
LangGraph
→
shipped ✦ latency first
Google Cloud STT latest_short
Gemini 2.0 Flash
Cartesia Sonic
Pipecat

kept from the first design ✦ BGE-M3 embeddings ✦ Milvus ✦ Silero VAD

max_tokens = 60
answers stay phone length
automatic punctuation off
one less transform in the loop
8 kHz PCM
matches Twilio native rate, nothing is resampled
temperature 0.3
phrasing stays repeatable
02 ✦ natural language to SQL ✦ facilities SaaS

₹95,000 a month for bad answers

The brief, in the client's own words, was to “optimise or find alternate to the existing Copilot functionality by improving contextual accuracy and interface responsiveness.” A facilities management platform was paying ₹95,000 a month to Microsoft Copilot Studio to answer questions about their own data, badly.

I built natural language to T-SQL against their Azure SQL schema. Intent classification routes each question to system docs, the database, or out-of-context. Semantic table selection puts only the relevant schemas in front of the model instead of the whole catalogue. Generated SQL carries correction logic, three cache tiers sit behind it: in-memory, disk and Redis. Every call is traced in LangSmith.

question→intent routing
docs · db · out of context
→semantic table selection
only relevant schemas
→T-SQL + correction→cache
memory · disk · redis
→answer

The figure on the right is a projection I built per layer before implementation. Spend after deployment tracked it. That is the part I care about: not the saving, but the fact that the model held.

before
₹95,000
licensing
after
₹35,000
inference + infra
the model, per 100k prompts
query routing · gpt-3.5 turbo$27.50
sql generation · gpt-4o mini$198.00
answer generation · gpt-4o mini$138.00
rag subtotal$363.50 · ₹30,898
doc intelligence · gpt-4.1-mini
per 10k pages
₹219
combined ₹31,116

the model totals ₹31,116; the balance to ₹35,000 is infrastructure and is unconfirmed. the two columns are not like-for-like. ₹95,000 was licensing, ₹35,000 is inference plus infra.

model tiering ✦ semantic table selection ✦ three-tier cache ✦ connection pooling

03 ✦ risk intelligence ✦ fintech

The metric was measuring the wrong thing

The reported distress rate was climbing, 21% to 27% year over year, and it read like a market signal. It was 187 spurious labels from companies that had simply not filed their accounts. I traced it to the label definition rather than to the model, and added a reportability gate that holds the rate inside a point.

It is XGBoost, and every prediction carries a SHAP attribution so an analyst can see which financials drove a flag before acting on it. The diagnostic engine sits at 92.155 percent on current-year distress flagging and the predictive engine at 87 percent on T+1 failure, both checked against company annual reports rather than a held-out split alone.

result ✦ 92.155% on current-year distress flagging ✦ 187 false labels removed

reported distress rate 21%→ 27% 187 spurious labels, not a market signal
diagnostic engine92.155%
current-year distress flagging
predictive engine87%
T+1 failure, checked against annual reports
04 ✦ calling product ✦ diagnostic

I built a diagnostic instead of an architecture

A Hinglish calling product for bus operators was hanging up mid-call, missing clear answers and talking over customers. The ask was for a better architecture. I took nine recordings and built a reusable diagnostic first, and scoped the paid deliverable as a labelled failure taxonomy with per-category rates rather than an architecture prescription.

That one-off script hardened into a repeatable method. It now runs on any Indian voice agent's real calls, and it answers what generic call-analytics cannot, which layer actually broke and what the cheapest fix is. Config first, then audio, then telephony, and the model is blamed last, only once the cheaper causes are ruled out.

Dual-channel capture was the first thing I asked for. The stall signature is an utterance under 1.2 s followed by more than 2.5 s of silence, and the format change mid-window means the June and July files are not strictly comparable.

Two of my own numbers moved while I worked. The stall count went from nine to eleven once I tightened the VAD merge window from 300 ms to 80 ms. And I withdrew a claim that one stall cluster sat just before the call ended, after checking the file was 117 seconds long with the cluster in the middle. I would rather show that than a clean story.

result ✦ the build was stopped before the money went in

the method, written up →
what the nine recordings said
recordings in the sample9
sample rate8 kHz mono, all nine
energy above 3.4 kHz0.003% to 0.6%
capture formatPCM in June, 32 kbps MP3 in July
stall signature4 of 9 recordings
true barge-innot measurable
which layer is it
caller→ telephony→ recording→ suppression→ VAD→ STT→ LLM→ TTS→ playback

every finding is pinned to one stage and ordered cheapest fix first

why India breaks it differently
narrowband wallmeasured per call, not a label
retroflex vs dentalrides F3, above 3.4 kHz
Hinglishcode-switch mid-sentence
name confusabilityMeenatchi to Minakshi
the discipline

Every reported cause is a measured claim carrying its value, unit, threshold, timestamp and uncertainty. A hypothesis is never dressed up as a definite cause. The audio is analysed on-device and never leaves the machine, and each measurement is licensed from DSP first principles and checked against Praat.

05 ✦ independent evaluation ✦ AI coding agent

An outside pass before it shipped

An AI coding agent in private alpha gave me access to evaluate the product before wider release. I ran it as six separate tracks, from a static review of the shipped build to a security pentest, black-box functional testing, LLM red-teaming, capability evaluation and a usability pass. Each track was scored against a rubric the industry already recognises, so the findings sit against OWASP LLM Top 10, CWE, MITRE ATLAS and SWE-bench rather than my own taste.

The code-level work tested what the product actually ships, not a paraphrase of it. Its own guard functions were lifted from the shipped build and run under a small offline harness, so every number here re-runs to the same result. Live behaviour was tested as fixed protocols, a set prompt with a prediction registered in advance, so a result is unambiguous rather than a matter of opinion.

the finding worth reporting

The attacks it was built to stop were held, cleanly. The gaps were architectural, the product exposing more of itself than it should and mishandling unusual input with no attacker present. That distinction is the part worth paying for, since it points at what to fix rather than only scoring the build.

result ✦ 18 findings ranked and reproducible, handed back with a remediation order

findings, ranked
critical3
high6
medium5
low4
verified strengths 8
six tracks, one rubric each
static reviewOWASP LLM · CWE
security pentestWSTG · ATLAS
black-box functionalISTQB
LLM red-teamOWASP LLM · ATLAS
capability evalSWE-bench · HumanEval
usabilityNielsen · NIST AI RMF
the discipline

Every code-level number re-runs from an offline harness against the shipped build, so anyone on the team can reproduce it. Live results were registered as predictions before the session, not written up to fit after.

Something in production behaving oddly?

Three questions and I will tell you whether it is worth a call.

Send me the brief → How I consult
home partners github linkedin x scholar weights & biases npm events whatsapp email
reviewed october 2026
WhatsApp me