Five of them, with the build decisions, the numbers, and the failures that were worth reporting.
Backend infrastructure so the client can stand up voice support for any company they onboard. My first plan was deliberately vendor-independent. I costed it, latency-tested it, then reversed my own position.
Independence lost to latency and reliability, and I would make that trade again.
The per-minute number came out of tuning, not procurement, and the four settings on the right are where it came from. Three profiles switch on an environment variable, so a faster stack can be A/B tested without a redeploy.
Retrieval was querying a Milvus collection before anything had been ingested. No crash, no error, just empty results and confident ungrounded answers. It now checks collection existence and row count at startup and, if empty, disables retrieval with a loud log line rather than failing mid-call.
result ✦ ₹2.29 per minute all in, at production quality
stack ✦ pipecat ✦ gemini 2.0 flash ✦ cartesia sonic ✦ bge-m3 ✦ milvus ✦ silero vad ✦ twilio ✦ redis
₹2.29 at ₹95/USD
kept from the first design ✦ BGE-M3 embeddings ✦ Milvus ✦ Silero VAD
The brief, in the client's own words, was to “optimise or find alternate to the existing Copilot functionality by improving contextual accuracy and interface responsiveness.” A facilities management platform was paying ₹95,000 a month to Microsoft Copilot Studio to answer questions about their own data, badly.
I built natural language to T-SQL against their Azure SQL schema. Intent classification routes each question to system docs, the database, or out-of-context. Semantic table selection puts only the relevant schemas in front of the model instead of the whole catalogue. Generated SQL carries correction logic, three cache tiers sit behind it: in-memory, disk and Redis. Every call is traced in LangSmith.
The figure on the right is a projection I built per layer before implementation. Spend after deployment tracked it. That is the part I care about: not the saving, but the fact that the model held.
the model totals ₹31,116; the balance to ₹35,000 is infrastructure and is unconfirmed. the two columns are not like-for-like. ₹95,000 was licensing, ₹35,000 is inference plus infra.
model tiering ✦ semantic table selection ✦ three-tier cache ✦ connection pooling
The reported distress rate was climbing, 21% to 27% year over year, and it read like a market signal. It was 187 spurious labels from companies that had simply not filed their accounts. I traced it to the label definition rather than to the model, and added a reportability gate that holds the rate inside a point.
It is XGBoost, and every prediction carries a SHAP attribution so an analyst can see which financials drove a flag before acting on it. The diagnostic engine sits at 92.155 percent on current-year distress flagging and the predictive engine at 87 percent on T+1 failure, both checked against company annual reports rather than a held-out split alone.
result ✦ 92.155% on current-year distress flagging ✦ 187 false labels removed
A Hinglish calling product for bus operators was hanging up mid-call, missing clear answers and talking over customers. The ask was for a better architecture. I took nine recordings and built a reusable diagnostic first, and scoped the paid deliverable as a labelled failure taxonomy with per-category rates rather than an architecture prescription.
That one-off script hardened into a repeatable method. It now runs on any Indian voice agent's real calls, and it answers what generic call-analytics cannot, which layer actually broke and what the cheapest fix is. Config first, then audio, then telephony, and the model is blamed last, only once the cheaper causes are ruled out.
Dual-channel capture was the first thing I asked for. The stall signature is an utterance under 1.2 s followed by more than 2.5 s of silence, and the format change mid-window means the June and July files are not strictly comparable.
Two of my own numbers moved while I worked. The stall count went from nine to eleven once I tightened the VAD merge window from 300 ms to 80 ms. And I withdrew a claim that one stall cluster sat just before the call ended, after checking the file was 117 seconds long with the cluster in the middle. I would rather show that than a clean story.
result ✦ the build was stopped before the money went in
the method, written up →every finding is pinned to one stage and ordered cheapest fix first
Every reported cause is a measured claim carrying its value, unit, threshold, timestamp and uncertainty. A hypothesis is never dressed up as a definite cause. The audio is analysed on-device and never leaves the machine, and each measurement is licensed from DSP first principles and checked against Praat.
An AI coding agent in private alpha gave me access to evaluate the product before wider release. I ran it as six separate tracks, from a static review of the shipped build to a security pentest, black-box functional testing, LLM red-teaming, capability evaluation and a usability pass. Each track was scored against a rubric the industry already recognises, so the findings sit against OWASP LLM Top 10, CWE, MITRE ATLAS and SWE-bench rather than my own taste.
The code-level work tested what the product actually ships, not a paraphrase of it. Its own guard functions were lifted from the shipped build and run under a small offline harness, so every number here re-runs to the same result. Live behaviour was tested as fixed protocols, a set prompt with a prediction registered in advance, so a result is unambiguous rather than a matter of opinion.
The attacks it was built to stop were held, cleanly. The gaps were architectural, the product exposing more of itself than it should and mishandling unusual input with no attacker present. That distinction is the part worth paying for, since it points at what to fix rather than only scoring the build.
result ✦ 18 findings ranked and reproducible, handed back with a remediation order
Every code-level number re-runs from an offline harness against the shipped build, so anyone on the team can reproduce it. Live results were registered as predictions before the session, not written up to fit after.
Three questions and I will tell you whether it is worth a call.