Eight ways to work together, from an independent evaluation before launch to a system built, shipped and put in front of buyers.
One system, one written report. What is going wrong, where the choke points and bottlenecks are, and the fixes I would make, in one to two weeks.
An outside pass over an AI product during private alpha, under authorised access. Static analysis of the shipped build, adversarial live sessions and functional oracles run together, so the report covers what the product does as well as where it gives. Findings are mapped to the OWASP Top 10 for LLM Applications, MITRE ATLAS, CWE and the OWASP Web Security Testing Guide.
I model the unit economics layer by layer, then the stack gets chosen against that model rather than against a benchmark. You see the per-unit cost of every layer before anyone commits to it.
I build the system, put it in production, and stay long enough to see what it does once real users arrive. Cost and behaviour are instrumented from the first commit, not added later.
For teams with their own engineers who want a second opinion on architecture, cost and evaluation choices before they commit to them.
Whether a system can be explained, audited and defended before it ships. Findings are mapped to the NIST AI Risk Management Framework and ISO/IEC 42001, and to the data protection regimes that apply to you: India's DPDP Act 2023 and the EU GDPR. The same framing sits behind my C-RAI policy draft and the responsible AI handbook.
Sessions on production AI, cost modelling and evaluation, run as workshops for your engineers rather than as lectures.
Positioning, pricing and the route to the first buyers who are not already your friends. I open the doors I can open, and I say plainly which ones I cannot.
An evaluation tells you what is true about the product today. What usually matters more is what that implies for the next two quarters: which defect is a roadmap item and which is a launch blocker, what the product can honestly claim in market, and which conversation moves it further than another release.
I work across that span, with founders and with the engineers who own the code, and I build go-to-market strategies for AI products around what they can credibly take to market. I keep a network of builders, operators and buyers, and make introductions where warranted.
My main ground is small and mid-sized businesses (SMBs), and the post-seed startups at Series A and B that sit alongside them. I go deepest here, evaluating the product, diagnosing what is holding it back, and settling what it can honestly take to market. I take on very small enterprises (VSEs) when the team is highly competent or already structured, and for larger, mid-market companies I come in for evaluation and strategy. Enterprise, with its field sales and procurement cycles, is a different game and not one I chase.
When a client needs more hands than mine, I bring in partners, engineering specialists, designers, GTM people and product founders. I do it to bring the product up to the standard it deserves and ready it for market, owning the whole process while the partner delivers their part under my brief.
The method does not change much between sectors. What changes is the data, the regulator and what counts as a wrong answer. These are the ones I have shipped in.
Voice agents that take bookings and handle support calls in Hinglish, around the clock. I diagnose them from real call recordings and hand back a labelled failure taxonomy with per-category rates, not an architecture diagram.
how I audit a voice agent →Agents that detect the caller's language, answer from your own knowledge base rather than the open internet, and escalate to a human the moment retrieval comes back thin.
Natural language to SQL over production schemas, with intent routing and semantic table selection so only the relevant tables ever reach the model. Costed per layer before a line of code is written.
Distress and failure signals, trade alerting, and transaction anomaly detection. Every prediction carries an attribution, so an analyst can see which inputs drove a flag before acting on it.
Clinical assistants that read a provider's curated material instead of general model knowledge, because general models hallucinate confidently on medical advice.
Vessel segmentation from X-ray angiography for cardiac care, and text-guided facial editing for aesthetic and reconstructive planning.
Analyst agents that reason across category hierarchies, customer profiles, transactions and purchase sessions together, rather than one flat exported table.
Workshops for engineering teams, university hackathon mentoring, and written technical material that practitioners actually use.
not on this list ? → the constraints usually rhyme. tell me the system and i will say plainly whether i am the right person.
brief → diagnostic → findings → fix → measure
3 questions: what the system does, what it is doing that you cannot explain, and how you would know it was fixed. From that I can tell you whether I am the right person to help.
I read the system, the logs and the data before forming an opinion. Often the reported problem and the actual problem are not the same thing.
The bottlenecks, the cause of each, and what I would change. Sometimes the finding is that the thing you asked for is not worth building, and I will tell you that.
Built by me, or handed to your engineers with enough detail that they can build it. Either way it ships behind a measurement, not behind a demo.
The projection gets checked against production spend and production behaviour, so you can see whether the fix held.
Three passes, run together rather than in sequence. Static analysis of the shipped build, where I deobfuscate what reaches the client and lift the guard functions verbatim, then run them offline under a harness rather than testing a paraphrase of them. Adversarial live sessions against the running product: extraction corpora, multi-turn and crescendo jailbreaks, prompt injection, and how the product handles safety and crisis signals. Functional oracles that establish whether the product does the thing it claims to do at all.
The first pass tells you what the product is. The second tells you where it gives. The third tells you whether the claim holds.
Written authorisation, and the level of access a real user would have. Private alpha credentials are usually enough. I do not need your source repository, though the work is sharper when the engineers who own the code are reachable for questions.
Without written authorisation I do not begin. This is not a formality. Unauthorised testing is indefensible for both of us, whatever the intent behind it.
Scope and timeline are agreed at the brief, before anything starts. A single product in private alpha usually runs in the same range as a production diagnostic, one to two weeks from brief to report. A broad surface, or a build that is heavily obfuscated, takes longer, and I say so at the brief rather than afterwards.
A severity ladder, where every finding carries steps to reproduce, so your engineers can confirm it themselves rather than take my word for it. A remediation priority order, which is deliberately not the same as the severity order, because some severe findings are cheap to close and some minor ones are not.
A framework crosswalk, mapping each finding to the standards that name it: the OWASP Top 10 for LLM Applications, MITRE ATLAS, CWE and the OWASP Web Security Testing Guide, and where governance is in scope, the NIST AI Risk Management Framework, ISO/IEC 42001, India's DPDP Act 2023 and the EU GDPR.
And a founder memo, the short version of what the findings mean for the launch date and for the claims you were planning to make.
I do not test what I have not been authorised in writing to test. I do not publish, distribute or reuse what I find; the engagement sits under a confidentiality and non-distribution covenant, and the report is yours. I do not hand over scanner output and call it an evaluation.
Priced per engagement, after the brief. What moves the number is the size of the surface, how much of the product is reachable without source access, and whether governance frameworks are in scope alongside the security ones. You get the scope and the price in writing before any testing begins.
Not to be difficult. Each of these makes the work impossible to judge afterwards, which is bad for both of us.
Every engagement is different enough that a number on a page would be dishonest.