AI that survives contact with production
Most AI programmes stall between a demo that impressed someone and a system anyone trusts. We build the measurement first, then the feature — and hand you the eval set that proves it works.
Three places AI work actually goes wrong
Find the real candidates
Most AI wish-lists are a mix of two or three genuine opportunities and a dozen expensive distractions. We separate them before anyone writes a prompt.
→Task inventory scored on volume, judgement and risk→A baseline of what the work costs you today→Candidates ranked by value, not novelty→The honest list of what should stay manual→A first slice small enough to finishBuild the eval set first
You cannot improve what you have not agreed how to measure. The eval set is written with the people who do the job, and it belongs to you.
→Real questions written by your own practitioners→An expected answer or accept rule for every one→Retrieval scored separately from generation→The whole set run in CI on every change→Owned by your team in your own repositoryShip it into the workflow
An AI feature nobody uses is a demo. We put it where the work already happens, with the refusal behaviour and review path designed in.
→Deployed inside the tool your team already opens→Citations to the source passage, every answer→Refusal when the corpus does not support an answer→Cost and latency budgets per feature→A feedback loop that files low-confidence casesBook a call →If a change cannot be shown to help, it does not ship
Every prompt, model or index change runs the full eval set before release. That rule is unglamorous, it slows the first fortnight, and it is the only reason the following three months are fast.
What we build it on
A short, deliberate list — each in production on a system we maintain. See the full partner stack.
Clients we work with








AI case studies
What we will not do
Most AI briefs we receive are three projects wearing one name. These are the ones we hand back.
$./ai-readiness-review→