$capabilities/ai

AI that survives contact with production

Most AI programmes stall between a demo that impressed someone and a system anyone trusts. We build the measurement first, then the feature — and hand you the eval set that proves it works.

$what we do here

Three places AI work actually goes wrong

Find the real candidates

Most AI wish-lists are a mix of two or three genuine opportunities and a dozen expensive distractions. We separate them before anyone writes a prompt.

Task inventory scored on volume, judgement and riskA baseline of what the work costs you todayCandidates ranked by value, not noveltyThe honest list of what should stay manualA first slice small enough to finish

Build the eval set first

You cannot improve what you have not agreed how to measure. The eval set is written with the people who do the job, and it belongs to you.

Real questions written by your own practitionersAn expected answer or accept rule for every oneRetrieval scored separately from generationThe whole set run in CI on every changeOwned by your team in your own repository

Ship it into the workflow

An AI feature nobody uses is a demo. We put it where the work already happens, with the refusal behaviour and review path designed in.

Deployed inside the tool your team already opensCitations to the source passage, every answerRefusal when the corpus does not support an answerCost and latency budgets per featureA feedback loop that files low-confidence casesBook a call
$evaluation

If a change cannot be shown to help, it does not ship

Every prompt, model or index change runs the full eval set before release. That rule is unglamorous, it slows the first fortnight, and it is the only reason the following three months are fast.

Eval-gatedEvery prompt, model and index change
CitedAnswers point at the passage, not the document
YoursThe eval set lives in your repository
4-10 wksAudit to a first supervised feature

What we build it on

A short, deliberate list — each in production on a system we maintain. See the full partner stack.

Claude (Anthropic)AI platform partner
OpenAIPlatform — production use
LangSmithEvals and tracing
PineconeVector search
SupabaseData platform partner
PostgreSQLCore data engine
Google CloudCertified — data & AI
CloudflareEdge, DNS and WAF

Clients we work with

Bayt Travel
Travel Secrets
DAX — Doha Express
Orangetheory Fitness
Nova Fertility
Octillion Global
Ornamint
Shubhra Krishan
IKISAKI
$case studies

AI case studies

Clinical research console
HealthcareNova FertilityRetrieval over a decade of clinical literature, with citations to the passage and an honest refusal when the corpus falls short.
Document pipeline
LogisticsDAX — Doha ExpressFreight documents classified and reconciled automatically, with every run traced and ambiguity queued for a person.

What we will not do

Most AI briefs we receive are three projects wearing one name. These are the ones we hand back.

$./ai-readiness-review
01Ship an AI feature with no agreed measure of correct.02Put a model in front of a corpus nobody has audited.03Let a confident summary stand in for a cited answer.04Sell a rebuild when a scoped pilot would answer the question.05Keep the eval set as our leverage instead of your asset.