Depth over demos
A demo proves a model can. Depth proves a system will — at your latency, your cost, your load.
Anyone can call an API. We work the ten layers underneath it — where latency, cost and ownership are actually decided.
The lower layers decide whether a system is fast, affordable, and yours.
Most AI work stops at the surface. A prompt, an API key, a demo that impresses in a meeting and buckles in production. We go further down — because the layers nobody shows you are the ones that decide whether a system is fast, affordable, and actually yours.
A demo proves a model can. Depth proves a system will — at your latency, your cost, your load.
The layers nobody shows you decide whether a system is fast, affordable, and actually yours.
Every improvement we report has an evaluation behind it that you can run yourself.
We are consultants. The engagement ends. The system should not notice.
Ten layers between a question and an answer. Most consultancies rent you the top one. Open any layer to see what we actually do down there.
The part everyone sees, and the only part most vendors touch. Interfaces that make a probabilistic system feel dependable — streaming, citations, graceful failure, and the affordances that let a person stay in control.
Engagements are shaped around the layer your problem actually lives in — not around a package we happen to sell.
Where to apply AI, what to build, and what to buy.
End-to-end design and construction of production AI systems.
Fine-tuning, distillation, and small language models.
Serving stacks engineered for throughput and cost.
AI that survives procurement, security review, and scale.
When the answer isn't in a paper yet.
We start with the constraint, not the technology. What must be true for this to be worth doing — and what happens if it isn't.
The system drawn before it is built. Layer by layer, with the trade-offs written down where they can be argued with.
The riskiest assumption, built first and measured against a real evaluation set. Fast enough to be wrong cheaply.
Down the stack. Quality, latency and cost pushed until the curve flattens and further effort stops paying.
Shipped with the unglamorous parts intact: monitoring, rollback, on-call notes and a team that can operate it without us.
Models move, data drifts, costs change. A cadence of evaluation and tuning that keeps the system from quietly decaying.
Three engagements, anonymised. Detailed case studies are available on request.
Layers of the stack, worked end to end
From first call to a measurable prototype
Typical inference cost reduction
Engagements handed over with a runbook
A first conversation costs nothing and usually saves a quarter. Bring the constraint you're stuck on — we'll tell you which layer it lives in.
Typical reply within one working day.