Applied AI
Retrieval pipelines, agent systems, structured extraction, fine-tuning — scoped against an eval set, not a vibe.
Services
Three practices, one standard: if we build it, we can run it — and we can hand you the runbook that proves it.
What we take on
Retrieval pipelines, agent systems, structured extraction, fine-tuning — scoped against an eval set, not a vibe.
The substrate models feed on: pipelines, vector stores, feature systems, governance that survives an audit.
Landing zones, Kubernetes, IaC, networking and the cost model — architected before the first workload lands.
Model CI, registries, canary deploys, GPU scheduling, token-level observability. Where AI meets pager.
We keep running what we built — SLOs, on-call, upgrades, cost reviews — with a named engineer, not a queue.
Practice · 01
We build AI systems against a written definition of "works": an eval set agreed before the first line of code, run on every change, reported without varnish.
Chunking, embedding, reranking, citation. Grounded answers over your corpus with the failure modes measured, not assumed.
Tool-using agents with permission boundaries, spend caps and audit trails — autonomy budgeted like any other resource.
The part most teams skip. Golden sets, regression gates in CI, drift monitoring in production.
When prompting stops being enough: data curation, training runs, and inference serving with real latency budgets.
Practice · 02
Everything below the model: accounts, networks, clusters, pipelines, dashboards. Defined in code, deployed by machine, explained in writing.
Multi-account foundations with identity, networking and guardrails designed before workload one.
Clusters sized for reality, GPU pools scheduled for utilisation, autoscaling tuned against your actual traffic.
Terraform end to end, progressive delivery, rollback measured in seconds. The console is for reading, not clicking.
Traces that cross the model boundary, SLOs with error budgets, and cost per request on the same pane of glass.
Practice · 03
The offer that keeps the other two honest: we stay on the pager for what we ship. Monthly, with a named engineer and a standing review.
Availability and latency targets in the contract, error budgets reported monthly, postmortems you can read.
Kubernetes versions, model migrations, dependency CVEs — handled on a cadence, not in a crisis.
Quarterly unit-economics reviews. When your bill drops because of our work, you hear it from us first.
How an engagement runs
A paid discovery sprint. Output: a written architecture, an eval plan, a cost model, and a fixed quote for the slice.
A working path through the whole system on your infrastructure — small, but real: real data, real deploys, real metrics.
Load, chaos, security review, runbooks, cost tuning. The slice becomes a platform.
Managed operations with SLOs, or a documented handover with your team trained on the pager. Your call.