Production AI systems · Enterprise AI architecture

I've shipped the AI system you're trying to build.

Multi-tenant retrieval and multi-step AI processes on event-driven AWS — more than 1,000 concurrent connections, thousands of requests a second and 99.9% uptime for 50+ organisations. Document-heavy workflows: extraction, reconciliation and drafting across contracts, reports and technical archives. Customer-facing assistants: a member-facing pension assistant running on-premises, and a five-language self-help assistant for a humanitarian programme. I design the architecture, write the processes and the evaluations, and lead the engineers who own it afterwards.

Fractional CTO & consulting partner · open to the right permanent CTO or Head of AI role

Reference system 99.9% uptime
Client API Gateway Lambda Fargate RAG Service Vector DB Graph DB OpenAI · Bedrock · Azure · Mistral
~3,200 req/sec illustrative throughput

Systems delivered for

As Co-founder & CTO of Jaden Data, and earlier as a consultant

  • Industrial technology
  • Venture investing
  • Development bank
  • Pharma
  • Retail banking
  • Pension fund
  • Central bank
1,000+
concurrent connections on the platform
thousands
requests a second, sustained
99.9%
uptime
50+
organisations, many of them regulated
ISO 27001 + SOC 2
certified platform, delivered in about three months
01

I design the architecture.

Event-driven services on AWS built to survive real load: more than 1,000 concurrent connections and thousands of requests a second at 99.9% uptime. The diagram on this page is the shape of a real production system — request in, retrieval and reasoning across vector and graph stores, model calls fanned out across OpenAI, Bedrock, Azure and Mistral, chosen per client on quality, latency, cost and data sovereignty. A graph store is not decoration: one pharma client's research platform is vector- and graph-based.

02

I build the processes and the evaluations.

Multi-step processes rather than one prompt: parallel extraction into structured JSON, a reconciliation step, then generation with source references preserved in every output. Retrieval with reranking and page-level citations. Deterministic test cases over the pipeline before anything model-graded. I took my own company through ISO 27001 and SOC 2 Type 2 in about three months, because in regulated markets compliance is part of the architecture, not an afterthought.

03

I lead the engineers who own it.

I built a company from zero to €1M+ revenue and engineering from 0 to 10. At idealo I was a key contributor to a six-country login rollout that created 1.4M accounts in three months, 700% above forecast. I hand systems over to teams that keep shipping after I have moved on.

Evaluation depth

How I know the system is good enough to ship

The difference between a demo and a production workflow is that somebody can say, with numbers, what it gets right and what a person still has to check.

  1. Eval sets built from real work

    Test cases come from the customer's own documents and the questions their experts actually ask, not from synthetic prompts. Across a seven-part proof-of-concept portfolio, every workflow was validated by subject-matter experts on real documents against one evaluation template: time saved, rework reduced, result quality, practical applicability, improvement potential.

  2. Deterministic checks before model-graded ones

    Testbed runs deterministic cases over AI processes and retrieval pipelines first — structure, required fields, output language, refusal behaviour — because those failures are cheap to catch and unambiguous. Model-graded scoring comes after, for the judgements a rule cannot express.

  3. Retrieval measured separately from generation

    Whether the right passage was retrieved is a different question from whether the answer reads well. Over technical archives of up to 300 pages per document, retrieval used top-k vector search, a reranking step with a stated relevance justification, and citations down to document, file and page — so a wrong answer can be traced to retrieval or to generation.

  4. An error taxonomy, not an error count

    Failures get classified: missing extraction, wrong attribute, fabricated value, wrong output language, unsupported claim. On contract deliverables, that taxonomy showed a single wide table schema hallucinating, so extraction moved to an n-step design — identify first, then extract attributes.

  5. Human-review thresholds written down

    Every workflow states what a person still checks and when. Partner-report pre-checks carry a three-star consistency rating per indicator with comments explaining each deviation; contract extraction keeps human-in-the-loop validation at the accuracy the customer accepted.

  6. Launch criteria agreed before the build

    A workflow ships against numbers set in the specification, not against a demo reaction: on contract deliverables the targets were at least 70% recall, over 75% precision and more than 50% time reduction, with the projected review time falling from five hours to one at the agreed accuracy.

  7. Quality, latency and cost as one trade-off

    Model choice follows quality, latency, cost and data sovereignty per client, on a provider-independent layer. For a multilingual self-help assistant, the per-interaction cost against a static decision tree was part of the evaluation rather than a surprise after rollout.

Send a project brief.

The workflow, the documents or the assistant, the constraints you cannot move, and who owns it afterwards. I will tell you where it breaks before you build it.

Fractional CTO & consulting partner · open to the right permanent CTO or Head of AI role