Case study · Development finance

A repeatable way to find, build and evaluate AI workflows with domain experts

Seven proofs of concept across ministry reporting, partner-report checks, contract analysis, evaluation reports, technical knowledge retrieval and humanitarian self-service, each built as a multi-step process and validated by subject-matter experts on real documents.

Client engagement · via Jaden Data / entAIngine Proof of concept

Mandate

Designed the architecture and the multi-step AI processes for a portfolio of seven proofs of concept, ran the expert-review cadence, and was accountable for each one meeting the customer's acceptance criteria.

Key decisions

  1. One standardised specification and one evaluation template for every proof of concept, so results were comparable and expert reviews stayed weekly.
  2. Fan-out / fan-in process design: parallel extraction into structured JSON, reconciliation, then generation with source references in every output.
  3. Data preparation over model selection: PDF and Excel pre-processing as the main lever against numeric errors, and simpler output schemas after a wide single-pass extraction hallucinated.

Outcome — customer estimates

~1 day saved per ministry monitoring report (out of ~90 per year)
20–30% less drafting effort on an ex-post evaluation chapter
~15 min saved per partner-report indicator pre-check (~180 per year)
5 h → 1 h per contract deliverables review at the targeted accuracy Projected
Stakeholders
  • Portfolio managers
  • Technical experts
  • Evaluators
  • An NGO's programme and legal-assistance teams
  • IT and data owners
Constraints
  • Strict German- or English-only structured outputs
  • Documents of 200–500 pages
  • Expert time available only in weekly reviews
  • Per-interaction cost versus static tools had to be justified
Reuse
  • Standardised PoC specification and evaluation templates
  • Multi-node process patterns (parallel extraction → JSON hand-off → reconciliation → generation)
  • RAG with reranking and page-level citations
  • Prompt-injected decision-tree chatbot pattern

Context

The development-finance teams of a German development bank, and a humanitarian NGO it funds, run on documents: annual ministry reports, partner reports, contracts, evaluation reports, technical design archives and self-help guides for refugees.

Business problem

Each workflow consumed expert hours in summarising, cross-checking and drafting. The question was not whether a language model could write a paragraph, but where generative AI removes manual effort reliably enough that experts would adopt it.

My mandate

I designed the solution architecture and the multi-step AI processes for every proof of concept, wrote the prompts for extraction, reconciliation, generation, retrieval and the chatbot, and was accountable for each result meeting the acceptance criteria the experts set.

Decisions

  • Every proof of concept used one specification template (background, problem, objective, scope, methodology, success metrics, dependencies) and one evaluation template (time savings, rework reduction, result quality, applicability, improvement potential).
  • Processes were designed as fan-out / fan-in: parallel extraction nodes producing JSON, a reconciliation step, then generation with source references preserved in every output. The ministry-report PESTLE draft alone was an eight-node process with six parallel research analyses.
  • Knowledge retrieval over 300-page technical documents used vector retrieval, a reranking step with relevance justification, and citations by document, file and page, exposed through a chat interface with follow-up questions.
  • Contract deliverables extraction moved from a single wide table schema (which hallucinated) to an n-step design: identify first, then extract attributes, with human-in-the-loop validation.
  • The refugee self-help guide became a stateful, multilingual assistant (English, Arabic, Ukrainian, Bengali, Russian) by injecting the exported decision tree into the prompt and tracking the step pointer in the conversation state.

Delivery

Rapid prototyping on the platform, agile iterations, weekly reviews with the domain experts, every proof of concept validated on real documents.

Delivery shape
  1. Spec template — One specification (scope, metrics, dependencies)
  2. 7 PoCs — Run in parallel (same terms, same evaluation template)
  3. Weekly — Expert review (domain experts score each output)
  4. Decision — Portfolio call (what to take further, what to stop)

Outcome

Experts estimated about a day saved per ministry report, a fifth to a third less drafting effort on an evaluation chapter, and a quarter of an hour per partner-report pre-check; contract review was projected to drop from five hours to one at the targeted accuracy. Retrieval over the technical archive enabled research that had been impractical, and the experts asked for a production rollout to a wider group.

Retrieval over the technical archive enabled research that had been impractical, and the experts asked for a production rollout to a wider group.

Reuse

The specification and evaluation templates, the process patterns, the cited-retrieval design and the decision-tree assistant pattern were reused across the portfolio and later deployments.

Evidence

Time savings are the experts’ own estimates from the evaluation template; the contract-review figure is a projection at the targeted recall and precision. Outputs carried source references so every claim could be traced.

Want the same thing done in your environment?

This case is one of several. If the shape looks like your problem, the fastest route is to send me the constraints you cannot move.

Remote-first, on-site when it matters; NDA on request