Case study · Automotive

Foundation Model Evaluator: winning a global challenge on choosing between models

A global challenge sponsored by Stellantis asked for a way to assess and evaluate the performance of foundation models and improve how they are selected for practical AI use cases. Jaden Data entered with Patronus AI and won, with a proof of concept on retrieval-augmented generation evaluation using Ragas.

Client engagement · via Jaden Data Proof of concept

Mandate

Delivered the winning entry through Jaden Data together with Patronus AI: a concept and a proof of concept for assessing foundation models and choosing between them, built as retrieval-augmented generation evaluation with Ragas.

Key decisions

  1. Enter with Patronus AI rather than alone, pairing platform delivery with an evaluation specialist.
  2. Build the proof of concept on Ragas, so retrieval-augmented answers are scored against reference data instead of judged by demo.
  3. Answer both tracks the brief named — assessment before deployment and evaluation after it — with one measurement approach rather than two disconnected tools.

Outcome — qualitative

Winner of the Foundation Model Evaluator challenge, entered with Patronus AI
2 tracks covered: selection before deployment, performance evaluation after it Measured
Global challenge scope, sponsored by Stellantis
Co-creation funded co-creation, strategic partnerships and long-term integration as the award
Stakeholders
  • Stellantis as the challenge sponsor
  • ekipa as the challenge platform
  • Patronus AI as co-entrant
  • Jaden Data engineering
Constraints
  • A detailed concept and, if possible, a first prototype by the initial deadline
  • A functional proof of concept in the six months after that deadline
  • Model selection had to hold for practical AI use cases, not for one benchmark
  • Global scope
Reuse
  • Ragas-based evaluation harness for retrieval-augmented generation
  • Model comparison method spanning pre- and post-deployment
  • Partnering pattern: platform delivery plus an evaluation specialist

Context

Stellantis sponsored an open challenge, run on the ekipa platform and titled “Foundation Model Evaluator”. The brief asked entrants to “Develop an innovative concept or prototype to assess & evaluate the performance of foundation models, optimizing their selection process for practical AI use cases”. Scope was global, and the challenge is now marked completed.

Business problem

An enterprise choosing a foundation model is choosing between options that all demo well. The brief split the problem into two tracks: assessing performance and selecting a model before deployment, and evaluating performance after deployment. Both need a measurement that survives contact with a real use case, not a leaderboard position.

My mandate

I delivered the entry through Jaden Data, with Patronus AI as co-entrant. The challenge page lists the winner as “JadenDate & Patronus AI” — a misspelling of Jaden Data.

Decisions

  • Enter as a pair. Jaden Data brought the platform and the delivery; Patronus AI brought evaluation as its specialism. A single-vendor entry would have had to argue both halves alone.
  • Build the proof of concept on Ragas, which scores retrieval-augmented answers against reference data, so retrieval quality and answer quality are separable rather than collapsed into one impression.
  • Treat pre-deployment selection and post-deployment monitoring as the same measurement applied at two moments, so a model chosen on evidence keeps being checked against it.

Delivery

The brief asked for “a detailed concept or if possible, a first prototype”, with “a functional POC in the next 6-month after the initial deadline”. What was built is a proof of concept on RAG evaluation using Ragas.

Outcome

The entry won. The award listed on the challenge page is “Funded co-creation, strategic partnerships & long-term integration”.

A global Foundation Model Evaluator challenge sponsored by Stellantis, won with Patronus AI on a proof of concept for evaluating retrieval-augmented generation with Ragas.

Reuse

The evaluation harness and the comparison method carry into any deployment where a model has to be picked and then kept honest — the same need that produced Testbed on the platform.

Evidence

The challenge title, the two tracks, the scope, the award and the winner are published on the challenge brief, linked in the facts panel. The Ragas proof of concept is a project record. No dates, contract value or post-challenge outcomes are published, so none are claimed here.

Want the same thing done in your environment?

This case is one of several. If the shape looks like your problem, the fastest route is to send me the constraints you cannot move.

Remote-first, on-site when it matters; NDA on request