PyData Amsterdam 2026

Building the
evaluation flywheel

How human judgment becomes an evaluator you can run on every change.

Show of hands
  1. 01Who has built an eval?
  2. 02Who runs an eval on every change, in CI or
    on production traffic?
  3. 03Who has checked an LLM judge against human labels?
  4. 04Who would let agents merge PRs autonomously?
The evaluation gap

Two correct answers.
Only one you want

“Where is the station?”

ANSWER A

52.3791° N, 4.9003° E.

ANSWER B

Two streets to your left, about five minutes on foot.

Where this started

We came for optimization.
We got stuck on the signal.

Self-improvement runs on the judge's critiques. A wrong judge, wrong improvements.

agent judge signal prompt update

Evaluation moved from correctness to alignment

Known answer COMPARE WITH GROUND TRUTH Human judgement COMPARE WITH AN EXPERT LLM judge COMPARE WITH EXPERT LABELS The evaluator now needs its own evaluation.
Agent Judge Expert
Tldr;

You cannot use evals to improve an agent
before the eval is aligned.

Bauke Brenninkmeijer
Bauke Brenninkmeijer
Applied AI Researcher · orq.ai
Red teaming · Evaluation · Agent simulation
Lead @ Agentic AI Foundation Amsterdam
Context

What is orq.ai?

Generative AI collaboration platform. One control plane to build, ship, and optimize AI products.

Build
Agents
Deploy agents with tools, memory, and knowledge bases.
think Letta · LangGraph · CrewAI
Ship
Router
One API for model routing, failovers, caching, and budget.
think LiteLLM · OpenRouter
Optimize
Observability
Traces, usage, evaluation, and annotation.
think Langfuse · LangSmith · Arize
The case
sphere.com

A B2B wholesaler of physical home appliances. The board wants to understand the quality of growth.

1Discounts and realized revenue
2Refunds and cancellations
3Regional and category mix
4Answers shaped for a business decision

Does the answer support the decision?

50distinct business situations
How do you create an eval?

Pass or fail

Clear enough to act on
No false precision from a 1 to 5 scale
Agreement and false passes become measurable
No space to stand
PASS
FAIL

The critique carries the nuance

The grey zone

Every evaluation has one.
All plausible. All different.

Question 02 · Trust the judge

Every failure has two suspects

The eval says FAILone case, one criterion
↗
System loop
The answer really was wrong
→
Fix
Update the agent
↘
Evaluator loop
The judge was wrong
→
Fix
Update the evaluator

Aligning the judge to humans is how you tell the two apart.

The quality-control process
has not changed

THEN An expert annotates the reference cases A non-expert annotates student or Mechanical Turk worker Compare the annotations expert / delegate agreement NOW An expert annotates the reference cases An LLM judge annotates the same cases Compare the annotations expert / judge agreement

In both cases, agreement with the expert decides whether to trust the delegate.

But we are lazy

We don't want to annotate all 50 cases.
Let the LLM judges find the ambiguous cases, then spend human time only there.

The most valuable thing ishuman attention.

Two ways of disagreement

Three judges · three repetitions

gpt-5.6-lunaqwen3.8-27bgemini-3.5-flash-lite
3 repetitions →
… × 50 cases

Judges disagree with each other, and one judge disagrees with itself.

Both signals exist for every case in the pool.

REVIEW FIRST 0 4 flagged CONTROL 0 4 sampled
judges disagreeone judge unstable

The judge
sorts the queue

Flagged cases go first. A random sample comes along to catch the cases the judges agreed on and got wrong.

Disagreement gives us
a question

CASES THE PANEL SPLIT ON ? ONE BOUNDARY QUESTION

The flagged cases produced no labels. They produced the question the criterion never answered.

One answer exposed another ambiguity

The human answered one boundary question: claims visible in the evidence must be valid. Then the rule went into the evaluator.

Before the rule

4cases the panel split on

After the rule

8cases the panel split on

We aligned the principle, but not what counts as unsupported.

The grey-zone loop

Disagreement shows where the evaluator still needs a human decision.

HELD-OUT TEST SET STAYS SEALED
01

Run the jury

02

Surface instability

03

Collaborator reads the reasons

04

Ask one boundary question

05

Update the evaluator

RERUN THE SAME FROZEN DEVELOPMENT CASES
PASS

Clarify the metric first

Valid evidence supports the scoped answer.

FAIL

Best month net

Visible evidence contradicts the stated definition.

PASS

Earlier context still counts

The response established the context earlier in the conversation.

Human labels reveal the judge limits

Orq experiment grid: three evaluator prompt versions scored by three evaluators over the frozen development cases

Jury signals helped us find the unresolved questions. Once the human decisions became labels, we could see which judges reproduced them.

What changes with agents?

The answer is only the endpoint

Agents require evaluating behavior, rather than final answers.

NR. OF TOKENS →
user turn assistant tool call tool result

Fifty Sphere.com cases. Each bar is one run, each block one message, sized by how much context it added.

EVALS ON THE BEHAVIOR
Tool-call efficiency
Error recovery
Instruction adherence

The Eval Lifecycle

DISCOVERY

Build with humans

TARGET ROUNDS OF HUMAN REVIEW →

Hard cases and edge cases. A low pass rate is the eval working; one that never fails is too easy.

REGRESSION

Guard every commit

TARGET CAUGHT EVERY COMMIT →

Known behavior and core paths. Everything passes, or something broke and the commit stops.

Scaling evaluation

When and where to run evaluations

RELEASE Offline The curated 50, when you ask. Online Sampled production traces, as traffic arrives. Continuous Both, on every change and then on a schedule.

The same criterion runs in all three, and it holds only while production stays inside the slice humans validated.

2026 Software Factory

Most software will ship
without a human reading it.

SHIPPED
READ

Evals in the software factory

Evals in the software factory

Same throughput.
Different factory.

Factory A

Most pull requests merge untouched.

Eval pass rateCommits →
Holdingthe criterion still passes
Factory B

Most pull requests merge untouched.

Eval pass rateCommits →
Rottingthe same criterion started failing

The factory reports how much moved. Only an eval tells you which of these you are running.

Conclusion

The evaluation flywheel

Judge disagreement directs attention.

Human judgement sets the boundary.

The resulting eval guards every change.

What to keep

Run this on your own agent

$ npx skills add orq-ai/assistant-plugins
$ pip install evaluatorq
QR code linking to the talk repository
Questions

Q&A

QR code linking to Bauke Brenninkmeijer on LinkedIn