PyData Amsterdam 2026
Building the
evaluation flywheel
How human judgment becomes an evaluator you can run on every change.
Bauke BrenninkmeijerOrq.aiSeptember 2026
Show of hands
- 01Who has built an eval?
- 02Who runs an eval on every change, in CI or
on production traffic?
- 03Who has checked an LLM judge against human labels?
- 04Who would let agents merge PRs autonomously?
The evaluation gap
Two correct answers.
Only one you want
“Where is the station?”
ANSWER A
52.3791° N, 4.9003° E.
ANSWER B
Two streets to your left, about five minutes on foot.
Where this started
We came for optimization.
We got stuck on the signal.
Self-improvement runs on the judge's critiques. A wrong judge, wrong improvements.
Evaluation moved from correctness to alignment
Tldr;
You cannot use evals to improve an agent
before the eval is aligned.
Bauke Brenninkmeijer
Applied AI Researcher · orq.ai
Red teaming · Evaluation · Agent simulation
Lead @ Agentic AI Foundation Amsterdam
Context
What is orq.ai?
Generative AI collaboration platform. One control plane to build, ship, and optimize AI products.
Build
Agents
Deploy agents with tools, memory, and knowledge bases.
think Letta · LangGraph · CrewAI
Ship
Router
One API for model routing, failovers, caching, and budget.
think LiteLLM · OpenRouter
Optimize
Observability
Traces, usage, evaluation, and annotation.
think Langfuse · LangSmith · Arize
The case
sphere.com
A B2B wholesaler of physical home appliances. The board wants to understand the quality of growth.
1Discounts and realized revenue
2Refunds and cancellations
3Regional and category mix
4Answers shaped for a business decision
Does the answer support the decision?
50distinct business situations
How do you create an eval?
Pass or fail
Clear enough to act on
No false precision from a 1 to 5 scale
Agreement and false passes become measurable
No space to stand
PASS
FAIL
The critique carries the nuance
The grey zone
Every evaluation has one.
All plausible. All different.
Question 02 · Trust the judge
Every failure has two suspects
The eval says FAILone case, one criterion
↗
System loop
The answer really was wrong
→
↘
Evaluator loop
The judge was wrong
→
Aligning the judge to humans is how you tell the two apart.
The quality-control process
has not changed
In both cases, agreement with the expert decides whether to trust the delegate.
But we are lazy
We don't want to annotate all 50 cases.
Let the LLM judges find the ambiguous cases, then spend human time only there.
The most valuable thing ishuman attention.
Two ways of disagreement
Three judges · three repetitions
gpt-5.6-lunaqwen3.8-27bgemini-3.5-flash-lite
3 repetitions →
… × 50 cases
Judges disagree with each other, and one judge disagrees with itself.
Both signals exist for every case in the pool.
judges disagreeone judge unstable
The judge
sorts the queue
Flagged cases go first. A random sample comes along to catch the cases the judges agreed on and got wrong.
Disagreement gives us
a question
The flagged cases produced no labels. They produced the question the criterion never answered.
One answer exposed another ambiguity
The human answered one boundary question: claims visible in the evidence must be valid. Then the rule went into the evaluator.
Before the rule
4cases the panel split on
After the rule
8cases the panel split on
We aligned the principle, but not what counts as unsupported.
The grey-zone loop
Disagreement shows where the evaluator still needs a human decision.
HELD-OUT TEST SET STAYS SEALED
01
Run the jury
02
Surface instability
03
Collaborator reads the reasons
04
Ask one boundary question
05
Update the evaluator
RERUN THE SAME FROZEN DEVELOPMENT CASES
PASSClarify the metric first
Valid evidence supports the scoped answer.
FAILBest month net
Visible evidence contradicts the stated definition.
PASSEarlier context still counts
The response established the context earlier in the conversation.
Human labels reveal the judge limits
Jury signals helped us find the unresolved questions. Once the human decisions became labels, we could see which judges reproduced them.
What changes with agents?
The answer is only the endpoint
Agents require evaluating behavior, rather than final answers.
NR. OF TOKENS →
user turn
assistant
tool call
tool result
Fifty Sphere.com cases. Each bar is one run, each block one message, sized by how much context it added.
EVALS ON THE BEHAVIOR
Tool-call efficiency
Error recovery
Instruction adherence
The Eval Lifecycle
DISCOVERY
Build with humans
Hard cases and edge cases. A low pass rate is the eval working; one that never fails is too easy.
REGRESSION
Guard every commit
Known behavior and core paths. Everything passes, or something broke and the commit stops.
Scaling evaluation
When and where to run evaluations
The same criterion runs in all three, and it holds only while production stays inside the slice humans validated.
Most software will ship
without a human reading it.
Evals in the software factory
Evals in the software factory
Same throughput.
Different factory.
Factory A
Most pull requests merge untouched.
Eval pass rateCommits →
Holdingthe criterion still passes
Factory B
Most pull requests merge untouched.
Eval pass rateCommits →
Rottingthe same criterion started failing
Conclusion
The evaluation flywheel
Judge disagreement directs attention.
Human judgement sets the boundary.
The resulting eval guards every change.
What to keep
Run this on your own agent
$ npx skills add orq-ai/assistant-plugins
$ pip install evaluatorq
Questions
Q&A
Bauke BrenninkmeijerOrq.ai