The Jetty Blog: Ground Truth
Subscribe
Sign in
Home
Archive
Latest
Top
Discussions
Looping into a Game About AI Evals
Agents play-testing a game about AI Evals
Jul 21
•
Jonathan Lebensold
1
June 2026
Tests are Boolean. Evals are a Gradient
Why a green suite is necessary but never sufficient
Jun 22
•
Jonathan Lebensold
1
1
Verifiable Agents Engender Trust in Systems
Turning confident guesses into measurable claims
Jun 13
•
Jonathan Lebensold
1
Valley-Dodging: The Half of Agent Reliability You're Not Optimizing
Why your best runbook lines were never about the score
Jun 8
•
Jonathan Lebensold
and
Ade Oshineye
1
1
May 2026
A Pelican Learns to Ride
What happens when every iteration of a benchmark is on disk
May 22
•
Jonathan Lebensold
1
Lights-Out Manufacturing Had a Brake Pedal
The lights are going off in software the same way they went off in manufacturing
May 16
•
Jonathan Lebensold
5
1
Research Closes the Loop. Production Keeps Us In It.
Why we kept the outer loop open by design
May 8
•
Jonathan Lebensold
1
April 2026
Patterns Were the Map in the Search for Beauty
Christopher Alexander told the patterns community they missed the point. Thirty years later, agents finally let us listen.
Apr 28
•
Jonathan Lebensold
14
6
2
My Backend is 442 Lines of Markdown
We shipped a web app whose entire backend is a structured document
Apr 21
•
Jonathan Lebensold
2
1
The Jagged Frontier Is an Evaluation Problem
Why your AI system breaks in ways your evals won't catch
Apr 13
•
Jonathan Lebensold
1
1
March 2026
Visual workflows are procedural programming in a costume
Why outcome specs beat node graphs in production
Mar 31
•
Jonathan Lebensold
2
1
Runbooks: what agents need to hill-climb
The Missing Layer Between “Call This API” and “Accomplish This Outcome”
Mar 27
•
Jonathan Lebensold
9
1
3
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts