Discovery & AI
AI Workflows That Catch Their Own Mistakes
Also called: Agentic reasoning systems
- Personal interest
- Philosophical question
- Working interpretation
AI tools can be arranged into a research process, with some parts proposing ideas and other parts criticizing and checking them. I am interested in AI helpers that change their plans, keep track of what they are unsure about and point out mistakes, instead of only sounding confident.
Which step needs a skeptical checker rather than one more source of ideas?
Why it attracts me
A fluent answer is not the same as a correct one. The AI systems that interest me most revise their plans, keep track of what they are unsure about and show their mistakes, rather than covering them with confident prose.
The idea
An agent works over several steps, using tools and reacting to results. Several agents can be arranged into a workflow: one proposes, another criticizes, a third checks. The key design choice is where to put a skeptical checker instead of one more generator of ideas. Each extra role adds cost and new ways to fail, so it should earn its place.
An example
In my write-up From Vibe Coding to Engineering, I argue that the request I type to a coding agent is not the product. The product is the loop around the model: its context, its tools, how success is checked and when a person steps in. The piece draws a sharp line: written guidance tells an agent what to do, a test is evidence, and a check that blocks bad work is enforcement. Splitting work between agents makes sense when they have truly separate jobs, such as one that builds and another that independently validates.
Where it connects
This links to Learning What to Do in Each Situation, where a learning agent improves its rules for choosing from rewards. The comparison helps, but human purpose cannot be reduced to a reward score. These workflows power AI Systems That Discover, Not Just Summarize, and they lean on Proofs a Computer Can Check for checks that need no one's judgment.
Questions I'm still exploring
- When does adding another AI role help, and when is it just complexity for show?
- How can an AI system report how unsure it is in a way people can trust?
- Which checks should be done by ordinary code rather than by another AI?
Sources and further reading
- Erik Schluntz and Barry Zhang, "Building effective agents", Anthropic (December 2024) — Recommends starting with the simplest workflow that works, and describes a pattern where one model call generates and another evaluates.
Working interpretation: drafted from my notes and interests for review. It is not a direct quotation, and I may still change it.