Discovery & AI · constellation

AI Systems That Discover, Not Just Summarize

Also called: Autonomous scientific discovery

  • Personal interest
  • Philosophical question
  • Working interpretation

I want to find out whether groups of AI agents, programs that can take steps on their own, can do more than summarize what is already known: whether they can suggest new ideas worth testing and find results that other people can check for themselves. The goal is a real new contribution, not just a lot of automated activity.

What would count as a truly new result, and who would be able to check it?

Why it attracts me

Most of what AI tools do today is rearrange knowledge that already exists. They summarize, translate and draft. That is useful, but it is not discovery. What I want to know is whether a group of AI agents, set up carefully, can go further: suggest a new idea, test it, and produce a result that someone else can check without trusting the system or me.

The question matters to me because it sits where several of my interests meet: mathematics, computer search, careful evidence and the wish to make a real contribution (Making a Meaningful Contribution).

The idea

Research has a rough cycle. You pick a precise question, propose an answer, try to break it, check it again independently, and compare it with what is already known. An autonomous system tries to run more of that cycle with less step-by-step human direction. The pieces have their own places on this map: AI workflows that plan and criticize (AI Workflows That Catch Their Own Mistakes), computer search (Computer Search and Rules of Thumb), proofs a machine can check (Proofs a Computer Can Check) and the search for earlier work (Checking Whether Someone Got There First).

The hard part is not producing output. It is producing a result that is both new and true.

What I think (and don't know)

I think the most promising designs keep finding and accepting separate. The program that proposes an answer should not be the only one allowed to approve it. I also think the strongest early results will come where a proposed answer is cheap to check exactly, even if finding it took a huge search.

What I do not know is how far this goes. A system can look impressively busy while producing nothing new. It can also produce something correct that turns out to be already known. I do not yet know where the line falls between AI doing research and AI being a very good research tool, or whether that line matters as much as the result itself.

Where it connects

This idea pulls against Curious Wandering and Lucky Discoveries. Directed research aims at a target, but some of the best connections turn up while wandering without one, and both have value. It also depends on Pairing AI's Reach With Human Judgment, because judging which results matter is still a human job. And it leads to Contributing Without Needing the Credit: the goal is the contribution, not the credit.

An example

My Agentic Maths side quest asks whether coding agents can find publishable mathematics if novelty, proof and attempts to break a claim are treated as separate gates. A candidate passes through steps that include reproducing known results, a hostile review, an independent rebuild, a search for earlier work and, finally, handoff to a human expert. The strongest result so far is a construction whose main theorem is checked by the Lean proof checker. It is marked ready for expert review, not peer reviewed, and the page is explicit that even a top-level result can still be wrong or already known. That honesty is the point. The system's job is to take a claim as far as it honestly can, then stop and hand it to people.

Questions I am still carrying

  • How do I measure progress when most attempts fail quietly?
  • How should credit and authorship work when AI does much of the search?
  • Will the best results come from fully automatic systems or from close teams of people and AI?

What this does not establish

Being interested in AI-led discovery does not show that current AI systems can make new, correct discoveries on their own. Each claimed result still needs independent checking, a search for earlier work and review by human experts.

Questions I'm still exploring

  • What is the smallest result that would count as a real discovery rather than a summary?
  • Which parts of research can be handed to AI, and which must stay with people?
  • How do I avoid mistaking a large volume of output for progress?

Sources and further reading

  • Bernardino Romera-Paredes et al., "Mathematical discoveries from program search with large language models", Nature 625 (2024): 468–475 — Describes FunSearch, which paired a language model with an automatic evaluator and found new mathematical constructions.
  • Alex Davies et al., "Advancing mathematics by guiding human intuition with AI", Nature 600 (2021): 70–74 — Machine learning pointed mathematicians toward patterns that they then proved themselves.

Working interpretation: drafted from my notes and interests for review. It is not a direct quotation, and I may still change it.