Discovery & AI

Results That Hold Up When Rerun

Also called: Reproducibility

  • Established idea
  • Observed in experiments
  • Working interpretation

A useful finding should still hold when someone runs it again, writes the code a different way or changes the assumptions a little. Keeping exact versions of the data, fixing the starting point of any randomness and writing down every step are what turn an exciting result into evidence.

Could someone recreate my result without having to trust me?

Why it attracts me

The question I keep asking is whether someone could recreate my result without trusting me. If the answer is no, I do not really have evidence yet. I have a story about what happened on my computer.

The idea

Reproducibility has layers. The first is rerunning the same steps on the same data. The next is rebuilding the method separately, in different code. The strongest is changing the assumptions and seeing whether the finding survives. Each layer rules out a different kind of accident: a lucky random draw, a bug in one program, or a result that only holds under one narrow setup.

An example

In my Agentic Solvers side quest, the certificates behind each result, checkable records of each answer, can be regenerated exactly. A separately written checker program confirms each one, and the number of trees checked is compared with the established count of distinct trees of each size. In my Kryptos K4 side quest, an independent rebuild of a public decoding method is what revealed that it could fit almost any message.

Where it connects

Making results easy to rerun is one way to invite others to try to break them (Trying Hard to Break My Ideas). It is also what lets collaborators (Doing Research With Critics and Experts) and AI-led discovery (AI Systems That Discover, Not Just Summarize) build on a result without trusting its author. Personal growth (Growth Through Feedback) benefits from honest records too, though a life should not be run like a lab.

What this does not establish

Getting the same result again shows that the procedure is consistent. It does not show that the result is correct, because a shared flaw in the data or the assumptions would be reproduced too.

Questions I'm still exploring

  • How different does a second method need to be to count as a real independent check?
  • Which assumptions should I deliberately change to see whether a result survives?
  • How do I make a rerun easy enough that someone will actually do it?

Sources and further reading

Working interpretation: drafted from my notes and interests for review. It is not a direct quotation, and I may still change it.