Choices, Games & Branching Paths
Try Something New or Use What Works
Also called: Exploration versus exploitation
- Personal interest
- Formal theory
- Personal metaphor
- Working interpretation
I am drawn to the pull between going deeper on a path I know and trying something completely new. In life, as in computer programs that learn by trial and error, trying new things can change what becomes possible later.
When should I stop fine-tuning what I have and start exploring?
The idea
Every learner faces the same choice: use the option that has worked so far, or try one that might be better. Using it gives a dependable payoff now. Exploring gives up some of that payoff in exchange for information. Researchers study this as the multi-armed bandit problem, and the theory backs a piece of common sense: explore more early, when there is lots of time left to benefit, and lean on what works later.
Why it attracts me
My interests are wide (Depth in One Field, Curiosity in Many). This tradeoff lets me treat that as a strategy instead of a flaw. A new field can change what an old skill is good for, and wandering can reveal connections that an efficient search would miss (Curious Wandering and Lucky Discoveries). It also keeps me honest: exploring has a real cost, so it should be a choice.
An example
Imagine choosing where to eat on a Friday. One restaurant is reliably good. A new one has opened nearby. If you have just moved to town, trying the new place is worth the risk, because you could eat there for years if it turns out to be great. If you are leaving next week, the reliable favorite makes more sense. The same choice changes with the time left to use what you learn.
Where it connects
This is the central puzzle in learning what to do in each situation (Learning What to Do in Each Situation). It also shapes which branches of the future I bother to explore at all (Mapping the Futures a Choice Opens).
Questions I'm still exploring
- How do I know when I have explored enough to commit?
- What would I try if exploring counted as part of the work rather than a distraction from it?
- How should the balance shift as the time left to use what I learn gets shorter?
Sources and further reading
- Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, 2nd edition (MIT Press, 2018) — Chapter 2 covers the multi-armed bandit problem.
- Brian Christian and Tom Griffiths, Algorithms to Live By (Henry Holt, 2016), chapter "Explore/Exploit"
Working interpretation: drafted from my notes and interests for review. It is not a direct quotation, and I may still change it.