Exploration Is a State, Not a Setting: A Markov-Switching Reinforcement-Learning Model of Strategy Transitions in Sequential Choice
DOI:
https://doi.org/10.62411/jcta.16818Keywords:
Computational behavioral modeling, Directed and random exploration, Exploration–exploitation, Hidden Markov model, Latent states, Markov switching, Reinforcement learning, Strategy dynamicsAbstract
Computational accounts of human exploration usually assume a stationary policy, in which a single set of parameters for value sensitivity and uncertainty seeking generates every choice in a task, so that all within-task variation is treated as decision noise. Neuroscience instead treats exploration and exploitation as dissociable modes between which the brain switches, which a stationary model cannot represent. We introduce the Markov-Switching Reinforcement-Learning (MS-RL) model, in which a first-order hidden Markov chain governs transitions among a small number of latent decision regimes, each with its own softmax policy over a shared value-learning process. The three regimes are exploitation, directed exploration, and random exploration. The model contains the standard stationary account (one regime) and a temporally unstructured mixture (memoryless transitions) as nested special cases, making the stationarity assumption testable. We estimate the model using Expectation-Maximization with Viterbi decoding and select the number of regimes using the Bayesian information criterion. A parameter- and state-recovery study confirmed that the generating parameters and latent regime paths are recoverable (parameter correlations 0.88–0.94; state accuracy 86 percent; Cohen’s kappa ≈ 0.79). Applied to an openly available two-armed bandit dataset (46 adults, 13,800 choices), a three-regime MS-RL model was preferred over the nested baselines, two- and four-regime variants, and a single-regime model with smoothly time-varying value sensitivity. An ablation analysis attributes the largest gains to the latent regimes and temporal switching. The decoded regime path revealed a systematic within-game shift from directed exploration toward exploitation. The estimated transition matrix provides a per-participant measure of strategy change that has no counterpart in stationary models. We conclude that exploration is better described as a dynamic state than as a fixed trait, and that modeling it as a switching process is both more accurate and useful.References
N. D. Daw, J. P. O’Doherty, P. Dayan, B. Seymour, and R. J. Dolan, “Cortical substrates for exploratory decisions in humans,” Nature, vol. 441, no. 7095, pp. 876–879, Jun. 2006, doi: 10.1038/nature04766.
R. C. Wilson, A. Geana, J. M. White, E. A. Ludvig, and J. D. Cohen, “Humans use directed and random exploration to solve the explore–exploit dilemma.,” J. Exp. Psychol. Gen., vol. 143, no. 6, pp. 2074–2081, 2014, doi: 10.1037/a0038199.
E. Schulz and S. J. Gershman, “The algorithmic architecture of exploration in the human brain,” Curr. Opin. Neurobiol., vol. 55, pp. 7–14, Apr. 2019, doi: 10.1016/j.conb.2018.11.003.
M. M. Botvinick, T. S. Braver, D. M. Barch, C. S. Carter, and J. D. Cohen, “Conflict monitoring and cognitive control.,” Psychol. Rev., vol. 108, no. 3, pp. 624–652, 2001, doi: 10.1037/0033-295X.108.3.624.
R. Ligneul, “Sequential exploration in the Iowa gambling task: Validation of a new computational model in a large dataset of young and old healthy participants,” PLOS Comput. Biol., vol. 15, no. 6, p. e1006989, Jun. 2019, doi: 10.1371/journal.pcbi.1006989.
D. Tuzsus, A. Brands, I. Pappas, and J. Peters, “Exploration–Exploitation Mechanisms in Recurrent Neural Networks and Human Learners in Restless Bandit Problems,” Comput. Brain Behav., vol. 7, no. 3, pp. 314–356, Sep. 2024, doi: 10.1007/s42113-024-00202-y.
S. J. Gershman, “Deconstructing the human algorithms for exploration,” Cognition, vol. 173, pp. 34–42, Apr. 2018, doi: 10.1016/j.cognition.2017.12.014.
M. Speekenbrink and E. Konstantinidis, “Uncertainty and Exploration in a Restless Bandit Problem,” Top. Cogn. Sci., vol. 7, no. 2, pp. 351–367, Apr. 2015, doi: 10.1111/tops.12145.
N. D. Daw, S. J. Gershman, B. Seymour, P. Dayan, and R. J. Dolan, “Model-Based Influences on Humans’ Choices and Striatal Prediction Errors,” Neuron, vol. 69, no. 6, pp. 1204–1215, Mar. 2011, doi: 10.1016/j.neuron.2011.02.027.
Y. Ger, E. Nachmani, L. Wolf, and N. Shahar, “Harnessing the flexibility of neural networks to predict dynamic theoretical parameters underlying human choice behavior,” PLOS Comput. Biol., vol. 20, no. 1, p. e1011678, Jan. 2024, doi: 10.1371/journal.pcbi.1011678.
R. Schurr, D. Reznik, H. Hillman, R. Bhui, and S. J. Gershman, “Dynamic computational phenotyping of human cognition,” Nat. Hum. Behav., vol. 8, no. 5, pp. 917–931, Feb. 2024, doi: 10.1038/s41562-024-01814-x.
W. K. Zajkowski, M. Kossut, and R. C. Wilson, “A causal role for right frontopolar cortex in directed, but not random, exploration,” Elife, vol. 6, p. e27430, Sep. 2017, doi: 10.7554/eLife.27430.
G. Aston-Jones and J. D. Cohen, “An integrative theory of locus coeruleus-norepinephrine function: Adaptive gain and optimal performance,” Annu. Rev. Neurosci., vol. 28, no. 1, pp. 403–450, Jul. 2005, doi: 10.1146/annurev.neuro.28.061604.135709.
M. Jepma and S. Nieuwenhuis, “Pupil Diameter Predicts Changes in the Exploration–Exploitation Trade-off: Evidence for the Adaptive Gain Theory,” J. Cogn. Neurosci., vol. 23, no. 7, pp. 1587–1596, Jul. 2011, doi: 10.1162/jocn.2010.21548.
C. S. Chen, D. Mueller, E. Knep, R. B. Ebitz, and N. M. Grissom, “Dopamine and Norepinephrine Differentially Mediate the Exploration–Exploitation Tradeoff,” J. Neurosci., vol. 44, no. 44, p. e1194232024, Oct. 2024, doi: 10.1523/JNEUROSCI.1194-23.2024.
J. S. B. T. Evans and K. E. Stanovich, “Dual-process theories of higher cognition: Advancing the debate,” Perspect. Psychol. Sci., vol. 8, no. 3, pp. 223–241, May 2013, doi: 10.1177/1745691612460685.
J. B. Hirsh, R. A. Mar, and J. B. Peterson, “Psychological entropy: A framework for understanding uncertainty-related anxiety.,” Psychol. Rev., vol. 119, no. 2, pp. 304–320, Apr. 2012, doi: 10.1037/a0026767.
A. Modirshanechi, W.-H. Lin, H. A. Xu, M. H. Herzog, and W. Gerstner, “Novelty as a drive of human exploration in complex stochastic environments,” Proc. Natl. Acad. Sci., vol. 122, no. 39, Sep. 2025, doi: 10.1073/pnas.2502193122.
M. S. Tomov, V. Q. Truong, R. A. Hundia, and S. J. Gershman, “Dissociable neural correlates of uncertainty underlie different exploration strategies,” Nat. Commun., vol. 11, no. 1, p. 2371, May 2020, doi: 10.1038/s41467-020-15766-z.
Z. C. Ashwood et al., “Mice alternate between discrete strategies during perceptual decision-making,” Nat. Neurosci., vol. 25, no. 2, pp. 201–212, Feb. 2022, doi: 10.1038/s41593-021-01007-z.
Š. Kucharský, N.-H. Tran, K. Veldkamp, M. Raijmakers, and I. Visser, “Hidden Markov Models of Evidence Accumulation in Speeded Decision Tasks,” Comput. Brain Behav., vol. 4, no. 4, pp. 416–441, Dec. 2021, doi: 10.1007/s42113-021-00115-0.
M. Jepma, J. V Schaaf, I. Visser, and H. M. Huizenga, “Uncertainty-driven regulation of learning and exploration in adolescents: A computational account,” PLOS Comput. Biol., vol. 16, no. 9, p. e1008276, Sep. 2020, doi: 10.1371/journal.pcbi.1008276.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction. MIT Press, 2018. [Online]. Available: http://incompleteideas.net/book/the-book-2nd.html
L. R. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,” Proc. IEEE, vol. 77, no. 2, pp. 257–286, 1989, doi: 10.1109/5.18626.
C. M. Bishop, Pattern Recognition and Machine Learning. Springer New York, NY, 2006. [Online]. Available: https://link.springer.com/book/9780387310732
R. C. Wilson and A. G. Collins, “Ten simple rules for the computational modeling of behavioral data,” Elife, vol. 8, p. e49547, Nov. 2019, doi: 10.7554/eLife.49547.
P. L. Lockwood and M. C. Klein-Flügge, “Computational modelling of social cognition and behaviour—a reinforcement learning primer,” Soc. Cogn. Affect. Neurosci., vol. 16, no. 8, pp. 761–771, Mar. 2020, doi: 10.1093/scan/nsaa040.
S. J. Gershman, “Uncertainty and exploration.,” Decision, vol. 6, no. 3, pp. 277–286, Jul. 2019, doi: 10.1037/dec0000101.
M. Binz and E. Schulz, “Modeling Human Exploration Through Resource-Rational Reinforcement Learning,” in Advances in Neural Information Processing Systems 35, 2022, vol. 35, pp. 31755–31768. doi: 10.52202/068431-2302.
T. S. Braver, “The variable nature of cognitive control: a dual mechanisms framework,” Trends Cogn. Sci., vol. 16, no. 2, pp. 106–113, Feb. 2012, doi: 10.1016/j.tics.2011.12.010.
A. Diamond, “Executive Functions,” Annu. Rev. Psychol., vol. 64, no. 1, pp. 135–168, Jan. 2013, doi: 10.1146/annurev-psych-113011-143750.
M. Mittner, G. E. Hawkins, W. Boekel, and B. U. Forstmann, “A Neural Model of Mind Wandering,” Trends Cogn. Sci., vol. 20, no. 8, pp. 570–578, Aug. 2016, doi: 10.1016/j.tics.2016.06.004.
H. Fan, T. Burke, D. C. Sambrano, E. Dial, E. A. Phelps, and S. J. Gershman, “Pupil Size Encodes Uncertainty during Exploration,” J. Cogn. Neurosci., vol. 35, no. 9, pp. 1508–1520, Sep. 2023, doi: 10.1162/jocn_a_02025.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Fathimah Al-Ma'shumah, Feneta Fidi Kirani, Nita Ratnawaty

This work is licensed under a Creative Commons Attribution 4.0 International License.














