<?xml version="1.0" ?><!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.3 20210610//EN"  "JATS-archivearticle1-mathml3.dtd"><article xmlns:ali="http://www.niso.org/schemas/ali/1.0/" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" dtd-version="1.3" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="nlm-ta">elife</journal-id>
<journal-id journal-id-type="publisher-id">eLife</journal-id>
<journal-title-group>
<journal-title>eLife</journal-title>
</journal-title-group>
<issn publication-format="electronic" pub-type="epub">2050-084X</issn>
<publisher>
<publisher-name>eLife Sciences Publications, Ltd</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">94231</article-id>
<article-id pub-id-type="doi">10.7554/eLife.94231</article-id>
<article-id pub-id-type="doi" specific-use="version">10.7554/eLife.94231.1</article-id>
<article-version-alternatives>
<article-version article-version-type="publication-state">reviewed preprint</article-version>
<article-version article-version-type="preprint-version">1.2</article-version>
</article-version-alternatives>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Human Exploration Strategically Balances Approaching and Avoiding Uncertainty</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Abir</surname>
<given-names>Yaniv</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<email xlink:href="mailto:yaniv.abir@columbia.edu">yaniv.abir@columbia.edu</email>
<xref ref-type="corresp" rid="cor1">*</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shadlen</surname>
<given-names>Michael N.</given-names>
</name>
<xref ref-type="aff" rid="a2">2</xref>
<xref ref-type="aff" rid="a3">3</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Shohamy</surname>
<given-names>Daphna</given-names>
</name>
<xref ref-type="aff" rid="a1">1</xref>
<xref ref-type="aff" rid="a2">2</xref>
<email xlink:href="mailto:ds2619@columbia.edu">ds2619@columbia.edu</email>
</contrib>
<aff id="a1"><label>1</label><institution>Department of Psychology, Columbia University</institution>, <city>New York</city>, <state>NY</state>, <country>USA</country></aff>
<aff id="a2"><label>2</label><institution>Zuckerman Mind Brain Behavior Institute, and Kavli Institute for Brain Science, Columbia University</institution>, <city>New York</city>, <state>NY</state>, <country>USA</country></aff>
<aff id="a3"><label>3</label><institution>Department of Neuroscience and Howard Hughes Medical Institute, Columbia University</institution>, <city>New York</city>, <state>NY</state>, <country>USA</country></aff>
</contrib-group>
<contrib-group content-type="section">
<contrib contrib-type="editor">
<name>
<surname>Gillan</surname>
<given-names>Claire M</given-names>
</name>
<role>Reviewing Editor</role>
<aff>
<institution-wrap>
<institution-id institution-id-type="ror">https://ror.org/02tyrky19</institution-id><institution>Trinity College Dublin</institution>
</institution-wrap>
<city>Dublin</city>
<country>Ireland</country>
</aff>
</contrib>
<contrib contrib-type="senior_editor">
<name>
<surname>Büchel</surname>
<given-names>Christian</given-names>
</name>
<role>Senior Editor</role>
<aff>
<institution-wrap>
<institution-id institution-id-type="ror">https://ror.org/01zgy1s35</institution-id><institution>University Medical Center Hamburg-Eppendorf</institution>
</institution-wrap>
<city>Hamburg</city>
<country>Germany</country>
</aff>
</contrib>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>*</label>For correspondence: <email xlink:href="mailto:yaniv.abir@columbia.edu">yaniv.abir@columbia.edu</email> (YA); <email xlink:href="mailto:ds2619@columbia.edu">ds2619@columbia.edu</email> (DS)</corresp>
</author-notes>
<pub-date date-type="original-publication" iso-8601-date="2024-02-15">
<day>15</day>
<month>02</month>
<year>2024</year>
</pub-date>
<volume>13</volume>
<elocation-id>RP94231</elocation-id>
<history><date date-type="sent-for-review" iso-8601-date="2023-12-08">
<day>08</day>
<month>12</month>
<year>2023</year>
</date>
</history>
<pub-history>
<event>
<event-desc>Preprint posted</event-desc>
<date date-type="preprint" iso-8601-date="2023-12-08">
<day>08</day>
<month>12</month>
<year>2023</year>
</date>
<self-uri content-type="preprint" xlink:href="https://doi.org/10.31234/osf.io/gtxam"/>
</event>
</pub-history>
<permissions>
<copyright-statement>© 2024, Abir et al</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Abir et al</copyright-holder>
<ali:free_to_read/>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<ali:license_ref>https://creativecommons.org/licenses/by/4.0/</ali:license_ref>
<license-p>This article is distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License</ext-link>, which permits unrestricted use and redistribution provided that the original author and source are credited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="elife-preprint-94231-v1.pdf"/>
<abstract>
<title>Abstract</title>
<p>A central purpose of exploration is to reduce goal-relevant uncertainty. Consequentially, individuals often explore by focusing on areas of uncertainty in the environment. However, people sometimes adopt the opposite strategy, one of avoiding uncertainty. How are the conflicting tendencies to approach and avoid uncertainty reconciled in human exploration? We hypothesized that the balance between avoiding and approaching uncertainty can be understood by considering capacity constraints. Accordingly, people are expected to approach uncertainty in most cases, but to avoid it when overall uncertainty is highest. To test this, we developed a new task and used modeling to compare human choices to a range of plausible policies. The task required participants to learn the statistics of a simulated environment by active exploration. On each trial, participants chose to explore a better-known or lesser-known option. Participants generally chose to approach uncertainty, however, when overall uncertainty about the choice options was highest, they instead avoided uncertainty and chose to sample better-known objects. This strategy was associated with faster decisions and, despite reducing the rate of observed information, it did not impair learning. We suggest that balancing approaching and avoiding uncertainty reduces the cognitive costs of exploration in a resource-rational manner.</p>
</abstract>

</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<title>Introduction</title>
<p>The purpose of exploration is to reduce uncertainty about the aspects of one’s environment that are goal relevant or otherwise important. Yet, devising an optimal strategy to reduce uncertainty is known to be very difficult (<xref ref-type="bibr" rid="c17">Cohen et al., 2007</xref>; <xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c63">Sutton and Barto, 2018</xref>), especially for agents with limited memory and processing capacities. A heuristic strategy that is often efficient for exploration is focusing on the parts of the environment that one is most uncertain about. This principle of approaching uncertainty has been applied in a range of fields, including statistics (<xref ref-type="bibr" rid="c38">MacKay, 1992</xref>; <xref ref-type="bibr" rid="c56">Sebastiani and Wynn, 2000</xref>), artificial intelligence (<xref ref-type="bibr" rid="c5">Badia et al., 2020</xref>; <xref ref-type="bibr" rid="c8">Bellemare et al., 2016</xref>; <xref ref-type="bibr" rid="c44">Pathak et al., 2017</xref>; <xref ref-type="bibr" rid="c48">Raposo et al., 2021</xref>), and cognitive theories of human exploration (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c55">Schwartenbeck et al., 2019</xref>). Indeed, humans have been shown to approach uncertainty when learning about rewards in the environment through trial and error (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c62">Speekenbrink and Konstantinidis, 2015</xref>; <xref ref-type="bibr" rid="c69">Wilson et al., 2014</xref>; <xref ref-type="bibr" rid="c70">Wu et al., 2022</xref>).</p>
<p>However, there are also many examples of uncertainty avoidance in the decision making of humans and animals. Uncertainty avoidance has been documented in situations where resolving uncertainty may reveal negative outcomes and news (<xref ref-type="bibr" rid="c2">Ahmadlou et al., 2021</xref>; <xref ref-type="bibr" rid="c9">Botta et al., 2020</xref>; <xref ref-type="bibr" rid="c24">Eilam and Golani, 1989</xref>; <xref ref-type="bibr" rid="c28">Glickman and Sroges, 1966</xref>; <xref ref-type="bibr" rid="c30">Gordon et al., 2014</xref>; <xref ref-type="bibr" rid="c27">Gigerenzer and Garcia-Retamero, 2017</xref>; <xref ref-type="bibr" rid="c29">Golman et al., 2017</xref>), or may make overcoming a conflict in motivation more difficult (<xref ref-type="bibr" rid="c14">Carrillo and Mariotti, 2000</xref>; <xref ref-type="bibr" rid="c29">Golman et al., 2017</xref>). When the goal is to maximize immediate rewards, choosing the most rewarding option often entails avoiding more uncertain options (<xref ref-type="bibr" rid="c64">Trudel et al., 2020</xref>; <xref ref-type="bibr" rid="c69">Wilson et al., 2014</xref>).</p>
<p>How are the two conflicting tendencies to approach and avoid uncertainty reconciled when exploring? To answer this question, we must address gaps in the literature about exploration at two levels of analysis. At the computational level, it is unclear what might compel individuals to avoid uncertainty instead of approaching it, bar holding goals other than attaining knowledge. Indeed, avoiding uncertainty reduces the rate of information intake, and so might result in poorer learning. At the algorithmic level, we lack an understanding of how individuals compute uncertainty to make exploratory choices. Tallying uncertainty in an exact manner is complicated and often intractable. Several candidate algorithms for approximating the computation of uncertainty have been suggested (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>), but evidence as to their use by humans is still preliminary.</p>
<p>It is the complexity of choosing based on uncertainty, set against the limited processing and memory capacities that are inherent to human cognition, that motivated our hypotheses regarding both the algorithmic and computational questions. First, we charted a hypothesis space of plausible algorithms for computing uncertainty and making exploratory choices (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>), starting with the optimal but complex, and ending with simple approximations. Second, we hypothesized that the complexity of choosing what to explore, even when using approximate algorithms, is the key factor explaining why and when individuals might avoid uncertainty in exploration. Adhering to the goal of approaching uncertainty may well be an efficient policy for an agent with unlimited cognitive resources. Since humans have finite memory systems, inference bandwidth, and time, it stands to reason that they would try to conserve these resources by regulating their exploration (<xref ref-type="bibr" rid="c36">Lieder and Griffiths, 2020</xref>), possibly by selectively avoiding uncertainty. Following this insight, we examined exploratory choices as a function of two factors affecting the difficulty of making an exploratory choice: participants’ overall uncertainty about choice options (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>), and forgetting.</p>
<p>We developed a task requiring participants to make multiple exploratory choices, incrementally building knowledge in the service of a distant goal (<xref ref-type="fig" rid="fig1">Figure 1</xref>). Importantly, participants were given reward feedback only at the end of a round and not after every trial, allowing us to focus on choices made to accumulate knowledge, rather than choices driven by the need to exploit available rewards. Seeking ecological validity, we designed a task that posed a challenging exploration problem for participants, requiring that they infer and remember the values of multiple latent parameters from repeated experience (<xref ref-type="bibr" rid="c32">Hartley, 2022</xref>; <xref ref-type="bibr" rid="c36">Lieder and Griffiths, 2020</xref>). The task could nonetheless be captured by a few mathematical expressions, allowing for the derivation of the optimal exploration policy. This optimal policy served as a basis for a quantitative analysis of participants’ choices and reaction times with the aim of identifying the algorithm driving their exploratory choices (<xref ref-type="bibr" rid="c3">Anderson, 1990</xref>; <xref ref-type="bibr" rid="c15">Chater and Oaksford, 1999</xref>; <xref ref-type="bibr" rid="c67">Waskom et al., 2019</xref>).</p>
</sec>
<sec id="s2" sec-type="results">
<title>Results</title>
<p>194 participants from a pre-registered (<xref ref-type="bibr" rid="c1">Abir et al., 2021</xref>) sample were recruited to complete up to 22 rounds of the exploration task over four online sessions. The task simulated a room with four tables, with two decks of cards on each table (<xref ref-type="fig" rid="fig1">Figure 1</xref>a-b). If a card was flipped, it was revealed to be, for example, either orange or blue (each round used a different pair of colors). The proportion of orange vs. blue cards, <italic>π</italic> differed between the two decks on each table. Participants’ goal was to learn <italic>sgn(π</italic><sub>1</sub><italic> – π</italic><sub>2</sub> or which deck had more orange (blue) cards on each table. We will denote this term, which serves as the learning desideratum for participants, as <italic>θ.</italic></p>
<fig id="fig1" position="float" fig-type="figure">
<label>Figure 1.</label>
<caption><title>Examining exploration strategy in relation to uncertainty in an incremental learning task.</title>
<p><bold>a</bold>, Structure of the task. Participants explored four tables, each containing two decks with different proportions of blue/orange cards. The goal was to learn the difference in proportions of the decks on each table. <bold>b</bold>, The two phases of the task - exploration and test. On a single exploration trial (left), participants chose between two tables, and then sampled a card from one of the decks on that table, observing its color. After a random number of exploration trials, participants were tested on their knowledge (right). A color was designated as rewarding, and participants then chose the deck with the highest proportion of the rewarding color on each table. They were rewarded for correct test-phase choices, and received no reward during exploration. <bold>c</bold>, Histogram of round lengths. Participants played 22 rounds. The length of exploration in each round followed a shifted geometric distribution, such that the test was equally likely to occur following any trial after the first 10. <bold>d</bold>, We considered a hierarchy of strategies for choosing which table to explore. The normatively prescribed strategy is to choose the table affording maximal expected information gain. This is the table for which the next card is expected to maximally decrease uncertainty (measured as entropy <italic>H</italic>) about the value of the goal-relevant latent parameter <italic>θ,</italic> given observations thus far <italic>x.</italic> A simpler strategy is to choose the table with the maximum uncertainty, as it does not necessitate computing an expectation over the next observation. An even simpler heuristic is to equate previous exposure and choose the table with the least previous observations <italic>n<sub>x</sub>.</italic> Even though these three strategies vary considerably in complexity, they are all uncertainty-approaching on average. Lastly, people may be random explorers.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig2" position="float" fig-type="figure">
<label>Figure 2.</label>
<caption><title>Hypothetical strategies make differing predictions for exploratory choice behavior.</title>
<p>We computed the three quantities hypothesized to drive exploratory choices using a Bayesian observer model. To illustrate this process, we plot the derivation of Bayesian belief on a single trial. (<bold>a</bold>) and across multiple trials (<bold>b</bold>, <bold>c</bold>). For visualization, we use a simplified version with two tables only. <bold>a</bold> depicts the Bayesian observer’s belief about a single table on a single trial. Given a sequence of previously observed cards (left), the Bayesian observer forms posterior beliefs about the proportion of orange cards in each deck (center). These beliefs are expressed as Beta distributions. From these, it is possible to derive a belief about the difference in the proportion of orange cards between the two decks π<sub>1</sub> – π<sub>2</sub> (right). The probability that π<sub>1</sub><italic>&lt;</italic> π<sub>2</sub> is given by the proportional size of the area marked in gray (0.74 in this example). <bold>b</bold> Depicts the same process over a series of20trials. The observed card sequence for each table is presented at the top of each panel. The matching belief state about π<sub>1</sub> – π<sub>2</sub> is plotted below it as an evolving posterior density in white (high) and black (low). The green arrows mark the true value of π<sub>1</sub> – π<sub>2</sub> for that round. As the round progresses, belief converges towards the true value, and becomes more certain. <bold>c</bold>, The three choice strategies prescribe different table choices on most trials. The difference between <xref ref-type="table" rid="app3-tbl1">table 1</xref> and <xref ref-type="table" rid="app3-tbl2">table 2</xref> in each of the three quantities (EIG, uncertainty and exposure) is plotted for each trial. This difference is the hypothesized decision variable for choosing between <xref ref-type="table" rid="app3-tbl1">tables 1</xref> and <xref ref-type="table" rid="app3-tbl2">2</xref>. A positive value indicates a preference for exploring table 1, and a negative value a preference for <xref ref-type="table" rid="app3-tbl2">table 2</xref>. The three variables are normalized to facilitate visual comparison.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig2.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<p>The task begins with an exploration phase, followed by a test phase. On each trial of the exploration phase participants chose which of two tables to explore, and then revealed one card from a deck on that table (<xref ref-type="fig" rid="fig1">Figure 1 b</xref>). Participants were instructed that the exploration phase would be followed by a test phase after a random number of trials (drawn from a geometric distribution to discourage pre-planning, <xref ref-type="fig" rid="fig1">Figure 1c</xref>). They were further instructed that one of the colors would be designated as rewarding at the beginning of the test phase. During the test phase, participants were asked to indicate which deck had more of the rewarding color on each table (<xref ref-type="fig" rid="fig1">Figure 1b</xref>). They also rated their confidence in the choice. For every correct test-phase choice they received $0.25. Crucially, they received no reward during exploration. Participants’ only incentive during the exploration phase was to maximize their confidence about the value of <italic>θ.</italic></p>
<sec id="s2-1">
<title>Three Hypothetical Strategies Derived by Rational Analysis</title>
<p>To explain how participants chose between tables in the exploration phase, we first asked how an optimal agent might solve the problem of choosing which table to explore on each trial of the task. We limited our consideration to strategies that optimize learning only for the next trial, since a globally optimal strategy is intractable for this task (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c63">Sutton and Barto, 2018</xref>). We started by deriving the optimal strategy and progressively simplified it to generate two additional strategies. While they differ in the level of complexity they assume, all three strategies direct an agent using them to approach the option they are more uncertain about.</p>
<p>The optimal strategy, given by the expression at the top of <xref ref-type="fig" rid="fig1">Figure 1d</xref>, is choosing the table affording maximal expected information gain (EIG; Gureckis and Markant (2012); MacKay (1992); Yang et al. (2016)). EIG is the difference between the uncertainty in the value of the learning desider-atum, <italic>θ,</italic> given observed cards <italic>x<sub>0: t</sub></italic> and the expected uncertainty after observing the next card on trial <italic>t</italic> +1. In other words, EIG is the amount of uncertainty resolvable on the next trial.</p>
<fig id="fig3" position="float" fig-type="figure">
<label>Figure 3.</label>
<caption><title>The Bayesian observer model is validated by participants’ accuracy and confidence on the test phase.</title>
<p><bold>a</bold>, Participants were accurate when an exploration phase ended with low uncertainty, and performed at chance level when the phase ended with high uncertainty. <bold>b</bold>, Participant’s confidence on correct choices fell with rising uncertainty. Confidence on error trials did not depend as much on Bayesian observer uncertainty. When a test question was unsolvable because no evidence was observed on each deck during exploration, participants had very low confidence. Data presented as mean values ±1 SE, n=194 participants.</p>
<p><bold>Figure 3—figure supplement 1.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig3.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS3-1" position="float" fig-type="figure">
<label>Figure 3—figure supplement 1.</label>
<caption><title>Reproducing the analysis using the preliminary sample: The Bayesian observer model is validated by participants’ accuracy and confidence on the test phase.</title>
<p><bold>a</bold>, Participants were accurate when an exploration phase ended with low uncertainty, and performed at chance level when the phase ended with high uncertainty. <bold>b</bold>, Participant’s confidence on correct choices fell with rising uncertainty. Confidence on errors did not depend as much on Bayesian observer uncertainty. Data presented as mean values ±1 SE, n=62 participants. Nats are the units of entropy, a mathematically convenient measure of uncertainty.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS3-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<p>Computing the second term in the EI expression requires averaging over future unseen outcomes, which may be beyond the ability of participants. As an alternative, they might avoid computing this term by simply choosing the table they were more uncertain about at the moment of making the choice (<xref ref-type="fig" rid="fig1">Figure 1d</xref>, second tier; Schulz and Gershman, 2019). While this strategy has intuitive appeal, computing uncertainties may still be too complicated for human participants. An even simpler heuristic is given on the third tier of <xref ref-type="fig" rid="fig1">Figure 1</xref>d: choosing the table with the least prior exposure (<xref ref-type="bibr" rid="c4">Auer, 2002</xref>; <xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>), measured as the number of already observed cards <italic>n<sub>x</sub>.</italic> Since on average additional observations result in lower uncertainty, this strategy is an approximate way to approach the more uncertain table. Finally, participants might explore at random, rather than in a directed manner (<xref ref-type="bibr" rid="c21">Daw et al., 2006</xref>; <xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c69">Wilson et al., 2014</xref>).</p>
</sec>
<sec id="s2-2">
<title>Test Phase Performance Validates Observation Model</title>
<p>To relate the three hypothesized strategies to participants’ behavior, we assumed a model of participants’ beliefs about the goal-relevant parameter θ and the mechanism by which they updated these beliefs. We used a Bayesian observer model which forms beliefs about θ based on the actual card sequence each participant observed, and updates these beliefs according to Bayes’ rule (<xref ref-type="fig" rid="fig2">Figure 2</xref>). On its own, the Bayesian observer does not predict participants’ exploration choices, but only models the process of inference from observation.</p>
<fig id="fig4" position="float" fig-type="figure">
<label>Figure 4.</label>
<caption><title>Uncertainty is the best predictor of choice.</title>
<p><bold>a</bold>, On each plot the difference in the hypothesized quantity between the two tables presented on each trial is plotted against actual choices of the table presented on the right. For each plot, the relevant hypothesis predicts a positive smooth curve. -uncertainty, plotted on the left, matches this prediction better than Δ (center). The relationship between Δ-exposure (right) and choice is negative, rather than the hypothesized positive correlation. <bold>b</bold> Quantitative model comparison confirms this observation. Out of the three hypothesized strategies uncertainty has the highest approximate expected log predictive density (using PSIS LOO; see Methods). Data presented as mean values ±1SE, n=194 participants.</p>
<p><bold>Figure 4—figure supplement 1.</bold> Fitting simulated data successfully recovers the underlying strategy.</p>
<p><bold>Figure 4—figure supplement 2.</bold> Uncertainty is a sufficient predictor of choice.</p>
<p><bold>Figure 4—figure supplement 3.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig4.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS4-1" position="float" fig-type="figure">
<label>Figure 4—figure supplement 1.</label>
<caption><title>Our analysis approach successfully recovers the strategy used by simulated agents.</title>
<p>We compared the actual data (top) to datasets generated by artificial agents. Each simulated dataset comprised a group of agents operating according to one of the hypothesized strategies. We fixed the effect size for each strategy in the simulations to the effect size we observed for uncertainty in the actual data. Each agent matched a single participant in the true dataset, choosing to observe cards from the same decks presented to the participant. Each agent chose the table on the right or the table on the left on each trial, with the probability of choosing the table on the right given by <italic>f (a</italic> + <italic>b</italic> × Δ<italic>x),</italic> where <italic>f</italic> is the logistic function, <italic>b</italic> is the degree to which the agent’s choices are dependent on <italic>Ax,</italic> standing for the relevant decision variable, and <italic>a</italic> is a general bias towards rightward or leftward choices. Coefficients <italic>a</italic> and <italic>b</italic> were extracted per participant from the uncertainty model described in <xref ref-type="fig" rid="fig4">Figure 4</xref>. For the sake of this analysis, we assumed the agents choose a random deck on the table of their choice. Here, the simulated data for each group of agents was plotted against each of the three decision variables and fit with the same models we used on the actual dataset (center). We tested whether our procedure for qualitative and quantitative model comparison used in <xref ref-type="fig" rid="fig4">Figure 4</xref> is potent at recovering the true strategy generating the data. For easy comparison, the actual data is re-plotted on the first row. For each of the three simulated strategies, we observe successful recovery: the decision variable matching the true strategy shows the strongest positive correlation with choice (center), and the correct strategy is indicated as best fitting the data (right). Furthermore, a negative correlation between behavior and Δ-exposure, as observed in the true data, is only evident in the uncertainty-based group of agents (second row). Data plotted as means ±1SE, n=194 participants/agents.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS4-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS4-2" position="float" fig-type="figure">
<label>Figure 4—figure supplement 2.</label>
<caption><title>Simulations confirm that uncertainty is a sufficient predictor of choice.</title>
<p>We further confirmed our conclusion that uncertainty is the best predictor of participants’ choices, by plotting the posterior predictive distribution for each of the models predicting choice from a hypothesized strategy. We simulated 500 datasets for each of the three models, and plotted the distribution of the simulated data (green lines with 50% and 95% posterior interval bands, n=500 iterations) against the observed dataset (means ±1SE plotted in black, n=194 participants). The simulation procedure was similar to that used in <xref ref-type="fig" rid="fig4">Figure 4—figure Supplement 1</xref>, with the exception that coefficients were extracted from the posterior distribution of each model fitted to the actual data. We expect that the posterior predictive distribution for each model would capture the relationship between the relevant decision variable and choice well. The extent to which the posterior predictive distribution can recreate the association with the other two decision variables is a test of model fit. <bold>a</bold>, The posterior predictive distribution for the EIG model does not match the observed data well: it does not reproduce the strong slope for Δ-uncertainty, nor the negative correlation with Δ-exposure. <bold>b</bold>, The posterior predictive distribution for uncertainty captures the data very well, matching the particular shape of the correspondence between choices and Δ-EIG, and the negative correlation between choices and. Δ-exposure. <bold>c</bold>, The posterior predictive distribution of exposure does not match observe data well: it fails to recreate the positive correlations between choice and EIG and uncertainty.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS4-2.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS4-3" position="float" fig-type="figure">
<label>Figure 4—figure supplement 3.</label>
<caption><title>Reproducing the analysis in <xref ref-type="fig" rid="fig4">Figure 4</xref> using the preliminary sample: Uncertainty is the best predictor of choice.</title>
<p><bold>a</bold>, On each plot the difference in the hypothesized quantity between the two tables presented on each trial is plotted against actual choices of the table presented on the right. For each plot, the relevant hypothesis predicts a positive smooth curve. Δ-uncertainty, plotted on the left, matches this prediction better than Δ-EIG (center). The relationship between Δ-exposure (right) and choice is negative, rather than the hypothesized positive correlation. <bold>b</bold> Quantitative model comparison confirms this observation. Out of the three hypothesized strategies uncertainty has the highest approximate expected log predictive density (PSIS LOO; see Methods). Data presented as mean values ±1 SE, n=62 participants.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS4-3.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<p>Before evaluating the hypothesized exploration strategies, we sought to validate the assumptions of the Bayesian observer model. To this end, we related the predictions of the Bayesian observer model to participants’ choices during the test phase. We predicted that test accuracy should be greater when the Bayesian observer model had low uncertainty about <italic>θ</italic> at the end of the learning phase. The data supported this prediction (<xref ref-type="fig" rid="fig3">Figure 3</xref>). Using a multilevel logistic regression model, we confirmed that test accuracy was strongly related to the Bayesian observer’s uncertainty b=-5.59, 95% posterior interval (PI)=[-6.25,-4.95] (all effect sizes given in original units, full model and coefficients reported in Appendix 3—<xref ref-type="table" rid="tbl1">Table 1</xref>). Participants’ reports of confidence after making a correct choice also followed the Bayesian observer’s uncertainty b=-4.04, 95% PI=[- 4.50,-3.56]. After committing errors, participants’ reported confidence was lower overall b=-1.09, 95% PI=[-1.27,-0.92], and considerably less dependent on Bayesian observer uncertainty, interaction b=-3.10, 95% PI=[-3.76,-2.46] (<xref ref-type="fig" rid="fig3">Figure 3b</xref>, Appendix 3—<xref ref-type="table" rid="app3-tbl2">Table 2</xref>).</p>
</sec>
<sec id="s2-3">
<title>Uncertainty is the Best Predictor of Exploratory Choice</title>
<p>To evaluate the three exploration strategies, we tested whether participants’ exploration-phase choices could be predicted from the difference between the two tables that were presented as choice options in each of the hypothesized quantities. We fit the data with a multilevel logistic regression model for each strategy (Appendix 3—<xref ref-type="table" rid="app3-tbl3">Tables 3</xref>-<xref ref-type="table" rid="app3-tbl5">5</xref>). In a formal comparison of the three models we found that uncertainty was the best predictor of exploratory choices, as indicated by a reliably better prediction metric (<xref ref-type="fig" rid="fig4">Figure 4</xref>). Accordingly, the difference in uncertainty for the table presented on the right versus the table presented on the left (Δ-uncertainty) predicts participants’ choices. Δ-EIG provides a poorer fit to choices, and Δ-exposure is anti-correlated with choice, in contradiction of the exposure hypothesis. We confirmed that our analysis approach can recover the true model generating a simulated dataset(<xref ref-type="fig" rid="fig4">Figure 4—figure Supplement 1</xref>). Furthermore, simulations showed that uncertainty is a sufficient predictor of choice. Simulated datasets generated by uncertainty-driven agents recreated the entire set of qualitative and quantitative results (<xref ref-type="fig" rid="fig4">Figure 4—figure Supplement 2</xref>). The simulations demonstrate that the surprising negative correlation between choice and Δ-exposure is an epiphenomenon of uncertainty-based exploration.</p>
</sec>
<sec id="s2-4">
<title>Participants Systematically Change Their Exploration Strategy According to Overall Uncertainty</title>
<p>We next asked whet her participants’ strategy of exploring by approaching uncertainty is modulated by the state of their knowledge when making an exploratory choice. Specifically, we examined how participants’ overall uncertainty about the two options they could choose to explore on a given trial changed the way the explored (<xref ref-type="fig" rid="fig5">Figure5a</xref>). Since table choice options were presented at random, participants sometimes had to choose between tables they already knew a lot about, and sometimes between tables they were very uncertain about. When overall uncertainty was high, the choice between tables had to be made with very little evidence. Note that from a normative perspective, choice should follow the difference in uncertainty between options and shouldn’t be influenced by overall uncertainty.</p>
<p>We found a systematic deviation in exploration strategy in relation to overall uncertainty. When overall uncertainty for the two choice options was below a certain threshold, participants chose the more uncertain table, as expected. However, when overall uncertainty was above the threshold, they chose the less uncertain table, thereby slowing the rate of information intake (<xref ref-type="fig" rid="fig5">Figure 5b, c</xref>).</p>
<p>We validated this observation using a multilevel piecewise-regression model, allowing for the influence of Δ-uncertainty on choice to differ below and above a fitted threshold of overall uncertainty. We observed a positive relationship between Δ-uncertainty and choice below the threshold b=0.97, 95% PI=[0.83, 1.11], but above the threshold we found that the influence of Δ-uncertainty on choice became strongly negative (interaction b=-4.3e+02, 95% PI=[-5.4e+02,-3.4e+02]]). The threshold was estimated to be 1.28 nats of overall uncertainty (95% PI=[1.27, 1.29]; Appendix 3— <xref ref-type="table" rid="app3-tbl6">Table 6</xref>), leaving 21.58% of trials in the high overall uncertainty range (95% PI=[20.12,24.45]). This bias in exploration cannot be viewed merely as a noisier version of optimal performance. Rather, it constitutes a systematic modulation of exploration strategy on about a fifth of the trials.</p>
</sec>
<sec id="s2-5">
<title>Costs and Benefits of Strategically Avoiding Uncertainty</title>
<p>What motivates participants to systematically avoid learning about more uncertain objects? By the standards of an ideal agent, uncertainty avoidance is clearly suboptimal, as it reduces the rate of observed information, and thus the potential capacity to learn. We hypothesized that the limited processing and memory capacities that are inherent to human cognition and set it apart from the optimal agent are the reason for uncertainty avoidance. To test this hypothesis, we conduct a cost benefit analysis of uncertainty avoidance in the following sections. We ask whether uncertainty avoidance is associated with costs to learning, and whether it affords any benefits in managing cognitive effort.</p>
<sec id="s2-1-1">
<title>Uncertainty Avoidance is not Associated with Learning Deficits</title>
<p>Since efficient learning is the purpose of exploration, we asked how the tendencies to approach uncertainty and avoid it when overall uncertainty is high affect learning as reflected in performance attest. If approaching uncertainty is the only rational exploration policy, then participants who tend to approach uncertainty to a greater degree should learn more and perform better at test, while participants with a strong tendency to avoid uncertainty should learn less and perform worse at test, since they are choosing to forgo valuable information as they explore.</p>
<p>To test these predictions we examined individual differences in exploration strategy in relation to test performance. We found that participants’ baseline tendency to approach uncertainty predicted better performance at test b=2.96, 95% PI=[2.67, 3.25] (<xref ref-type="fig" rid="fig6">Figure 6b</xref>; Appendix 3—<xref ref-type="table" rid="app3-tbl7">Table 7</xref>). In contrast, we found no evidence that participants with a strong tendency to avoid uncertainty performed worse at test. Indeed, a stronger tendency to avoid uncertainty when overall uncertainty is high was associated with a small improvement in test performance b=1.18, 95% PI=[0.80, 1.58] (<xref ref-type="fig" rid="fig6">Figure 6b</xref>; Appendix 3—<xref ref-type="table" rid="app3-tbl8">Table 8</xref>). Thus, modulating exploration according to overall uncertainty was not maladaptive, resulting in no decrement to learning. This result suggests that the rate of information intake is not the limiting factor for the efficiency of exploration and learning.</p>
<fig id="fig5" position="float" fig-type="figure">
<label>Figure 5.</label>
<caption><title>Participants approach vs. avoid Δ-uncertainty as a function of overall uncertainty.</title>
<p><bold>a</bold>, While the Δ-uncertainty is the decision variable identified above, overall uncertainty, defined as the sum of uncertainty for both tables, is a measure of decision difficulty. <bold>b</bold>, The influence of Δ-uncertainty on choice differed markedly below and above a threshold of overall uncertainty. Below an estimated threshold of overall uncertainty, Δ-uncertainty had a significant positive effect on choice. Above this threshold of overall uncertainty, the influence of Δ-uncertainty became strongly negative. Points denote mean posterior estimate from regression models fitted to binned data, error bars mark 50% PI. The solid line depicts the prediction from a piecewise regression model capturing the non-linear relationship and estimating the threshold, with darker ribbon marking50%PI and light ribbon marking95% PI. Data from three regions of overall uncertainty marked in color are plotted in <bold>c</bold>. For low overall uncertainty (blue) participants tend to choose the table they are more uncertain about, as normatively prescribed. But that relationship is broken for medium levels of overall uncertainty (purple). For high overall uncertainty (red), participants strongly prefer to choose the table they are less uncertain about, thereby slowing down the rate of information intake. Data plotted as mean ±SE, n=194 participants.</p>
<p><bold>Figure 5—figure supplement 1.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig5.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS5-1" position="float" fig-type="figure">
<label>Figure 5—figure supplement 1.</label>
<caption><title>Reproducing the analysis using the preliminary sample: Participants approach vs. avoid Δ-uncertainty as a function of overall uncertainty.</title>
<p>a, The influence of Δ-uncertainty on choice differed markedly below and above a threshold of overall uncertainty. Below a certain estimated threshold of overall uncertainty, Δ-uncertainty had a significant positive effect on choice. Above this threshold of overall uncertainty, the influence of Δ-uncertainty decreased significantly. Points denote mean posterior estimate from regression models fitted to binned data, error bars mark 50% PI. The solid line depicts the prediction from a piecewise regression model capturing the non-linear relationship and estimating the threshold, with the darker ribbon marking 50% PI and the light ribbon marking 95% PI. Data from three regions of overall uncertainty marked in color are plotted in b. For low overall uncertainty (blue) participants tend to choose the table they are more uncertain about, as normatively prescribed. But that relationship is broken for medium levels of overall uncertainty (purple). For high overall uncertainty (red), participants strongly prefer to choose the table they are less uncertain about, thereby slowing down the rate of information intake. Data plotted as mean ±SE, n=62 participants.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS5-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig6" position="float" fig-type="figure">
<label>Figure 6.</label>
<caption><title>Approaching uncertainty benefits learning while avoiding uncertainty does not hurt it.</title>
<p><bold>a</bold>, We observe substantial individual differences in strategy. Replotting <xref ref-type="fig" rid="fig5">Figure 5</xref>e for each individual reveals differences in the baseline tendency to approach uncertainty, and differences in the interaction with overall uncertainty, which captures uncertainty avoidance when overall uncertainty is high. <bold>b</bold>, Associations between test performance and the parameters describing approaching and avoiding uncertainty. The baseline tendency to approach uncertainty (left) is strongly associated with performance at test, such that participants who are unable to approach uncertainty also perform poorly at test. There is a weak and positive correlation between test performance and the tendency to avoid uncertainty when overall uncertainty is high (right), indicating that uncertainty avoidance does not hinder learning. Uncertainty avoidance is quantified based on the individual lines plotted in panel <bold>a</bold> as the triangular area charted by the piecewise regression line.</p>
<p><bold>Figure 6—figure supplement 1.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig6.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS6-1" position="float" fig-type="figure">
<label>Figure 6—figure supplement 1.</label>
<caption><title>Reproducing the analysis using the preliminary sample: Approaching uncertainty benefits learning while avoiding uncertainty does not hurt it.</title>
<p><bold>a</bold>, We observe substantial individual differences in strategy. Replotting <xref ref-type="fig" rid="fig5">Figure 5e</xref> for each individual reveals differences in the baseline tendency to approach uncertainty, and differences in the interaction with overall uncertainty, which captures uncertainty avoidance when overall uncertainty is high. <bold>b</bold>, Associations between test performance and the parameters describing approaching and avoiding uncertainty. The baseline tendency to approach uncertainty (left) is strongly associated with performance at test, such that participants who are unable to approach uncertainty also perform poorly at test. There is a weak and positive correlation between test performance and the tendency to avoid uncertainty when overall uncertainty is high (right), indicating that uncertainty avoidance does not hinder learning. Uncertainty avoidance is quantified based on the individual lines plotted in panel <bold>a</bold> as the triangular area charted by the piecewise regression line.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS6-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
</sec>
<sec id="s2-1-2">
<title>Strategic Exploration Involves Costly Deliberation</title>
<p>To understand the costs involved in exploration, we asked whether making exploratory choices in this task involves prolonged deliberation. If that is the case, and exploratory choices are guided by ΔΔ-uncertainty, we reasoned that decisions should require longer deliberation when the absolute value of Δ-uncertainty is small (<xref ref-type="bibr" rid="c43">Palmer et al., 2005</xref>; <xref ref-type="bibr" rid="c60">Shushruth et al., 2022</xref>). To test this prediction, we fit the data with a generative model of choice and RTs. We used a sequential sampling model, which explains decisions as the outcome of a process of sequential sampling that stops when the accumulation of evidence satisfies a bound. This model explains RTs as jointly influenced by participant’s efficacy in deliberating about Δ-uncertainty, and their tendency to deliberate longer vs. make quick responses (<xref ref-type="bibr" rid="c49">Ratcliff and McKoon, 2008</xref>; <xref ref-type="bibr" rid="c57">Shadlen and Kiani, 2013</xref>; <xref ref-type="bibr" rid="c58">Shadlen and Shohamy, 2016</xref>). One prediction of sequential sampling theory is that greater deliberation efficacy should be manifested as as greater dependence of RT on absolute Δ-uncertainty (<xref ref-type="bibr" rid="c43">Palmer et al., 2005</xref>).</p>
<p>We found that RTs indeed varied in relation to the absolute value of Δ-uncertainty as expected b=0.69, 95% PI=[0.58,0.78] (Appendix 3—<xref ref-type="table" rid="app3-tbl9">Table 9</xref>). Crucially, a strong dependence of RT on the absolute value of Δ-uncertainty predicted better performance at test b=0.81, 95% PI=[0.58,1.07]. We further found that participants who tended to deliberate longer for the sake of accuracy also tended to perform better at test b=1.46, 95% PI=[0.58,2.34] (<xref ref-type="fig" rid="fig7">Figure 7c</xref>, Appendix 3—<xref ref-type="table" rid="app3-tbl10">Table 10</xref>). In summary, participants who were better at deliberating about uncertainty during exploration, and who deliberated for longer, performed better at test. Thus, making good exploratory choices that lead to efficient learning involves prolonged deliberation.</p>
</sec>
<sec id="s2-1-3">
<title>Deliberation is Reduced by Choice Repetition</title>
<p>We have shown that when overall uncertainty is high participants avoid uncertainty rather than approach it, and that they do not pay a learning cost as a result. It remains to be shown that participants’ alternative strategy is able to shorten the time spent deliberating. Unfortunately, we could not test for such a benefit by directly comparing RTs as a function of overall uncertainty, as overall uncertainty is related to the difficulty of making an exploratory choice. With a single independent variable, deconfounding the effect of difficulty from the strategies used to ameliorate it is impossible. Fortunately, we could take advantage of a conceptually-related but independent tendency we observed in our dataset to examine the benefits of reduced deliberation times.</p>
<p>As in many learning tasks, participants in our task tended to repeat their previous choice (<xref ref-type="bibr" rid="c70">Wu etal.,2022</xref>),a tendency that was in dependent of Δ-uncertainty or overall uncertainty. We observed that participants generally preferred to re-choose the table they had last chosen (<xref ref-type="fig" rid="fig8">Figure 8b</xref>). We corroborated this with a multilevel regression model controlling for the effects of Δ-uncertainty and overall uncertainty b=0.50, 95% PI=[0.42,0.59] (Appendix 3—<xref ref-type="table" rid="app3-tbl11">Table 11</xref>). Crucially, the tendency to repeat choices was also reflected in RTs, which for repeat choices were less related to Δ-uncertainty (b=-0.32, 95% PI=[-0.43,-0.22]). We also found that participants tended to make repeat choices more quickly rather than deliberate longer (b=-0.05, 95% PI=[-0.05,-0.04]; <xref ref-type="fig" rid="fig8">Figure 8c</xref>, Appendix 3— <xref ref-type="table" rid="app3-tbl12">Table 12</xref>).</p>
<p>As in other aspects of exploration strategy, we observed considerable individual differences in the tendency to repeat previous choices. These differences were associated with the uncertainty based aspects of exploration discussed above (<xref ref-type="fig" rid="fig8">Figure 8d</xref>). Participants with a general tendency to repeat choices show stronger uncertainty avoidance when overall uncertainty is high, indicating that these two conceptually related strategies also co-occur in the population r=-0.60, 95% PI=[-0.74,-0.43] (Appendix 3—<xref ref-type="table" rid="app3-tbl11">Table 11</xref>). Furthermore, the tendency to repeat previous choices is associated with better test performance, logistic regression b=0.09, 95% PI=[0.07,0.11] (Appendix 3—<xref ref-type="table" rid="app3-tbl13">Table 13</xref>). The tendency to repeat is also correlated with a stronger baseline tendency to approach uncertainty r=0.32, 95% PI=[0.17,0.46] (Appendix 3—<xref ref-type="table" rid="app3-tbl11">Table 11</xref>), which was shown above to be correlated with test performance. Thus, while from a normative point of view repeating the previous choice appears to be a context-insensitive heuristic, in practice participants who use this strategy do not learn any worse.</p>
</sec>
</sec>
<sec id="s2-6">
<title>Forgetting as a Conceptual Control</title>
<p>Explaining participants’ deviation from the optimal exploration strategy as rational is interesting only to the extent that rationality is not a forgone conclusion. Is the alternative hypothesis of a failure in decision making also a-priori plausible?. We turned to forgetting as a second source of difficulty in our task and a conceptual control condition. Due to the random presentation of choice options, there was variability in the number of trials passed since either of the presented tables was last explored. We assumed that choosing between tables that had not been explored fora long time is more difficult than between tables for which evidence has been recently observed. Indeed, we found that RTs were longer with a larger lag, indicating greater difficulty of making a choice (log normal regression b=0.02, 95% PI=[0.02, 0.03]; <xref ref-type="fig" rid="fig9">Figure 9a</xref>, Appendix 3—<xref ref-type="table" rid="app3-tbl14">Table 14</xref>). Furthermore, we observed that exploration choices on trials with a greater lag depended less on Δ-uncertainty b=- 0.08,95% PI=[-0.11,-0.04], and that the tendency to repeat the last chosen table on these trials was also diminished b=-0.13, 95% PI=[-0.15, -0.11] (<xref ref-type="fig" rid="fig9">Figure 9b</xref>, Appendix 3—<xref ref-type="table" rid="app3-tbl15">Table 15</xref>). Finally, on trials with a large lag the difference in RTs between making a repeat and a switch choice disappeared, interaction b=0.02, 95% PI=[0.02,0.03] (<xref ref-type="fig" rid="fig9">Figure 9a</xref>, Appendix 3—<xref ref-type="table" rid="app3-tbl14">Table 14</xref>). These patterns suggest that prior evidence is forgotten with increasing lag and that as a consequence exploration becomes more random. Hence, in contrast to the systematic effect of overall uncertainty, forgetting results in a failure to make principled exploratory choices.</p>
<fig id="fig7" position="float" fig-type="figure">
<label>Figure 7.</label>
<caption><title>Individuals who spend time deliberation during exploration make strategic choices and learn well.</title>
<p>Participants varied not only in the pattern of their choices, but also in their RTs. <bold>a</bold>, Data from three example participants. The relationship of choice and RTs with Δ-uncertainty weakens from left to right. Data plotted as mean ±SE. <bold>b</bold>, These individual differences were captured by a sequential sampling model, explaining choices and RTs as the interaction between participant’s efficacy of deliberating about Δ-uncertainty and their tendency to deliberate longer vs. make quick responses. Plotting model predictions, we observe a u-shaped dependence of RTs on Δ-uncertainty for participants whose performance at test was in the top accuracy tertile. This characteristic u-shape is indicative of decisions made by prolonged deliberation. This relationship is weaker for participants in the bottom two test accuracy tertiles. Such participants also exhibit shorter RTs overall. Lines mark mean predictions from a sequential sampling model fit by tertiles for visualization, ribbons denote 50% PIs. <bold>c</bold>, Correlating the sequential sampling model parameters with test performance confirms these observations. Participants with a stronger dependence of RT on Δ–uncertainty perform better at test, as do participants who deliberate longer for the sake of accuracy. Example participants from <bold>a</bold> are marked in red. Lines are mean predictions from a logistic regression model.</p>
<p><bold>Figure 7—figure supplement 1.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig7.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS7-1" position="float" fig-type="figure">
<label>Figure 7—figure supplement 1.</label>
<caption><title>Reproducing the analysis using the preliminary sample: Individuals who spend time deliberation during exploration make strategic choices and learn well.</title>
<p>Participants varied not only in the pattern of their choices, but also in their RTs. <bold>a</bold>, Data from three example participants. The relationship of choice and RTs with Δ-uncertainty weakens from left to right. Data plotted as mean ±SE. <bold>b</bold>, These individual differences were captured by a sequential sampling model, explaining choices and RTs as the interaction between participant’s efficacy of deliberating about Δ-uncertainty and their tendency to deliberate longer vs. make quick responses. Plotting model predictions, we observe a u-shaped dependence of RTs on Δ-uncertainty for participants whose performance attest was in the top accuracy tertile. This characteristic u-shape is indicative of decisions made by prolonged deliberation. This relationship is weaker for participants in the bottom two test accuracy tertiles. Such participants also exhibit shorter RTs overall. Lines mark mean predictions from a sequential sampling model fit by tertiles for visualization, ribbons denote 50% PIs. <bold>c</bold>, Correlating the sequential sampling model parameters with test performance confirms these observations. Participants with a stronger dependence of RT on Δ-uncertainty perform better at test, as do participants who deliberate longer for the sake of accuracy. Example participants from a are marked in red. Lines are mean predictions from a logistic regression model.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS7-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig8" position="float" fig-type="figure">
<label>Figure 8.</label>
<caption><title>Participants tend to repeat previous choices instead of deliberating over uncertainty.</title>
<p><bold>a</bold>, On a given trial one table has been chosen more recently than the other (frames denote previous choices). In the example the green table had been chosen more recently, hence it is designated the repeat option and the other table the switch option. <bold>b</bold>, Participants tend to choose the table displayed on the right more often when it is the repeat option than when it is the switch option. Data plotted as mean ±SE, n=194 participants. <bold>c</bold>, When choosing a repeat option, participants’ RTs are shorter and less dependent on -uncertainty. Lines mark mean predictions from a sequential sampling model, ribbons denote 50% PIs. <bold>d</bold>, Participants who tended to repeat their previous choice also tended to perform better at test (left), were more likely to have a stronger baseline tendency to approach uncertainty (middle), and a stronger tendency to avoid uncertainty when overall uncertainty is high (right). Regression lines are plotted for visualization.</p>
<p><bold>Figure 8—figure supplement 1.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig8.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS8-1" position="float" fig-type="figure">
<label>Figure 8—figure supplement 1.</label>
<caption><title>Reproducing the analysis using the preliminary sample: Participants tend to repeat previous choices instead of deliberating over uncertainty.</title>
<p><bold>a</bold>, On a given trial one table has been chosen more recently than the other (frames denote previous choices). In the example the green table had been chosen more recently, hence it is designated the repeat option and the other table the switch option. <bold>b</bold>, Participants tend to choose the table displayed on the right more often when it is the repeat option than when it is the switch option. Data plotted as mean ±SE, n=194participants. <bold>c</bold>, When choosing a repeat option, participants’ RTs are shorter and less dependent on Δ-uncertainty. Lines mark mean predictions from a sequential sampling model, ribbons denote 50% PIs. <bold>d</bold>, Participants who tended to repeat their previous choice also tended to perform better at test (left), were more likely to have a stronger baseline tendency to approach uncertainty(middle),and a stronger tendency to avoid uncertainty when overall uncertainty is high (right). Regression lines are plotted for visualization.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS8-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="fig9" position="float" fig-type="figure">
<label>Figure 9.</label>
<caption><title>Forgetting is associated with random choice rather than a systematic bias.</title>
<p><bold>a</bold>, Memory lag, defined as trials since last choice, serves as a proxy for forgetting and contributes to the difficulty of making an exploratory choice. RTs rise with memory lag. The RT advantage for repeat choices disappears with higher memory lag. <bold>b</bold>, With higher memory lag choices become less dependent on Δ-uncertainty, as indicated by flatter curves. The tendency to repeat the last choice is also diminished with memory lag. Both effects amount to choice becoming more random due to forgetting. Data plotted as mean ±SE, n=194 participants.</p>
<p><bold>Figure 9—figure supplement 1.</bold> Matching results in the preliminary sample.</p>
</caption>
<graphic xlink:href="31234.gtxam_fig9.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
<fig id="figS9-1" position="float" fig-type="figure">
<label>Figure 9—figure supplement 1.</label>
<caption><title>Reproduction the analysis using the preliminary sample.</title>
<p>Forgetting is associated with random choice rather than a systematic bias. <bold>a</bold>, Memory lag, defined as trials since last choice, serves as a proxy for forgetting and contributes to the difficulty of making an exploratory choice. RTs rise with memory lag. The RT advantage for repeat choices disappears with higher memory lag. <bold>b</bold>, With higher memory lag choices become less dependent on Δ-uncertainty, as indicated by flatter curves. The tendency to repeat the last choice is also diminished with memory lag. Both effects amount to choice becoming more random due to forgetting. Data plotted as mean ±SE, n=194 participants.</p>
</caption>
<graphic xlink:href="31234.gtxam_figS9-1.tif" mime-subtype="tiff" mimetype="image"/>
</fig>
</sec>
</sec>
<sec id="s3" sec-type="discussion">
<title>Discussion</title>
<p>We examined the cognitive computations behind exploratory choices using a paradigm that encourages incremental learning in the service of a distant goal. We found that uncertainty played an important role in guiding participants’ choices about how to sample their environment for learning. In general, participants chose to learn more about the options they were more uncertain about. However, when overall uncertainty was especially high, participants instead avoided the more uncertain options and sampled the options they already knew more about. In addition, we found that participants tended to repeat previous choices. Together, this pattern suggests that participants systematically balance approaching and avoiding uncertainty while exploring.</p>
<p>Examining individual differences in exploration and learning revealed the costs and benefits of avoiding uncertainty when exploring. We found that strategically avoiding uncertainty is not associated with a detriment to learning, even though it slows down the rate of information intake. We also found an association between the length of deliberation and learning efficiency. Participants who deliberated longer also learned better, and deliberation time could be shortened by repeating previous choices. Based on these results, we conclude that balancing approaching and avoiding uncertainty is a way to manage cognitive resources by regulating deliberation costs. In this sense, our results serve as an example of how human cognition is adapted to the inherent constraints of the human mind, consistent with the resource rationality framework (<xref ref-type="bibr" rid="c36">Lieder and Griffiths, 2020</xref>).</p>
<p>While the literature on exploration is expansive, the paradigm presented here extends it in important ways. Researchers of reinforcement learning have previously examined how exploration manifests when agents learn incrementally about their environment. Crucially, this literature has focused on cases where reward can be gained on each trial (<xref ref-type="bibr" rid="c10">Brown et al., 2022</xref>; <xref ref-type="bibr" rid="c17">Cohen et al., 2007</xref>; <xref ref-type="bibr" rid="c21">Daw et al., 2006</xref>; <xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c61">Song et al., 2019</xref>; <xref ref-type="bibr" rid="c65">Tversky and Edwards, 1966</xref>; <xref ref-type="bibr" rid="c69">Wilson et al., 2014</xref>; <xref ref-type="bibr" rid="c70">Wu et al., 2022</xref>). In contrast, our task was designed to remove the impetus to exploit current knowledge immediately, a motivation that predominates exploration in tasks with immediate reward. Accordingly, we were able to observe many exploratory choices and had greater experimental power to describe in detail how participants approach uncertainty and when they avoid it instead. Secondly, exploration has been studied in the information search literature (<xref ref-type="bibr" rid="c31">Gureckis and Markant, 2012</xref>; <xref ref-type="bibr" rid="c39">Markant and Gureckis, 2014</xref>; <xref ref-type="bibr" rid="c41">Oaksford and Chater, 1994</xref>; <xref ref-type="bibr" rid="c45">Petitet et al., 2021</xref>; <xref ref-type="bibr" rid="c50">Rothe et al., 2018</xref>; <xref ref-type="bibr" rid="c51">Ruggeri et al., 2017</xref>). In most studies of this field participants make decisions without relying on their memory, as the entire history of learning is displayed to them on screen (cf. related work in active sensing; Yang et al., 2016). This differs from our task, which places heavy demands on memory. Rather than treating capacity limitations as a source of noise and a nuisance to measurement, we find that the rational use of limited resources is central for successful exploration.</p>
<p>Several previous studies of exploration inspired us to design a task with separate exploration and test phases. Using the “observe or bet” paradigm, Tversky and Edwards (1966) examined how participants trade off exploration and exploitation on a trial-by-trial basis. By using a short block of exploration followed by a test, Wilson et al. (2014) achieved a first reliable demonstration of directed exploration in humans. Finally, the expansive literature on the description-experience gap (<xref ref-type="bibr" rid="c71">Wulff et al., 2018</xref>) has used a similar paradigm to examine when participants choose to self terminate their exploration, and how that affects their learning. The paradigm presented here extends these approaches, as it is crafted to reveal the strategy driving each exploration choice.</p>
<p>We observed considerable individual differences in exploration strategy, as would be expected in a complex task requiring memory-based learning and inference. In the face of such variability, one may question the prudence of drawing conclusions about the population, since the average might be a poor summary of a plurality of idiosyncratic strategies. However, the strong correlation we observed between individual differences in exploration and test performance mitigates this concern. The correlation suggests that participants who were engaged with this task and able to learn from observation can well be described as exploring by a strategy of combining approaching and avoiding uncertainty. The relationship between test performance and RTs lends additional mechanistic support to this idea.</p>
<p>Our theoretical analysis and experiments leave several questions open. First, overall uncertainty in our task was correlated with the number of cards observed. While our results hold when trial number is added as a covariate to the regression models (see Appendix 3—<xref ref-type="table" rid="app3-tbl16">Table 16</xref>), future work orthogonalizing overall uncertainty and time on task would help to fully disentangle the contribution of each factor to uncertainty avoidance.</p>
<p>Another open question is the nature of the limitation driving participants to avoid uncertainty when overall uncertainty is high. This could be due to limitations in committing prior experiences to memory, inferring latent parameters from disparate experiences, retrieving prior knowledge, or estimating the uncertainty of existent knowledge. While the idea that decisions based on high overall uncertainty are more difficult has been raised previously (<xref ref-type="bibr" rid="c54">Schulz and Gershman, 2019</xref>; <xref ref-type="bibr" rid="c59">Shafir, 1994</xref>), an explanation grounded in cognitive mechanisms is still needed. Accordingly, the mechanism by which uncertainty avoidance ameliorates choice difficulty remains unknown.</p>
<p>One intriguing explanation for the source of difficulty and the way it is managed lies in the distinction between strategies dependent on remembering single experiences and those dependent on the incremental acumulation of knowledge in the form of summary statistics (<xref ref-type="bibr" rid="c19">Collins et al., 2017</xref>; <xref ref-type="bibr" rid="c18">Collins and Frank, 2012</xref>; <xref ref-type="bibr" rid="c20">Daw et al., 2005</xref>; <xref ref-type="bibr" rid="c35">Knowlton et al., 1996</xref>; <xref ref-type="bibr" rid="c46">Plonsky et al., 2015</xref>; <xref ref-type="bibr" rid="c47">Poldrack et al., 2001</xref>). Both strategies could contribute to performance in tasks such as ours. A participant may be encoding prior observations as single instances, or summarizing them into a central tendency with a margin of uncertainty around it. Crucially, each strategy is associated with a different profile of cognitive resource use. Keeping track of individual experiences is much costlier than tracking a single expectation and a confidence interval around it (<xref ref-type="bibr" rid="c20">Daw et al., 2005</xref>; <xref ref-type="bibr" rid="c40">Nicholas et al., 2022</xref>) and more likely to incur costs when switching between exploring different tables. Prior work suggests individuals use single experiences or summary statistics according to the reliability of each strategy, and the cost of using it (<xref ref-type="bibr" rid="c20">Daw et al., 2005</xref>; <xref ref-type="bibr" rid="c40">Nicholas et al., 2022</xref>). In our case, summary statistics may be perceived as unreliable when overall uncertainty is high, compelling participants to rely on committing individual experiences to working memory (<xref ref-type="bibr" rid="c6">Bavard et al., 2021</xref>; <xref ref-type="bibr" rid="c19">Collins et al., 2017</xref>; <xref ref-type="bibr" rid="c20">Daw et al., 2005</xref>; <xref ref-type="bibr" rid="c23">Duncan et al., 2019</xref>; <xref ref-type="bibr" rid="c47">Poldrack et al., 2001</xref>). Furthermore, recent work examining how humans make a series of dependent decisions demonstrates that the tension between remembering single experiences and discarding them in favor of summary statistics is accompanied by a tendency to revisit previous choices instead of switching to new alternatives (<xref ref-type="bibr" rid="c73">Zylberberg, 2021</xref>).</p>
<p>The questions we addressed here were partly motivated by the well-established observation that humans and animals often avoid uncertainty in various situations. Two broad categories of explanation for such avoidance have been proposed (<xref ref-type="bibr" rid="c29">Golman et al., 2017</xref>). First, individuals avoid resolving uncertainty when it could lead to negative news, for example by avoiding ambiguous prospects when making economic choices (<xref ref-type="bibr" rid="c25">Ellsberg, 1961</xref>; <xref ref-type="bibr" rid="c26">Fox and Tversky, 1995</xref>). An extension of this idea is dread avoidance (<xref ref-type="bibr" rid="c27">Gigerenzer and Garcia-Retamero, 2017</xref>; <xref ref-type="bibr" rid="c29">Golman et al., 2017</xref>). One might avoid resolving the uncertainty about a medical diagnosis to avoid the unpleasant affective response to the news, even if the information could be very useful in determining treatment. Relatedly, humans might avoid uncertainty as a by-product of pursuing a goal other than exploration. More uncertain options are often avoided for the sake of choosing immediately rewarding options (<xref ref-type="bibr" rid="c64">Trudel et al., 2020</xref>; <xref ref-type="bibr" rid="c69">Wilson et al., 2014</xref>) — why try an unknown dish, when your absolute favorite is on the menu? Lastly, uncertainty avoidance may be a strategy for managing conflict between different motivations, or different mechanisms of action selection (<xref ref-type="bibr" rid="c14">Carrillo and Mariotti, 2000</xref>; <xref ref-type="bibr" rid="c29">Golman et al., 2017</xref>). For example, to maintain their diet, an individual might choose to avoid resolving the uncertainty about what snacks can be found in the office kitchen. Our findings highlight a different kind of strategic uncertainty avoidance. In our tasks there were no negative consequences to learning about the color proportions of card decks, and no conflicting motivations. Rather, we explain participants’ tendency to avoid uncertainty in terms of managing their limited cognitive resources.</p>
<p>The idea of a balance between approaching and avoiding uncertainty has conceptual parallels in other literatures. A group of relevant findings concern how animals explore their proximal environment. A classic finding in rats is that when placed in a novel open arena, they alternate between the exploration strategy of walking around the arena (uncertainty approaching) and a strategy of returning to their initial position and pausing there (termed “home base” behavior, which is uncertainty avoiding; Eilam and Golani, 1989). Relatedly, by using computational models to understand how rats use their whiskers to explore near objects, researchers have identified an alternation between uncertainty approaching and avoiding strategies (<xref ref-type="bibr" rid="c30">Gordon et al., 2014</xref>). Recent work in mice and primates has uncovered neural circuits driving exploration by framing the problem of exploration as striking a balance between approach and avoidance (<xref ref-type="bibr" rid="c2">Ahmadlou et al., 2021</xref>; <xref ref-type="bibr" rid="c9">Botta et al., 2020</xref>; <xref ref-type="bibr" rid="c42">Ogasawara et al., 2022</xref>). Our findings highlight the shared computational principles between human exploration in symbolic space and animal exploration of the physical environment and suggest that mechanisms involved in avoidance responses may also play a part in knowledge acquisition.</p>
<p>Finally, planning (<xref ref-type="bibr" rid="c34">Hunt et al., 2021</xref>), learning (<xref ref-type="bibr" rid="c31">Gureckis and Markant, 2012</xref>), and sensing (<xref ref-type="bibr" rid="c30">Gordon et al., 2014</xref>; <xref ref-type="bibr" rid="c72">Yang et al., 2016</xref>) are increasingly studied as active processes, situated within our environment and interacting with it. Understanding the complicated dynamics between agent and environment has been greatly facilitated by comparing behavior against the computational ideal of maximizing the amount of information observed (<xref ref-type="bibr" rid="c41">Oaksford and Chater, 1994</xref>; shman, 2019; <xref ref-type="bibr" rid="c55">Schwartenbeck et al., 2019</xref>). The findings we present here suggest a modification to this computational premise. Rather than trying to uncover as much information as possible, the goal of human exploration may be to maximize the amount of information retained in memory, by modulating the rate and order of observed information.</p>
</sec>
<sec id="s4" sec-type="methods">
<title>Methods</title>
<sec id="s4-1">
<title>Data Collection and Participants</title>
<p>A sample of 298 participants was recruited via Amazon MTurk to participate in four sessions of the exploration task. They were paid $3.60 for each session and earned a bonus contingent on their test phase performance, adding up to $4.50 for the first session and $6 for later sessions. Additionally, a $2 bonus was paid out for completion of the fourth session. Participants were asked to complete the four sessions over the course of a week and were invited by email to each session after the first, as long as the data from their last session was not excluded according to the criteria we had specified (see below). All participants provided informed consent; all protocols were approved by the Columbia University Institutional Review Board.</p>
<p>The first session was terminated early for 89 participants due to recorded interactions with other applications during the experiment or failure to comply with instructions. An additional 32 sessions played by participants who had successfully completed the first session were excluded for the same reasons. One participant was excluded after reporting technical problems with stimulus presentation in the second session. Twenty-seven further sessions were excluded for failure to sample cards from both decks, a prerequisite for learning on which participants were instructed as part of the training. Altogether, data from 194 participants was included in the analyzed sample (120 female, 72 male, 2 other gender, average age 29.63, range 20-48). This sample included 194 first sessions, 156 second sessions, 129 third sessions, and 116 fourth sessions.</p>
<p>Before running this experiment, we pre-registered (<xref ref-type="bibr" rid="c1">Abir et al., 2021</xref>) a sample size of 190 participants satisfying our exclusion criteria. We chose this number to be three times larger than a preliminary sample of N=62 participants, which provided the dataset we used to develop our analysis approach and pipeline, and first identify exploration strategies as described above. Results for the preliminary sample are provided in figure supplements.</p>
</sec>
<sec id="s4-2">
<title>Task Design and Procedure</title>
<p>On each round of the exploration task participants were presented with a simple environment of four tables with two decks of cards on each table. Tables were distinguished by unique colorful patterns and decks by geometric symbols that did not repeat within an experimental session. The hidden side of each card was painted in one of two colors, with a unique color pair for each round. The proportion of colors in each deck were determined pseudo-randomly (see Supplementary Information), resulting in variability in the difference in proportion between each deck pair - the learning desideratum of this task.</p>
<p>At the beginning of each round, participants were first presented with the color pair for the round, and then with the table-deck assignments. Participants then had to pass a multiple-choice test on the table-deck assignment, making sure they remembered the structure of the task before proceeding to explore. Failing to get a perfect score on this test resulted in repeating this phase. The exploration phase then commenced. Trial structure for the exploration phase is depicted in <xref ref-type="fig" rid="fig1">Figure 1b</xref>. The lengths of the exploration phases varied from round to round. They were sampled from a geometric distribution with rate, shifted by 10 trials. The same list of round lengths was used for all participants, but their order was randomized.</p>
<p>Following the exploration phase, participants were tested on their learning. They were presented with the rewarding color for this round, and then had to indicate which deck had a greater proportion of that color on each table (<xref ref-type="fig" rid="fig1">Figure 1c</xref>). After answering this question for each of the four tables, they rated their confidence in each of the four choices on a 1-5 Likert scale. Participants were then told whether each of the test choices were correct, and the true color proportions for the two decks on each table were presented to them as 10 open cards.</p>
<p>The first session started with extensive instructions explaining the structure of each of the two phases of the task and clearly stating the learning goal. Participants were also instructed on the independence of color proportion within each deck pair, necessitating sampling from both decks to succeed in the task. The instructions also included training on how to make the relevant choices in each of the two stages. A quiz followed the instruction phase, and participants had to repeat reading the instructions if they had given the wrong response to any question on this quiz.</p>
<p>Each session started with a short practice round (12-19 trials). Data from this round was excluded from analysis. In the first session participants then played three more rounds and in later sessions five more rounds, for a total of 18 experimental rounds.</p>
</sec>
<sec id="s4-3">
<title>Data Analysis</title>
<p>Analysis was performed using Julia 1.4.2. Hierarchical regression models were fitted using the Stan probabilistic programming language 2.30.1 (<xref ref-type="bibr" rid="c13">Carpenter et al., 2017</xref>), using the interface supplied by the brims package version 2.16.1 (<xref ref-type="bibr" rid="c11">Burkner, 2017</xref>), running on top of R 4.1.2. The complete computing environment was packaged as a Docker image, which can be used to reproduce the entire analysis pipeline. Sequential sampling models were fitted on a separate Docker image (<xref ref-type="bibr" rid="c16">Chuan- Peng et al., 2022</xref>) containing HDDM 0.8 (<xref ref-type="bibr" rid="c68">Wiecki et al., 2013</xref>) running on top of python 3.8.8.</p>
<sec id="s4-1-1">
<title>Bayesian Observer</title>
<p>Each of the three hypothesized strategies for exploration postulates a different summary statistic of prior learning as the driver of exploratory choice. To derive these summary statistics, we first had to construct a model of prior learning. We chose a simple Bayesian observer model (<xref ref-type="bibr" rid="c7">Behrens et al., 2007</xref>; <xref ref-type="bibr" rid="c72">Yang et al., 2016</xref>). Like our participants, this model’s goal was to learn <italic>θ = sgn(π<sub>1</sub> – π<sub>2</sub>)</italic> from observed outcomes <italic>x<sub>0:t</sub>.</italic> It did so by placing a probabilistic prior over the value of each updating it after every observation according to Bayes’ rule, and solving for <italic>θ</italic> using the rules of probability. The result is a posterior distribution capturing the agent’s expectation of the value of <italic>θ,</italic> and their uncertainty about the expectation. This process is depicted in <xref ref-type="fig" rid="fig2">Figure 2</xref> for two tables and their matching pairs of decks.</p>
<p>This computation can be put into formulaic form as follows. At the beginning of a round, the Bayesian observer places a flat Beta distribution prior on the proportion of colors in each of the eight decks:</p>
<disp-formula id="FD1">
<alternatives>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>π</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>~</mml:mo><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow>
</mml:math>
<graphic xlink:href="31234.gtxam_eqn1.tif" mime-subtype="tiff" mimetype="image"/>
</alternatives>
</disp-formula>
<p>After observing a card, this prior would be updated to form a posterior distribution. Since the posterior of a Beta prior and a Bernoulli observation likelihood is also a Beta distribution, the posterior has a simple analytic form: after completing t trials, observing <italic>c<sub>i</sub></italic> cards of one color and <italic>t – c<sub>i</sub></italic> cards of the other color, the posterior would</p>
<disp-formula id="FD2">
<alternatives>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>π</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
<mml:mo>~</mml:mo><mml:mi>B</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mfenced>
<mml:mrow>
<mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>t</mml:mi><mml:mo>−</mml:mo><mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced></mml:mrow>
</mml:math>
<graphic xlink:href="31234.gtxam_eqn2.tif" mime-subtype="tiff" mimetype="image"/>
</alternatives>
</disp-formula>
<p>We can then find the probability that θ = 1, i.e. that <italic>π</italic><sub>1</sub><italic> &gt; π</italic><sub>2</sub><italic>,</italic> by calculative the probability that π<sub>2</sub> is smaller than a given π<sub>1</sub><italic> = z,</italic> and integrating over <italic>z</italic>, the possible values of π<sub>1</sub>:</p>
<disp-formula id="FD3">
<alternatives>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>P</mml:mi><mml:mfenced>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced><mml:mo>=</mml:mo><mml:mstyle displaystyle="true">
<mml:mrow>
<mml:msubsup>
<mml:mo>∫</mml:mo>
<mml:mn>0</mml:mn>
<mml:mn>1</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>π</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mstyle><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>π</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:mrow>
</mml:math>
<graphic xlink:href="31234.gtxam_eqn3.tif" mime-subtype="tiff" mimetype="image"/>
</alternatives>
</disp-formula>
<p>Where <italic>f</italic> is the Beta probability density function, <italic>F</italic> is the Beta cumulative density function, and <italic>x<sub>0:t</sub></italic> are observations thus far. We computed the value of this integral numerically using the Julia package QuadGK.jl. Finally, <italic>θ</italic> can only take two values, and so</p>
<disp-formula id="FD4">
<alternatives>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>P</mml:mi><mml:mfenced>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>=</mml:mo><mml:mo>−</mml:mo><mml:mn>1</mml:mn><mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>−</mml:mo><mml:mi>P</mml:mi><mml:mfenced>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced></mml:mrow>
</mml:math>
<graphic xlink:href="31234.gtxam_eqn4.tif" mime-subtype="tiff" mimetype="image"/>
</alternatives>
</disp-formula>
</sec>
<sec id="s4-1-2">
<title>Computing Hypothesized Decision Variables</title>
<p>The theory of decision making defines a decision variable as the quantity evaluated by the decision maker in order to choose between two choice options (<xref ref-type="bibr" rid="c57">Shadlen and Kiani, 2013</xref>). The difficulty of the decision should scale with the absolute value of the decision variable. Each of the three hypothesized strategies is defined by a specific summary statistic of prior learning that might serve as the decision variable for an exploratory choice. The three summary statistics are given in <xref ref-type="fig" rid="fig1">Figure 1e</xref>.</p>
<p>Both EIG and uncertainty are derived from the uncertainty of the posterior for <italic>θ</italic> as defined above. We quantified uncertainty as the entropy of the posterior belief (<xref ref-type="bibr" rid="c38">MacKay, 1992</xref>; <xref ref-type="bibr" rid="c41">Oaksford and Chater, 1994</xref>; <xref ref-type="bibr" rid="c72">Yang et al., 2016</xref>).</p>
<disp-formula id="FD5">
<alternatives>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:mi>H</mml:mi><mml:mfenced>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced><mml:mo>=</mml:mo><mml:mo>−</mml:mo><mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>∑</mml:mo>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>=</mml:mo><mml:mo>−</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow>
</mml:munder>
<mml:mi>P</mml:mi>
</mml:mstyle><mml:mfenced>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced><mml:mi>ln</mml:mi><mml:mi>P</mml:mi><mml:mfenced>
<mml:mrow>
<mml:mi>θ</mml:mi><mml:mo>∣</mml:mo><mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>0</mml:mn><mml:mo>:</mml:mo><mml:mi>t</mml:mi></mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced></mml:mrow>
</mml:math>
<graphic xlink:href="31234.gtxam_eqn5.tif" mime-subtype="tiff" mimetype="image"/>
</alternatives>
</disp-formula>
<p>Entropy takes the unit of nats, ranging from 0 should the participant be absolutely sure about the value of <italic>θ</italic> for both table choice options, to 0.69 when they know nothing about a table. This is the equivalentof1bitofinformation, were we to replace the natural logarithm with a base 2logarithm.</p>
</sec>
<sec id="s4-1-3">
<title>Estimating Multilevel Bayesian Models for Inference</title>
<p>The regression coefficients and PIs reported here were all estimated using multilevel regression models accounting for individual differences in behavior. We used regularizing priors for all coefficients of interest to facilitate robust estimation (<xref ref-type="table" rid="app3-tbl1">Table 1</xref>). For RT data we selected informative priors for the interceptterm in the regression (capturing the grand average of RTs)following established recommendations (<xref ref-type="bibr" rid="c53">Schad et al., 2021</xref>). For predicting choices, we used logistic regression, for confidence ratings we used ordinal-logistic regression, and for average RTs we used log normal regression. We estimated these models with Hamiltonian Monte Carlo implemented in the Stan probabilistic programming language using the R package brms. Three Monte-Carlo chains were run for each model, collecting 1000 samples each after a warm up period of at least 1000 samples (warm up was extended if convergence had not been reached). Sequential sampling models were estimated using slice sampling, implemented in the python package HDDM. Four Monte-Carlo chains were run for each model, collecting 2000 samples each after a warm up period of at least 2000 samples. Convergence for both model types was assessed using the Ȓ metric, and visual inspection of trace plots. R syntax formulae and coefficients for covariates for all models mentioned in the main text are reported in Supplementary Information.</p>
</sec>
<sec id="s4-1-4">
<title>Sequential Sampling Model of Reaction Times</title>
<p>To draw inference from participants’ RTs we turned to the sequential sampling theory of deliberation and choice. This theory encompasses a family of models in which decisions arise through a process of sequential sampling that stops when the accumulation of evidence satisfies a threshold or bound (<xref ref-type="bibr" rid="c43">Palmer et al., 2005</xref>; <xref ref-type="bibr" rid="c57">Shadlen and Kiani, 2013</xref>). From this family of models we chose to use the drift diffusion model (DDM) to fit our data, as it is very well described and extensively studied (<xref ref-type="bibr" rid="c49">Ratcliff and McKoon, 2008</xref>; <xref ref-type="bibr" rid="c57">Shadlen and Kiani, 2013</xref>). The DDM explains RTs as the culmination of three interpretable terms. The first is the efficacy of a participant’s thought process in furnishing relevant evidence for the decision - in our case the efficacy of calculating Δ-uncertainty (the drift rate in DDM parlance). The second term governs the participant’s speed-accuracy tradeoff by determining how much evidence they require to commit to a decision. This can also be thought of as how long a participant is willing to deliberate when a decision is difficult(bound height). Finally, the portion of the RT not linked to the deliberation process is captured by a third term (non-decision time). Since behavior was considerably different when overall uncertainty was high, DDM models were fit excluding trials with total uncertainty above the participant’s estimated threshold.</p>
<table-wrap id="tbl1" orientation="portrait" position="float">
<label>Table 1.</label>
<caption><p>Regularizing priors used in regression models.</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_tbl1.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Type of coefficient</th>
<th align="left" valign="top">Prior for logistic and ordered-logistic regression</th>
<th align="left" valign="top">Prior for lognormal regression (RTs; following <xref ref-type="bibr" rid="c52">Schad et al. 2019</xref>)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Intercept</td>
<td align="left" valign="top">normal(0,1) (not applicable for ordered logistic models; <xref ref-type="bibr" rid="c12">Bürkner and Vuorre 2019</xref>)</td>
<td align="left" valign="top">normal(-0.25, 0.5)</td>
</tr>
<tr>
<td align="left" valign="top">Group-level effects of predictors</td>
<td align="left" valign="top">normal(0,1)</td>
<td align="left" valign="top">normal(0, 0.5)</td>
</tr>
<tr>
<td align="left" valign="top">Scale of by-participant terms</td>
<td align="left" valign="top">normal(0,1)</td>
<td align="left" valign="top">normal(0, 0.01)</td>
</tr>
<tr>
<td align="left" valign="top">Correlation matrices for by-participant terms</td>
<td align="left" valign="top">LKJ(2)</td>
<td align="left" valign="top">LKJ(2)</td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="tbl1-fn1"><p>Prior distributions are given in Stan syntax. All predictors used in models were centered and scaled prior to fitting, so that the same priors can apply to all parameters.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-1-5">
<title>Model Evaluation</title>
<p>We compared the models of choice and RTs to alternative models, either reduced or expanded (see Supplementary Information). We used the LOO R package to perform approximate leave- one-out cross validation for models implemented in Stan. This method uses pareto-smoothed importance sampling to approximate cross validation in an efficient manner (<xref ref-type="bibr" rid="c66">Vehtari et al., 2017</xref>). Models implemented in HDDM were compared using the DIC metric. We also performed recovery analysis and posterior predictive checks for our models, making sure they capture the theoretically important qualitative features of the data.</p>
</sec>
</sec>
</sec>
</body>
<back>
<ack>
<title>Acknowledgments</title>
<p>We thank the Shohamy lab, Christopher A. Baldassano, and Gabriel M. Stine for their insightful discussion of the project. We are thankful for the support of the Stan user community, especially Matti Vuorre and Paul Burkner. We are grateful for funding support from the NSF (award #1822619 to D.S.), NIMH/NIH (#MH121093to D.S.)and the Templeton Foundation (#60844 to D.S.).</p>
</ack>
<sec id="s5">
<title>Additional Information</title>
<sec id="s5-1">
<title>Author Contributions</title>
<p>Y.A., M.N.S, and D.S. designed research; Y.A. collected and analyzed data; M.N.S and D.S supervised.</p>
</sec>
</sec>
<ref-list>
<title>References</title>
<ref id="c1"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y</given-names> <surname>Abir</surname></string-name>, <string-name><given-names>MN</given-names> <surname>Shadlen</surname></string-name>, <string-name><given-names>D</given-names> <surname>Shohamy</surname></string-name></person-group>, <source>Memory-based incremental exploration in a stochastic environment</source>; <year>2021</year>. <ext-link ext-link-type="uri" xlink:href="https://aspredicted.org/hx6gj.pdf">https://aspredicted.org/hx6gj.pdf</ext-link>.</mixed-citation></ref>
<ref id="c2"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M</given-names> <surname>Ahmadlou</surname></string-name>, <string-name><given-names>JH</given-names> <surname>Houba</surname></string-name>, <string-name><surname>van Vierbergen</surname> <given-names>JF</given-names></string-name>, <string-name><given-names>M</given-names> <surname>Giannouli</surname></string-name>, <string-name><given-names>GA</given-names> <surname>Gimenez</surname></string-name>, <string-name><given-names>C</given-names> <surname>van Weeghel</surname></string-name>, <string-name><given-names>M</given-names> <surname>Darbanfouladi</surname></string-name>, <string-name><given-names>MY</given-names> <surname>Shirazi</surname></string-name>, <string-name><given-names>J</given-names> <surname>Dziubek</surname></string-name>, <string-name><given-names>M</given-names> <surname>Kacem</surname></string-name>, <string-name><surname>others.</surname> <given-names>A</given-names></string-name></person-group> <chapter-title>cell type-specific cortico-subcortical brain circuit for investigatory and novelty-seeking behavior.</chapter-title> <source>Science</source>. <year>2021</year>;<volume>372</volume>(<issue>6543</issue>):<comment>eabe9681</comment>. <publisher-name>Publisher: American Association for the Advancement of Science</publisher-name>.</mixed-citation></ref>
<ref id="c3"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>JR.</given-names> <surname>Anderson</surname></string-name></person-group> <article-title>The adaptive character of thought.</article-title> <source>Psychology Press;</source> <year>1990</year>.</mixed-citation></ref>
<ref id="c4"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Auer</surname></string-name></person-group> <article-title>Using confidence bounds for exploitation-exploration trade-offs.</article-title> <source>Journal of Machine Learning Research</source>. <year>2002</year>; <day>3</day>(<month>Nov</month>):<fpage>397</fpage>–<lpage>422</lpage>.</mixed-citation></ref>
<ref id="c5"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>AP</given-names> <surname>Badia</surname></string-name>, <string-name><given-names>B</given-names> <surname>Piot</surname></string-name>, <string-name><given-names>S</given-names> <surname>Kapturowski</surname></string-name>, <string-name><given-names>P</given-names> <surname>Sprechmann</surname></string-name>, <string-name><given-names>A</given-names> <surname>Vitvitskyi</surname></string-name>, <string-name><given-names>ZD</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Blundell</surname></string-name></person-group> <article-title>Agent57: Outperforming the Atari Human Benchmark</article-title>. In: <source>Proceedings of the 37th International Conference on Machine Learning PMLR;</source> <year>2020</year>. p. <fpage>507</fpage>–<lpage>517</lpage>. <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v119/badia20a.html">https://proceedings.mlr.press/v119/badia20a.html</ext-link>, <comment>iSSN: 2640-3498</comment>.</mixed-citation></ref>
<ref id="c6"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S</given-names> <surname>Bavard</surname></string-name>, <string-name><given-names>A</given-names> <surname>Rustichini</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Palminteri</surname></string-name></person-group> <chapter-title>Two sides of the same coin: Beneficial and detrimental consequences of range adaptation in human reinforcement learning</chapter-title>. <source>Science Advances</source>. <year>2021</year> <month>Apr</month>; <volume>7</volume>(<issue>14</issue>):<comment>eabe0340</comment>. <ext-link ext-link-type="uri" xlink:href="https://www.science.org/doi/full/10.1126/sciadv.abe0340">https://www.science.org/doi/full/10.1126/sciadv.abe0340</ext-link>, doi: <pub-id pub-id-type="doi">10.1126/sciadv.abe0340,</pub-id> <publisher-name>publisher: American Association for the Advancement of Science</publisher-name>.</mixed-citation></ref>
<ref id="c7"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>TEJ</given-names> <surname>Behrens</surname></string-name>, <string-name><given-names>MW</given-names> <surname>Woolrich</surname></string-name>, <string-name><given-names>ME</given-names> <surname>Walton</surname></string-name></person-group>, <person-group person-group-type="author"><string-name><given-names>MFS</given-names> <surname>Rushworth</surname></string-name></person-group>. <chapter-title>Learning the value of information in an uncertain world</chapter-title>. <source>Nature Neuroscience</source>. <year>2007</year> <month>Sep</month>; <volume>10</volume>(<issue>9</issue>):<fpage>1214</fpage>–<lpage>1221</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/nn1954">https://www.nature.com/articles/nn1954</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/nn1954,</pub-id> <comment>number: 9</comment> <publisher-name>Publisher: Nature Publishing Group</publisher-name>.</mixed-citation></ref>
<ref id="c8"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M</given-names> <surname>Bellemare</surname></string-name>, <string-name><given-names>S</given-names> <surname>Srinivasan</surname></string-name>, <string-name><given-names>G</given-names> <surname>Ostrovski</surname></string-name>, <string-name><given-names>T</given-names> <surname>Schaul</surname></string-name>, <string-name><given-names>D</given-names> <surname>Saxton</surname></string-name>, <string-name><surname>Munos</surname> <given-names>R.</given-names></string-name></person-group> <article-title>Unifying count-based exploration and intrinsic motivation.</article-title> <source>Advances in neural information processing systems</source>. <year>2016</year>; <day>29</day>.</mixed-citation></ref>
<ref id="c9"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P</given-names> <surname>Botta</surname></string-name>, <string-name><given-names>A</given-names> <surname>Fushiki</surname></string-name>, <string-name><given-names>AM</given-names> <surname>Vicente</surname></string-name>, <string-name><given-names>LA</given-names> <surname>Hammond</surname></string-name>, <string-name><given-names>AC</given-names> <surname>Mosberger</surname></string-name>, <string-name><given-names>CR</given-names> <surname>Gerfen</surname></string-name>, <string-name><given-names>D</given-names> <surname>Peterka</surname></string-name>, <string-name><surname>Costa</surname> <given-names>RM.</given-names></string-name></person-group> <article-title>An Amygdala Circuit Mediates Experience-Dependent Momentary Arrests during Exploration</article-title>. <source>Cell</source>. <year>2020</year>; <volume>183</volume>(<issue>3</issue>):<fpage>605</fpage>–<lpage>619</lpage>.<comment>e22</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.cell.2020.09.023.</pub-id></mixed-citation></ref>
<ref id="c10"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>VM</given-names> <surname>Brown</surname></string-name>, <string-name><given-names>MN</given-names> <surname>Hallquist</surname></string-name>, <string-name><given-names>MJ</given-names> <surname>Frank</surname></string-name>, <string-name><surname>Dombrovski</surname> <given-names>AY.</given-names></string-name></person-group> <article-title>Humans adaptively resolve the explore-exploit dilemma under cognitive constraints: Evidence from a multi-armed bandit task</article-title>. <source>Cognition</source>. <year>2022</year> <month>Dec</month>; <volume>229</volume>:<issue>105233</issue>. <ext-link ext-link-type="uri" xlink:href="https://www.sciencedirect.com/science/article/pii/S0010027722002219">https://www.sciencedirect.com/science/article/pii/S0010027722002219</ext-link>, doi: <pub-id pub-id-type="doi">10.1016/j.cognition.2022.105233.</pub-id></mixed-citation></ref>
<ref id="c11"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>PC</given-names> <surname>Bürkner</surname></string-name></person-group>. <article-title>sbrims: An R Package for Bayesian Multilevel Models Using Stan</article-title>. <source>Journal of Statistical Software</source>. <year>2017</year> <month>Aug</month>; <volume>80</volume>:<fpage>1</fpage>–<lpage>28</lpage>. <ext-link ext-link-type="uri" xlink:href="https:///doi.org/10.18637/jss.v080.i01">https://doi.org/10.18637/jss.v080.i01</ext-link>, doi: <pub-id pub-id-type="doi">10.18637/jss.v080.i01.</pub-id></mixed-citation></ref>
<ref id="c12"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>PC</given-names> <surname>Bürkner</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Vuorre</surname></string-name></person-group> <article-title>Ordinal regression models in psychology: A tutorial</article-title>. <source>Advances in Methods and Practices in Psychological Science</source>. <year>2019</year>; <volume>2</volume>(<issue>1</issue>):<fpage>77</fpage>–<lpage>101</lpage>.</mixed-citation></ref>
<ref id="c13"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>B</given-names> <surname>Carpenter</surname></string-name>, <string-name><given-names>A</given-names> <surname>Gelman</surname></string-name>, <string-name><given-names>MD</given-names> <surname>Hoffman</surname></string-name>, <string-name><given-names>D</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>B</given-names> <surname>Goodrich</surname></string-name>, <string-name><given-names>M</given-names> <surname>Betancourt</surname></string-name>, <string-name><given-names>M</given-names> <surname>Brubaker</surname></string-name>, <string-name><given-names>J</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>P</given-names> <surname>Li</surname></string-name>, <string-name><surname>Riddell</surname> <given-names>A.</given-names></string-name></person-group> <chapter-title>Stan: A Probabilistic Programming Language</chapter-title>. <source>Journal of Statistical Software</source>. <year>2017</year> <month>Jan</month>; <volume>76</volume>(<issue>1</issue>). <ext-link ext-link-type="uri" xlink:href="https://www.osti.gov/pages/biblio/1430202">https://www.osti.gov/pages/biblio/1430202</ext-link>, doi: <pub-id pub-id-type="doi">10.18637/jss.v076.i01,</pub-id> <publisher-name>institution: Columbia Univ., New York, NY (United States); Harvard Univ., Cambridge, MA (United States)</publisher-name>.</mixed-citation></ref>
<ref id="c14"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>JD</given-names> <surname>Carrillo</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Mariotti</surname></string-name></person-group> <chapter-title>Strategic ignorance as a self-disciplining device</chapter-title>. <source>The Review of Economic Studies</source>. <year>2000</year>; <volume>67</volume>(<issue>3</issue>):<fpage>529</fpage>–<lpage>544</lpage>. <publisher-name>Publisher: Wiley-Blackwell</publisher-name>.</mixed-citation></ref>
<ref id="c15"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N</given-names> <surname>Chater</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Oaksford</surname></string-name></person-group> <article-title>Ten years of the rational analysis of cognition</article-title>. <source>Trends in Cognitive Sciences</source>. <year>1999</year> <month>Feb</month>; <volume>3</volume>(<issue>2</issue>):<fpage>57</fpage>–<lpage>65</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.sciencedirect.com/science/article/pii/S136466139801273X">https://www.sciencedirect.com/science/article/pii/S136466139801273X</ext-link>, doi: <pub-id pub-id-type="doi">10.1016/S1364-6613(98)01273-X</pub-id>.</mixed-citation></ref>
<ref id="c16"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chuan-Peng</surname> <given-names>H</given-names></string-name>, <string-name><given-names>H</given-names> <surname>Geng</surname></string-name>, <string-name><given-names>L</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>A</given-names> <surname>Fengler</surname></string-name>, <string-name><given-names>M</given-names> <surname>Frank</surname></string-name>, <string-name><given-names>RY</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>A</given-names> <surname>Hitchhiker’s</surname></string-name></person-group> <article-title>Guide to Bayesian Hierarchical Drift-Diffusion Modeling with docker HDDM</article-title>. <source>PsyArXiv;</source> <year>2022</year>. <ext-link ext-link-type="uri" xlink:href="https://psyarxiv.com/6uzga/">https://psyarxiv.com/6uzga/</ext-link>, doi: <pub-id pub-id-type="doi">10.31234/osf.io/6uzga</pub-id>.</mixed-citation></ref>
<ref id="c17"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>JD</given-names> <surname>Cohen</surname></string-name>, <string-name><given-names>SM</given-names> <surname>McClure</surname></string-name>, <string-name><given-names>AJ</given-names> <surname>Yu</surname></string-name></person-group>. <article-title>Should I stay or should I go? How the human brain manages the trade-off between exploitation and exploration</article-title>. <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>. <year>2007</year>; <volume>362</volume>(<issue>1481</issue>):<fpage>933</fpage>–<lpage>942</lpage>. doi: <pub-id pub-id-type="doi">10.1098/rstb.2007.2098.</pub-id></mixed-citation></ref>
<ref id="c18"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>AGE</given-names> <surname>Collins</surname></string-name>, <string-name><given-names>MJ</given-names> <surname>Frank</surname></string-name></person-group>. <article-title>How much of reinforcement learning is working memory, not reinforcement learning? A behavioral, computational, and neurogenetic analysis</article-title>. <source>European Journal of Neuroscience</source>. <year>2012</year>; <volume>35</volume>(<issue>7</issue>):<fpage>1024</fpage>–<lpage>1035</lpage>. <ext-link ext-link-type="uri" xlink:href="https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1460-9568.2011.07980.x">https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1460-9568.2011.07980.x</ext-link>, doi: <pub-id pub-id-type="doi">10.1111/j.1460-9568.2011.07980.x</pub-id>, _eprint:<ext-link ext-link-type="uri" xlink:href="https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1460-9568.2011.07980.x">https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1460-9568.2011.07980.x</ext-link>.</mixed-citation></ref>
<ref id="c19"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>AGE</given-names> <surname>Collins</surname></string-name>, <string-name><given-names>B</given-names> <surname>Ciullo</surname></string-name>, <string-name><given-names>MJ</given-names> <surname>Frank</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Badre</surname></string-name></person-group> <article-title>Working memory load strengthens reward prediction errors</article-title>. <source>Journal of Neuroscience</source>. <year>2017</year>; <volume>37</volume>(<issue>16</issue>):<fpage>4332</fpage>–<lpage>4342</lpage>. doi: <pub-id pub-id-type="doi">10.1523/JNEUROSCI.2700-16.2017.</pub-id></mixed-citation></ref>
<ref id="c20"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>ND</given-names> <surname>Daw</surname></string-name>, <string-name><given-names>Y</given-names> <surname>Niv</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Dayan</surname></string-name></person-group> <article-title>Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control</article-title>. <source>Nature neuroscience</source>. <year>2005</year>; <volume>8</volume>(<issue>12</issue>):<fpage>1704</fpage>–<lpage>1711</lpage>.</mixed-citation></ref>
<ref id="c21"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>ND</given-names> <surname>Daw</surname></string-name>, <string-name><surname>O’doherty</surname> <given-names>JP</given-names></string-name>, <string-name><given-names>P</given-names> <surname>Dayan</surname></string-name>, <string-name><given-names>B</given-names> <surname>Seymour</surname></string-name>, <string-name><given-names>RJ.</given-names> <surname>Dolan</surname></string-name></person-group> <chapter-title>Cortical substrates for exploratory decisions in humans</chapter-title>. <source>Nature</source>. <year>2006</year>; <volume>441</volume>(<issue>7095</issue>):<fpage>876</fpage>–<lpage>879</lpage>. <publisher-name>Publisher: Nature Publishing Group</publisher-name>.</mixed-citation></ref>
<ref id="c22"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>JR.</given-names> <surname>De Leeuw</surname></string-name></person-group> <article-title>jsPsych: A JavaScript library for creating behavioral experiments in a Web browser</article-title>. <source>Behavior research methods</source>. <year>2015</year>; <volume>47</volume>(<issue>1</issue>):<fpage>1</fpage>–<lpage>12</lpage>.</mixed-citation></ref>
<ref id="c23"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K</given-names> <surname>Duncan</surname></string-name>, <string-name><given-names>A</given-names> <surname>Semmler</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Shohamy</surname></string-name></person-group> <article-title>Modulating the Use of Multiple Memory Systems in Value-based Decisions with Contextual Novelty</article-title>. <source>Journal of Cognitive Neuroscience</source>. <year>2019</year> Oct; <volume>31</volume>(<issue>10</issue>):<fpage>1455</fpage>–<lpage>1467</lpage>. <pub-id pub-id-type="doi">10.1162/jocn_a_01447</pub-id>, doi: <pub-id pub-id-type="doi">10.1162/jocn_a_01447.</pub-id></mixed-citation></ref>
<ref id="c24"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D</given-names> <surname>Eilam</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Golani</surname></string-name></person-group> <article-title>Home base behavior of rats (Rattus norvegicus) exploring a novel environment</article-title>. <source>Behavioural Brain Research</source>. <year>1989</year>; <volume>34</volume>(<issue>3</issue>):<fpage>199</fpage>–<lpage>211</lpage>. doi: <pub-id pub-id-type="doi">10.1016/S0166-4328</pub-id>(89)80102-0.</mixed-citation></ref>
<ref id="c25"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Ellsberg</surname></string-name></person-group> <chapter-title>Risk, ambiguity, and the Savage axioms</chapter-title>. <source>The quarterly journal of economics</source>. <year>1961</year>; p. <fpage>643</fpage>–<lpage>669</lpage>. <publisher-name>Publisher: JSTOR</publisher-name>.</mixed-citation></ref>
<ref id="c26"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>CR</given-names> <surname>Fox</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Tversky</surname></string-name></person-group> <chapter-title>Ambiguity a version and comparative ignorance</chapter-title>. <source>The quarterly journal of economics</source>. <year>1995</year>; <volume>110</volume>(<issue>3</issue>):<fpage>585</fpage>–<lpage>603</lpage>. <publisher-name>Publisher: MIT Press</publisher-name>.</mixed-citation></ref>
<ref id="c27"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>G</given-names> <surname>Gigerenzer</surname></string-name>, <string-name><surname>Garcia-Retamero</surname> <given-names>R.</given-names></string-name></person-group> <chapter-title>Cassandra’s regret: The psychology of not wanting to know</chapter-title>. <source>Psychological review</source>. <year>2017</year>; <volume>124</volume>(<issue>2</issue>):<comment>179</comment>. <publisher-name>Publisher: American Psychological Association</publisher-name>.</mixed-citation></ref>
<ref id="c28"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>SE</given-names> <surname>Glickman</surname></string-name>, <string-name><given-names>RW.</given-names> <surname>Sroges</surname></string-name></person-group> <chapter-title>Curiosity in zoo animals</chapter-title>. <source>Behaviour</source>. <year>1966</year>; <volume>26</volume>(<issue>1-2</issue>):<fpage>151</fpage>–<lpage>187</lpage>. <publisher-name>Publisher: Brill</publisher-name>.</mixed-citation></ref>
<ref id="c29"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R</given-names> <surname>Golman</surname></string-name>, <string-name><given-names>D</given-names> <surname>Hagmann</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Loewenstein</surname></string-name></person-group> <article-title>Information avoidance</article-title>. <source>Journal of economic literature</source>. <year>2017</year>; <volume>55</volume>(<issue>1</issue>):<fpage>96</fpage>–<lpage>135</lpage>.</mixed-citation></ref>
<ref id="c30"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G</given-names> <surname>Gordon</surname></string-name>, <string-name><given-names>E</given-names> <surname>Fonio</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Ahissar</surname></string-name></person-group> <article-title>Emergent Exploration via Novelty Management</article-title>. <source>Journal of Neuroscience</source>. <year>2014</year>; <volume>34</volume>(<issue>38</issue>):<fpage>12646</fpage>–<lpage>12661</lpage>. doi: <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1872-14.2014.</pub-id></mixed-citation></ref>
<ref id="c31"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>TM</given-names> <surname>Gureckis</surname></string-name>, <string-name><surname>Markant</surname> <given-names>DB.</given-names></string-name></person-group> <article-title>Self-directed learning: A cognitive and computational perspective</article-title>. <source>Perspectives on Psychological Science</source>. <year>2012</year>; <volume>7</volume>(<issue>5</issue>):<fpage>464</fpage>–<lpage>481</lpage>.</mixed-citation></ref>
<ref id="c32"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>CA</given-names> <surname>Hartley</surname></string-name></person-group>. <article-title>How do natural environments shape adaptive cognition across the lifespan?</article-title> <source>Trends in Cognitive Sciences</source>. <year>2022</year> <month>Dec</month>; <volume>26</volume>(<issue>12</issue>):<fpage>1029</fpage>–<lpage>1030</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.sciencedirect.com/science/article/pii/S1364661322002601">https://www.sciencedirect.com/science/article/pii/S1364661322002601</ext-link>, doi: <pub-id pub-id-type="doi">10.1016/j.tics.2022.10.002.</pub-id></mixed-citation></ref>
<ref id="c33"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>DJ</given-names> <surname>Hauser</surname></string-name>, <string-name><given-names>AJ</given-names> <surname>Moss</surname></string-name>, <string-name><given-names>C</given-names> <surname>Rosenzweig</surname></string-name>, <string-name><given-names>SN</given-names> <surname>Jaffe</surname></string-name>, <string-name><given-names>J</given-names> <surname>Robinson</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Litman</surname></string-name></person-group> <article-title>Evaluating Cloud Research’s Approved Group as a solution for problematic data quality on MTurk</article-title>. <source>Behavior Research Methods</source>. <year>2022</year> <month>Nov</month>; <pub-id pub-id-type="doi">10.3758/s13428-022-01999-x</pub-id>, doi: <pub-id pub-id-type="doi">10.3758/s13428-022-01999-x.</pub-id></mixed-citation></ref>
<ref id="c34"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>LT</given-names> <surname>Hunt</surname></string-name>, <string-name><given-names>ND</given-names> <surname>Daw</surname></string-name>, <string-name><given-names>P</given-names> <surname>Kaanders</surname></string-name>, <string-name><given-names>MA</given-names> <surname>MacIver</surname></string-name>, <string-name><given-names>U</given-names> <surname>Mugan</surname></string-name>, <string-name><given-names>E</given-names> <surname>Procyk</surname></string-name>, <string-name><given-names>AD</given-names> <surname>Redish</surname></string-name>, <string-name><given-names>E</given-names> <surname>Russo</surname></string-name>, <string-name><given-names>J</given-names> <surname>Scholl</surname></string-name>, <string-name><given-names>K</given-names> <surname>Stachenfeld</surname></string-name>, <string-name><given-names>CRE</given-names> <surname>Wilson</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Kolling</surname></string-name></person-group> <article-title>Formalizing planning and information search in naturalistic decision-making</article-title>. <source>Nature Neuroscience</source>. <year>2021</year> <month>Aug</month>; <volume>24</volume>(<issue>8</issue>):<fpage>1051</fpage>–<lpage>1064</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/s41593-021-00866-w">https://www.nature.com/articles/s41593-021-00866-w</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/s41593-021-00866-w,</pub-id> number: 8 Publisher: Nature Publishing Group.</mixed-citation></ref>
<ref id="c35"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>BJ</given-names> <surname>Knowlton</surname></string-name>, <string-name><given-names>JA</given-names> <surname>Mangels</surname></string-name>, <string-name><given-names>LR.</given-names> <surname>Squire</surname></string-name></person-group> <article-title>A Neostriatal Habit Learning System in Humans</article-title>. <source>Science</source>. <year>1996</year> <month>Sep</month>; <volume>273</volume>(<issue>5280</issue>):<fpage>1399</fpage>–<lpage>1402</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.science.org/doi/10.1126/science.273.5280.1399">https://www.science.org/doi/10.1126/science.273.5280.1399</ext-link>, doi: <pub-id pub-id-type="doi">10.1126/science.273.5280.1399</pub-id>.</mixed-citation></ref>
<ref id="c36"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>F</given-names> <surname>Lieder</surname></string-name>, <string-name><surname>Griffiths</surname> <given-names>TL.</given-names></string-name></person-group> <chapter-title>Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources</chapter-title>. <source>Behavioral and Brain Sciences</source>. <year>2020</year>; <fpage>43</fpage>. doi: <pub-id pub-id-type="doi">10.1017/S0140525X1900061X,</pub-id> <publisher-name>publisher: Cambridge University Press</publisher-name>.</mixed-citation></ref>
<ref id="c37"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L</given-names> <surname>Litman</surname></string-name>, <string-name><given-names>J</given-names> <surname>Robinson</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Abberbock</surname></string-name></person-group> <article-title>TurkPrime.com: A versatile crowdsourcing data acquisition platform for the behavioral sciences</article-title>. <source>Behavior Research Methods</source>. <year>2017</year> <month>Apr</month>; <volume>49</volume>(<issue>2</issue>):<fpage>433</fpage>–<lpage>442</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-016-0727-z</pub-id>, doi: <pub-id pub-id-type="doi">10.3758/s13428-016-0727-z.</pub-id></mixed-citation></ref>
<ref id="c38"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>DJC.</given-names> <surname>MacKay</surname></string-name></person-group> <article-title>Information-based objective functions for active data selection</article-title>. <source>Neural computation</source>. <year>1992</year>; <volume>4</volume>(<issue>4</issue>):<fpage>590</fpage>–<lpage>604</lpage>.</mixed-citation></ref>
<ref id="c39"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>DB</given-names> <surname>Markant</surname></string-name>, <string-name><given-names>TM.</given-names> <surname>Gureckis</surname></string-name></person-group> <article-title>A preference for the unpredictable over the informative during self-directed learning</article-title>. <source>Proceedings of the 36th Annual Conference of the Cognitive Science Society</source>. <year>2014</year>;.</mixed-citation></ref>
<ref id="c40"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>J</given-names> <surname>Nicholas</surname></string-name>, <string-name><given-names>ND</given-names> <surname>Daw</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Shohamy</surname></string-name></person-group> <chapter-title>Uncertainty alters the balance between incremental learning and episodic memory</chapter-title>. <source>eLife</source>. <year>2022</year> <month>Dec</month>; <volume>11</volume>:<comment>e81679</comment>. <pub-id pub-id-type="doi">10.7554/eLife.81679</pub-id>, doi: <pub-id pub-id-type="doi">10.7554/eLife.81679,</pub-id> <publisher-name>publisher: eLife Sciences Publications, Ltd</publisher-name>.</mixed-citation></ref>
<ref id="c41"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M</given-names> <surname>Oaksford</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Chater</surname></string-name></person-group> <article-title>A Rational Analysis of the Selection Task as Optimal Data Selection</article-title>. <source>Psychological Review</source>. <year>1994</year>; <volume>101</volume>(<issue>4</issue>):<fpage>608</fpage>–<lpage>631</lpage>. doi: <pub-id pub-id-type="doi">10.1037/0033-295X.101.4.608.</pub-id></mixed-citation></ref>
<ref id="c42"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>T</given-names> <surname>Ogasawara</surname></string-name>, <string-name><given-names>F</given-names> <surname>Sogukpinar</surname></string-name>, <string-name><given-names>K</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>YY</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>J</given-names> <surname>Pai</surname></string-name>, <string-name><given-names>A</given-names> <surname>Jezzini</surname></string-name>, <string-name><given-names>IE.</given-names> <surname>Monosov</surname></string-name></person-group> <chapter-title>A primate temporal cortex-zona incerta pathway for novelty seeking</chapter-title>. <source>Nature Neuroscience</source>. <year>2022</year> <month>Jan</month>; <volume>25</volume>(<issue>1</issue>):<fpage>50</fpage>–<lpage>60</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/s41593-021-00950-1">https://www.nature.com/articles/s41593-021-00950-1</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/s41593-021-00950-1,</pub-id> <comment>number: 1</comment> <publisher-name>Publisher: Nature Publishing Group</publisher-name>.</mixed-citation></ref>
<ref id="c43"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J</given-names> <surname>Palmer</surname></string-name>, <string-name><given-names>AC</given-names> <surname>Huk</surname></string-name>, <string-name><given-names>MN.</given-names> <surname>Shadlen</surname></string-name></person-group> <article-title>The effect of stimulus strength on the speed and accuracy of a perceptual decision</article-title>. <source>Journal of vision</source>. <year>2005</year>; <volume>5</volume>(<issue>5</issue>):<fpage>1</fpage>.</mixed-citation></ref>
<ref id="c44"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D</given-names> <surname>Pathak</surname></string-name>, <string-name><given-names>P</given-names> <surname>Agrawal</surname></string-name>, <string-name><given-names>AA</given-names> <surname>Efros</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Darrell</surname></string-name></person-group> <article-title>Curiosity-driven Exploration by Self-supervised Prediction</article-title>. In: <source>Proceedings of the 34th International Conference on Machine Learning PMLR</source>; <year>2017</year>. p. <fpage>2778</fpage>–<lpage>2787</lpage>. <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v70/pathak17a.html">https://proceedings.mlr.press/v70/pathak17a.html</ext-link>, <comment>iSSN: 2640-3498</comment>.</mixed-citation></ref>
<ref id="c45"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>P</given-names> <surname>Petitet</surname></string-name>, <string-name><given-names>B</given-names> <surname>Attaallah</surname></string-name>, <string-name><given-names>SG</given-names> <surname>Manohar</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Husain</surname></string-name></person-group> <chapter-title>The computational cost of active information sampling before decision-making under uncertainty</chapter-title>. <source>Nature Human Behaviour</source>. <year>2021</year> <month>Jul</month>; <volume>5</volume>(<issue>7</issue>):<fpage>935</fpage>–<lpage>946</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/s41562-021-01116-6">https://www.nature.com/articles/s41562-021-01116-6</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/s41562-021-01116-6,</pub-id> <comment>number: 7</comment> <publisher-name>Publisher: Nature Publishing Group</publisher-name>.</mixed-citation></ref>
<ref id="c46"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>O</given-names> <surname>Plonsky</surname></string-name>, <string-name><given-names>K</given-names> <surname>Teodorescu</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Erev</surname></string-name></person-group> <article-title>Reliance on small samples, the wavy recency effect, and similarity-based learning</article-title>. <source>Psychological review</source>. <year>2015</year>; <volume>122</volume>(<issue>4</issue>):<fpage>621</fpage>.</mixed-citation></ref>
<ref id="c47"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>RA</given-names> <surname>Poldrack</surname></string-name>, <string-name><given-names>J</given-names> <surname>Clark</surname></string-name>, <string-name><surname>Paré-Blagoev</surname> <given-names>EJ</given-names></string-name>, <string-name><given-names>D</given-names> <surname>Shohamy</surname></string-name>, <string-name><surname>Creso</surname> <given-names>Moyano J</given-names></string-name>, <string-name><given-names>C</given-names> <surname>Myers</surname></string-name>, <string-name><given-names>MA.</given-names> <surname>Gluck</surname></string-name></person-group> <article-title>Interactive memory systemsinthehumanbrain</article-title>. <source>Nature</source>. <year>2001</year> <month>Nov</month>; <comment>41</comment> <volume>4</volume>(<issue>6863</issue>):<fpage>546</fpage>–<lpage>550</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/35107080">http://www.nature.com/articles/35107080</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/35107080.</pub-id></mixed-citation></ref>
<ref id="c48"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D</given-names> <surname>Raposo</surname></string-name>, <string-name><given-names>S</given-names> <surname>Ritter</surname></string-name>, <string-name><given-names>A</given-names> <surname>Santoro</surname></string-name>, <string-name><given-names>G</given-names> <surname>Wayne</surname></string-name>, <string-name><given-names>T</given-names> <surname>Weber</surname></string-name>, <string-name><given-names>M</given-names> <surname>Botvinick</surname></string-name>, <string-name><surname>van Hasselt</surname> <given-names>H</given-names></string-name>, <string-name><given-names>F</given-names> <surname>Song</surname></string-name></person-group>, <article-title>Synthetic Returns for Long-Term Credit Assignment</article-title>. <source>arXiv;</source> <year>2021</year>. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2102.12425">http://arxiv.org/abs/2102.12425</ext-link>, <comment>arXiv: 2102.12425 [cs]</comment>.</mixed-citation></ref>
<ref id="c49"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R</given-names> <surname>Ratcliff</surname></string-name>, <string-name><surname>McKoon</surname> <given-names>G.</given-names></string-name></person-group> <article-title>The Diffusion Decision Model: Theory and Data for Two-Choice Decision Tasks</article-title>. <source>Neural Computation</source>. <year>2008</year> <month>Apr</month>; <volume>20</volume>(<issue>4</issue>):<fpage>873</fpage>–<lpage>922</lpage>. <ext-link ext-link-type="uri" xlink:href="https://direct.mit.edu/neco/article/20/4/873-922/7299">https://direct.mit.edu/neco/article/20/4/873-922/7299</ext-link>, doi: <pub-id pub-id-type="doi">10.1162/neco.2008.12-06-420.</pub-id></mixed-citation></ref>
<ref id="c50"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A</given-names> <surname>Rothe</surname></string-name>, <string-name><given-names>BM</given-names> <surname>Lake</surname></string-name>, <string-name><given-names>TM.</given-names> <surname>Gureckis</surname></string-name></person-group> <article-title>Do People Ask Good Questions?</article-title> <source>Computational Brain &amp; Behavior</source>. <year>2018</year>; <volume>1</volume>(<issue>1</issue>):<fpage>69</fpage>–<lpage>89</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s42113-018-0005-5.</pub-id></mixed-citation></ref>
<ref id="c51"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A</given-names> <surname>Ruggeri</surname></string-name>, <string-name><given-names>ZL</given-names> <surname>Sim</surname></string-name>, <string-name><surname>Xu</surname> <given-names>F.</given-names></string-name></person-group> “<chapter-title>Whyis Toma late to school again?” Preschoolers identify the most informative questions</chapter-title>. <source>Developmental psychology</source>. <year>2017</year>; <volume>53</volume>(<issue>9</issue>):<fpage>1620</fpage>. <publisher-name>Publisher: American Psychological Association</publisher-name>.</mixed-citation></ref>
<ref id="c52"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>DJ</given-names> <surname>Schad</surname></string-name>, <string-name><given-names>M</given-names> <surname>Betancourt</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Vasishth</surname></string-name></person-group> <article-title>Toward a principled Bayesian workflow in cognitive science</article-title>. <source>arXiv preprint arXiv:190412765</source>. <year>2019</year>;.</mixed-citation></ref>
<ref id="c53"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>DJ</given-names> <surname>Schad</surname></string-name>, <string-name><given-names>M</given-names> <surname>Betancourt</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Vasishth</surname></string-name></person-group> <chapter-title>Toward a principled Bayesian workflow in cognitive science</chapter-title>. <source>Psychological methods</source>. <year>2021</year>; <volume>26</volume>(<issue>1</issue>):<fpage>103</fpage>. <publisher-name>Publisher: American Psychological Association</publisher-name>.</mixed-citation></ref>
<ref id="c54"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E</given-names> <surname>Schulz</surname></string-name>, <string-name><given-names>SJ.</given-names> <surname>Gershman</surname></string-name></person-group> <article-title>The algorithmic architecture of exploration in the human brain</article-title>. <source>Current Opinion in Neurobiology</source>. <year>2019</year>; <volume>55</volume>:<fpage>7</fpage>–<lpage>14</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.conb.2018.11.003.</pub-id></mixed-citation></ref>
<ref id="c55"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P</given-names> <surname>Schwartenbeck</surname></string-name>, <string-name><given-names>J</given-names> <surname>Passecker</surname></string-name>, <string-name><given-names>TU</given-names> <surname>Hauser</surname></string-name>, <string-name><given-names>TH</given-names> <surname>FitzGerald</surname></string-name>, <string-name><given-names>M</given-names> <surname>Kronbichler</surname></string-name>, <string-name><given-names>KJ.</given-names> <surname>Friston</surname></string-name></person-group> <article-title>Computational mechanisms of curiosity and goal-directed exploration</article-title>. <source>eLife</source>. <year>2019</year>; <volume>8</volume>:<fpage>1</fpage>–<lpage>45</lpage>. doi: <pub-id pub-id-type="doi">10.7554/eLife.41703.</pub-id></mixed-citation></ref>
<ref id="c56"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P</given-names> <surname>Sebastiani</surname></string-name>, <string-name><given-names>HP.</given-names> <surname>Wynn</surname></string-name></person-group> <article-title>Maximum entropy sampling and optimal Bayesian experimental design</article-title>. <source>Journal of the Royal Statistical Society: Series B (Statistical Methodology)</source>. <year>2000</year>; <volume>62</volume>(<issue>1</issue>):<fpage>145</fpage>–<lpage>157</lpage>.</mixed-citation></ref>
<ref id="c57"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>MN</given-names> <surname>Shadlen</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kiani</surname></string-name></person-group> <article-title>Decision making as a window on cognition</article-title>. <source>Neuron</source>. <year>2013</year>; <volume>80</volume>(<issue>3</issue>):<fpage>791</fpage>–<lpage>806</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2013.10.047</pub-id>, doi: <pub-id pub-id-type="doi">10.1016/j.neuron.2013.10.047.</pub-id></mixed-citation></ref>
<ref id="c58"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>MN</given-names> <surname>Shadlen</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Shohamy</surname></string-name></person-group> <article-title>Decision Making and Sequential Sampling from Memory</article-title>. <source>Neuron</source>. <year>2016</year> <month>Jun</month>; <volume>90</volume>(<issue>5</issue>):<fpage>927</fpage>–<lpage>939</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.sciencedirect.com/science/article/pii/S0896627316301234">https://www.sciencedirect.com/science/article/pii/S0896627316301234</ext-link>, doi: <pub-id pub-id-type="doi">10.1016/j.neuron.2016.04.036.</pub-id></mixed-citation></ref>
<ref id="c59"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Shafir</surname></string-name></person-group> <article-title>Uncertainty and the difficulty of thinking through disjunctions</article-title>. <source>Cognition</source>. <year>1994</year> <month>Apr</month>; <volume>50</volume>(<issue>1</issue>):<fpage>403</fpage>–<lpage>430</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.sciencedirect.com/science/article/pii/0010027794900388">https://www.sciencedirect.com/science/article/pii/0010027794900388</ext-link>, doi: <pub-id pub-id-type="doi">10.1016/0010-0277(94)90038-8</pub-id>.</mixed-citation></ref>
<ref id="c60"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S</given-names> <surname>Shushruth</surname></string-name>, <string-name><given-names>A</given-names> <surname>Zylberberg</surname></string-name>, <string-name><given-names>MN.</given-names> <surname>Shadlen</surname></string-name></person-group> <article-title>Sequential sampling from memory underlies action selection during abstract decision-making</article-title>. <source>Current Biology</source>. <year>2022</year>; <volume>32</volume>(<issue>9</issue>):<fpage>1949</fpage>–<lpage>1960</lpage>.</mixed-citation></ref>
<ref id="c61"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Z</given-names> <surname>Bnaya</surname></string-name>, <string-name><given-names>WJ.</given-names> <surname>Ma</surname></string-name></person-group> <chapter-title>Sources of suboptimality in a minimalistic explore-exploit task</chapter-title>. <source>Nature Human Behaviour</source>. <year>2019</year> <month>Apr</month>; <volume>3</volume>(<issue>4</issue>):<fpage>361</fpage>–<lpage>368</lpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/s41562-018-0526-x">https://www.nature.com/articles/s41562-018-0526-x</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/s41562-018-0526-x,</pub-id> <comment>number: 4</comment> <publisher-name>Publisher: Nature Publishing Group</publisher-name>.</mixed-citation></ref>
<ref id="c62"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M</given-names> <surname>Speekenbrink</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Konstantinidis</surname></string-name></person-group> <article-title>Uncertainty and Exploration in a Restless Bandit Problem</article-title>. <source>Topics in Cognitive Science</source>. <year>2015</year>; <volume>7</volume>(<issue>2</issue>):<fpage>351</fpage>–<lpage>367</lpage>. <ext-link ext-link-type="uri" xlink:href="https://onlinelibrary.wiley.com/doi/abs/10.1111/tops.12145">https://onlinelibrary.wiley.com/doi/abs/10.1111/tops.12145</ext-link>, doi: <pub-id pub-id-type="doi">10.1111/tops.12145</pub-id>, _eprint: <ext-link ext-link-type="uri" xlink:href="https://onlinelibrary.wiley.com/doi/pdf/10.1111/tops.12145">https://onlinelibrary.wiley.com/doi/pdf/10.1111/tops.12145</ext-link>.</mixed-citation></ref>
<ref id="c63"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>RS</given-names> <surname>Sutton</surname></string-name>, <string-name><given-names>AG.</given-names> <surname>Barto</surname></string-name></person-group> <source>Reinforcement learning: An introduction, 2nd ed. Reinforcement learning: An introduction, 2nd ed</source>, <publisher-loc>Cambridge, MA, US</publisher-loc>: <publisher-name>The MIT Press</publisher-name>; <year>2018</year>. Pages: <fpage>xxii</fpage>, <lpage>526</lpage>.</mixed-citation></ref>
<ref id="c64"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N</given-names> <surname>Trudel</surname></string-name>, <string-name><given-names>J</given-names> <surname>Scholl</surname></string-name>, <string-name><surname>Klein-Flügge</surname> <given-names>MC</given-names></string-name>, <string-name><given-names>E</given-names> <surname>Fouragnan</surname></string-name>, <string-name><given-names>L</given-names> <surname>Tankelevitch</surname></string-name>, <string-name><given-names>MK</given-names> <surname>Wittmann</surname></string-name>, <string-name><given-names>MFS.</given-names> <surname>Rushworth</surname></string-name></person-group> <article-title>Polarity of uncertainty representation during exploration and exploitation in ventromedial prefrontal cortex</article-title>. <source>Nature Human Behaviour</source>. <year>2020</year>; <day>5</day>(<month>January</month>). <pub-id pub-id-type="doi">10.1038/s41562-020-0929-3</pub-id>, doi: <pub-id pub-id-type="doi">10.1038/s41562-020-0929-3.</pub-id></mixed-citation></ref>
<ref id="c65"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A</given-names> <surname>Tversky</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Edwards</surname></string-name></person-group> <chapter-title>Information versus reward in binary choices</chapter-title>. <source>Journal of Experimental Psychology</source>. <year>1966</year>; <volume>71</volume>(<issue>5</issue>):<fpage>680</fpage>–<lpage>683</lpage>. doi: <pub-id pub-id-type="doi">10.1037/h0023123,</pub-id> <publisher-name>place: US Publisher: American Psychological Association</publisher-name>.</mixed-citation></ref>
<ref id="c66"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A</given-names> <surname>Vehtari</surname></string-name>, <string-name><given-names>A</given-names> <surname>Gelman</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Gabry</surname></string-name></person-group> <article-title>Practical Bayesian model evaluation using leave-one-out cross-validation and WAIC</article-title>. <source>Statistics and Computing</source>. <year>2017</year> <month>Sep</month>; <volume>27</volume>(<issue>5</issue>):<fpage>1413</fpage>–<lpage>1432</lpage>. <pub-id pub-id-type="doi">10.1007/s11222-016-9696-4</pub-id>, doi: <pub-id pub-id-type="doi">10.1007/s11222-016-9696-4.</pub-id></mixed-citation></ref>
<ref id="c67"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>ML</given-names> <surname>Waskom</surname></string-name>, <string-name><given-names>G</given-names> <surname>Okazawa</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kiani</surname></string-name></person-group> <chapter-title>Designing and Interpreting Psychophysical Investigations of Cognition</chapter-title>. <source>Neuron</source>. <year>2019</year>; <volume>104</volume>(<issue>1</issue>):<fpage>100</fpage>–<lpage>112</lpage>. <publisher-name>Publisher: Elsevier</publisher-name>.</mixed-citation></ref>
<ref id="c68"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T</given-names> <surname>Wiecki</surname></string-name>, <string-name><given-names>I</given-names> <surname>Sofer</surname></string-name>, <string-name><surname>Frank</surname> <given-names>M.</given-names></string-name></person-group> <article-title>HDDM: Hierarchical Bayesian estimation of the Drift-Diffusion Model in Python</article-title>. <source>Frontiers in Neuroinformatics</source>. <year>2013</year>; <volume>7</volume>. <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fninf.2013.00014">https://www.frontiersin.org/articles/10.3389/fninf.2013.00014</ext-link>.</mixed-citation></ref>
<ref id="c69"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>RC</given-names> <surname>Wilson</surname></string-name>, <string-name><given-names>A</given-names> <surname>Geana</surname></string-name>, <string-name><given-names>JM</given-names> <surname>White</surname></string-name>, <string-name><given-names>EA</given-names> <surname>Ludvig</surname></string-name>, <string-name><given-names>JD.</given-names> <surname>Cohen</surname></string-name></person-group> <chapter-title>Humans use directed and random exploration to solve the explore-exploit dilemma</chapter-title>. <source>Journal of Experimental Psychology: General</source>. <year>2014</year>; <volume>143</volume>(<issue>6</issue>):<fpage>2074</fpage>. <publisher-name>Publisher: American Psychological Association</publisher-name>.</mixed-citation></ref>
<ref id="c70"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>CM</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>E</given-names> <surname>Schulz</surname></string-name>, <string-name><given-names>TJ</given-names> <surname>Pleskac</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Speekenbrink</surname></string-name></person-group> <chapter-title>Time pressure changes how people explore and respond to uncertainty</chapter-title>. <source>Scientific Reports</source>. <year>2022</year> <month>Mar</month>; <volume>12</volume>(<issue>1</issue>):<fpage>4122</fpage>. <ext-link ext-link-type="uri" xlink:href="https://www.nature.com/articles/s41598-022-07901-1">https://www.nature.com/articles/s41598-022-07901-1</ext-link>, doi: <pub-id pub-id-type="doi">10.1038/s41598-022-07901-1,</pub-id> <comment>number: 1</comment> <publisher-name>Publisher: Nature Publishing Group</publisher-name>.</mixed-citation></ref>
<ref id="c71"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>DU</given-names> <surname>Wulff</surname></string-name>, <string-name><surname>Mergenthaler-Canseco</surname> <given-names>M</given-names></string-name>, <string-name><given-names>R.</given-names> <surname>Hertwig</surname></string-name></person-group> <article-title>A meta-analytic review of two modes of learning and the description-experience gap</article-title>. <source>Psychological bulletin</source>. <year>2018</year>; <volume>144</volume>(<issue>2</issue>):<fpage>140</fpage>.</mixed-citation></ref>
<ref id="c72"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>SCH</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>M</given-names> <surname>Lengyel</surname></string-name>, <string-name><given-names>DM.</given-names> <surname>Wolpert</surname></string-name></person-group> <article-title>Active sensing in the categorization of visual patterns</article-title>. <source>eLife</source>. <year>2016</year>; <volume>5</volume>:<fpage>1</fpage>–<lpage>22</lpage>. doi: <pub-id pub-id-type="doi">10.7554/elife.12215.</pub-id></mixed-citation></ref>
<ref id="c73"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Zylberberg</surname></string-name></person-group> <chapter-title>Decision prioritization and causal reasoning in decision hierarchies</chapter-title>. <source>PLOS Computational Biology</source>. <year>2021</year> <month>Dec</month>; <volume>17</volume>(<issue>12</issue>):<comment>e1009688</comment>. <ext-link ext-link-type="uri" xlink:href="https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1009688">https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1009688</ext-link>, doi: <pub-id pub-id-type="doi">10.1371/journal.pcbi.1009688,</pub-id> <publisher-name>publisher: Public Library of Science</publisher-name>.</mixed-citation></ref>
</ref-list>
<app-group>
<app id="app1">
<title>Appendix 1</title>
<sec id="s6">
<title>Full Details of Task and Procedure</title>
<sec id="s6-1">
<title>Recruitment</title>
<p>Participants were recruited from the pool of Amazon Mechanical Turk (MTurk) vetted by cloudresearch.com (<xref ref-type="bibr" rid="c33">Hauser et al., 2022</xref>; <xref ref-type="bibr" rid="c37">Litman et al., 2017</xref>). We further restricted enrollment to participants with an approval rating higher than 95%, and at least 100 prior jobs completed. Enrollment was also restricted to participants registered as being 18-35 years old. Participants were presented with an ad for a multi-session study, describing base pay and the performance-dependent bonus. The ad stated that we were only looking for participants who were willing to complete all four sessions, and that the task required undivided attention.</p>
<p>After accepting the task, participants were directed to a website running the experiment, which was coded using jsPsych 6.0.4 (<xref ref-type="bibr" rid="c22">De Leeuw, 2015</xref>).</p>
</sec>
<sec id="s6-2">
<title>Instructions and Training</title>
<p>At the beginning of the first session, participants were thoroughly instructed about the task, and practiced each of its phases. First, the two-phase structure and the learning goal were introduced. Participants then were shown an example table with two decks on it. They practiced choosing a deck on a single table. Participants were instructed that only the symbol on the deck determined its identity, while its location (which was randomly determined) did not matter. Participants then completed ten practice trials limited to choosing decks on a single table, with the goal of figuring out the difference in proportions of colors between the decks. In this practice, one deck had 7/10 cards of the rewarding color, and the other 9/10 cards. After practicing choosing decks, participants were told which had more of the rewarding color, and it was explained that the differences could be more difficult to figure out in the game itself. Next, participants were told that while both decks may have a majority of cards of the same color. As such, they would have to sample from both decks to learn which had more of each color. This point was demonstrated by presenting ten cards from two decks -and asking (i) which deck had a majority of color 1 cards (both did), (ii) which had a majority of color 2 cards (none did), (iii) which deck had more of color 1 than the other (one of the decks did), and (iv) which deck had more of color 2 than the other (the other deck did). A failure to give a correct answer on all four questions prompted a repetition of this section.</p>
<p>Participants were then introduced to choosing a table before making a deck choice. As practice, they were tasked with revealing a single card from each deck on two tables over 4 trials. Failing to do so resulted in participants having to repeat this section.</p>
<p>Next, the test phase was introduced. Participants were instructed that a particular deck had more of the rewarding color in it. They then had to choose that deck on the following test screen. Having successfully done so, they received the same visual feedback they would receive in the actual task - ten cards of each deck were presented to them, demonstrating the true proportion of colors in each deck. A message was displayed alongside this demonstration, stating the accuracy of their choice.</p>
<p>Before starting work on the main task, participants were reminded that the test phase could commence on any given trial. They then answered six multiple-choice questions about the structure of the task and their goal. If they got any of these wrong, they were instructed on the correct answer, and then had to repeat the quiz. After successfully completing the quiz, they played the whole task over a 20-trials-long practice round. For this practice, only two tables were included in the set of stimuli. Finally, after completing this practice round, they continued to play four more rounds of the full task with four tables.</p>
<p>Each of the following three sessions began with a short reminder of the instructions. Participants then completed five rounds of the task. The first round was always a short one (12-19 trials, these are the four lowest numbers drawn from the geometric distribution of round lengths), and was treated as a practice round in analysis, i.e. data from this round was discarded.</p>
</sec>
<sec id="s6-3">
<title>Procedure for Single Round</title>
<sec id="s6-3-1">
<title>Familiarization</title>
<p>At the beginning of each round, participants were presented with each of the tables included in the round, and the two decks associated with each table. They were then tested on these associations: they were shown a table, and two decks next to it, and had to indicate which of the two belonged to the table. Failure to answer correctly on all eight trials resulted in a repetition of the introduction. Following the introductions of tables and decks, participants were shown the two card colors used in this round. They were reminded about their goal, and then commenced to play the learning phase.</p>
</sec>
<sec id="s6-3-2">
<title>Learning Phase</title>
<p>Each trial of the learning phase began with a 500ms period during which a fixation cross was displayed at the centre of the screen. Two tables were then presented as choice options. Choice options were chosen for each trial at random, and the left-right presentation of the two choice options was also determined randomly. Participants indicated their choice of table by pressing either the ‘d’ or ‘k’ keys. Next, the unchosen table was removed from the screen, and a frame with the same pattern as the chosen table was presented for a period of 1000 ms. This was followed by a 500ms presentation of a fixation cross at the center of this frame. Then, the two decks associated with the table appeared in the frame (deck location was determined randomly). Using the same keys, participants chose to reveal a card on one of the decks. A short animation was played showing the deck being shuffled, and a single card was flipped to reveal its color. The colored card remained on screen for 2400 ms. A 1700 ms inter-trial-interval followed each trial.</p>
</sec>
<sec id="s6-3-3">
<title>Test Phase</title>
<p>At the end of the learning phase, participants were instructed that they will now be tested on their learning, and were shown the color designated as the rewarding color for this round. They then chose the deck they believed had more of the rewarding color on each table. After making all four choices, participants were shown their choice on each table and asked to rate their confidence that their choice was indeed the correct one.</p>
</sec>
<sec id="s6-3-4">
<title>Memory Test</title>
<p>Following the confidence ratings, participants were tested on their memory of table and deck associations. They were shown each deck participating in the round, and had to indicate which of the four tables it belonged to. Participants received no feedback for this memory test.</p>
</sec>
<sec id="s6-3-5">
<title>Feedback</title>
<p>Next, participants received feedback for their test-phase choices. For each table they were shown ten cards drawn from each deck. The cards represented the true proportion of colors in the deck. They were reminded which deck they had chosen, and were told whether that was the correct choice. After observing this for each of the four tables, participants were told how much bonus money they earned in this round.</p>
</sec>
</sec>
<sec id="s6-4">
<title>Debriefing</title>
<p>At the end of the experiment, after collecting demographic information, participants were asked about any strategy they may have implemented to remember the card colors better, to choose between decks and between tables in the learning phase, and between decks in the test phase. Lastly they were asked if anything in the instructions remained unclear. These responses were evaluated for any use of external aids, such as pen and paper, and for any technical difficulties.</p>
</sec>
<sec id="s6-5">
<title>Compliance and Attention Checks</title>
<p>We implemented several measures to incentivize participants to devote their undivided attention to the task. First, the task ran in full screen mode, and if participants chose to exit it, a warning message was shown explaining that this study only runs in fullscreen mode. Additionally, refreshing the webpage and right-clicking on it were disabled. Instances where the participant interacted with another application on their computer were recorded by jsPsych, and when this occurred a warning message was displayed on screen.</p>
<p>Participants had to make each learning-phase choice within 3000 ms. Failing to do so resulted in the display of a warning message asking them to choose more quickly. Additionally, if a reaction time (RT) of less than 250 ms was recorded on three consecutive choices, a warning message asking participants to comply with instructions was displayed on screen.</p>
<p>When more than ten warning messages of any kind had been displayed, the session terminated, and participants were asked to return the job to MTurk. Participants were paid for terminated sessions, but were not invited to following sessions.</p>
</sec>
</sec>
</app>
<app id="app2">
<title>Appendix 2</title>
<sec id="s7">
<title>Preliminary Sample</title>
<sec id="s7-1">
<title>Data Collection and Participants</title>
<p>A preliminary sample of 70 participants was recruited via MTurk to participate in four sessions of the task. Task design was identical to that later used for the pre-registered sample. Recruitment was not limited to the cloudresearch.com approved sample, since data collection predated the proliferation of fake worker accounts on MTurk (<xref ref-type="bibr" rid="c33">Hauser et al.,2022</xref>).</p>
<p>The first session was terminated early for 7 participants due to recorded interactions with other applications during the experiment, or failure to comply with instructions. An additional 8 sessions played by participants who had successfully completed the first session were excluded for the same reasons. Five further sessions were excluded for failure to sample cards from both decks, a prerequisite for learning on which participants were instructed as part of the training. Altogether, data from 62 participants was included in the analyzed sample (33 female, 28 male, 1 other gender, average age 29.10, range 20-38). This sample included 62 first sessions, 45 second sessions, 33 third sessions, and 29 fourth sessions.</p>
</sec>
<sec id="s7-2">
<title>Results</title>
<p>Results from the pre-registered sample largely replicate the results we first observed in the preliminary sample (see matching figure supplement for each figure). Two points of divergence are observed. In the preliminary sample, we didn’t see a significant correlation between the tendency to avoid uncertainty when overall uncertainty is high and testperformance, while in the larger pre-registered sample we observed a positive significant correlation. This difference could be due to sample size, as the correlation is weak, we shouldn’t over interpret it.</p>
<p>A second point of divergence regards the plotting of predicted RTs by test performance tertile. While both in the preliminary sample and the pre-registered sample we find that the bound height estimated by the DDM is predictive of test performance(Supplementary <xref ref-type="table" rid="app3-tbl12">Table 12</xref>), in the preliminary sample this relationship does not translate to a monotonic rising of bound height when comparing participants by test performance tertile (Supplementary <xref ref-type="fig" rid="fig4">Figure 4b</xref>). We hesitate to interpret any non-linear relationship between the bound height and test performance, given the small size of the sample, the divergence from the pre-registered sample, and the complexity of the models involved.</p>
</sec>
</sec>
</app>
<app id="app3">
<title>Appendix 3</title>
<table-wrap id="app3-tbl1" orientation="portrait" position="float">
<label>Table 1.</label>
<caption><p>Test Accuracy as a Function of the Final Uncertainty in the Exploration-Phase.</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl1.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (12379 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (3482 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p>Predictors</p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>0.86</p></td>
<td valign="top" align="center"><p>[0.61, 1.11]</p></td>
<td valign="top" align="center"><p>0.90</p></td>
<td valign="top" align="center"><p>[0.47, 1.32]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Final uncertainty</p></td>
<td valign="top" align="center"><p>–5.59</p></td>
<td valign="top" align="center"><p>[-6.25, -4.95]</p></td>
<td valign="top" align="center"><p>–5.69</p></td>
<td valign="top" align="center"><p>[-6.86, -4.68]</p></td>
<td valign="top" align="center"><p>logit/nats</p></td>
</tr>
<tr>
<td colspan="6" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>1.58</p></td>
<td valign="top" align="center"><p>[1.38, 1.82]</p></td>
<td valign="top" align="center"><p>1.45</p></td>
<td valign="top" align="center"><p>[1.13, 1.89]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of final</p></td>
<td valign="top" align="center"><p>7.49</p></td>
<td valign="top" align="center"><p>[6.54, 8.61]</p></td>
<td valign="top" align="center"><p>6.83</p></td>
<td valign="top" align="center"><p>[5.33, 8.88]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>uncertainty</p><p>Correlation of intercept and final uncertainty</p></td>
<td valign="top" align="center"><p>–0.68</p></td>
<td valign="top" align="center"><p>[-0.84, -0.46]</p></td>
<td valign="top" align="center"><p>–0.81</p></td>
<td valign="top" align="center"><p>[-0.97, -0.40]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl1-fn1"><p>The model can be summarized with the following R syntax formula: <italic>accuracy</italic> ~ 0.5 + 0.5 * <italic>inv_logit(1</italic> + <italic>final uncertainty</italic> + (1 + <italic>final uncertainty\participant)), where inv_logit</italic> is the inverse logit function. This functional form limits predicted accuracy between 0.5 and 1.0, since guessing-level accuracy on a two-alternative forced-choice test is 0.5. Since accuracy is a binary variable, this regression was fit with a Bernoulli likelihood.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl2" orientation="portrait" position="float">
<label>Table 2.</label>
<caption><p>Test Confidence as a Function of the Final Uncertainty in the Exploration-Phase.</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl2.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (12007 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (3362 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Threshold 1</p></td>
<td valign="top" align="center"><p>–2.22</p></td>
<td valign="top" align="center"><p>[-2.44, -2.00]</p></td>
<td valign="top" align="center"><p>–2.11</p></td>
<td valign="top" align="center"><p>[-2.49, -1.72]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Threshold 2</p></td>
<td valign="top" align="center"><p>–0.44</p></td>
<td valign="top" align="center"><p>[-0.67, -0.23]</p></td>
<td valign="top" align="center"><p>–0.20</p></td>
<td valign="top" align="center"><p>[-0.58, 0.17]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Threshold 3</p></td>
<td valign="top" align="center"><p>1.31</p></td>
<td valign="top" align="center"><p>[1.08, 1.53]</p></td>
<td valign="top" align="center"><p>1.60</p></td>
<td valign="top" align="center"><p>[1.22, 1.98]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Threshold 4</p></td>
<td valign="top" align="center"><p>3.11</p></td>
<td valign="top" align="center"><p>[2.88, 3.34]</p></td>
<td valign="top" align="center"><p>3.58</p></td>
<td valign="top" align="center"><p>[3.18, 3.98]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Final uncertainty</p></td>
<td valign="top" align="center"><p>–2.48</p></td>
<td valign="top" align="center"><p>[-2.93, -2.06]</p></td>
<td valign="top" align="center"><p>–2.32</p></td>
<td valign="top" align="center"><p>[-3.11, -1.56]</p></td>
<td valign="top" align="center"><p>logit/nats</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Choice accuracy</p></td>
<td valign="top" align="center"><p>1.09</p></td>
<td valign="top" align="center"><p>[0.92, 1.27]</p></td>
<td valign="top" align="center"><p>1.05</p></td>
<td valign="top" align="center"><p>[0.71, 1.39]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Final uncertainty x choice accuracy</p></td>
<td valign="top" align="center"><p>–3.10</p></td>
<td valign="top" align="center"><p>[-3.76, -2.46]</p></td>
<td valign="top" align="center"><p>–3.20</p></td>
<td valign="top" align="center"><p>[-4.60, -1.93]</p></td>
<td valign="top" align="center"><p>logit/nats</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>1.45</p></td>
<td valign="top" align="center"><p>[1.31, 1.63]</p></td>
<td valign="top" align="center"><p>1.43</p></td>
<td valign="top" align="center"><p>[1.19, 1.73]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of final</p></td>
<td valign="top" align="center"><p>2.10</p></td>
<td valign="top" align="center"><p>[1.73, 2.52]</p></td>
<td valign="top" align="center"><p>2.12</p></td>
<td valign="top" align="center"><p>[1.39, 2.97]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>uncertainty SD of choice accuracy</p></td>
<td valign="top" align="center"><p>0.81</p></td>
<td valign="top" align="center"><p>[0.64, 0.99]</p></td>
<td valign="top" align="center"><p>0.84</p></td>
<td valign="top" align="center"><p>[0.52, 1.25]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of uncertainty × accuracy</p></td>
<td valign="top" align="center"><p>2.11</p></td>
<td valign="top" align="center"><p>[1.46, 2.82]</p></td>
<td valign="top" align="center"><p>3.02</p></td>
<td valign="top" align="center"><p>[1.58, 4.59]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and uncertainty</p></td>
<td valign="top" align="center"><p>0.23</p></td>
<td valign="top" align="center"><p>[0.03, 0.41]</p></td>
<td valign="top" align="center"><p>0.07</p></td>
<td valign="top" align="center"><p>[-0.26, 0.39]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and accuracy</p></td>
<td valign="top" align="center"><p>–0.20</p></td>
<td valign="top" align="center"><p>[-0.39, 0.01]</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[-0.33, 0.40]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of uncertainty and accuracy</p></td>
<td valign="top" align="center"><p>–0.57</p></td>
<td valign="top" align="center"><p>[-0.81, -0.30]</p></td>
<td valign="top" align="center"><p>–0.51</p></td>
<td valign="top" align="center"><p>[-0.83, -0.03]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and uncertainty × accuracy</p></td>
<td valign="top" align="center"><p>0.27</p></td>
<td valign="top" align="center"><p>[-0.01, 0.52]</p></td>
<td valign="top" align="center"><p>0.28</p></td>
<td valign="top" align="center"><p>[-0.11, 0.62]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of uncertainty and uncertainty × accuracy</p></td>
<td valign="top" align="center"><p>0.68</p></td>
<td valign="top" align="center"><p>[0.34, 0.91]</p></td>
<td valign="top" align="center"><p>0.26</p></td>
<td valign="top" align="center"><p>[-0.24, 0.70]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of accuracy and uncertainty × accuracy</p></td>
<td valign="top" align="center"><p>–0.84</p></td>
<td valign="top" align="center"><p>[-0.95, -0.61]</p></td>
<td valign="top" align="center"><p>–0.23</p></td>
<td valign="top" align="center"><p>[-0.66, 0.35]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl2-fn1"><p>The model can be summarized with the following R syntax formula: <italic>confidence</italic> ~ 0 + <italic>final uncertainty</italic> * <italic>accuracy</italic> + (1 + <italic>final uncertainty</italic> * <italic>accuracy |PID).</italic> This model was fit as an ordered-logisticregression,with four threshold variables since confidence was rated on a 5-point likert scale (<bold>?</bold>).</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl3" orientation="portrait" position="float">
<label>Table 3.</label>
<caption><p>Exploration-Phase Choices as a Function of Δ-Uncertainty</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl3.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (146766 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (41009 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.10, 0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.11, 0.06]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>?-uncertainty</p></td>
<td valign="top" align="center"><p>0.89</p></td>
<td valign="top" align="center"><p>[0.76, 1.03]</p></td>
<td valign="top" align="center"><p>1.01</p></td>
<td valign="top" align="center"><p>[0.70, 1.30]</p></td>
<td valign="top" align="center"><p>logit / nat</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.39</p></td>
<td valign="top" align="center"><p>[0.35, 0.43]</p></td>
<td valign="top" align="center"><p>0.36</p></td>
<td valign="top" align="center"><p>[0.30, 0.45]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of</p></td>
<td valign="top" align="center"><p>0.89</p></td>
<td valign="top" align="center"><p>[0.79, 1.00]</p></td>
<td valign="top" align="center"><p>1.11</p></td>
<td valign="top" align="center"><p>[0.91, 1.38]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>?-uncertainty Correlation of intercept and ?-uncertainty</p></td>
<td valign="top" align="center"><p>–0.05</p></td>
<td valign="top" align="center"><p>[-0.20, 0.12]</p></td>
<td valign="top" align="center"><p>0.07</p></td>
<td valign="top" align="center"><p>[-0.21, 0.31]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl3-fn1"><p>The model can be summarized with the following R syntax formula: <italic>table on right chosen</italic> ~ 1 + Δ <italic>uncertainty</italic> + (1 + Δ <italic>uncertainty\participant).</italic> This model was fit as a logistic regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl4" orientation="portrait" position="float">
<label>Table 4.</label>
<caption><p>Exploration-Phase Choices as a Function of Δ-EIG</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl4.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (146766 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (41009 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.09, 0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.12, 0.07]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-EIG</p></td>
<td valign="top" align="center"><p>10.12</p></td>
<td valign="top" align="center"><p>[8.00, 12.47]</p></td>
<td valign="top" align="center"><p>12.50</p></td>
<td valign="top" align="center"><p>[7.97, 16.91]</p></td>
<td valign="top" align="center"><p>logit / nat</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.39</p></td>
<td valign="top" align="center"><p>[0.35, 0.43]</p></td>
<td valign="top" align="center"><p>0.35</p></td>
<td valign="top" align="center"><p>[0.29, 0.43]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-EIG</p></td>
<td valign="top" align="center"><p>1.21</p></td>
<td valign="top" align="center"><p>[1.07, 1.37]</p></td>
<td valign="top" align="center"><p>1.48</p></td>
<td valign="top" align="center"><p>[1.21, 1.82]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and Δ-EIG</p></td>
<td valign="top" align="center"><p>–0.02</p></td>
<td valign="top" align="center"><p>[-0.17, 0.14]</p></td>
<td valign="top" align="center"><p>–0.11</p></td>
<td valign="top" align="center"><p>[-0.35, 0.16]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl4-1fn1"><p>The model can be summarized with the following R syntax formula: <italic>table onright chosen</italic> ~ <italic>1+</italic>Δ<italic>EIG</italic>+ (1 + Δ <italic>EIG\participant).</italic> This model was fit as a logistic regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl5" orientation="portrait" position="float">
<label>Table 5.</label>
<caption><p>Exploration-Phase Choices as a Function of Δ-Exposure</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl5.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (146766 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminarysample (41009 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.10, 0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.13, 0.07]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-exposure</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.04, -0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.05, -0.02]</p></td>
<td valign="top" align="center"><p>logit / trial</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.39</p></td>
<td valign="top" align="center"><p>[0.35, 0.43]</p></td>
<td valign="top" align="center"><p>0.36</p></td>
<td valign="top" align="center"><p>[0.29, 0.44]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-exposure</p></td>
<td valign="top" align="center"><p>0.05</p></td>
<td valign="top" align="center"><p>[0.04, 0.06]</p></td>
<td valign="top" align="center"><p>0.05</p></td>
<td valign="top" align="center"><p>[0.04, 0.06]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and Δ-exposure</p></td>
<td valign="top" align="center"><p>0.16</p></td>
<td valign="top" align="center"><p>[-0.01, 0.32]</p></td>
<td valign="top" align="center"><p>–0.08</p></td>
<td valign="top" align="center"><p>[-0.36, 0.21]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl5-fn1"><p>The model can be summarized with the following R syntax formula: <italic>table on right chosen</italic> ~ 1 + Δ <italic>exposure</italic> + (1 + A <italic>exposure\participant).</italic> This model was fit as a logistic regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl6" orientation="portrait" position="float">
<label>Table 6.</label>
<caption><p>Exploration-Phase Choices as a Function of ΔΔ-uncertainty and Overall Uncertainty</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl6.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (146766 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (41009 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.10, 0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.13, 0.07]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty</p></td>
<td valign="top" align="center"><p>0.97</p></td>
<td valign="top" align="center"><p>[0.83, 1.11]</p></td>
<td valign="top" align="center"><p>1.12</p></td>
<td valign="top" align="center"><p>[0.83, 1.42]</p></td>
<td valign="top" align="center"><p>logit /</p></td>
</tr>
<tr>
<td valign="top" align="left"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p>nat</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty × overall</p></td>
<td valign="top" align="center"><p>–428.44</p></td>
<td valign="top" align="center"><p>[-536.60, -339.27]</p></td>
<td valign="top" align="center"><p>–444.27</p></td>
<td valign="top" align="center"><p>[-559.73, -353.23]</p></td>
<td valign="top" align="center"><p>logit / nat<sup>2</sup></p></td>
</tr>
<tr>
<td valign="top" align="left"><p>uncertainty Transformed threshold α</p></td>
<td valign="top" align="center"><p>2.52</p></td>
<td valign="top" align="center"><p>[2.40, 2.64]</p></td>
<td valign="top" align="center"><p>2.33</p></td>
<td valign="top" align="center"><p>[2.17, 2.49]</p></td>
<td valign="top" align="center"><p>a.u.</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.39</p></td>
<td valign="top" align="center"><p>[0.35, 0.43]</p></td>
<td valign="top" align="center"><p>0.36</p></td>
<td valign="top" align="center"><p>[0.30, 0.44]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-uncertainty</p></td>
<td valign="top" align="center"><p>0.92</p></td>
<td valign="top" align="center"><p>[0.81, 1.04]</p></td>
<td valign="top" align="center"><p>1.09</p></td>
<td valign="top" align="center"><p>[0.89, 1.35]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>12.85</p></td>
<td valign="top" align="center"><p>[0.56, 41.41]</p></td>
<td valign="top" align="center"><p>8.97</p></td>
<td valign="top" align="center"><p>[0.46, 28.35]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of transformed threshold</p></td>
<td valign="top" align="center"><p>0.45</p></td>
<td valign="top" align="center"><p>[0.38, 0.54]</p></td>
<td valign="top" align="center"><p>0.35</p></td>
<td valign="top" align="center"><p>[0.26, 0.47]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl6-fn1"><p>This model can be summarized with the following formula: <italic>logit(P(tableonrightchosen))</italic> = <italic>Intercept+ b</italic><sub>1</sub> * Δ <italic>uncertainty</italic> + <italic>b<sub>2</sub></italic> * <italic>(overall uncertainty</italic> — ω) * <italic>step(overall uncertainty</italic> — ω) * Δ <italic>uncertainty, where</italic> step is the step function, ω = —2ln(0.5) * <italic>inv_logit(a).</italic> The intercept and parameters <italic>b<sub>1</sub>, b<sub>2</sub>,</italic> and α all vary by participant. This model was fit as a logistic regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl7" orientation="portrait" position="float">
<label>Table 7.</label>
<caption><p>Test Performance as a Function of the Tendency to Approach Uncertainty in Exploration</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl7.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>1. 56</p></td>
<td valign="top" align="center"><p>[1.51, 1.61]</p></td>
<td valign="top" align="center"><p>1. 62</p></td>
<td valign="top" align="center"><p>[1.53, 1.72]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Approach tendency</p></td>
<td valign="top" align="center"><p>2.96</p></td>
<td valign="top" align="center"><p>[2.67, 3.25]</p></td>
<td valign="top" align="center"><p>3.09</p></td>
<td valign="top" align="center"><p>[2.65, 3.57]</p></td>
<td valign="top" align="center"><p>logit<sup>2</sup> / nat</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl7-fn1"><p>The model can be summarized with the following R syntax formula: <italic>test accuracy ~</italic> 1 + <italic>approach tendency.</italic> For the tendency to approach uncertainty we computed the mean posterior approach parameter for each participant in the model described in <xref ref-type="table" rid="app3-tbl6">Table 6</xref>. The model described here was fit as a logistic regression with binomial likelihood.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl8" orientation="portrait" position="float">
<label>Table 8.</label>
<caption><p>Test Performance as a Function of Tendency to Approach Uncertainty in Exploration when Overall Uncertainty is High</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl8.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th colspan="2" valign="top" align="center"><p>95% PI Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>1. 52</p></td>
<td colspan="2" valign="top" align="center"><p>[1.48, 1.57] 1.59</p></td>
<td valign="top" align="center"><p>[1.50, 1.68]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Avoid tendency</p></td>
<td valign="top" align="center"><p>1. 18</p></td>
<td colspan="2" valign="top" align="center"><p>[0.80, 1.58] -0.52</p></td>
<td valign="top" align="center"><p>[-1.20, 0.21]</p></td>
<td valign="top" align="center"><p>nat2</p></td>
</tr>
<tr>
<td valign="top" align="left"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td colspan="3" valign="top" align="center"><p>Participant-wise variability</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl8-fn1"><p>The model can be summarized with the following R syntax formula: <italic>testaccuracy ~</italic> 1+ <italic>avoidtendency.</italic> For the tendency to avoid uncertainty when overall uncertainty is high a we computed for each participant the area under the curve of the uncertainty approach / avoid graph, averaging across the posterior of the model described in <xref ref-type="table" rid="app3-tbl6">Table 6</xref>. The model described here was fit as a logistic regression with binomial likelihood.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl9" orientation="portrait" position="float">
<label>Table 9.</label>
<caption><p>Drift Diffusion Model of Exploration-Phase Choice and RTs</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl9.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (113746 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (31205 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI Units</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>B - bound height</p></td>
<td valign="top" align="center"><p>0.74</p></td>
<td valign="top" align="center"><p>[0.72, 0.76]</p></td>
<td valign="top" align="center"><p>0.74</p></td>
<td valign="top" align="center"><p>[0.71, 0.76]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td valign="top" align="left"><p>μ<sub>0</sub> - drift rate offset</p></td>
<td valign="top" align="center"><p>–0.01</p></td>
<td valign="top" align="center"><p>[-0.06, 0.03]</p></td>
<td valign="top" align="center"><p>–0.01</p></td>
<td valign="top" align="center"><p>[-0.08, 0.06]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td valign="top" align="left"><p><italic>k</italic>- dependence of drift rate on</p></td>
<td valign="top" align="center"><p>0.69</p></td>
<td valign="top" align="center"><p>[0.58, 0.78]</p></td>
<td valign="top" align="center"><p>0.78</p></td>
<td valign="top" align="center"><p>[0.58, 0.99]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td valign="top" align="left"><p>uncertainty <bold><italic>t<sub>ND</sub></italic></bold> non-decision time</p></td>
<td valign="top" align="center"><p>0.28</p></td>
<td valign="top" align="center"><p>[0.26, 0.31]</p></td>
<td valign="top" align="center"><p>0.27</p></td>
<td valign="top" align="center"><p>[0.25, 0.31]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td colspan="6" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of B</p></td>
<td valign="top" align="center"><p>0.12</p></td>
<td valign="top" align="center"><p>[0.11, 0.14]</p></td>
<td valign="top" align="center"><p>0.10</p></td>
<td valign="top" align="center"><p>[0.08, 0.12]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td valign="top" align="left"><p>SD of μ<italic><sub>0</sub></italic></p></td>
<td valign="top" align="center"><p>0.30</p></td>
<td valign="top" align="center"><p>[0.27, 0.33]</p></td>
<td valign="top" align="center"><p>0.25</p></td>
<td valign="top" align="center"><p>[0.21, 0.31]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td valign="top" align="left"><p>SD of K</p></td>
<td valign="top" align="center"><p>0.66</p></td>
<td valign="top" align="center"><p>[0.58, 0.74]</p></td>
<td valign="top" align="center"><p>0.76</p></td>
<td valign="top" align="center"><p>[0.62, 0.94]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
<tr>
<td valign="top" align="left"><p>SD of t<sub>ND</sub></p></td>
<td valign="top" align="center"><p>0.16</p></td>
<td valign="top" align="center"><p>[0.14, 0.19]</p></td>
<td valign="top" align="center"><p>0.12</p></td>
<td valign="top" align="center"><p>[0.10, 0.15]</p></td>
<td valign="top" align="center"><p> </p></td></tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl9-fn1"><p>We used a drift-diffusion model to formalize the dependence of RTs and choice on evidence. A drift-diffusion model is one variant in the sequential sampling family of models. The model posits that samples of momentary evidence are integrated over time. The expectation of the momentary evidence distribution is termed the drift rate <italic>μ,</italic> and its standard deviation is termed the diffusion coefficient. The decision is made when integrated evidence reaches an upper or lower bound (±<italic>B</italic>), whose sign determines the choice. Processes external to decision making are modelled by <italic>t<sub>ND</sub>,</italic> a constant added to the RT. In this model μ is allowed to depend linearly on Δ-uncertainty, μ<italic> = μ<sub>0</sub> + <sub>K</sub> ‧ Δ - uncertainty,</italic> such that <sc><italic>k</italic></sc> captures the dependence of drift rate on Δ-uncertainty, and μ<sub><italic>0</italic></sub> is a general bias to make rightward or leftward choices.</p></fn>
<fn id="app3-tbl9-fn2"><p>Prior to fitting the model to the data, we excluded trials for which overall uncertainty was above the threshold estimated for each participant by the model described in <xref ref-type="table" rid="app3-tbl6">Table 6</xref>. As we find qualitatively different choice behavior above the threshold, we couldn’t justify modelling these trials together with the majority of trials. Fitting a piecewise regression DDM model was beyond the capabilities of current software.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl10" orientation="portrait" position="float">
<label>Table 10.</label>
<caption><p>Test Performance as a Function of Drift Diffusion Model Parameters for Exploration Phase</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl10.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (113746 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (31205 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI Units</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>0.33</p></td>
<td valign="top" align="center"><p>[-0.34, 1.01]</p></td>
<td valign="top" align="center"><p>0.23</p></td>
<td valign="top" align="center"><p>[-0.86, 1.28]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Final uncertainty</p></td>
<td valign="top" align="center"><p>–0.87</p></td>
<td valign="top" align="center"><p>[-0.98, -0.76]</p></td>
<td valign="top" align="center"><p>–0.93</p></td>
<td valign="top" align="center"><p>[-1.13, -0.74]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>B - bound height</p></td>
<td valign="top" align="center"><p>1.46</p></td>
<td valign="top" align="center"><p>[0.58, 2.34]</p></td>
<td valign="top" align="center"><p>1.49</p></td>
<td valign="top" align="center"><p>[0.05, 2.90]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>k</italic>- dependence of drift rate on uncertainty</p></td>
<td valign="top" align="center"><p>0.81</p></td>
<td valign="top" align="center"><p>[0.58, 1.07]</p></td>
<td valign="top" align="center"><p>0.88</p></td>
<td valign="top" align="center"><p>[0.57, 1.21]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td colspan="6" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.87</p></td>
<td valign="top" align="center"><p>[0.74, 1.01]</p></td>
<td valign="top" align="center"><p>0.68</p></td>
<td valign="top" align="center"><p>[0.47, 0.96]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of final uncertainty</p></td>
<td valign="top" align="center"><p>0.49</p></td>
<td valign="top" align="center"><p>[0.39, 0.60]</p></td>
<td valign="top" align="center"><p>0.44</p></td>
<td valign="top" align="center"><p>[0.24, 0.67]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and final uncertainty</p></td>
<td valign="top" align="center"><p>-0.81</p></td>
<td valign="top" align="center"><p>[-0.92, -0.66]</p></td>
<td valign="top" align="center"><p>-0.83</p></td>
<td valign="top" align="center"><p>[-0.98, -0.44]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl10-fn1"><p>This model can be summarized with the following R syntax formula: <italic>test accuracy ~</italic> 1 + <italic>final uncertainty</italic> + <italic>B</italic> + <italic>K</italic> + (1 + <italic>final uncertainty\participant).</italic> This model was fit as a logistic regression. As B and <sc><italic>k</italic></sc> are parameters estimated from the model described in <xref ref-type="table" rid="app3-tbl9">Table 9</xref>, we took into account our error in measuring them when using them as predictors in this model. Thus, the posterior distribution for each participant’s B and <sc><italic>k</italic></sc> parameters was summarized as a mean and standard deviation. These summary statistics were used to approximate the posterior as a normal distribution from which a latent variable was drawn during the estimation of this model. This method propagates the uncertainty in the values of B and <sc><italic>k</italic></sc> into the estimates reported here. Prior to using this method, we inspected the posteriors from the model summarized in <xref ref-type="table" rid="app3-tbl9">Table 9</xref>, and made sure the normal distribution is an adequate approximation for these posteriors.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl11" orientation="portrait" position="float">
<label>Table 11.</label>
<caption><p>Exploration-phase Choices as a Function of Δ-Uncertainty, Overall Uncertainty, and Side of Repeat Option</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl11.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (146766 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (41009 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.10, 0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.14, 0.07]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Δ-uncertainty</p></td>
<td valign="top" align="center"><p>1.01</p></td>
<td valign="top" align="center"><p>[0.86, 1.14]</p></td>
<td valign="top" align="center"><p>1.16</p></td>
<td valign="top" align="center"><p>[0.87, 1.44]</p></td>
<td valign="top" align="center"><p>logit / nat</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>–326.14</p></td>
<td valign="top" align="center"><p>[-468.61, -245.18]</p></td>
<td valign="top" align="center"><p>–349.54</p></td>
<td valign="top" align="center"><p>[-691.47, -251.13]</p></td>
<td valign="top" align="center"><p>logit / nat<sup>2</sup></p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Transformed threshold</p></td>
<td valign="top" align="center"><p>2.55</p></td>
<td valign="top" align="center"><p>[2.40, 2.72]</p></td>
<td valign="top" align="center"><p>2.36</p></td>
<td valign="top" align="center"><p>[2.16, 2.67]</p></td>
<td valign="top" align="center"><p>a.u.</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Repeat choice on right</p></td>
<td valign="top" align="center"><p>0.50</p></td>
<td valign="top" align="center"><p>[0.42, 0.59]</p></td>
<td valign="top" align="center"><p>0.57</p></td>
<td valign="top" align="center"><p>[0.43, 0.72]</p></td>
<td valign="top" align="center"><p>logit difference</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="center"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.40</p></td>
<td valign="top" align="center"><p>[0.36, 0.44]</p></td>
<td valign="top" align="center"><p>0.38</p></td>
<td valign="top" align="center"><p>[0.31, 0.46]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>SD of Δ-uncertainty</p></td>
<td valign="top" align="center"><p>0.90</p></td>
<td valign="top" align="center"><p>[0.80, 1.02]</p></td>
<td valign="top" align="center"><p>1.05</p></td>
<td valign="top" align="center"><p>[0.86, 1.31]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>SD of Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>14.21</p></td>
<td valign="top" align="center"><p>[0.98, 41.45]</p></td>
<td valign="top" align="center"><p>9.23</p></td>
<td valign="top" align="center"><p>[0.46, 29.28]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>SD of transformed threshold</p></td>
<td valign="top" align="center"><p>0.44</p></td>
<td valign="top" align="center"><p>[0.36, 0.54]</p></td>
<td valign="top" align="center"><p>0.35</p></td>
<td valign="top" align="center"><p>[0.22, 0.50]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>SD of repeat choice</p></td>
<td valign="top" align="center"><p>0.57</p></td>
<td valign="top" align="center"><p>[0.51, 0.64]</p></td>
<td valign="top" align="center"><p>0.53</p></td>
<td valign="top" align="center"><p>[0.44, 0.66]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of intercept and Δ-uncertainty</p></td>
<td valign="top" align="center"><p>-0.06</p></td>
<td valign="top" align="center"><p>[-0.21, 0.10]</p></td>
<td valign="top" align="center"><p>0.05</p></td>
<td valign="top" align="center"><p>[-0.21, 0.32]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of intercept and Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>-0.08</p></td>
<td valign="top" align="center"><p>[-0.76, 0.72]</p></td>
<td valign="top" align="center"><p>-0.01</p></td>
<td valign="top" align="center"><p>[-0.76, 0.75]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of Δ-uncertainty and Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>0.01</p></td>
<td valign="top" align="center"><p>[-0.69, 0.69]</p></td>
<td valign="top" align="center"><p>-0.10</p></td>
<td valign="top" align="center"><p>[-0.78, 0.70]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of intercept and repeat choice</p></td>
<td valign="top" align="center"><p>-0.16</p></td>
<td valign="top" align="center"><p>[-0.30, -0.00]</p></td>
<td valign="top" align="center"><p>-0.06</p></td>
<td valign="top" align="center"><p>[-0.31, 0.22]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of Δ-uncertainty and repeat choice</p></td>
<td valign="top" align="center"><p>0.32</p></td>
<td valign="top" align="center"><p>[0.17, 0.46]</p></td>
<td valign="top" align="center"><p>0.11</p></td>
<td valign="top" align="center"><p>[-0.18, 0.38]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of Δ-uncertainty × overall uncertainty and repeat choice</p></td>
<td valign="top" align="center"><p>0.38</p></td>
<td valign="top" align="center"><p>[-0.67, 0.87]</p></td>
<td valign="top" align="center"><p>0.00</p></td>
<td valign="top" align="center"><p>[-0.76, 0.77]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of intercept and threshold</p></td>
<td valign="top" align="center"><p>-0.03</p></td>
<td valign="top" align="center"><p>[-0.26, 0.22]</p></td>
<td valign="top" align="center"><p>0.31</p></td>
<td valign="top" align="center"><p>[-0.09, 0.64]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of Δ-uncertainty and threshold</p></td>
<td valign="top" align="center"><p>-0.35</p></td>
<td valign="top" align="center"><p>[-0.52, -0.16]</p></td>
<td valign="top" align="center"><p>0.09</p></td>
<td valign="top" align="center"><p>[-0.27, 0.43]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of Δ-uncertainty × overall uncertainty and threshold</p></td>
<td valign="top" align="center"><p>-0.42</p></td>
<td valign="top" align="center"><p>[-0.88, 0.61]</p></td>
<td valign="top" align="center"><p>-0.10</p></td>
<td valign="top" align="center"><p>[-0.79, 0.70]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="center"><p>Correlation of repeat choice and threshold</p></td>
<td valign="top" align="center"><p>-0.60</p></td>
<td valign="top" align="center"><p>[-0.74, -0.43]</p></td>
<td valign="top" align="center"><p>-0.38</p></td>
<td valign="top" align="center"><p>[-0.64, -0.05]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl11-fn1"><p>This model can be summarized with the following formula: <italic>logit(P(table on right chosen)) = Intercept</italic> + <italic>b1 ‧</italic> Δ — <italic>uncertainty</italic> + b2 ‧ <italic>(overall uncertainty — ω</italic>) ‧ <italic>step(overall uncertainty — ω) ‧ Δ — uncertainty</italic> + <italic>b3 ‧ repeat on right</italic>, where step is the step function, ω<italic> = —2ln(0.5) ‧ inv_logit(a).</italic> The intercept and parameters b1, b2, b3, and α all vary by participant, and their correlations across participants are modelled. This model was fit as a logistic regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl12" orientation="portrait" position="float">
<label>Table 12.</label>
<caption><p>Drift Diffusion Model of Exploration-Phase Choice and RTs, Differentiating between Repeat and Switch Choices</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl12.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (113746 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (31205 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI Units</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>B<sub>0</sub></italic> – average bound height</p></td>
<td valign="top" align="center"><p>0.75</p></td>
<td valign="top" align="center"><p>[0.73, 0.76]</p></td>
<td valign="top" align="center"><p>0.74</p></td>
<td valign="top" align="center"><p>[0.72, 0.77]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>B<sub>repeat</sub></italic> - difference in bound height between repeat and switch chosen</p></td>
<td valign="top" align="center"><p>–0.05</p></td>
<td valign="top" align="center"><p>[-0.05, -0.04]</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.06, -0.02]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>μ<sub>0</sub></italic> - drift rate offset</p></td>
<td valign="top" align="center"><p>–0.01</p></td>
<td valign="top" align="center"><p>[-0.06, 0.03]</p></td>
<td valign="top" align="center"><p>–0.01</p></td>
<td valign="top" align="center"><p>[-0.08, 0.05]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>k</italic>- dependence of drift rate on uncertainty</p></td>
<td valign="top" align="center"><p>0.70</p></td>
<td valign="top" align="center"><p>[0.61, 0.80]</p></td>
<td valign="top" align="center"><p>0.81</p></td>
<td valign="top" align="center"><p>[0.60, 1.01]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>K<bold><sub>repeat</sub></bold></italic> - difference in dependence between repeat and switch chosen</p></td>
<td valign="top" align="center"><p>–0.32</p></td>
<td valign="top" align="center"><p>[-0.43, -0.22]</p></td>
<td valign="top" align="center"><p>–0.28</p></td>
<td valign="top" align="center"><p>[-0.49, -0.08]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p><bold><italic>t<sub>ND</sub></italic></bold> - non-decision time</p></td>
<td valign="top" align="center"><p>0.28</p></td>
<td valign="top" align="center"><p>[0.26, 0.31]</p></td>
<td valign="top" align="center"><p>0.28</p></td>
<td valign="top" align="center"><p>[0.25, 0.31]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of <italic>B<sub>0</sub></italic></p></td>
<td valign="top" align="center"><p>0.12</p></td>
<td valign="top" align="center"><p>[0.11, 0.14]</p></td>
<td valign="top" align="center"><p>0.10</p></td>
<td valign="top" align="center"><p>[0.08, 0.12]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of <italic>B<sub>repeat</sub></italic></p></td>
<td valign="top" align="center"><p>0.05</p></td>
<td valign="top" align="center"><p>[0.04, 0.05]</p></td>
<td valign="top" align="center"><p>0.05</p></td>
<td valign="top" align="center"><p>[0.04, 0.07]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of μ<italic><sub>0</sub></italic></p></td>
<td valign="top" align="center"><p>0.30</p></td>
<td valign="top" align="center"><p>[0.27, 0.33]</p></td>
<td valign="top" align="center"><p>0.26</p></td>
<td valign="top" align="center"><p>[0.21, 0.31]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of K</p></td>
<td valign="top" align="center"><p>0.64</p></td>
<td valign="top" align="center"><p>[0.57, 0.72]</p></td>
<td valign="top" align="center"><p>0.74</p></td>
<td valign="top" align="center"><p>[0.60, 0.92]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of <italic>K<sub>repeai</sub></italic></p></td>
<td valign="top" align="center"><p>0.50</p></td>
<td valign="top" align="center"><p>[0.40, 0.60]</p></td>
<td valign="top" align="center"><p>0.58</p></td>
<td valign="top" align="center"><p>[0.39, 0.80]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of <italic>t<sub>ND</sub></italic></p></td>
<td valign="top" align="center"><p>0.16</p></td>
<td valign="top" align="center"><p>[0.14, 0.19]</p></td>
<td valign="top" align="center"><p>0.12</p></td>
<td valign="top" align="center"><p>[0.10, 0.15]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl12-fn1"><p>We refit the DDM described in <xref ref-type="table" rid="app3-tbl9">Table 9</xref>, allowing both B and <sc><italic>k</italic></sc> to vary by whether the choice was a repeat or switch choice (see main text for definition).</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl13" orientation="portrait" position="float">
<label>Table 13.</label>
<caption><p>Test Performance as a Function of the Tendency to Repeat Exploration-Phase Choices</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl13.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th colspan="2" valign="top" align="center"><p>95% PI Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
<td colspan="2" valign="top" align="center"><p>Predictors</p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>1. 52</p></td>
<td colspan="2" valign="top" align="center"><p>[1.47, 1.56] 1.59</p></td>
<td valign="top" align="center"><p>[1.50, 1.67]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Tendency to repeat</p></td>
<td valign="top" align="center"><p>0.09</p></td>
<td colspan="2" valign="top" align="center"><p>[0.07, 0.11] 0.11</p></td>
<td valign="top" align="center"><p>[0.06, 0.16]</p></td>
<td valign="top" align="center"><p>logit / logit</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl13-fn1"><p>The model can be summarized with the following R syntax formula: <italic>test accuracy ~</italic> 1 + <italic>tendency to repeat.</italic> For tendency to repeat we computed the mean posterior parameter for each participant in the model described in <xref ref-type="table" rid="app3-tbl11">Table 11</xref>. This model was fit as a logistic regression with binomial likelihood.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl14" orientation="portrait" position="float">
<label>Table 14.</label>
<caption><p>Exploration-Phase RTs as a Function of Memory Lag and Side of Repeat Option</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl14.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (126848 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (35264 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.34</p></td>
<td valign="top" align="center"><p>[-0.38, -0.30]</p></td>
<td valign="top" align="center"><p>–0.35</p></td>
<td valign="top" align="center"><p>[-0.42, -0.30]</p></td>
<td valign="top" align="center"><p>logs</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Memory lag</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.02, 0.03]</p></td>
<td valign="top" align="center"><p>0.03</p></td>
<td valign="top" align="center"><p>[0.02, 0.03]</p></td>
<td valign="top" align="center"><p>log s / trial</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Repeat choice on right</p></td>
<td valign="top" align="center"><p>-0.05</p></td>
<td valign="top" align="center"><p>[-0.06, -0.04]</p></td>
<td valign="top" align="center"><p>–0.05</p></td>
<td valign="top" align="center"><p>[-0.07, -0.03]</p></td>
<td valign="top" align="center"><p>log s difference</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Memory lag x repeat</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.02, 0.03]</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.02, 0.03]</p></td>
<td valign="top" align="center"><p>1 / trial</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.29</p></td>
<td valign="top" align="center"><p>[0.26, 0.31]</p></td>
<td valign="top" align="center"><p>0.23</p></td>
<td valign="top" align="center"><p>[0.19, 0.27]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of memory lag</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.01, 0.02]</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.01, 0.02]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of repeat choice on right</p></td>
<td valign="top" align="center"><p>0.06</p></td>
<td valign="top" align="center"><p>[0.05, 0.07]</p></td>
<td valign="top" align="center"><p>0.07</p></td>
<td valign="top" align="center"><p>[0.06, 0.09]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of memory lag × repeat</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.01, 0.02]</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.01, 0.03]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and memory lag</p></td>
<td valign="top" align="center"><p>0.41</p></td>
<td valign="top" align="center"><p>[0.24, 0.56]</p></td>
<td valign="top" align="center"><p>0.19</p></td>
<td valign="top" align="center"><p>[-0.11, 0.46]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and repeat</p></td>
<td valign="top" align="center"><p>–0.27</p></td>
<td valign="top" align="center"><p>[-0.42, -0.11]</p></td>
<td valign="top" align="center"><p>–0.35</p></td>
<td valign="top" align="center"><p>[-0.57, -0.07]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of memory lag and repeat</p></td>
<td valign="top" align="center"><p>–0.70</p></td>
<td valign="top" align="center"><p>[-0.83, -0.53]</p></td>
<td valign="top" align="center"><p>–0.27</p></td>
<td valign="top" align="center"><p>[-0.57, 0.07]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of intercept and memory lag × repeat</p></td>
<td valign="top" align="center"><p>0.27</p></td>
<td valign="top" align="center"><p>[0.04, 0.48]</p></td>
<td valign="top" align="center"><p>0.26</p></td>
<td valign="top" align="center"><p>[-0.17, 0.65]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of memory lag and memory lag x repeat</p></td>
<td valign="top" align="center"><p>0.75</p></td>
<td valign="top" align="center"><p>[0.51, 0.90]</p></td>
<td valign="top" align="center"><p>0.21</p></td>
<td valign="top" align="center"><p>[-0.28, 0.65]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Correlation of repeat and memory lag × repeat</p></td>
<td valign="top" align="center"><p>–0.91</p></td>
<td valign="top" align="center"><p>[-0.98, -0.77]</p></td>
<td valign="top" align="center"><p>–0.74</p></td>
<td valign="top" align="center"><p>[-0.94, -0.36]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl14-fn1"><p>The model can be summarized with the following R syntax formula: <italic>log RT</italic> ~ 1 + <italic>memory lag</italic> * <italic>repeatonright+(1 +memorylag -repeatonright\participant).</italic> This model was fit as a lognormal regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl15" orientation="portrait" position="float">
<label>Table 15.</label>
<caption><p>Exploration-Phase Choices as a Function of ΔΔ-Uncertainty, Memory Lag, and Side of Repeat Option</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl15.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (126973 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (35304 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.09, 0.02]</p></td>
<td valign="top" align="center"><p>–0.02</p></td>
<td valign="top" align="center"><p>[-0.12, 0.08]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty</p></td>
<td valign="top" align="center"><p>1.03</p></td>
<td valign="top" align="center"><p>[0.89, 1.17]</p></td>
<td valign="top" align="center"><p>1.16</p></td>
<td valign="top" align="center"><p>[0.87, 1.43]</p></td>
<td valign="top" align="center"><p>logit / nat</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Memory lag</p></td>
<td valign="top" align="center"><p>–0.01</p></td>
<td valign="top" align="center"><p>[-0.02, 0.00]</p></td>
<td valign="top" align="center"><p>0.01</p></td>
<td valign="top" align="center"><p>[-0.00, 0.02]</p></td>
<td valign="top" align="center"><p>logit / trial</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Repeat choice on right</p></td>
<td valign="top" align="center"><p>0.45</p></td>
<td valign="top" align="center"><p>[0.37, 0.52]</p></td>
<td valign="top" align="center"><p>0.50</p></td>
<td valign="top" align="center"><p>[0.37, 0.63]</p></td>
<td valign="top" align="center"><p>logit difference</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty × memory lag</p></td>
<td valign="top" align="center"><p>–0.08</p></td>
<td valign="top" align="center"><p>[-0.11, -0.04]</p></td>
<td valign="top" align="center"><p>–0.14</p></td>
<td valign="top" align="center"><p>[-0.20, -0.07]</p></td>
<td valign="top" align="center"><p>logit / nat * trial</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Memory lag x repeat</p></td>
<td valign="top" align="center"><p>–0.13</p></td>
<td valign="top" align="center"><p>[-0.15, -0.11]</p></td>
<td valign="top" align="center"><p>–0.08</p></td>
<td valign="top" align="center"><p>[-0.12, -0.04]</p></td>
<td valign="top" align="center"><p>1 / trial</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.40</p></td>
<td valign="top" align="center"><p>[0.36, 0.45]</p></td>
<td valign="top" align="center"><p>0.37</p></td>
<td valign="top" align="center"><p>[0.31, 0.45]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-uncertainty</p></td>
<td valign="top" align="center"><p>0.92</p></td>
<td valign="top" align="center"><p>[0.81, 1.04]</p></td>
<td valign="top" align="center"><p>1.06</p></td>
<td valign="top" align="center"><p>[0.87, 1.32]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of memory lag</p></td>
<td valign="top" align="center"><p>0.03</p></td>
<td valign="top" align="center"><p>[0.02, 0.04]</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.00, 0.04]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of repeat choice on right</p></td>
<td valign="top" align="center"><p>0.49</p></td>
<td valign="top" align="center"><p>[0.44, 0.55]</p></td>
<td valign="top" align="center"><p>0.45</p></td>
<td valign="top" align="center"><p>[0.36, 0.57]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-uncertainty × memory lag</p></td>
<td valign="top" align="center"><p>0.13</p></td>
<td valign="top" align="center"><p>[0.08, 0.17]</p></td>
<td valign="top" align="center"><p>0.07</p></td>
<td valign="top" align="center"><p>[0.00, 0.19]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of memory lag × repeat</p></td>
<td valign="top" align="center"><p>0.10</p></td>
<td valign="top" align="center"><p>[0.08, 0.13]</p></td>
<td valign="top" align="center"><p>0.11</p></td>
<td valign="top" align="center"><p>[0.07, 0.16]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl15-fn1"><p>The model can be summarized with the following R syntax formula: <italic>table on right chosen</italic> ~ 1 + ΔΔ<italic>-uncertainty ‧ memory lag</italic> + <italic>memory lag</italic>: <italic>repeat on right</italic> + (1 + Δ<italic>-uncertainty ‧ memory lag</italic> + <italic>memory lag</italic>: <italic>repeat on right\participant).</italic> This model was fit as a logistic regression. For brevity, the correlations in participant-wise variability are emitted from this table.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl16" orientation="portrait" position="float">
<label>Table 16.</label>
<caption><p>Exploration-Phase Choices as a Function of Δ-Uncertainty, Overall Uncertainty, and Trial Number</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl16.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th rowspan="2" valign="top" align="center"><p>Term</p></th>
<th colspan="2" valign="top" align="center"><p>Pre-registered sample (146766 trials, 194 participants)</p></th>
<th colspan="2" valign="top" align="center"><p>Preliminary sample (41009 trials, 62 participants)</p></th>
<th rowspan="2" valign="top" align="center"><p>Units</p></th>
</tr>
<tr>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
<th valign="top" align="center"><p>Median</p></th>
<th valign="top" align="center"><p>95% PI</p></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="6" valign="top" align="center"><p>Predictors</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Intercept</p></td>
<td valign="top" align="center"><p>–0.04</p></td>
<td valign="top" align="center"><p>[-0.09, 0.02]</p></td>
<td valign="top" align="center"><p>–0.03</p></td>
<td valign="top" align="center"><p>[-0.12, 0.06]</p></td>
<td valign="top" align="center"><p>logit</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty</p></td>
<td valign="top" align="center"><p>1.00</p></td>
<td valign="top" align="center"><p>[0.85, 1.14]</p></td>
<td valign="top" align="center"><p>1.16</p></td>
<td valign="top" align="center"><p>[0.85, 1.46]</p></td>
<td valign="top" align="center"><p>logit / nat</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>–480.86</p></td>
<td valign="top" align="center"><p>[-628.28, -379.29]</p></td>
<td valign="top" align="center"><p>–450.20</p></td>
<td valign="top" align="center"><p>[-568.82, -356.27]</p></td>
<td valign="top" align="center"><p>logit / nat<sup>2</sup></p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Transformed threshold</p></td>
<td valign="top" align="center"><p>2.56</p></td>
<td valign="top" align="center"><p>[2.43, 2.69]</p></td>
<td valign="top" align="center"><p>2.32</p></td>
<td valign="top" align="center"><p>[2.18, 2.48]</p></td>
<td valign="top" align="center"><p>a.u.</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty × trial #</p></td>
<td valign="top" align="center"><p>0.00</p></td>
<td valign="top" align="center"><p>[-0.00, 0.00]</p></td>
<td valign="top" align="center"><p>0.00</p></td>
<td valign="top" align="center"><p>[-0.01, 0.00]</p></td>
<td valign="top" align="center"><p>logit / nat × trial</p></td>
</tr>
<tr>
<td colspan="6" valign="top" align="center"><p>Participant-wise variability</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of intercept</p></td>
<td valign="top" align="center"><p>0.39</p></td>
<td valign="top" align="center"><p>[0.35, 0.43]</p></td>
<td valign="top" align="center"><p>0.36</p></td>
<td valign="top" align="center"><p>[0.30, 0.45]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-uncertainty</p></td>
<td valign="top" align="center"><p>0.94</p></td>
<td valign="top" align="center"><p>[0.83, 1.06]</p></td>
<td valign="top" align="center"><p>1.13</p></td>
<td valign="top" align="center"><p>[0.93, 1.39]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of Δ-uncertainty × overall uncertainty</p></td>
<td valign="top" align="center"><p>10.68</p></td>
<td valign="top" align="center"><p>[0.52, 36.54]</p></td>
<td valign="top" align="center"><p>8.87</p></td>
<td valign="top" align="center"><p>[0.45, 28.51]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD of transformed threshold</p></td>
<td valign="top" align="center"><p>0.43</p></td>
<td valign="top" align="center"><p>[0.36, 0.51]</p></td>
<td valign="top" align="center"><p>0.34</p></td>
<td valign="top" align="center"><p>[0.25, 0.46]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>SD Δ-uncertainty × trial #</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.01, 0.02]</p></td>
<td valign="top" align="center"><p>0.02</p></td>
<td valign="top" align="center"><p>[0.01, 0.02]</p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl16-fn1"><p>We refit the piece ise regression odel described in <xref ref-type="table" rid="app3-tbl6">Table 6</xref>,accounting fora possible interaction between ΔΔ-uncertainty and trial number. We find no significant interaction in the pre-registered sample, nor the preliminary sample. All other terms in the model remained practically the same. The model can be summarized with the following formula: <italic>logit(P(table on right chosen)) = Intercept</italic>+ <italic>b1 ‧ Δ-uncertainty</italic> + b2 ‧ <italic>(overall uncertainty — ®</italic>) ‧ <italic>step(overall uncertainty — ω</italic>) ‧ Δ<italic>-uncertainty</italic> + b4 ‧ <italic>trial # ‧ Δ-uncertainty,</italic> where step is the step function, ω = <italic>—2ln(0.5) ‧ inv_logit(a).</italic> The intercept and parameters b1, b2, b4, and <italic>a</italic> all vary by participant. This model was fit as a logistic regression.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="app3-tbl17" orientation="portrait" position="float">
<label>Table 17.</label>
<caption><p>Model Comparison for Sequential Sampling Models of the Tendency to Repeat Previous Choices</p></caption>
<alternatives>
<graphic xlink:href="31234.gtxam_app1-tabl17.tif" mime-subtype="tiff" mimetype="image"/>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><p>Parameters varying by repeat / switch choice</p></th>
<th valign="top" align="center"><p>Pre-registered sample DIC</p></th>
<th valign="top" align="center"><p>Preliminary sample DIC</p></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><p>None</p></td>
<td valign="top" align="center"><p>218,877.63</p></td>
<td valign="top" align="center"><p>59,365.40</p></td>
</tr>
<tr>
<td valign="top" align="left"><p><italic>K</italic> - the dependence of RT on</p></td>
<td valign="top" align="center"><p>218,666.99</p></td>
<td valign="top" align="center"><p>59,311.25</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Δ-uncertainty</p></td>
<td valign="top" align="center"><p> </p></td>
<td valign="top" align="center"><p> </p></td>
</tr>
<tr>
<td valign="top" align="left"><p>B - Bound height</p></td>
<td valign="top" align="center"><p>217,928.89</p></td>
<td valign="top" align="center"><p>59,096.38</p></td>
</tr>
<tr>
<td valign="top" align="left"><p>Both <italic>K</italic> and B</p></td>
<td valign="top" align="center"><p>217,669.43</p></td>
<td valign="top" align="center"><p>59,046.67</p></td>
</tr>
</tbody>
</table>
</alternatives>
<table-wrap-foot>
<fn id="app3-tbl17-fn1"><p>The model reported in <xref ref-type="table" rid="app3-tbl12">Table 12</xref> captures the tendency to repeat previous choices by allowing both the dependence of RT on Δ-uncertainty and the bound height parameters to vary by whether the choice was a repeat or switch choice (last row in this table). Here, it is compared against the simpler models nested within it. For both samples, the full model is favored over the partial models, as is indicated by lower deviance information criterion (DIC) values. DIC values are derived from the likelihood of the data given estimated parameters, and the effective number of parameters in the model. Absolute values of DIC depend on sample size and the attributes of the noise distribution. Accordingly, DIC values should only be compared between models fit to the same dataset.</p></fn>
</table-wrap-foot>
</table-wrap>
</app>
</app-group>
</back>
<sub-article id="sa0" article-type="editor-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.94231.1.sa2</article-id>
<title-group>
<article-title>eLife Assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Gillan</surname>
<given-names>Claire M</given-names>
</name>
<role specific-use="editor">Reviewing Editor</role>
<aff>
<institution-wrap>
<institution-id institution-id-type="ror">https://ror.org/02tyrky19</institution-id><institution>Trinity College Dublin</institution>
</institution-wrap>
<city>Dublin</city>
<country>Ireland</country>
</aff>
</contrib>
</contrib-group>
<kwd-group kwd-group-type="evidence-strength">
<kwd>Convincing</kwd>
</kwd-group>
<kwd-group kwd-group-type="claim-importance">
<kwd>Valuable</kwd>
</kwd-group>
</front-stub>
<body>
<p>This study presents a <bold>valuable</bold> investigation of how people approach and avoid uncertainty, with a particular focus on the effects of overall uncertainty. They find that individuals approach uncertainty to a point, but when uncertainty is particularly high, they avoid it. The results are interpreted under a cognitive cost-resource rational framework. The methods are <bold>convincing</bold>, using appropriate and current methodologies, but more details on analyses and placing the work more fully in the context of the existing literature would make the contribution more significant.</p>
</body>
</sub-article>
<sub-article id="sa1" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.94231.1.sa1</article-id>
<title-group>
<article-title>Reviewer #1 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>This manuscript reports on the behavior of participants playing a game to measure exploration. Specifically, participants completed a task with blocks of exploratory choices (choosing between two 'tables', and within each table, two 'card decks', each of which had a specific probability of showing cards with one color versus another) and test choices, where participants were asked to choose which of the two decks per table had a higher likelihood of one color. Blocks differed on how long (how many trials) the exploration phase lasted. Participants' choices were fit to increasingly complex models of next-trial exploration. Participants' choices were best fit by an intermediate model where the difference in uncertainty between tables influenced the choice. Next, the authors investigated factors affecting whether participants sought out or avoided uncertainty, their choice reaction times, and the relationship of these measures with performance during the test phase of each block. Participants were uncertainty-seeking (exploratory) under most levels of overall uncertainty but became less uncertainty-seeking at high levels of total uncertainty. Participants with a stronger tendency to approach uncertainty at lower levels of total uncertainty were more accurate in the test phase, while the tendency to avoid uncertainty when total uncertainty was high was also weakly positively related to test accuracy. In terms of reaction times, participants whose reaction times were more related to the level of uncertainty, and who deliberated longer, performed better. The individual tendency to repeat choices was related to avoidance of uncertainty under high total uncertainty and better test performance. Lastly, choices made after a longer lag were less affected by these measures.</p>
<p>The authors note that their paradigm, which does not provide immediate rewarding feedback, is novel. However, the resulting behavior appears similar to other exploratory learning tasks, so it's unclear what this task design adds - besides perhaps showing that exploratory behavior is similar across types of reward environments. Several papers have shown that cognitive constraints modulate exploration (PMIDs: 30667262, 24664860, 35917612, 35260717); although this paper provides novel insights, it does not situate its findings in the context of this prior literature. As a result, what it adds to the literature is difficult to discern.</p>
<p>Other methodological questions include whether the same model provides the best fit for all participants and whether possible individual differences in models used relate to individual differences in exploration and performance; how some analyses were carried out that currently lack sufficient detail in the manuscript; and how the two stages of choice behavior (tables versus card decks) were accounted for in the analyses.</p>
</body>
</sub-article>
<sub-article id="sa2" article-type="referee-report">
<front-stub>
<article-id pub-id-type="doi">10.7554/eLife.94231.1.sa0</article-id>
<title-group>
<article-title>Reviewer #2 (Public Review):</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<anonymous/>
<role specific-use="referee">Reviewer</role>
</contrib>
</contrib-group>
</front-stub>
<body>
<p>Summary:</p>
<p>
This paper focuses on an interesting question that has puzzled psychologists for decades, that is, why do people demonstrate a mix of uncertainty approach and avoidance behavior, given the fact that reducing uncertainty could always gain information and seems beneficial? This paper designed a novel task to demonstrate behavioral signatures of uncertainty approaching and avoidance during the exploration phase within the same task at both a within-subject and between-subject level. On the algorithmic level, this paper compared four different implementations of uncertainty-guided exploration and found that the model sensitive to relative uncertainty provides the best fit for human behavior compared to its counterparts using expected information gain or past exposure. This paper then links people's uncertainty attitude with accuracy and finds that uncertainty avoidance during exploration does not impair task performance, implying that uncertainty avoidance may be the output of a resource-rational decision-making process. To examine this account, this paper uses reaction time as an independent proxy of costly deliberation and shows that people deliberate shorter when engaging in repetitive choice, which presumably saves cognitive resources. Finally, the paper shows that people's tendency to engage in repetitive choice correlates with their tendency to avoid uncertainty, which supports the argument that avoiding uncertainty could be a strategy developed under the constraint of limited cognitive resources.</p>
<p>Strengths:</p>
<p>
One of the highlights of this paper, as mentioned in the previous paragraph, is that the authors can establish the existence of the uncertainty approach and avoidance behavior within the same task whereas previous work usually focuses on one of them. This dissociation allows the authors to examine what situational factor is related to the emergence of the act of avoiding uncertainty, and extract parameters describing participants' attitude towards uncertainty during baseline as well as during situations where uncertainty avoidance is more common. Besides documenting the existence of uncertainty avoidance behavior, this paper also tried to explain this behavior by proposing under the resource rational framework and has carefully quantified different aspects (e.g., accuracy; choice speed) of participants' behavior as well as examined their relationships. Though more experiments are needed to fully understand human uncertainty avoidance behavior, this paper has provided both empirical and theoretical contributions toward a mechanistic understanding of how people balance approaching and avoiding uncertainty.</p>
<p>Weaknesses:</p>
<p>
I have a couple of concerns related to this paper. First, there seems to exist an anti-correlation between total uncertainty and absolute relative uncertainty (Figure 5 panel C, \delta uncertainty is restricted to a small range when total uncertainty is high). It seems to be a natural product of the exploration process since the high total uncertainty phase is usually the period where the participant knows little about either option, leading to a less distinguishable relative uncertainty. However, it remains unknown whether the documented uncertainty avoidance still applies when extrapolating to larger absolute relative uncertainty. It would be great if the experiment allows for a manipulation of uncertainty in the middle of the experiment (e.g., introducing a new deck/informing that one deck has been updated). Relatedly, the current 'threshold' of uncertainty avoidance behavior, if I understand correctly, is found by empirically fitting participants' data. This brings the question: can we predict when people will demonstrate uncertainty avoidance behavior before collecting any data? Or, is it possible that by measuring some metrics related to cognitive cost sensitivity, we could predict the proportion of choices that participants will show uncertainty-avoidant behavior? Finally, regarding the analysis of different behavior patterns in the game, it seems that the authors try to link repetitive behavior, uncertainty attitude, and accuracy together by testing the correlation between the two of them. I wonder whether other multivariate statistical methods e.g., mediation analysis, will be better suited for this purpose.</p>
</body>
</sub-article>
</article>