Two-Stage Decision Task¶
HED task ID: hedtsk_two_stage_decision
Family: Conditioning, reinforcement and implicit learning tasks
Also known as: Two-Step Task, Daw Task, MB/MF Task
Sequential two-choice task with probabilistic transitions to second-stage states and drifting rewards; choice patterns dissociate model-based from model-free control.
Description¶
Participants make two sequential choices per trial. The first-stage choice leads (with fixed transition probabilities — 70% common, 30% rare) to one of two second-stage states, each containing its own pair of options. Second-stage options yield rewards with slowly drifting probabilities. Daw et al. (2011) designed this task to dissociate model-based reinforcement learning (using knowledge of the transition structure to plan) from model-free reinforcement learning (repeating previously rewarded actions regardless of transition type). The diagnostic signature is the interaction between reward and transition type on subsequent first-stage choices: model-based agents show opposite stay/switch patterns after common vs. rare transitions, while model-free agents show the same pattern regardless of transition type.
Inclusion test¶
An experiment is an instance of this task when its procedure matches, it manipulates at least one of the listed variables, and it records at least one of the listed measures.
Procedure |
First stage: choose between two options that lead probabilistically (70/30) to one of two second-stage choice sets. Second stage: choose between two options with drifting reward probabilities. This separates model-based from model-free learning. |
Manipulation |
Transition structure (common vs. rare); reward probability drift rate; reward magnitude. |
Measurement |
Stay probability as a function of previous trial outcome × transition type; model-based index; computational model fits (hybrid MB/MF learning rates, mixing weight). |
Variations¶
Named versions that change what the participant experiences or does. The identifier
of a variation is hedvar_<task>__<variation>.
Variation |
Description |
Justification |
|---|---|---|
Standard Daw et al. (2011) Version
|
Two first-stage options, two second-stage states, each with two options; 70/30 transition probabilities. |
Canonical two-step Markov decision task; measures model-based vs. model-free |
Shortened/Simplified Versions
|
Fewer trials or simplified transition structure for clinical or developmental populations. |
Fewer trials or simplified state structure; recognized efficient version |
Enhanced Model-Based Version
|
More complex transition structures (three or more stages) to increase model-based demands. |
Design features that enhance model-based learning; different incentive structure |
Devaluation Manipulation
|
Changing reward magnitudes mid-task to test sensitivity to outcome value (model-based predicts rapid adjustment). |
Reward devalued after training; tests habitual vs. goal-directed control |
Two-Stage with Instructed Knowledge
|
Explicitly teaching the transition structure before the task; tests whether model-based behavior increases with explicit knowledge. |
Transition structure explicitly taught; tests instructed vs. learned model |
Cognitive processes¶
This task is designed to engage the following processes:
Key references¶
Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P., & Dolan, R. J. (2011). Model-based influences on humans’ choices and striatal prediction errors. Neuron, 69(6), 1204–1215. (DOI, PubMed)
Dolan, R. J., & Dayan, P. (2013). Goals and habits in the brain. Neuron, 80(2), 312–325. (DOI, PubMed)
Glascher, J., Daw, N., Dayan, P., & O’Doherty, J. P. (2010). States versus rewards: Dissociable neural prediction error signals underlying model-based and model-free reinforcement learning. Neuron, 66(4), 585–595. (DOI, PubMed)
Further references¶
Kool, W., Cushman, F. A., & Gershman, S. J. (2016). When does model-based control pay off? PLoS Computational Biology, 12(8), e1005090. (DOI, PubMed)
Gillan, C. M., Kosinski, M., Whelan, R., Phelps, E. A., & Daw, N. D. (2016). Characterizing a psychiatric symptom dimension related to deficits in goal-directed control. eLife, 5, e11305. (DOI, PubMed)
da Silva, C. F., & Hare, T. A. (2020). Humans primarily use model-based inference in the two-stage task. Nature Human Behaviour, 4, 1053–1066. (DOI, PubMed)
Feher da Silva, C., & Hare, T. A. (2018). A note on the analysis of two-stage task results: How changes in task structure affect what model-free and model-based strategies predict about the effects of reward and transition on the stay probability. PLoS ONE, 13(4), e0195328. (DOI, PubMed)
External links¶
Cognitive Atlas: 2-stage decision task