The HED Task Catalog is under development. IDs are not stable until formal release. Comments are welcome at github.com/hed-standard/hed-task/issues.

Two-Stage Decision Task

HED task ID: hedtsk_two_stage_decision

Family: Conditioning, reinforcement and implicit learning tasks

Also known as: Two-Step Task, Daw Task, MB/MF Task

Sequential two-choice task with probabilistic transitions to second-stage states and drifting rewards; choice patterns dissociate model-based from model-free control.

Description

Participants make two sequential choices per trial. The first-stage choice leads (with fixed transition probabilities — 70% common, 30% rare) to one of two second-stage states, each containing its own pair of options. Second-stage options yield rewards with slowly drifting probabilities. Daw et al. (2011) designed this task to dissociate model-based reinforcement learning (using knowledge of the transition structure to plan) from model-free reinforcement learning (repeating previously rewarded actions regardless of transition type). The diagnostic signature is the interaction between reward and transition type on subsequent first-stage choices: model-based agents show opposite stay/switch patterns after common vs. rare transitions, while model-free agents show the same pattern regardless of transition type.

Inclusion test

An experiment is an instance of this task when its procedure matches, it manipulates at least one of the listed variables, and it records at least one of the listed measures.

Procedure

First stage: choose between two options that lead probabilistically (70/30) to one of two second-stage choice sets. Second stage: choose between two options with drifting reward probabilities. This separates model-based from model-free learning.

Manipulation

Transition structure (common vs. rare); reward probability drift rate; reward magnitude.

Measurement

Stay probability as a function of previous trial outcome × transition type; model-based index; computational model fits (hybrid MB/MF learning rates, mixing weight).

Variations

Named versions that change what the participant experiences or does. The identifier of a variation is hedvar_<task>__<variation>.

Variation

Description

Justification

Standard Daw et al. (2011) Version

hedvar_two_stage_decision__standard_daw_et_al_2011_version

Two first-stage options, two second-stage states, each with two options; 70/30 transition probabilities.

Canonical two-step Markov decision task; measures model-based vs. model-free

Shortened/Simplified Versions

hedvar_two_stage_decision__shortened_simplified_versions

Fewer trials or simplified transition structure for clinical or developmental populations.

Fewer trials or simplified state structure; recognized efficient version

Enhanced Model-Based Version

hedvar_two_stage_decision__enhanced_model_based_version

More complex transition structures (three or more stages) to increase model-based demands.

Design features that enhance model-based learning; different incentive structure

Devaluation Manipulation

hedvar_two_stage_decision__devaluation_manipulation

Changing reward magnitudes mid-task to test sensitivity to outcome value (model-based predicts rapid adjustment).

Reward devalued after training; tests habitual vs. goal-directed control

Two-Stage with Instructed Knowledge

hedvar_two_stage_decision__two_stage_with_instructed_knowledge

Explicitly teaching the transition structure before the task; tests whether model-based behavior increases with explicit knowledge.

Transition structure explicitly taught; tests instructed vs. learned model

Cognitive processes

This task is designed to engage the following processes:

Key references

  • Daw, N. D., Gershman, S. J., Seymour, B., Dayan, P., & Dolan, R. J. (2011). Model-based influences on humans’ choices and striatal prediction errors. Neuron, 69(6), 1204–1215. (DOI, PubMed)

  • Dolan, R. J., & Dayan, P. (2013). Goals and habits in the brain. Neuron, 80(2), 312–325. (DOI, PubMed)

  • Glascher, J., Daw, N., Dayan, P., & O’Doherty, J. P. (2010). States versus rewards: Dissociable neural prediction error signals underlying model-based and model-free reinforcement learning. Neuron, 66(4), 585–595. (DOI, PubMed)

Further references

  • Kool, W., Cushman, F. A., & Gershman, S. J. (2016). When does model-based control pay off? PLoS Computational Biology, 12(8), e1005090. (DOI, PubMed)

  • Gillan, C. M., Kosinski, M., Whelan, R., Phelps, E. A., & Daw, N. D. (2016). Characterizing a psychiatric symptom dimension related to deficits in goal-directed control. eLife, 5, e11305. (DOI, PubMed)

  • da Silva, C. F., & Hare, T. A. (2020). Humans primarily use model-based inference in the two-stage task. Nature Human Behaviour, 4, 1053–1066. (DOI, PubMed)

  • Feher da Silva, C., & Hare, T. A. (2018). A note on the analysis of two-stage task results: How changes in task structure affect what model-free and model-based strategies predict about the effects of reward and transition on the stay probability. PLoS ONE, 13(4), e0195328. (DOI, PubMed)