The relationship between the model and potential neurological function has produced research attempting to use TD to explain many aspects of behavioral research.[15][16] It has also been used to study conditions such as schizophrenia or the consequences of pharmacological manipulations of dopamine on learning.[17]
12Sutton, Richard S. (1 August 1988). "Learning to predict by the methods of temporal differences". Machine Learning. 3 (1): 9–44. doi:10.1007/BF00115009. ISSN1573-0565. S2CID207771194.
12Schultz, W, Dayan, P & Montague, PR. (1997). "A neural substrate of prediction and reward". Science. 275 (5306): 1593–1599. CiteSeerX10.1.1.133.6176. doi:10.1126/science.275.5306.1593. PMID9054347. S2CID220093382.{{cite journal}}: CS1 maint: multiple names: authors list (link)
12Montague, P. R.; Dayan, P.; Sejnowski, T. J. (1996-03-01). "A framework for mesencephalic dopamine systems based on predictive Hebbian learning"(PDF). The Journal of Neuroscience. 16 (5): 1936–1947. doi:10.1523/JNEUROSCI.16-05-01936.1996. ISSN0270-6474. PMC6578666. PMID8774460.
12Montague, P.R.; Dayan, P.; Nowlan, S.J.; Pouget, A.; Sejnowski, T.J. (1993). "Using aperiodic reinforcement for directed self-organization"(PDF). Advances in Neural Information Processing Systems. 5: 969–976.
↑ Schultz, W. (1998). "ドーパミンニューロンの予測報酬シグナル". Journal of Neurophysiology . 80 (1): 1– 27. CiteSeerX 10.1.1.408.5994 . doi : 10.1152/jn.1998.80.1.1 . PMID 9658025. S2CID 52857162 .
↑ Dayan, P. (2001). "動機づけられた強化学習" (PDF) . Advances in Neural Information Processing Systems . 14 . MIT Press: 11– 18. 2012-05-25 のオリジナル(PDF)からアーカイブ済み。2009-03-03に取得。