↑ Xie, Zhaoming; Hung Yu Ling; Nam Hee Kim; Michiel van de Panne (2020). "ALLSTEPS: カリキュラム主導型ステップストーンスキル学習". arXiv : 2005.04323 [ cs.GR ].
↑ Vergara, Pedro P.; Salazar, Mauricio; Giraldo, Juan S.; Palensky, Peter (2022). "強化学習を用いた不平衡配電システムにおけるPVインバータの最適ディスパッチ" . International Journal of Electrical Power & Energy Systems . 136 107628. Bibcode : 2022IJEPE.13607628V . doi : 10.1016/j.ijepes.2021.107628 . S2CID 244099841 .
↑ Tokic, Michel; Palm, Günther (2011), "Value-Difference Based Exploration: Adaptive Control Between Epsilon-Greedy and Softmax" (PDF) , KI 2011: Advances in Artificial Intelligence , Lecture Notes in Computer Science, vol. 7006, Springer, pp. 335– 346, ISBN978-3-642-24455-1
↑ Singh, Satinder P.; Sutton, Richard S. (1996-03-01). "Reinforcement learning with replacement eligibility traces" . Machine Learning . 22 (1): 123– 158. doi : 10.1007/BF00114726 . ISSN 1573-0565 .
↑ Sutton, Richard S. (1984). Temporal Credit Assignment in Reinforcement Learning (PhD thesis). University of Massachusetts, Amherst, MA. 2017-03-30 のオリジナルからアーカイブ済み。2017-03-29 に取得。
↑ Matzliach, Barouch; Ben-Gal, Irad; Kagan, Evgeny (2022). "Detection of Static and Mobile Targets by an Autonomous Agent with Deep Q-Learning Abilities" . Entropy . 24 ( 8): 1168. Bibcode : 2022Entrp..24.1168M . doi : 10.3390/e24081168 . PMC 9407070. PMID 36010832 .
↑ Williams, Ronald J. (1987). "ニューラルネットワークにおける強化学習のための勾配推定アルゴリズムのクラス". Proceedings of the IEEE First International Conference on Neural Networks . CiteSeerX 10.1.1.129.8871 .
↑ Peters, Jan ; Vijayakumar, Sethu ; Schaal, Stefan (2003). Reinforcement Learning for Humanoid Robotics (PDF) . IEEE-RAS International Conference on Humanoid Robots. 2013-05-12 のオリジナル(PDF)からアーカイブ済み。2006-05-08に取得。
↑ Juliani, Arthur (2016-12-17). "Tensorflow を使用したシンプルな強化学習 パート 8: 非同期アクタークリティックエージェント (A3C)" . Medium . 2018-02-22に取得.
↑ Deisenroth, Marc Peter ; Neumann, Gerhard; Peters, Jan (2013). A Survey on Policy Search for Robotics (PDF) . Foundations and Trends in Robotics. Vol. 2. NOW Publishers. pp. 1– 142. doi : 10.1561/2300000021 . hdl : 10044/1/12051 .
↑ Zou, Lan (2023-01-01), Zou, Lan (編), "第7章 - メタ強化学習" , Meta-Learning , Academic Press, pp. 267–297 , doi : 10.1016/b978-0-323-89931-4.00011-0 , ISBN978-0-323-89931-42023年11月8日取得{{citation}}: CS1メンテナンス: ISBNを使用した作業パラメータ (リンク)
↑ van Hasselt, Hado; Hessel, Matteo; Aslanides, John (2019). 「強化学習でパラメトリックモデルを使用するタイミングは?」(PDF) . Advances in Neural Information Processing Systems . Vol. 32.
↑ 「ゲームメカニクスのテストにおける強化学習の使用について:ACM - Computers in Entertainment」 . cie.acm.org . 2018年11月27日取得。
↑ Li, Xiao; Vasile, Cristian-Ioan; Belta, Calin (2017). "Reinforcement Learning with Temporal Logic Rewards" . 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . pp. 3834– 3839. doi : 10.1109/IROS.2017.8206234 .
↑ Toro Icarte, Rodrigo; Klassen, Toryn Q.; Valenzano, Richard; McIlraith, Sheila A. (2022). "報酬マシン: 強化学習における報酬関数構造の活用" . Journal of Artificial Intelligence Research . 73 : 173– 208. arXiv : 2010.03950 . doi : 10.1613/jair.1.12440 .
↑ Riveret, Régis; Gao, Yang; Governatori, Guido; Rotolo, Antonino; Pitt, Jeremy; Sartor, Giovanni (2019). "強化学習エージェントのための確率的議論フレームワーク" . Autonomous Agents and Multi-Agent Systems . 33 ( 1– 2): 216– 274. doi : 10.1007/s10458-019-09404-2 .
↑ Dey, Somdip; Singh, Amit Kumar; Wang, Xiaohang; McDonald-Maier, Klaus (2020年3月) 「CPU-GPUモバイルMPSoCの電力および熱効率のためのユーザーインタラクション認識型強化学習」 2020 Design, Automation & Test in Europe Conference & Exhibition (DATE) (PDF) pp. 1728–1733 . doi : 10.23919/DATE48585.2020.9116294 . ISBN978-3-9819263-4-7. S2CID 219858480 .
↑ウィリアムズ、リアノン(2020年7月21日)「未来のスマートフォンは、所有者の行動を監視することでバッテリー寿命を延ばすだろう」" . i . 2021-06-17に取得.
↑ Kaplan, F.; Oudeyer, P. (2004). "学習進捗の最大化:発達のための内部報酬システム". Iida, F.; Pfeifer, R.; Steels, L.; Kuniyoshi, Y. (編)『Embodied Artificial Intelligence』Lecture Notes in Computer Science. Vol. 3139. Berlin; Heidelberg: Springer. pp. 259–270 . doi : 10.1007/978-3-540-27833-7_19 . ISBN978-3-540-22484-6. S2CID 9781221 .
↑ Klyubin, A.; Polani, D.; Nehaniv, C. (2008). "選択肢を広げておく:感覚運動系のための情報に基づく運転原理" . PLOS ONE . 3 (12) e4018. Bibcode : 2008PLoSO...3.4018K . doi : 10.1371/journal.pone.0004018 . PMC 2607028 . PMID 19107219 .
↑ Barto, AG (2013). 「内発的動機づけと強化学習」。自然システムと人工システムにおける内発的動機づけ学習(PDF)。ベルリン、ハイデルベルク:Springer。pp. 17–47。
↑ Dabérius, Kevin; Granat, Elvin; Karlsson, Patrik (2020). "Deep Execution - Value and Policy Based Reinforcement Learning for Trading and Beating Market Benchmarks". The Journal of Machine Learning in Finance . 1 . SSRN 3374766 .
↑ Mnih, Volodymyr; Badia, Adrià; Mirza, Mehdi; Graves, Alex; Lillicrap, Timothy; Harley, Tim; Silver, David; Kavukcuoglu, Koray (2016年6月16日). " Asynchronous Methods for Deep Reinforcement Learning" . Proceedings of Machine Learning Research (PMLR) . 48. PMLR: 1928–1937 . 2026年6月15日取得。
↑ J Duan; Y Guan; S Li ( 2021). "Distributional Soft Actor-Critic: Off-policy reinforcement learning for addressing value estimation errors". IEEE Transactions on Neural Networks and Learning Systems . 33 (11): 6584–6598 . arXiv : 2001.02811 . doi : 10.1109/TNNLS.2021.3082568 . PMID 34101599. S2CID 211259373 .
↑ Y Ren; J Duan; S Li (2020). "Improving Generalization of Reinforcement Learning with Minimax Distributional Soft Actor-Critic". 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC) . pp. 1– 6. arXiv : 2002.05502 . doi : 10.1109/ITSC45102.2020.9294300 . ISBN978-1-7281-4149-7. S2CID 211096594 .
↑ Duan, J; Wang, W; Xiao, L (2025). "3つの改良を加えた分布型ソフトアクタークリティック". IEEE Transactions on Pattern Analysis and Machine Intelligence . PP (5): 3935– 3946. arXiv : 2310.05858 . Bibcode : 2025ITPAM..47.3935D . doi : 10.1109/TPAMI.2025.3537087 . PMID 40031258 .
↑ Berenji, HR (1994). "ファジーQ学習:ファジー動的計画法への新しいアプローチ". 1994 IEEE 第3回国際ファジーシステム会議議事録. オーランド、フロリダ州、米国:IEEE. pp. 486–491 . doi : 10.1109/FUZZY.1994.343737 . ISBN0-7803-1896-X. S2CID 56694947 .
↑ Vincze, David (2017). "ファジールール補間と強化学習" (PDF) . 2017 IEEE 15th International Symposium on Applied Machine Intelligence and Informatics (SAMI) . IEEE. pp. 173–178 . doi : 10.1109/SAMI.2017.7880298 . ISBN978-1-5090-5655-2. S2CID 17590120 .
↑ Ng, AY; Russell, SJ (2000). "逆強化学習のためのアルゴリズム" (PDF) . Proceeding ICML '00 Proceedings of the Seventeenth International Conference on Machine Learning . Morgan Kaufmann Publishers. pp. 663–670 . ISBN1-55860-707-2。
↑ Ziebart, Brian D.; Maas, Andrew; Bagnell, J. Andrew; Dey, Anind K. (2008-07-13). "最大エントロピー逆強化学習" . Proceedings of the 23rd National Conference on Artificial Intelligence - Volume 3 . AAAI'08. Chicago, Illinois: AAAI Press: 1433–1438 . ISBN978-1-57735-368-3. S2CID 336219 .
↑ Tamar, Aviv; Glassner, Yonatan; Mannor, Shie (2015-02-21). "サンプリングによるCVaRの最適化" . Proceedings of the AAAI Conference on Artificial Intelligence . 29 (1). arXiv : 1404.3862 . doi : 10.1609/aaai.v29i1.9561 . ISSN 2374-3468 .
↑ Greenberg, Ido; Chow, Yinlam; Ghavamzadeh, Mohammad; Mannor, Shie (2022-12-06). "効率的なリスク回避型強化学習" . Advances in Neural Information Processing Systems . 35 : 32639– 32652. arXiv : 2205.05138 .
↑ Bozinovski, S. (1982). 「二次強化を用いた自己学習システム」。Trappl, Robert (編)『サイバネティクスとシステム研究:第6回ヨーロッパサイバネティクスとシステム研究会議議事録』North-Holland、pp. 397–402。ISBN 978-0-444-86488-8
↑ Bozinovski S. (1995)「神経遺伝学的因子と自己強化学習システムの構造理論」CMPSCIテクニカルレポート95-107、マサチューセッツ大学アマースト校
↑ Bozinovski, S. (2014) 「1981年以降の人工ニューラルネットワークにおける認知と感情の相互作用メカニズムのモデリング」 Procedia Computer Science p. 255–263
Sutton, Richard S. (1988). "Learning to predict by the method of temporal differences" . Machine Learning . 3 (1): 9–44 . Bibcode : 1988MLear...3....9S . doi : 10.1007/BF00115009 .