↑ Katehakis, Michael N.; Veinott, Jr., Arthur F. (1987). "多腕バンディット問題: 分解と計算". Mathematics of Operations Research . 12 (2): 262– 268. doi : 10.1287/moor.12.2.262 . S2CID 656323 .
↑ Bubeck, Sébastien (2012). "Regret Analysis of Stochastic and Nonstochastic Multi-armed Bandit Problems". Foundations and Trends in Machine Learning . 5 : 1–122 . arXiv : 1204.5721 . doi : 10.1561/2200000024 (2026年3月30日非アクティブ)。{{cite journal}}: CS1メンテナンス: DOIは2026年3月現在非アクティブです(リンク)
1 2 3 4 Gittins, JC (1989), Multi-armed bandit allocation indices , Wiley-Interscience Series in Systems and Optimization., Chichester: John Wiley & Sons, Ltd., ISBN978-0-471-92059-5
1 2 3 4ベリー、ドナルド A. ; フリステット、バート (1985)、バンディット問題:実験の逐次割り当て、統計学および応用確率論モノグラフ、ロンドン:チャップマン&ホール、ISBN978-0-412-24810-8
↑ Robbins, H. (1952). "実験の逐次設計のいくつかの側面" . Bulletin of the American Mathematical Society . 58 (5): 527– 535. Bibcode : 1952BAMaS..58..527R . doi : 10.1090/S0002-9904-1952-09620-8 .
↑ JC Gittins (1979). "バンディット過程と動的配分指標". Journal of the Royal Statistical Society. Series B (Methodological) . 41 (2): 148– 177. doi : 10.1111/j.2517-6161.1979.tb01068.x . JSTOR 2985029 . S2CID 17724147 .
↑ Press, William H. (2009)、「Bandit ソリューションは、ランダム化臨床試験および比較有効性研究のための統一された倫理モデルを提供する」、米国科学アカデミー紀要、106 (52): 22387–22392、Bibcode : 2009PNAS..10622387P、doi : 10.1073/pnas.0912378106、PMC 2793317、PMID 20018711。
1 2 Vermorel, Joannes; Mohri, Mehryar (2005), Multi-armed bandit algorithms and empirical evaluation (PDF) , In European Conference on Machine Learning, Springer, pp . 437–448
↑ Whittle, Peter (1988)、「落ち着きのない盗賊:変化する世界における活動配分」、Journal of Applied Probability、25A:287–298、doi:10.2307/3214163、JSTOR 3214163、MR 0974588、S2CID 202109695
↑ Auer, P.; Cesa-Bianchi, N.; Freund, Y.; Schapire, RE (2002). "The Nonstochastic Multiarmed Bandit Problem". SIAM J. Comput. 32 (1): 48– 77. CiteSeerX 10.1.1.130.158 . doi : 10.1137/S0097539701398375 . S2CID 13209702 .
↑ Aurelien Garivier; Emilie Kaufmann (2016). "Optimal Best Arm Identification with Fixed Confidence". arXiv : 1602.04589 [ math.ST ].
↑ Lai, TL; Robbins, H. (1985). "漸近的に効率的な適応的割り当てルール" . Advances in Applied Mathematics . 6 (1): 4– 22. Bibcode : 1985AdApM...6....4L . doi : 10.1016/0196-8858(85)90002-8 .
↑ Katehakis, MN; Robbins, H. (1995). "複数の集団からの逐次選択" . Proceedings of the National Academy of Sciences of the United States of America . 92 (19): 8584– 5. Bibcode : 1995PNAS...92.8584K . doi : 10.1073/pnas.92.19.8584 . PMC 41010 . PMID 11607577 .
↑ Burnetas, AN; Katehakis, MN (1996). "Optimal adaptive policies for sequential allocation problems" . Advances in Applied Mathematics . 17 (2): 122– 142. doi : 10.1006/aama.1996.0007 .
↑ Sutton, RS & Barto, AG 1998 強化学習入門。ケンブリッジ、マサチューセッツ州:MIT Press。
↑ Tokic, Michel (2010), "Adaptive ε-greedy exploration in reinforcement learning based on value differences" (PDF) , KI 2010: Advances in Artificial Intelligence , Lecture Notes in Computer Science, vol. 6359, Springer-Verlag, pp. 203– 210, CiteSeerX 10.1.1.458.464 , doi : 10.1007/978-3-642-16111-7_23 , ISBN978-3-642-16110-0。
↑ Tokic, Michel; Palm, Günther (2011), "Value-Difference Based Exploration: Adaptive Control Between Epsilon-Greedy and Softmax" (PDF) , KI 2011: Advances in Artificial Intelligence , Lecture Notes in Computer Science, vol. 7006, Springer-Verlag, pp. 335– 346, ISBN978-3-642-24455-1。
↑ Gimelfarb, Michel; Sanner, Scott; Lee, Chi-Guhn (2019), "ε-BMC: モデルフリー強化学習におけるイプシロン貪欲探索へのベイズアンサンブルアプローチ" (PDF) , Proceedings of the Thirty-Fifth Conference on Uncertainty in Artificial Intelligence , AUAI Press, p. 162。
1 2 Scott, SL (2010)、「多腕バンディットに対する現代ベイズ的アプローチ」、Applied Stochastic Models in Business and Industry、26 (2): 639–658、doi : 10.1002/asmb.874、S2CID 573750
↑ Olivier Chapelle; Lihong Li (2011)、「Thompsonサンプリングの経験的評価」、Advances in Neural Information Processing Systems、24、Curran Associates: 2249–2257
↑ Langford, John; Zhang, Tong ( 2008)、「コンテキストマルチアームバンディットのためのエポックグリーディアルゴリズム」、Advances in Neural Information Processing Systems、第20巻、Curran Associates, Inc.、pp. 817–824
↑ Auer, P. (2000). 「オンライン学習における信頼区間上限値の利用」.第41回コンピュータサイエンス基礎に関する年次シンポジウム議事録. IEEE Comput. Soc. pp. 270–279 . doi : 10.1109/sfcs.2000.892116 . ISBN978-0-7695-0850-4. S2CID 28713091 .
↑ Hong, Tzung-Pei; Song, Wei-Ping; Chiu, Chu-Tien (2011年11月). "進化的複合属性クラスタリング". 2011 International Conference on Technologies and Applications of Artificial Intelligence . IEEE. pp. 305–308 . doi : 10.1109 /taai.2011.59 . ISBN978-1-4577-2174-8. S2CID 14125100 .
↑ Kwang-Sung Jun; Aniruddha Bhargava; Robert D. Nowak; Rebecca Willett (2017)、「スケーラブルな一般化線形バンディット:オンライン計算とハッシュ化」、Advances in Neural Information Processing Systems、30、Curran Associates: 99–109、arXiv : 1706.00136、Bibcode : 2017arXiv170600136J
↑ Branislav Kveton; Manzil Zaheer; Csaba Szepesvári; Lihong Li; Mohammad Ghavamzadeh; Craig Boutilier (2020), "Randomized exploration in generalized linear bandits", Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS) , arXiv : 1906.08947 , Bibcode : 2019arXiv190608947K
↑ Michal Valko; Nathan Korda; Rémi Munos; Ilias Flaounas; Nello Cristianini (2013), Finite-Time Analysis of Kernelised Contextual Bandits , 29th Conference on Uncertainty in Artificial Intelligence (UAI 2013) and (JFPDA 2013)., arXiv : 1309.6869 , Bibcode : 2013arXiv1309.6869V
↑ Alekh Agarwal; Daniel J. Hsu; Satyen Kale; John Langford; Lihong Li; Robert E. Schapire (2014)、「Taming the monster: A fast and simple algorithm for contextual bandits」、第31回国際機械学習会議議事録:1638–1646、arXiv:1402.0555、Bibcode:2014arXiv1402.0555A
↑ Badanidiyuru, Ashwinkumar; Langford, John; Slivkins, Aleksandrs (2014), "Resourceful contextual bandits" , Balcan, Maria-Florina; Feldman, Vitaly; Szepesvári, Csaba (eds.), Proceedings of The 27th Conference on Learning Theory, COLT 2014, Barcelona, Spain, June 13–15, 2014 , JMLR Workshop and Conference Proceedings, vol. 35, JMLR.org, pp . 1109–1134
↑ Burtini, Giuseppe; Loeppky, Jason; Lawrence, Ramon (2015). "A Survey of Online Experiment Design with the Stochastic Multi-Armed Bandit". arXiv : 1510.00757 [ stat.ML ].
↑ Gajane, Pratik; Urvoy, Tanguy; Clérot, Fabrice (2015), "A Relative Exponential Weighing Algorithm for Adversarial Utility-based Dueling Bandits" (PDF) , Proceedings of the 32nd International Conference on Machine Learning (ICML-15) , archived from the original (PDF) on 2015-09-08 , retrieved 2016-04-29
↑ Zoghi, Masrour; Karnin, Zohar S; Whiteson, Shimon; Rijke, Maarten D (2015), "Copeland Dueling Bandits", Advances in Neural Information Processing Systems, NIPS'15 , arXiv : 1506.00312 , Bibcode : 2015arXiv150600312Z
↑ Wu, Huasen; Liu, Xin (2016), "Double Thompson Sampling for Dueling Bandits", The 30th Annual Conference on Neural Information Processing Systems (NIPS) , arXiv : 1604.07101 , Bibcode : 2016arXiv160407101W
↑チェサ・ビアンキ、ニコロ。ジェンティーレ、クラウディオ。 Zappella、Giovanni (2013)、A Gang of Bandits、Advances in Neural Information Processing Systems 26、NIPS 2013、arXiv : 1306.0811
↑ Gentile, Claudio; Li, Shuai; Zappella, Giovanni (2014), "Online Clustering of Bandits", The 31st International Conference on Machine Learning, Journal of Machine Learning Research (ICML 2014) , arXiv : 1401.8257 , Bibcode : 2014arXiv1401.8257G
↑ Li, Shuai; Alexandros, Karatzoglou; Gentile, Claudio (2016), "Collaborative Filtering Bandits", The 39th International ACM SIGIR Conference on Information Retrieval (SIGIR 2016) , arXiv : 1502.03473 , Bibcode : 2015arXiv150203473L
↑ Gai, Yi; Krishnamachari, Bhaskar; Jain, Rahul (2010年4月)、「認知無線ネットワークにおけるマルチユーザーチャネル割り当ての学習:組み合わせマルチアームバンディット定式化」(PDF)、2010 IEEE Symposium on New Frontiers in Dynamic Spectrum (DySPAN)、IEEE、pp. 1–9、doi : 10.1109/DYSPAN.2010.5457857、ISBN978-1-4244-5189-0
1 2 Chen, Wei; Wang, Yajun; Yuan, Yang (2013), "Combinatorial multi-armed bandit: General framework and applications", Proceedings of the 30th International Conference on Machine Learning (ICML 2013) (PDF) , pp. 151– 159, 2016-11-19 にオリジナル(PDF)からアーカイブ済み、2019-06-14 に取得
1 2 Santiago Ontañón (2017)、「リアルタイム戦略ゲームのための組み合わせ型マルチアームバンディット」、Journal of Artificial Intelligence Research、58 : 665–702、arXiv : 1710.04805、Bibcode : 2017arXiv171004805O、doi : 10.1613/jair.5398、S2CID 8517525
さらに読む
Guha, S.; Munagala, K.; Shi, P. (2010)、「不安定なバンディット問題に対する近似アルゴリズム」、Journal of the ACM、58 : 1–50、arXiv : 0711.3861、doi : 10.1145/1870103.1870106、S2CID 1654066
Dayanik, S.; Powell, W.; Yamazaki, K. (2008)、「利用可能性制約付き割引バンディット問題に対するインデックスポリシー」、Advances in Applied Probability、40 (2): 377–400、doi : 10.1239/aap/1214950209。
パウエル、ウォーレン B. (2007)、「第 10 章」、近似動的計画法: 次元の呪いの解決、ニューヨーク: ジョン・ワイリー・アンド・サンズ、ISBN978-0-470-17155-4。
Allesiardo, Robin (2014)、「コンテキストバンディット問題のためのニューラルネットワーク委員会」、ニューラル情報処理 – 第21回国際会議、ICONIP 2014、マレーシア、2014年11月3日~6日、議事録、Lecture Notes in Computer Science、vol. 8834、Springer、pp. 374–381、arXiv : 1409.8191、doi : 10.1007/978-3-319-12637-1_47、ISBN978-3-319-12636-4S2CID 14155718。
Weber, Richard (1992)、「多腕バンディット問題におけるギティンス指数について」、応用確率年報、2 (4): 1024–1033、doi : 10.1214/aoap/1177005588、JSTOR 2959678。
Katehakis, M. ; C. Derman (1986)、「臨床試験における最適な逐次割り当てルールの計算」、適応統計的手法と関連トピック、数理統計学研究所講義ノート - モノグラフシリーズ、第8巻、 29~ 39ページ、doi : 10.1214/lnms/1215540286、ISBN978-0-940600-09-6JSTOR 4355518。