Microsoft reported reaching 94.9% recognition accuracy on the Switchboard corpus, incorporating a vocabulary of 165,000 words. The approach used "dialog session-based long-short-term memory".[62]
2018:OpenAI used LSTM trained by policy gradients to beat humans in the complex video game of Dota 2,[15] and to control a human-like robot hand that manipulates physical objects with unprecedented dexterity.[14][63]
Aspects of LSTM were anticipated by "focused back-propagation",[64] cited by the LSTM paper.[1]
Sepp Hochreiter's 1991 German diploma thesis analyzed the vanishing gradient problem and developed principles of the method.[2] His supervisor, Jürgen Schmidhuber, considered the thesis highly significant.[65]
The most commonly used reference point for LSTM was published in 1997 in the journal Neural Computation.[1] By introducing Constant Error Carousel (CEC) units, LSTM deals with the vanishing gradient problem. The initial version of LSTM block included cells, input and output gates.[67]
Felix Gers, Jürgen Schmidhuber, and Fred Cummins introduced the forget gate (also called "keep gate") into the LSTM architecture in 1999,[68] enabling the LSTM to reset its own state.[67] This is the most commonly used version of LSTM nowadays. They added peephole connections in 2000.[21][22] Additionally, the output activation function was omitted.[67]
2009: An LSTM trained by CTC won the ICDAR connected handwriting recognition competition. Three such models were submitted by a team led by Alex Graves.[80] One was the most accurate model in the competition and another was the fastest.[81] This was the first time an RNN won international competitions.[63]
2013: Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton used LSTM networks as a major component of a network that achieved a record 17.7% phoneme error rate on the classic TIMIT natural speech dataset.[28]
2020: Kaplan and McCandlish et al. explored the scaling law of LSTM architecture in language modeling and compared it with Transformers, showing that Transformers scale better than LSTM on text data. [83]
123Hochreiter, Sepp (1991). Untersuchungen zu dynamischen neuronalen Netzen(PDF) (diploma thesis). Technical University of Munich, Institute of Computer Science.
12Hochreiter, Sepp; Schmidhuber, Jürgen (1996-12-03). "LSTM can solve hard long time lag problems". Proceedings of the 9th International Conference on Neural Information Processing Systems. NIPS'96. Cambridge, MA, USA: MIT Press: 473–479.
1 2 3 Felix A. Gers; Jürgen Schmidhuber; Fred Cummins (2000). "Learning to Forget: Continual Prediction with LSTM" . Neural Computation . 12 (10): 2451–2471 . CiteSeerX 10.1.1.55.5709 . doi : 10.1162/089976600300015015 . PMID 11032042. S2CID 11598600. 2019年4月7日にオリジナルからアーカイブ済み。 2017年4月15日に取得。
1 2 3 Graves, Alex; Fernández, Santiago; Gomez, Faustino; Schmidhuber, Jürgen (2006). "Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks". In Proceedings of the International Conference on Machine Learning, ICML 2006 : 369– 376. CiteSeerX 10.1.1.75.6306 .
1 2 Rodriguez, Jesus (2018年7月2日). 「OpenAI Fiveの背後にある科学:AI史上最大のブレークスルーの1つを生み出した」 . Towards Data Science . 2019年12月26日のオリジナルからアーカイブ済み。 2019年1月15日取得。
1 2スタンフォード、ステイシー(2019年1月25日)。「DeepMindのAI、AlphaStarがAGIに向けて大きな進歩を示す」。Medium ML Memoirs 。2019年1月15日取得。
↑ Schmidhuber, Jürgen (2021). "2010年代:ディープラーニングの10年/2020年代の展望" . AI Blog . IDSIA、スイス. 2022年4月30日取得.
12Maity, Abhishek; Tukarul, Viraj (23 January 2026). "Forecasting Energy Consumption using Recurrent Neural Networks: A Comparative Analysis". arXiv:2601.17110 [cs.CY].
↑Calin, Ovidiu (14 February 2020). Deep Learning Architectures. Cham, Switzerland: Springer Nature. p.555. ISBN978-3-030-36720-6.
↑Lakretz, Yair; Kruszewski, German; Desbordes, Theo; Hupkes, Dieuwke; Dehaene, Stanislas; Baroni, Marco (2019), "The emergence of number and syntax units in", The emergence of number and syntax units(PDF), Association for Computational Linguistics, pp.11–20, doi:10.18653/v1/N19-1002, hdl:11245.1/16cb6800-e10d-4166-8e0b-fed61ca6ebb4, S2CID81978369
123456Gers, F. A.; Schmidhuber, J. (2001). "LSTM Recurrent Networks Learn Simple Context Free and Context Sensitive Languages"(PDF). IEEE Transactions on Neural Networks. 12 (6): 1333–1340. Bibcode:2001ITNN...12.1333G. doi:10.1109/72.963769. PMID18249962. S2CID10192330. Archived from the original(PDF) on 2017-07-06.
1234Gers, F.; Schraudolph, N.; Schmidhuber, J. (2002). "Learning precise timing with LSTM recurrent networks"(PDF). Journal of Machine Learning Research. 3: 115–143. Archived from the original(PDF) on 2017-07-28. Retrieved 2017-04-15.
↑Xingjian Shi; Zhourong Chen; Hao Wang; Dit-Yan Yeung; Wai-kin Wong; Wang-chun Woo (2015). "Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting". Proceedings of the 28th International Conference on Neural Information Processing Systems: 802–810. arXiv:1506.04214. Bibcode:2015arXiv150604214S.
↑ Hochreiter, S.; Bengio, Y.; Frasconi, P.; Schmidhuber, J. (2001). "Gradient Flow in Recurrent Nets: the Difficulty of Learning Long-Term Dependencies (PDF Download Available)" . In Kremer and, SC; Kolen, JF (eds.). A Field Guide to Dynamical Recurrent Neural Networks . IEEE Press.
↑ Fernández, Santiago; Graves, Alex; Schmidhuber, Jürgen (2007). "階層型リカレントニューラルネットワークを用いた構造化ドメインにおけるシーケンスラベリング". Proc. 20th Int. Joint Conf. On Artificial Intelligence, Ijcai 2007 : 774– 779. CiteSeerX 10.1.1.79.1887 .
↑ A. Graves、J. Schmidhuber。「多次元リカレントニューラルネットワークによるオフライン手書き認識」。Advances in Neural Information Processing Systems 22、NIPS'22、pp 545–552、バンクーバー、MIT Press、2009年。
↑ Baccouche, M.; Mamalet, F.; Wolf, C.; Garcia, C.; Baskurt, A. (2011). "Sequential Deep Learning for Human Action Recognition". In Salah, AA; Lepri, B. (eds.). 2nd International Workshop on Human Behavior Understanding (HBU) . Lecture Notes in Computer Science. Vol. 7065. Amsterdam, Netherlands: Springer. pp. 29–39 . doi : 10.1007/978-3-642-25446-8_4 . ISBN978-3-642-25445-1。
↑ Tax, N.; Verenich, I.; La Rosa, M.; Dumas, M. (2017). "LSTMニューラルネットワークを用いた予測型ビジネスプロセスモニタリング". Advanced Information Systems Engineering . Lecture Notes in Computer Science. Vol. 10253. pp. 477–492 . arXiv : 1612.02130 . doi : 10.1007/978-3-319-59536-8_30 . ISBN978-3-319-59535-1. S2CID 2192354 .
↑Martin, Abbey; Hill, Andrew J.; Seiler, Konstantin M.; Balamurali, Mehala (2024-05-27). "Automatic excavator action recognition and localisation for untrimmed video using hybrid LSTM-Transformer networks". International Journal of Mining, Reclamation and Environment. 38 (5): 353–372. Bibcode:2024IJMRE..38..353M. doi:10.1080/17480930.2023.2290364. ISSN1748-0930.
↑Beaufays, Françoise (August 11, 2015). "The neural networks behind Google Voice transcription". Research Blog. Retrieved 2017-06-27.
↑Sak, Haşim; Senior, Andrew; Rao, Kanishka; Beaufays, Françoise; Schalkwyk, Johan (September 24, 2015). "Google voice search: faster and more accurate". Research Blog. Retrieved 2017-06-27.
↑"Neon prescription... or rather, New transcription for Google Voice". Official Google Blog. 23 July 2015. Retrieved 2020-04-25.
↑Khaitan, Pranav (May 18, 2016). "Chat Smarter with Allo". Research Blog. Retrieved 2017-06-27.
↑Metz, Cade (September 27, 2016). "An Infusion of AI Makes Google Translate More Powerful Than Ever | WIRED". Wired. Retrieved 2017-06-27.
↑"A Neural Network for Machine Translation, at Production Scale". Google AI Blog. 27 September 2016. Retrieved 2020-04-25.
↑エフラティ、アミール(2016年6月13日)。「アップルのマシンも学習できる」。The Information 。 2017年6月27日取得。
↑ Ranger, Steve (2016年6月14日). 「iPhone、AI、ビッグデータ:Appleがあなたのプライバシーを保護するために計画していること」 . ZDNet . 2017年6月27日閲覧。
1 2 3 Klaus Greff; Rupesh Kumar Srivastava; Jan Koutník; Bas R. Steunebrink; Jürgen Schmidhuber (2015). "LSTM: A Search Space Odyssey". IEEE Transactions on Neural Networks and Learning Systems . 28 (10): 2222– 2232. arXiv : 1503.04069 . Bibcode : 2015arXiv150304069G . doi : 10.1109/TNNLS.2016.2582924 . PMID 27411231. S2CID 3356463 .
1 2 3 Gers, Felix; Schmidhuber, Jürgen; Cummins, Fred (1999). "Learning to forget: Continual prediction with LSTM".第9回人工ニューラルネットワーク国際会議: ICANN '99 . Vol. 1999. pp. 850–855 . doi : 10.1049/cp:19991218 . ISBN0-85296-721-7。
↑ Srivastava, Rupesh K; Greff, Klaus; Schmidhuber, Juergen (2015 年 12 月). カナダ、モントリオールにて執筆。Training Very Deep Networks . NIPS'15: Proceedings of the 29th International Conference on Neural Information Processing Systems. Vol. 2. Cambridge, MA, United States: MIT Press. pp. 2377–2385 .
↑ Schmidhuber, Jürgen (2021). "最も引用されているニューラルネットワークはすべて、私の研究室で行われた研究に基づいています" . AI Blog . IDSIA、スイス. 2022年4月30日取得.
↑ He, Kaiming; Zhang, Xiangyu; Ren, Shaoqing; Sun, Jian (2016). "Deep Residual Learning for Image Recognition". 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE. pp. 770–778 . arXiv : 1512.03385 . doi : 10.1109/CVPR.2016.90 . ISBN978-1-4673-8851-1。