Computational methods have been developed to assist with information retrieval from scientific literature. Published approaches include methods for searching,[41] determining novelty,[42] and clarifying homonyms[43] among technical reports.
Digital humanities and computational sociology
The automatic analysis of vast textual corpora has created the possibility for scholars to analyze millions of documents in multiple languages with very limited manual intervention. Key enabling technologies have been parsing, machine translation, topic categorization, and machine learning.
The automatic parsing of textual corpora has enabled the extraction of actors and their relational networks on a vast scale, turning textual data into network data. The resulting networks, which can contain thousands of nodes, are then analyzed by using tools from network theory to identify the key actors, the key communities or parties, and general properties such as robustness or structural stability of the overall network, or centrality of certain nodes.[45] This automates the approach introduced by quantitative narrative analysis,[46] whereby subject-verb-object triplets are identified with pairs of actors linked by an action, or pairs formed by actor-object.[44]
Content analysis has been a traditional part of social sciences and media studies for a long time. The automation of content analysis has allowed a "big data" revolution to take place in that field, with studies in social media and newspaper content that include millions of news items. Gender bias, readability, content similarity, reader preferences, and even mood have been analyzed based on text mining methods over millions of documents.[47][48][49][50][51] The analysis of readability, gender bias and topic bias was demonstrated in Flaounas et al.[52] showing how different topics have different gender biases and levels of readability; the possibility to detect mood patterns in a vast population by analyzing Twitter content was demonstrated as well.[53][54]
Software
Text mining computer programs are available from many commercial and open source companies and sources.
Intellectual property law
Situation in the Europe Union
Video by Fix Copyright campaign explaining TDM and its copyright issues in the EU, 2016 [3:51
Under European copyright and database laws, the mining of in-copyright works (such as by web mining) without the permission of the copyright owner is permitted under Articles 3 and 4 of the 2019 Directive on Copyright in the Digital Single Market. A specific TDM exception for scientific research is described in article 3, whereas a more general exception described in article 4 only applies if the copyright holder has not opted out.[55]
↑ Hobbs, Jerry R.; Walker, Donald E.; Amsler, Robert A. (1982). "構造化テキストへの自然言語アクセス".第 9 回計算言語学会議議事録. 第1 巻. pp. 127–32 . doi : 10.3115/991813.991833 . S2CID 6433117 .
↑アントゥネス、ジョアン (2018-11-14)。Exploração de informationações contextuais para enriquecimento semântico em representações de textos (Mestrado em Ciências de Computação e Matemática Computacional thesis) (ポルトガル語)。サンカルロス: サンパウロ大学。土井:10.11606/d.55.2019.tde-03012019-103253。
↑ Moro, Andrea; Raganato, Alessandro; Navigli, Roberto (2014 年 12 月). "エンティティリンキングと単語意味曖昧性解消の統合的アプローチ" . Transactions of the Association for Computational Linguistics . 2 : 231– 244. doi : 10.1162/tacl_a_00179 . ISSN 2307-387X .
↑ Mehl, Matthias R. (2006). "Quantitative Text Analysis". Handbook of multimethod measurement in psychology . p. 141. doi : 10.1037/11383-011 . ISBN978-1-59147-318-3。
↑ Pang, Bo; Lee, Lillian (2008). "Opinion Mining and Sentiment Analysis". Foundations and Trends in Information Retrieval . 2 ( 1–2 ): 1–135 . CiteSeerX 10.1.1.147.2755 . doi : 10.1561/1500000011 . ISSN 1554-0669 . S2CID 207178694 .
↑ Paltoglou, Georgios; Thelwall, Mike (2012-09-01). "Twitter、MySpace、Digg: ソーシャルメディアにおける教師なし感情分析". ACM Transactions on Intelligent Systems and Technology . 3 (4): 66. doi : 10.1145/2337542.2337551 . ISSN 2157-6904 . S2CID 16600444 .
↑ Zanasi, Alessandro (2009). "仮想兵器による現実の戦争:国家安全保障のためのテキストマイニング". Proceedings of the International Workshop on Computational Intelligence in Security for Information Systems CISIS'08 . Advances in Soft Computing. Vol. 53. p. 53. doi : 10.1007/978-3-540-88181-0_7 . ISBN978-3-540-88180-3。
↑ Badal, Varsha D.; Kundrotas, Petras J.; Vakser, Ilya A. (2015-12-09). "Text Mining for Protein Docking" . PLOS Computational Biology . 11 (12) e1004630. Bibcode : 2015PLSCB..11E4630B . doi : 10.1371/journal.pcbi.1004630 . ISSN 1553-7358 . PMC 4674139 . PMID 26650466 .
↑ Cohen, K. Bretonnel; Hunter, Lawrence (2008). "Getting Started in Text Mining" . PLOS Computational Biology . 4 (1) e20. Bibcode : 2008PLSCB...4...20C . doi : 10.1371/journal.pcbi.0040020 . PMC 2217579 . PMID 18225946 .
↑Badal, V. D; Kundrotas, P. J; Vakser, I. A (2015). "Text mining for protein docking". PLOS Computational Biology. 11 (12) e1004630. Bibcode:2015PLSCB..11E4630B. doi:10.1371/journal.pcbi.1004630. PMC4674139. PMID26650466.
↑Papanikolaou, Nikolas; Pavlopoulos, Georgios A.; Theodosiou, Theodosios; Iliopoulos, Ioannis (2015). "Protein–protein interaction predictions using text mining methods". Methods. 74: 47–53. doi:10.1016/j.ymeth.2014.10.026. ISSN1046-2023. PMID25448298.
↑Szklarczyk, Damian; Morris, John H; Cook, Helen; Kuhn, Michael; Wyder, Stefan; Simonovic, Milan; Santos, Alberto; Doncheva, Nadezhda T; Roth, Alexander (2016-10-18). "The STRING database in 2017: quality-controlled protein–protein association networks, made broadly accessible". Nucleic Acids Research. 45 (D1): D362–D368. doi:10.1093/nar/gkw937. ISSN0305-1048. PMC5210637. PMID27924014.
↑Liem, David A.; Murali, Sanjana; Sigdel, Dibakar; Shi, Yu; Wang, Xuan; Shen, Jiaming; Choi, Howard; Caufield, John H.; Wang, Wei; Ping, Peipei; Han, Jiawei (2018-10-01). "Phrase mining of textual data to analyze extracellular matrix protein patterns across cardiovascular disease". American Journal of Physiology. Heart and Circulatory Physiology. 315 (4): H910–H924. doi:10.1152/ajpheart.00175.2018. ISSN1522-1539. PMC6230912. PMID29775406.
↑Van Le, D; Montgomery, J; Kirkby, KC; Scanlan, J (10 August 2018). "Risk Prediction using Natural Language Processing of Electronic Mental Health Records in an Inpatient Forensic Psychiatry Setting". Journal of Biomedical Informatics. 86: 49–58. doi:10.1016/j.jbi.2018.08.007. PMID30118855.
1 2 Coussement, Kristof; Van Den Poel, Dirk (2008). "コールセンターの電子メールを通じて顧客の声を統合し、顧客離脱予測のための意思決定支援システムに組み込む" . Information & Management . 45 (3): 164–74 . CiteSeerX 10.1.1.113.3238 . doi : 10.1016/j.im.2008.01.005 .
↑ Coussement, Kristof; Van Den Poel, Dirk (2008). "言語スタイルの特徴を予測因子として用いた自動メール分類による顧客苦情管理の改善" . Decision Support Systems . 44 (4): 870–82 . doi : 10.1016/j.dss.2007.10.010 .
↑ Ramiro H. Gálvez; Agustín Gravano (2017). "自動株価予測システムにおけるオンライン掲示板マイニングの有用性の評価". Journal of Computational Science . 19 : 1877– 7503. doi : 10.1016/j.jocs.2017.01.001 . hdl : 11336/60065 .
↑Pang, Bo; Lee, Lillian; Vaithyanathan, Shivakumar (2002). "Thumbs up?". Proceedings of the ACL-02 conference on Empirical methods in natural language processing. Vol.10. pp.79–86. doi:10.3115/1118693.1118704. S2CID7105713.
↑Alessandro Valitutti; Carlo Strapparava; Oliviero Stock (2005). "Developing Affective Lexical Resources"(PDF). PsychNology Journal. 2 (1): 61–83. Archived from the original(PDF) on 2018-09-20. Retrieved 2007-08-21.
↑Erik Cambria; Robert Speer; Catherine Havasi; Amir Hussain (2010). "SenticNet: a Publicly Available Semantic Resource for Opinion Mining"(PDF). Proceedings of AAAI CSK. pp.14–18.
↑Calvo, Rafael A; d'Mello, Sidney (2010). "Affect Detection: An Interdisciplinary Review of Models, Methods, and Their Applications". IEEE Transactions on Affective Computing. 1 (1): 18–37. doi:10.1109/T-AFFC.2010.1. S2CID753606.
↑"The University of Manchester". Manchester.ac.uk. Retrieved 2015-02-23.
↑"Tsujii Laboratory". Tsujii.is.s.u-tokyo.ac.jp. Archived from the original on 2012-03-07. Retrieved 2015-02-23.
↑"The University of Tokyo". UTokyo. Retrieved 2015-02-23.
↑Shen, Jiaming; Xiao, Jinfeng; He, Xinwei; Shang, Jingbo; Sinha, Saurabh; Han, Jiawei (2018-06-27). Entity Set Search of Scientific Literature: An Unsupervised Ranking Approach. ACM. pp.565–574. doi:10.1145/3209978.3210055. ISBN978-1-4503-5657-2. S2CID13748283.
↑Walter, Lothar; Radauer, Alfred; Moehrle, Martin G. (2017-02-06). "The beauty of brimstone butterfly: novelty of patents identified by near environment analysis based on text mining". Scientometrics. 111 (1): 103–115. doi:10.1007/s11192-017-2267-4. ISSN0138-9130. S2CID11174676.
↑Roll, Uri; Correia, Ricardo A.; Berger-Tal, Oded (2018-03-10). "Using machine learning to disentangle homonyms in large text corpora". Conservation Biology. 32 (3): 716–724. Bibcode:2018ConBi..32..716R. doi:10.1111/cobi.13044. ISSN0888-8892. PMID29086438. S2CID3783779.
12Automated analysis of the US presidential elections using Big Data and network analysis; S Sudhahar, GA Veltri, N Cristianini; Big Data & Society 2 (1), 1-28, 2015
↑Network analysis of narrative content in large corpora; S Sudhahar, G De Fazio, R Franzosi, N Cristianini; Natural Language Engineering, 1-32, 2013
↑Lansdall-Welfare, Thomas; Sudhahar, Saatviga; Thompson, James; Lewis, Justin; Team, FindMyPast Newspaper; Cristianini, Nello (2017-01-09). "Content analysis of 150 years of British periodicals". Proceedings of the National Academy of Sciences. 114 (4): E457–E465. Bibcode:2017PNAS..114E.457L. doi:10.1073/pnas.1606380114. ISSN0027-8424. PMC5278459. PMID28069962.
↑I. Flaounas, M. Turchi, O. Ali, N. Fyson, T. De Bie, N. Mosdell, J. Lewis, N. Cristianini, The Structure of EU Mediasphere, PLoS ONE, Vol. 5(12), pp. e14243, 2010.
↑Nowcasting Events from the Social Web with Statistical Learning V Lampos, N Cristianini; ACM Transactions on Intelligent Systems and Technology (TIST) 3 (4), 72
↑NOAM: news outlets analysis and monitoring system; I Flaounas, O Ali, M Turchi, T Snowsill, F Nicart, T De Bie, N Cristianini Proc. of the 2011 ACM SIGMOD international conference on Management of data
↑Automatic discovery of patterns in media content, N Cristianini, Combinatorial Pattern Matching, 2-13, 2011
↑I. Flaounas, O. Ali, T. Lansdall-Welfare, T. De Bie, N. Mosdell, J. Lewis, N. Cristianini, RESEARCH METHODS IN THE AGE OF DIGITAL JOURNALISM, Digital Journalism, Routledge, 2012
↑Circadian Mood Variations in Twitter Content; Fabon Dzogang, Stafford Lightman, Nello Cristianini. Brain and Neuroscience Advances, 1, 2398212817744501.