{"id":1064,"date":"2015-03-25T17:35:48","date_gmt":"2015-03-26T00:35:48","guid":{"rendered":"https:\/\/blogs.ubc.ca\/coetoolbox\/?p=1064"},"modified":"2015-05-05T13:03:49","modified_gmt":"2015-05-05T20:03:49","slug":"lets-talk-about-text-analytics","status":"publish","type":"post","link":"https:\/\/blogs.ubc.ca\/coetoolbox\/2015\/03\/25\/lets-talk-about-text-analytics\/","title":{"rendered":"Let\u2019s talk about Text Analytics"},"content":{"rendered":"<p><img loading=\"lazy\" decoding=\"async\" class=\"alignright\" src=\"http:\/\/www.cambridgesemantics.com\/sites\/default\/files\/solutions\/Unstructured%20Text%20Integration%20%28Custom%29.jpg\" alt=\"\" width=\"280\" height=\"210\" \/>Text analytics has been increasingly attracting\u00a0 the attention of the OR-analytics community as the amount of information stored as textual data increases. Just think about the amount of data stored in emails, news articles and social media, not to mention contact-center notes, surveys, feedback forms, and so forth. Some estimate that text analytics market will grow between 25% and 40% during the next 5 years.[<a title=\"Global Text Analytics Market to see 25.2% CAGR by 2020, Says a Research Report Available at Big Market Research [accessed Mar 25, 2015]\" href=\"http:\/\/www.prnewswire.com\/news-releases\/global-text-analytics-market-to-see-252-cagr-by-2020-says-a-research-report-available-at-big-market-research-296870731.html\" target=\"_blank\">1<\/a>][<a title=\"New report explores the North America text analytics market that is expected to reach $1,995.8 million in 2019 [accessed Mar 25, 2015]\" href=\"http:\/\/www.whatech.com\/market-research-reports\/press-release\/it\/44494-new-report-explores-the-north-america-text-analytics-market-that-is-expected-to-reach-1-995-8-million-in-2019\" target=\"_blank\">2<\/a>] But let\u2019s make sure we are talking about the same thing\u00a0when we talk about text analytics.<\/p>\n<p>So, what is text analytics? <!--more--><\/p>\n<p>If <a title=\"What is analytics? by E. Andrew Boyd, March\/April 2011 issue of Analytics magazine\" href=\"http:\/\/www.analytics-magazine.org\/march-april-2011\/275-profit-center-what-is-analytics.html\" target=\"_blank\">defining analytics is a bit of a challenge<\/a>, defining text analytics is an even greater one. But, basically, text analytics transforms unstructured text into data that can be analyzed with traditional techniques to obtain business insights. [<a title=\"Text mining - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Text_mining\" target=\"_blank\">3<\/a>][<a title=\"Text Mining (Big Data, Unstructured Data) - StatSoft.com\" href=\"http:\/\/www.statsoft.com\/textbook\/text-mining\" target=\"_blank\">4<\/a>]<\/p>\n<p>The distinctive characteristic of text analytics is that it works with <a title=\"Unstructured data - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Unstructured_data\" target=\"_blank\">unstructured text<\/a>. This kind of data does not have a predefined model of organization, resulting in irregularities and ambiguities that make it difficult to extract meaning from. (Although many times unstructured text can be found combined with structured data commonly found in records and surveys, such as dates, demographic information, etc.)<\/p>\n<p>The transformation of unstructured text into useful insights is done by applying methods from fields such as linguistics, statistics and computer science (most notably <a title=\"Natural language processing - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Natural_language_processing\">Natural Language Processing<\/a>). Some of these techniques are:<\/p>\n<ul>\n<li><a title=\"Tokenization (lexical analysis) - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Tokenization_%28lexical_analysis%29\" target=\"_blank\">Tokenizing<\/a>: decomposing sentences into individual words or phrases.<\/li>\n<li>Filtering: removing from the analysis useless words such as \u201cthe\u201d, \u201ca\u201d or others specific to the domain of the analysis.<\/li>\n<li><a title=\"Stemming - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Stemming\" target=\"_blank\">Stemming<\/a>: reducing words to their stem to identify different grammatical forms of words.<\/li>\n<li><a title=\"Tag (metadata) - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Tag_%28metadata%29\" target=\"_blank\">Tagging<\/a>: adding keywords that help describe the data, making searching it easier.<\/li>\n<li>Indexing: using indexes to quickly locate keywords without having to search every row in a database.<\/li>\n<li>Word frequency analysis: creating a matrix of frequencies that enumerates the number of times that each word occurs in each entry or document.<\/li>\n<\/ul>\n<p>Once text is indexed and quantified traditional statistical and data mining methods kick in. Exploratory data analysis, statistical classification, cluster analysis or predictive modelling techniques can help to make sense out of the text. For instance, topic modelling [<a title=\"Topic Model - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Topic_model\" target=\"_blank\">5<\/a>][<a title=\"Probabilistic topic models\" href=\"https:\/\/www.cs.princeton.edu\/~blei\/papers\/Blei2012.pdf\">6<\/a>] tries to discover the text topics by using probabilistic models based on the statistics of the words in the data. <a title=\"Sentiment analysis - Wikipedia.org\" href=\"http:\/\/en.wikipedia.org\/wiki\/Sentiment_analysis\" target=\"_blank\">Sentiment analysis<\/a> aims to determine the attitude of text comments towards a topic, for example to determine the public opinion on political topics based on Twitter entries. Which is the appropriate technique will depend basically on the objective of the analysis.<\/p>\n<p>In a recent <a title=\"Predicting  bill shocks at TELUS Mobility - COE website\" href=\"http:\/\/www.sauder.ubc.ca\/Faculty\/Research_Centres\/Centre_for_Operations_Excellence\/Industry_Projects\/Completed_Projects\/~\/media\/Files\/COE\/Projects\/2014\/Predicting%20Bill%20Shocks%20at%20Telus%20Mobility.ashx\" target=\"_blank\">COE project<\/a> we used agent notes from a call center to identify customers complaining about their phone bills and then analyzed the characteristics of these type of customers. Text analytics allowed us to interpret unstructured text from agent notes and to a massive amount of data to better understand customers&#8217; behavior.<\/p>\n<p>Interpreting unstructured text is a challenging problem that Text Analytics tackles. It&#8217;s not easy work but its potential rewards are worthwhile.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Text analytics has been increasingly attracting\u00a0 the attention of the OR-analytics community as the amount of information stored as textual data increases. Just think about the amount of data stored in emails, news articles and social media, not to mention contact-center notes, surveys, feedback forms, and so forth. Some estimate that text analytics market will [&hellip;]<\/p>\n","protected":false},"author":22982,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1054137],"tags":[5015],"class_list":["post-1064","post","type-post","status-publish","format-standard","hentry","category-text-analytics","tag-analytics"],"_links":{"self":[{"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/posts\/1064","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/users\/22982"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/comments?post=1064"}],"version-history":[{"count":15,"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/posts\/1064\/revisions"}],"predecessor-version":[{"id":1112,"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/posts\/1064\/revisions\/1112"}],"wp:attachment":[{"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/media?parent=1064"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/categories?post=1064"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.ubc.ca\/coetoolbox\/wp-json\/wp\/v2\/tags?post=1064"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}