{"id":24,"date":"2016-03-14T03:48:15","date_gmt":"2016-03-14T10:48:15","guid":{"rendered":"https:\/\/blogs.ubc.ca\/hamaker\/?p=24"},"modified":"2016-03-24T04:27:24","modified_gmt":"2016-03-24T11:27:24","slug":"look-whos-talking","status":"publish","type":"post","link":"https:\/\/blogs.ubc.ca\/hamaker\/2016\/03\/14\/look-whos-talking\/","title":{"rendered":"Look who&#8217;s talking"},"content":{"rendered":"<p>I feel remiss that, in a blog called &#8220;socially speaking&#8221; with a computational linguistics joke up there in the banner, that I haven&#8217;t actually touched on any of the linguistics of social media. Let&#8217;s fix that:<\/p>\n<p>&nbsp;<\/p>\n<p>A major field in the study of language is <em>corpus linguistics.\u00a0<\/em>Its methodology revolves around the creation and use of large databases, called\u00a0<em>corpora<\/em>, containing thousands if not millions of transcribed utterances and passages of written material. Copora are typically indexed down to the word and heavily encoded with metadata to allow researchers to search for subsets in the data that can be used to test a hypothesis about the use of language.<\/p>\n<p>One of the largest handcrafted corpora is COCA, the <em>Corpus of Contemporary America English<\/em>. COCA was developed around 2008 by researches at Brigham Young University, and continues to grow. The size of COCA is only possible because of the volume of American English text available online &#8212; it was originally built with, of all things, Internet Explorer &#8212; but COCA doesn&#8217;t actually include any natively online content. The corpus was built as a retrospective, balanced, and American corpus. The corpus archives data from back to 1990, and splits the data in each year evenly between the five genres it includes. In 1990 there simply wasn&#8217;t enough internet communication to make up an equal percentage of the data, especially if you limited to American sources (if you could even tell), so it was declared out-of-scope for the project.<\/p>\n<p>Still, the COCA is a behemoth.\u00a0 It has 520 <strong>million<\/strong> words from sources spanning 25 years, divided evenly between transcribed speech, fiction, popular magazines, newspapers, and academic journals. The corpus comprises some <strong>190,000 <\/strong>texts in total. The use of the data is free to the public, you can check out their search interface <a href=\"http:\/\/corpus.byu.edu\/coca\/\" target=\"_blank\">here<\/a>. For most of my linguistic training, it was one of the best &#8212; if not <strong>the<\/strong> best &#8212; English-language corpora.<\/p>\n<p>Compare that to this Facebook corpus a group of researchers generated <a href=\"http:\/\/journals.plos.org\/plosone\/article?id=10.1371\/journal.pone.0073791\" target=\"_blank\">just for their own research<\/a>. It comprises <strong>700 <\/strong>million words in it contributed from 75,000 volunteers (15.4 <strong>million\u00a0<\/strong>facebook status updates). They also got every volunteer to take a personality test. <a href=\"http:\/\/www.slate.com\/blogs\/lexicon_valley\/2014\/03\/12\/language_i_can_t_even_is_just_the_newest_example_of_an_old_greek_rhetorical.html\">I can&#8217;t even<\/a>.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-26 alignnone\" src=\"https:\/\/blogs.ubc.ca\/hamaker\/files\/2016\/03\/age-283x300.png\" alt=\"age\" width=\"283\" height=\"300\" srcset=\"https:\/\/blogs.ubc.ca\/hamaker\/files\/2016\/03\/age-283x300.png 283w, https:\/\/blogs.ubc.ca\/hamaker\/files\/2016\/03\/age.png 523w\" sizes=\"auto, (max-width: 283px) 100vw, 283px\" \/><img loading=\"lazy\" decoding=\"async\" class=\"size-medium wp-image-27 alignnone\" src=\"https:\/\/blogs.ubc.ca\/hamaker\/files\/2016\/03\/gender-282x300.png\" alt=\"gender\" width=\"282\" height=\"300\" srcset=\"https:\/\/blogs.ubc.ca\/hamaker\/files\/2016\/03\/gender-282x300.png 282w, https:\/\/blogs.ubc.ca\/hamaker\/files\/2016\/03\/gender.png 521w\" sizes=\"auto, (max-width: 282px) 100vw, 282px\" \/><\/p>\n<p>They&#8217;ve published some <a href=\"http:\/\/www.wwbp.org\/langTopics.html\" target=\"_blank\">neat visualizations<\/a> for their data on the links between word use, personality, age, and gender. It brings new meaning to &#8220;word cloud.&#8221;The power in these corpora is how easily they can be produced, and how easily their contents can be statistically manipulated and compared. Researchers are not only distributing their data sets, they&#8217;re sharing the code they used to collect them!\u00a0 (one such code release amusingly attempts to coin the term &#8216;<a href=\"http:\/\/link.springer.com\/chapter\/10.1007%2F978-3-642-40722-2_3\" target=\"_blank\">tworpus&#8217; for a twitter corpus<\/a>)<\/p>\n<p>&nbsp;<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>I feel remiss that, in a blog called &#8220;socially speaking&#8221; with a computational linguistics joke up there in the banner, that I haven&#8217;t actually touched on any of the linguistics of social media. Let&#8217;s fix that: &nbsp; A major field in the study of language is corpus linguistics.\u00a0Its methodology revolves around the creation and use [&hellip;]<\/p>\n","protected":false},"author":39408,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-24","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/posts\/24","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/users\/39408"}],"replies":[{"embeddable":true,"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/comments?post=24"}],"version-history":[{"count":5,"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/posts\/24\/revisions"}],"predecessor-version":[{"id":31,"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/posts\/24\/revisions\/31"}],"wp:attachment":[{"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/media?parent=24"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/categories?post=24"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/blogs.ubc.ca\/hamaker\/wp-json\/wp\/v2\/tags?post=24"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}