Tuesday, November 3, 2015

Continental Language Diversity

Since language data provides for the demonstration of many visualization techniques, I thought of using another set showing official languages spoken across continents using a new visualization in the "UpSetR" package.  The graph can be used for comparing sets of data numerically.  It provides an easier way to understand data for something like a Venn diagram more quantitatively.  Whereas the appeal of Venn diagrams is in their aesthetic, they do not provide understanding for the numeric value of sets when the sets reach a number that is visually challenging to interpret. 

Displaying the grouping of languages by continent was one way I thought to illustrate the use of the "UpSetR" package.  The data limits each country to one primary language.  The countries are then grouped into continents based on their location.  Below we can see the different continents and respective bar graphs associated with a continent or a set of continents.  The bar graph on the left side (y-axis) measures the number of languages in each continent.  The bar graph at the top (x-axis) measures the number of occurrences for each set of languages seen in the filled circles.  So each filled dot or set of dots represent a set or grouping of languages. 



For instance the red dot indicates 27 different languages in Europe that only occur in Europe, making it seemingly the most diverse in terms of official languages spoken.  The yellow dot set represents 2 languages that are spoken in countries located on 5 different continents (any guesses?*).  This graph does not capture the many different languages spoke in different countries but only the "lingua franca" associated with a continent.  For instance, Africa though linguistically diverse has as official status languages from the colonial-era, thus showing above a lack of diversity.

Alternatively we can see the transposition of this graph below.  Here we see the size of the intersects is small (because we are now considering continent intersects of which there are 1 for each continent).  Guessing the continents for this graph is perhaps a bit easier than guess the different languages for each continent above.



For those having data whose organization is in sets, visualizing sets in this way allows for various dimensions of data to be understood in a way not captured by other visualizations.  This tool for this particular language data set is an interesting view on official language use across continents.  The data used for these graphs is shown below.

The very few lines of code it took to make these graphs is available here.  Much thanks to the developers of this package!

*English and French


rows Africa Antarctic Asia Europe North.America Oceania South.America Countries
1 Albanian 0 0 0 1 0 0 0 1
2 Arabic 1 0 1 0 0 0 0 17
3 Armenian 0 0 0 1 0 0 0 1
4 Azerbaijani 0 0 1 0 0 0 0 1
5 Belarusian 0 0 0 1 0 0 0 1
6 Bosnian 0 0 0 1 0 0 0 1
7 Bulgarian 0 0 0 1 0 0 0 1
8 Catalan 0 0 0 1 0 0 0 1
9 Croatian 0 0 0 1 0 0 0 1
10 Czech 0 0 0 1 0 0 0 1
11 Danish 0 0 0 1 0 0 0 1
12 Dutch 0 0 0 1 1 0 0 3
13 English 1 0 0 1 1 1 1 39
14 Estonian 0 0 0 1 0 0 0 1
15 Filipino 0 0 1 0 0 0 0 1
16 Finnish 0 0 0 1 0 0 0 1
17 French 1 0 0 1 1 1 1 22
18 Georgian 0 0 0 1 0 0 0 1
19 German 0 0 0 1 0 0 0 4
20 Greek 0 0 0 1 0 0 0 2
21 Heard 0 1 0 0 0 0 0 1
22 Hebrew 0 0 1 0 0 0 0 1
23 Hindi 0 0 1 0 0 0 0 1
24 Hungarian 0 0 0 1 0 0 0 1
25 Icelandic 0 0 0 1 0 0 0 1
26 Indonesian 0 0 1 0 0 0 0 1
27 Italian 0 0 0 1 0 0 0 2
28 Japanese 0 0 1 0 0 0 0 1
29 Khmer 1 0 0 0 0 0 0 1
30 Korean 0 0 1 0 0 0 0 2
31 Lao 0 0 1 0 0 0 0 1
32 Latvian 0 0 0 1 0 0 0 1
33 Lithuanian 0 0 0 1 0 0 0 1
34 Malay 0 0 1 0 0 0 0 1
35 Maltese 0 0 0 1 0 0 0 1
36 Mandarin 0 0 1 0 0 0 0 1
37 Norwegian 0 0 0 1 0 0 0 1
38 Persian 0 0 1 0 0 0 0 1
39 Polish 0 0 0 1 0 0 0 1
40 Portuguese 1 0 0 1 0 0 1 7
41 Romanian 0 0 0 1 0 0 0 1
42 Russian 0 0 1 0 0 0 0 2
43 Slovak 0 0 0 1 0 0 0 1
44 Slovenian 0 0 0 1 0 0 0 1
45 Spanish 1 0 0 1 1 0 1 22
46 Swahili 1 0 0 0 0 0 0 1
47 Swedish 0 0 0 1 0 0 0 1
48 Thai 0 0 1 0 0 0 0 1
49 Turkish 0 0 1 0 0 0 0 1
50 Ukranian 0 0 0 1 0 0 0 1
51 Vietnamese 0 0 1 0 0 0 0 1

Wednesday, June 17, 2015

Language Difficulty and Diversity

*For R users not interested in the post but the code, a markdown file is available on github.  Thanks to Zuguang Gu and Bob Rudis for the 'circlize' and 'waffle' packages respectively.

I've been studying Arabic for about 10 months now and had some thoughts that I wanted to post about.  It's challenging, but I didn't know how challenging it was exactly when compared to other languages even though I knew it was on the more challenging end of the spectrum.  Turns out someone has measured (or attempted to) the amount of "class time" a native English speaker would need in order to learn a language.  In general, I've found my language ability improves most when I complement time in the classroom with time practicing with native speakers (which I think would be a more useful measure of time in tandem with "class time").

The data for the chord diagram below was retrieved from a language wiki site that used a study from the Foreign Service Institute.  The number of class hours for each of these languages communicates more a scale of difficulty than the exact number of hours it would take to speak a language (as not all learners are equal).  In general, this seemed to be a pretty comprehensive list of world languages so I thought it could look nice in a chord diagram.



From my own experience, I've been studying for about 10 months.  Not intensively perse, but about 6 hours per week along with conversation practice I have on my own.  Which means if I miss a few weeks I'll have logged about 300 class hours at a year.  Which is kind of disappointing considering I supposedly need 2,200 class hours!  I think these numbers are actually REALLY conservative but it does set some sort of benchmark for difficulty when comparing languages for English speakers.  I'm conversational now and feel comfortable with the language (though by no means fluent) after about 300 hours.  Which as a side note is why I think being immersed would decrease the amount of class time above dramatically.  

So, in learning Arabic and spending a supposed 2,200 hours studying it, how many more people can I actually communicate with?  Well, a lot more.  But in terms of world population I was surprised at the percentage of people who accounted for the top 5 most spoken languages (this includes Arabic).  I honestly had no idea language was quite this diverse (in that the top 5 languages comprise 35% of the world's languages, thought it would be more but that's just me).  Furthermore, if we get into dialects these percentages decrease further.  

Link to data

In real terms though, being able to speak with millions of additional people is fantastic and I would encourage all to pursue such an endeavor.  As for measuring my ability to communicate globally in percentage terms perhaps viewing language learning in the context of world language use is a scale reserved for those with unique skills in language acquisition.


Sunday, April 19, 2015

Boston Elite Field 2015

Last year I posted about how chances of a non-African country winning the Boston Marathon seemed to be good because of the widening interval of winning times (more recently there had been some historically "slower" races and some historically "faster" ones) and this actually happened.   Meb Kflezighi ran a remarkable race and was widely celebrated as he represented the US in a race more recently dominated by African countries.  His time for winning the race was obviously the fastest, but others in the field had faster PRs.  Because of the variation in winning times my conclusion has been that this provides opportunities for certain runners representing non-African countries to contest the race well.


The amount of participants from Africa in the elite field clearly increases the likelihood that the winner represents an African country.  The runners in the elite field mostly fall into or below the confidence interval shown in the graph above with the slight exception of Matt Tegenkamp whose PR for the marathon is 2:12 ish, just above where this statistical measurement would encompass.  It is clear that once again the elite field is dominated by African runners who are putting up some really impressive PRs.



And yet, with the difference in PRs, last year there was a similar dynamic.  Dennis Kimetto comes to the race with a 2:03 PR and Meb Kflezighi wins the Boston Marathon having run a 2:09 PR previously.  Thus we have another great story this year.  Incredible athletes, some of whom have in the past run much faster than others.  And yet, who can tell what will happen race day.

But why try?  Why did Meb think he could beat someone who in marathon terms could go somewhere he could not?  More broadly, why do we love these events?  Why should Matt Tegankamp attempt to rival someone who would be 2 miles ahead of him on each of their best days?  Variance.  Within these elite athletes there is the notion that on any given day, the guy next to you could be at his best or worst.  As spectators, we're drawn to variance...we love possibilities of things not turning out predictably, or that there is variation in what we assume to be true.  Athletes place their hopes in this, that they could run their absolute best and others may not.  Confidence intervals tell the story of variance, that statistically we can't know for certain.  I think this year yet again, we could see this same variance play out.  The athlete that doesn't have the fastest PR runs their best despite the odds.  This is what makes a great race and what we could see again tomorrow.