
How Similarly or Not Were We Grouped?
From just the class section data, I was grouped with Carla, Tanya, Emily, and Rebecca.

I noticed that all of Tanya’s choices are also choices that at least one other person selected. I had the most singularities.
Exploring the Rationale For Our Choices
I think that Carla, Tanya, Emily, and Rebecca are more closely aligned in their selection criteria than I was. However, in the end, we all made similar decisions based on the perceptions of sounds and feelings.
Carla: I was impressed that Carla’s selection was able to cover each of the continents. She chose songs based on how evocative they are and the sorted them based on her personal preference. I think if I used this method, I would have a similar experience to Brian where I’m surprised by the choices.
Tanya: Tanya’s overarching criteria were diversity of emotion and cultural inclusion. Similar to me, she also had to cull songs she liked for multi-cultural representation. I think my approach ends up being very similar to Tanya’s, but with the order switched. I started by organizing the data into continents and then started selecting within each of these pools.
Emily: Emily’s approach is similar to Carla’s in that she chose songs with a personal connection, memory, connections, and desires to be happy and present. She went into more in-depth research for some of the songs (e.g., Queen of the Night) which influenced her decisions. Like Emily, I also noticed that researching some of the songs was a rather selective experience because we can only really search for information in languages that we’re familiar with. Thinking back to Boroditsky’s comment that language shapes thought, I don’t know if the meaning or background of some of the songs can be effectively translated or if this was even recorded.
Rebecca: Rebecca’s approach is like a mix of everyone else’s in this group. She made selections based on title, story, culture, and uniqueness to represent emotions, and variety in music. A difficult task is picking between different songs! Rebecca and I both had a process of selecting songs based on what was previously selected. In this case, determining exclusion based on what was already included was a personally and culturally value laden process.
I think what really stands out about Emily, Rebecca, and Tanya’s approach is that they chose to curate based on human emotions and stories. This concept ended up trickling into my selection process towards the end when I realized that a “friendly” soundtrack wouldn’t necessarily capture the human experience.
Ungrouping Ourselves
When comparing the entire class, Carla, Tanya, Emily, and Rebecca remain grouped together! I get put into a new group with Valerie, Matthew, Allison, Patricia, Janice, and Heidi.

Initially, I thought that my original assumptions that there was a stronger link between my section quartet which is why I got put into a different group when there is a larger sample size.
With this larger group, I noticed that Matthew does have a similar curation approach to me. He did it based on populous representation (immediately including India–Jaat Kahan Ho and China–Flowing Streams) and then spread out across the world. However, I’m only observing that the visualization is not able to capture the reasons for our curation.
Grouping Ourselves Based on Strategy
I might group Andrew, Carla, Daniella, and Brian because they essentially chose songs that they like. Emily and Valerie could be put together because their criteria was related to happiness and dancing.
I think my strategy group would be:
- Matthew –grouping based on population size and then spreading across the globe
- Shaun –also used a hierarchy based on continent, music variation within continent/country, type of instrument, Western historical significance/personal preference
- Helen — filtered by continent and then emotional response
All of us used geography in some way to curate our lists. Shaun and Helen’s approaches are almost identical to mine. Although our strategies overlap, our curations are not the same. I think the differences are a result of both our filtering order and the personal selection elements.
We could also create groups based on curator. Taking the 23 people in our class section, there are 17 apparent females. The majority of people in our class have some sort of education work background, we are all financially able to pay for this course, many of us live in Canada, we can speak/write in English.
Thinking Like a Computer?
I’m not sure how Palladio made the grouping decisions, but it was not successful in visualizing the curation rationales.
At most, it captures the end results based on its algorithm. I tried thinking about this combinatorially even though it results in treating the songs as unique and unrelated elements.
Using random chance: given 27 songs, choose 10:

This leads to 8,436,285 possible combinations.
Out of curiousity, I want to determine the probability that at least two people in our class section got the exact same curation. If I treat each curation as an entity, this becomes analogous to the birthday problem. It makes my life easier from the calculation end, but it doesn’t capture the nuance that each entity is actually based on a selection.
Essentially, the probability of everyone choosing unique curations plus the probability of at least two people choosing the same curation is equal to 1. So I can just calculate 1 minus the probability of all unique curations to get what I want.



From my calculations, the probability that at least two people in our class section have the same curation is 0.003%.
With this tiny probability, I’m still unsure of how Palladio was able to make its groups. From how our class has done its curation criteria, the process are non-random and maybe similar criteria have led to the specific selection of songs that are higher than what could be predicted by random chance?
Implications of Data Visualization
From doing my calculations, I noticed that I had to simplify my calculations to make them easier to do (how did Past Linda do combinatorics in undergrad again???)
My calculations highlight that there is missing and assumed data. I am hoping that we all remember that within the 8,436,285 curations, there are some that are almost the same. Palladio visualizes this, but my calculation looks at exact similarity in the curation because of my calculation strategy.
The visualization captures the final results but doesn’t go into depth on any of the inclusion or exclusion criteria. This makes me think that the data visualization is very much another text: We judge it based on what is said rather than its intent.
The path to get to the same results may be very different. As I noticed with Carla’s curation, she ended up with songs from every continent but did it based on her emotional experience and preference while I did it with specific filtering criteria. This makes me realize that politically dissimilar groups may reach a similar conclusion but due to their differing philosophies, the context, and the task. A data visualization may tout this as unity, alignment, and conformity, when the underlying intent shows a different narrative. I can imagine some controversial topic (e.g., euthanasia, abortion) having what we might qualitatively identify as groups identifying with varying views. This doesn’t mean that one side has become “progressive” or “out-dated”, it just means that is the result but doesn’t necessarily mirror the intent.
With a large data set, I’m not sure how we could begin to salvage qualitative data and visualize it. How would you clean up and organize inclusion/exclusion data without compromising the creator’s intent? I think this is another case where data gets lost in translation and we treat it as equivalent exchange, when there will be deletion, substitution, and modification to the information.
Wow Linda, I was very impressed by your analytical approach to both how you chose your curated list and the analysis into the network connections. I too wondered about how Palladio made the grouping decisions and I tried digging through the documentation on the site but still don’t have a clear understanding of how this happens. Your work with calculating and analyzing the probability was impressive and it really opened my eyes to the vastness of the data even in such a small scale experiment such as ours. It then makes me sit in awe thinking about how data scientists handle ‘Big Data’ at all.
I was lucky enough to see and meet Data Visualization Designer Nadieh Bremer at last year’s Creative Technologies conference, CAMP Festival, in Calgary, Alberta. If you are interested in data visualizations and analyzing data I highly recommend her work. She has a background in Astronomy and Data Science, and loves not only data analysis but developing visualizations and insights. Her work is visually stunning and has the ability to help the viewer visually make connections that wouldn’t be possible in a table of numbers. She has done work for Scientific American, Physics Today, Greenpeace, and a litany of others. In her talk, she went through her journey for some of her biggest projects. It was fascinating to hear about the analysis behind the data, experimentation with different visualization strategies and algorithms, and why she chose the final visualization to support some very interesting insights.
I am a huge space nerd so I love her personal project ‘Figures in the Sky’, http://www.datasketch.es/may/code/nadieh/, where you can explore the constellations from 28 cultures around the world throughout time and see the similarities and differences, and learn some of the stories behind those constellations. Another fun project of hers is, ‘Google News Labs – Why do Cats and Dogs…?’, an interactive visualization where you can explore the most popular pet-related questions asked on Google found here: https://whydocatsanddogs.com/. If you are interested, you can see the rest of her portfolio at: https://www.visualcinnamon.com/portfolio/.