The data that I have cleaned is song lyrics. Luckily, this process hasn’t been too difficult, thanks to Taylor Swift’s huge fan base. I was able to find all of the lyrics to each album already on one big document, as well as lyrics separated by album as downloadable files. This made it easy for me to get a document for each album that includes every lyric from that specific album. Cleaning this data was satisfying, as I felt like I was very easily able to compile a large amount of data ready for analysis. One issue that did come up was that the downloadable files that were separated by song had credits at the top of each of them in different languages, which sometimes made a random word high in frequency. But, I was able to troubleshoot this by adding some of those words to the stop list on voyant. Similarly, on these files, the words “chorus” “verse” and “bridge” were included in each song denoting each portion of the song. So, as I have been beginning my analysis these words, especially chorus, have been coming up as high in frequency. Again, I just have added them to the stop list on voyant to fix this. I am going to visualize this data by having a different link/section for each album on my website, complete with song by song analysis of frequent words and an overall word cloud and voyant image for the album as a whole. This will provide a lot of visuals for each album. I also hope to upload all albums to into voyant, and then do one visual that compares them all to eachother.
Leave a Reply