Thursday, June 08, 2023
"Detecting Political Biases of Named Entities and Hashtags on Twitter"
Title: Detecting Political Biases of Named Entities and Hashtags on Twitter
Authors: Zhiping Xiao, Jeffrey Zhu, Yining Wang, Pei Zhou, Wen Hong Lam, Mason A. Porter, and Yizhou Sun
Abstract: Ideological divisions in the United States have become increasingly prominent in daily communication. Accordingly, there has been much research on political polarization, including many recent efforts that take a computational perspective. By detecting political biases in a text document, one can attempt to discern and describe its polarity. Intuitively, the named entities (i.e., the nouns and the phrases that act as nouns) and hashtags in text often carry information about political views. For example, people who use the term “pro-choice” are likely to be liberal and people who use the term “pro-life” are likely to be conservative. In this paper, we seek to reveal political polarities in social-media text data and to quantify these polarities by explicitly assigning a polarity score to entities and hashtags. Although this idea is straightforward, it is difficult to perform such inference in a trustworthy quantitative way. Key challenges include the small number of known labels, the continuous spectrum of political views, and the preservation of both a polarity score and a polarity-neutral semantic meaning in an embedding vector of words. To attempt to overcome these challenges, we propose the Polarity-aware Embedding Multi-task learning (PEM) model. This model consists of (1) a self-supervised context-preservation task, (2) an attention-based tweet-level polarity-inference task, and (3) an adversarial learning task that promotes independence between an embedding’s polarity component and its semantic component. Our experimental results demonstrate that our PEM model can successfully learn polarity-aware embeddings that perform well at tweet-level and account-level classification tasks. We examine a variety of applications—including a study of spatial and temporal distributions of polarities and a comparison between tweets from Twitter and posts from Parler—and we thereby demonstrate the effectiveness of our PEM model. We also discuss important limitations of our work and encourage caution when applying the PEM model to real-world scenarios.
Wednesday, October 26, 2022
"Analysis of Spatial and Spatiotemporal Anomalies Using Persistent Homology: Case Studies with COVID-19 Data"
Title: Analysis of Spatial and Spatiotemporal Anomalies Using Persistent Homology: Case Studies with COVID-19 Data
Authors: Abigail Hickok, Deanna Needell, and Mason A. Porter
Abstract: We develop a method for analyzing spatial and spatiotemporal anomalies in geospatial data using topological data analysis (TDA). To do this, we use persistent homology (PH), which allows one to algorithmically detect geometric voids in a data set and quantify the persistence of such voids. We construct an efficient filtered simplicial complex (FSC) such that the voids in our FSC are in one- to-one correspondence with the anomalies. Our approach goes beyond simply identifying anomalies; it also encodes information about the relationships between anomalies. We use vineyards, which one can interpret as time-varying persistence diagrams (which are an approach for visualizing PH), to track how the locations of the anomalies change with time. We conduct two case studies using spatially heterogeneous COVID-19 data. First, we examine vaccination rates in New York City by zip code at a single point in time. Second, we study a year-long data set of COVID-19 case rates in neighborhoods of the city of Los Angeles.
Thursday, April 30, 2020
"Community Matters"
Title: Community Matters
Authors: Anna C. F. Lewis, Nick S. Jones, Mason A. Porter, and Charlotte M. Deane
Wednesday, November 27, 2019
Facial Recognition Software for Sheep
Okay... pic.twitter.com/Qbsm6yptuH
— Thomas Lumley (@tslumley) November 27, 2019
(Tip of the cap to Richard Parker.)
Wednesday, July 31, 2019
UCLA's New Undergraduate Major in "Data Theory"
At @UCLA (joint between the Math & Stats departments), I helped design a brand new, very exciting Data Theory undergraduate major.
— Mason Porter (@masonporter) August 1, 2019
Courses: https://t.co/z27ZJybKQj
To look up course identities: https://t.co/cHcYrQzVcH
Catalog listing: https://t.co/MMOfGX4Cf4 pic.twitter.com/ZZA4E1G8tE
Tuesday, April 09, 2019
Monday, September 24, 2018
Will the Authors of this Paper Look Back in Anger?
"Can high-density human collective motion be forecasted by spatiotemporal fluctuations?"
— Chris Danforth (@ChrisDanforth) September 24, 2018
Eigenmodes of crowd dynamics at an Oasis concert. #moshmathhttps://t.co/srSG8Biyui pic.twitter.com/XxnlBQuDFJ
Monday, July 30, 2018
A Cartoon Depiction of Deep Learning Versus Traditional Machine Learning
great meme pic.twitter.com/d5AW6WliTy
— Lex Flagel (@flagelbagel) July 29, 2018
(Tip of the cap to Michael Stumpf.)
Wednesday, May 09, 2018
The Pulse of Manhattan
The population of #Manhattan, hour-by-hour. #NYC #datavizhttps://t.co/KATB4uq9va pic.twitter.com/ov3xTQlVm0
— Randy Olson (@randal_olson) May 9, 2018
Tuesday, May 08, 2018
Important Paper on Seussian Configurations of Random-Graph Models
Title: Configuring Random Graph Models with Fixed Degree Sequences
Authors: Bailey K. Fosdick, Daniel B. Larremore, Joel Nishimura, and Johan Ugander
In 2014, Aaron Clauset, David Kempe, and I (with help from Dan Larremore) organized a Mathematics Research Community in Network Science.
In addition to creating a network of network scientists from diverse backgrounds, some work was started there, and today the published version of what is in my opinion an extremely important paper has come out in final form in SIAM Review's 'Research Spotlights' section. I am, of course, talking about the aforementioned paper.
I'm very happy indeed for such excellent work to arise from this.
Congratulations to authors Bailey Fosdick, Daniel Larremore, Joel Nishimura, and Johan Ugander for creating this awesome paper!
I am posting this with an absolutely lovely Seussian picture from an arXiv version of the paper. This picture doesn't appear to have made the cut for the published piece. In addition to its wit and whimsy, a really great thing about the picture and its accompanying verse is that it also encodes the main message of the paper.
Note: I am on the editorial board of the Research Spotlights section of SIAM Review, but I had nothing whatsoever to do with the handling of this paper.
Tuesday, March 27, 2018
"Interval Signatures of Celebrities and Researchers on Twitter"
It's always nice to be analyzed (for my tweeting behavior) alongside luminaries like Katy Perry, Shaquille O'Neal, and Steve Martin. Lots of my peeps from network and data science are also put under the data-analytic microscope (or perhaps I should write 'mesoscope') on this page.
Friday, March 02, 2018
Proposal: Use Computational Topology to Study Aversion to Pictures of Holes
For science!
Sunday, October 08, 2017
Visualizations of Music
Data Visualisation of Music https://t.co/45mrRmaTW1 & https://t.co/xFiy9Gs5ay #music #dataviz pic.twitter.com/fLc071JyOG
— Benjamin Hennig (@geoviews) October 7, 2017
(Tip of the cap to Lucas Lacasa.
Wednesday, August 09, 2017
"A Roadmap for the Computation of Persistent Homology"
Title: A Roadmap for the Computation of Persistent Homology
Authors: Nina Otter, Mason A. Porter, Ulrike Tillmann, Peter Grindrod, and Heather A. Harrington
Abstract: Persistent homology (PH) is a method used in topological data analysis (TDA) to study qualitative features of data that persist across multiple scales. It is robust to perturbations of input data, independent of dimensions and coordinates, and provides a compact representation of the qualitative features of the input. The computation of PH is an open area with numerous important and fascinating challenges. The field of PH computation is evolving rapidly, and new algorithms and software implementations are being updated and released at a rapid pace. The purposes of our article are to (1) introduce theory and computational methods for PH to a broad range of computational scientists and (2) provide benchmarks of state-of-the-art implementations for the computation of PH. We give a friendly introduction to PH, navigate the pipeline for the computation of PH with an eye towards applications, and use a range of synthetic and real-world data sets to evaluate currently available open-source implementations for the computation of PH. Based on our benchmarking, we indicate which algorithms and implementations are best suited to different types of data sets. In an accompanying tutorial, we provide guidelines for the computation of PH. We make publicly available all scripts that we wrote for the tutorial, and we make available the processed version of the data sets used in the benchmarking.
Friday, July 21, 2017
Math Profs are the Least Boring (And Other Observations from Data Analysis)
You can see from this data analysis that math professors appear to be the least boring among all professors (yay!) and there are some expected (and perhaps more surprising) gender gaps (boo!) in the reviews.
(Tip of the cap to Nicholas Christakis.)
Tuesday, June 27, 2017
What Happens in London Stays in London
Friday, June 16, 2017
How Do People from Different Cultures Draw Circles?
I was once — well, at least once — told in elementary school that I was drawing circles the wrong way (because I was using the wrong sense around the clock). I think I responded with the definition of a circle and that the definition doesn't depend on the sense in which one draws it, and I think my teacher did not appreciate that.
On a similar note, in high school, I once lost 50% of the point total on an answer for misspelling ellipse (by using one 'l' instead of two), which was the correct answer. I called bullshit (on the grounds that it was my mathematics knowledge that was being tested), but unfortunately I lost.
More closely related to the article, one thing I noticed in the UK is that the most common way to write an 'x' there is with two arcs, so that they won't always cross if one writes quickly. In contrast, I write two attempts at lines that explicitly cross. (I haven't checked if this is US versus UK convention.)
And, indeed, most Americans drew their circles counterclockwise in this data set, and I draw mine clockwise.
(Tip of the cap to Improbable Research for their Facebook post.)
Update: Here is a lovely quote from the article: In a 1977 paper Theodore Blau, then-president of the American Psychological Association and creator of the torque test, argued that drawing clockwise circles was a sign of learning and behavioral aberrance.
Friday, May 26, 2017
Data Analysis of Gender in Film Dialog
The Largest Ever Analysis of Film Dialogue by Gender: 2,000 scripts, 25,000 actors, 4 million lines https://t.co/QV7NzluIpL pic.twitter.com/ampVQskeAu
— Sinan Aral (@sinanaral) May 26, 2017
Friday, May 12, 2017
Lyrical Repetition in Pop Music
I of course decided to look at how Depeche Mode stacks up, and I zoomed up on them as an individual artist.
From the song and artist library they used, Depeche Mode is listed as the band from the 1990s with the least repetitive lyrics, though it does only use a subset of their songs and it listed them in the 1990s instead of the 1980s. (Naturally, the employed songs span multiple decades.)
The most lyrically repetitive Depeche Mode song is very obvious.
(Tip of the cap to Taha Yasseri and I Fucking Love Data.)
Sunday, February 19, 2017
My Slides on Data Ethics for Mathematicians (and Others)
My slides on "Data Ethics for Mathematicians" (n the era of Big Data) #appliedmathematics #bigdata https://t.co/JBqOcMyf6f via @SlideShare
— Mason Porter (@masonporter) February 19, 2017

