Showing posts with label data science. Show all posts
Showing posts with label data science. Show all posts

Thursday, June 08, 2023

"Detecting Political Biases of Named Entities and Hashtags on Twitter"

One of my papers came out in final form earlier today. Here are some details. (This is in collaboration with computer scientists, and stylistically it is rather different from much of my work. However, you'll still notice my hand in it. :P)

Title: Detecting Political Biases of Named Entities and Hashtags on Twitter

Authors: Zhiping Xiao, Jeffrey Zhu, Yining Wang, Pei Zhou, Wen Hong Lam, Mason A. Porter, and Yizhou Sun

Abstract: Ideological divisions in the United States have become increasingly prominent in daily communication. Accordingly, there has been much research on political polarization, including many recent efforts that take a computational perspective. By detecting political biases in a text document, one can attempt to discern and describe its polarity. Intuitively, the named entities (i.e., the nouns and the phrases that act as nouns) and hashtags in text often carry information about political views. For example, people who use the term “pro-choice” are likely to be liberal and people who use the term “pro-life” are likely to be conservative. In this paper, we seek to reveal political polarities in social-media text data and to quantify these polarities by explicitly assigning a polarity score to entities and hashtags. Although this idea is straightforward, it is difficult to perform such inference in a trustworthy quantitative way. Key challenges include the small number of known labels, the continuous spectrum of political views, and the preservation of both a polarity score and a polarity-neutral semantic meaning in an embedding vector of words. To attempt to overcome these challenges, we propose the Polarity-aware Embedding Multi-task learning (PEM) model. This model consists of (1) a self-supervised context-preservation task, (2) an attention-based tweet-level polarity-inference task, and (3) an adversarial learning task that promotes independence between an embedding’s polarity component and its semantic component. Our experimental results demonstrate that our PEM model can successfully learn polarity-aware embeddings that perform well at tweet-level and account-level classification tasks. We examine a variety of applications—including a study of spatial and temporal distributions of polarities and a comparison between tweets from Twitter and posts from Parler—and we thereby demonstrate the effectiveness of our PEM model. We also discuss important limitations of our work and encourage caution when applying the PEM model to real-world scenarios.

Wednesday, October 26, 2022

"Analysis of Spatial and Spatiotemporal Anomalies Using Persistent Homology: Case Studies with COVID-19 Data"

I'm posting about one of my papers that was published in journal form a couple of months ago. I waited for a while because the journal made surprise, unwanted changes after the galley-proof stage — and unsurprisingly I objected very strongly to what they did — and I tried and failed to get those surprise changes addressed. They are very small, but they annoy me (and, as a matter of principle, they should not have made surprise wording changes between the version that we approved and the version that we published). Anyway, here are some details about the article.

Title: Analysis of Spatial and Spatiotemporal Anomalies Using Persistent Homology: Case Studies with COVID-19 Data

Authors: Abigail Hickok, Deanna Needell, and Mason A. Porter

Abstract: We develop a method for analyzing spatial and spatiotemporal anomalies in geospatial data using topological data analysis (TDA). To do this, we use persistent homology (PH), which allows one to algorithmically detect geometric voids in a data set and quantify the persistence of such voids. We construct an efficient filtered simplicial complex (FSC) such that the voids in our FSC are in one- to-one correspondence with the anomalies. Our approach goes beyond simply identifying anomalies; it also encodes information about the relationships between anomalies. We use vineyards, which one can interpret as time-varying persistence diagrams (which are an approach for visualizing PH), to track how the locations of the anomalies change with time. We conduct two case studies using spatially heterogeneous COVID-19 data. First, we examine vaccination rates in New York City by zip code at a single point in time. Second, we study a year-long data set of COVID-19 case rates in neighborhoods of the city of Los Angeles.

Thursday, April 30, 2020

"Community Matters"

Some art that arose from our research was published recently in the collection The Art of Theoretical Biology. Here are some details about our contribution.

Title: Community Matters

Authors: Anna C. F. Lewis, Nick S. Jones, Mason A. Porter, and Charlotte M. Deane

Wednesday, November 27, 2019

Facial Recognition Software for Sheep

This is classical improbable research. This is extremely useful for many things, so first it may make you laugh, but then it makes you think.


(Tip of the cap to Richard Parker.)

Wednesday, July 31, 2019

UCLA's New Undergraduate Major in "Data Theory"

Tuesday, April 09, 2019

A Data Scientist is as Lucky as Lucky Can Be?

They have created an amazing opportunity for song lyrics. :)


Here is their link.

Monday, September 24, 2018

Will the Authors of this Paper Look Back in Anger?

Key question: Will the authors look back in anger when they read their referee reports on this paper?

Tuesday, May 08, 2018

Important Paper on Seussian Configurations of Random-Graph Models

In network science, it is with very good reason that one should speak of "a" configuration model rather than "the" configuration model. To see an excellent discussion of this and related issues, be sure to go through this paper by Bailey Fosdick, Daniel Larremore, Joel Nishimura, and Johan Ugander.

Title: Configuring Random Graph Models with Fixed Degree Sequences

Authors: Bailey K. Fosdick, Daniel B. Larremore, Joel Nishimura, and Johan Ugander


In 2014, Aaron Clauset, David Kempe, and I (with help from Dan Larremore) organized a Mathematics Research Community in Network Science.

In addition to creating a network of network scientists from diverse backgrounds, some work was started there, and today the published version of what is in my opinion an extremely important paper has come out in final form in SIAM Review's 'Research Spotlights' section. I am, of course, talking about the aforementioned paper.

I'm very happy indeed for such excellent work to arise from this.

Congratulations to authors Bailey Fosdick, Daniel Larremore, Joel Nishimura, and Johan Ugander for creating this awesome paper!

I am posting this with an absolutely lovely Seussian picture from an arXiv version of the paper. This picture doesn't appear to have made the cut for the published piece. In addition to its wit and whimsy, a really great thing about the picture and its accompanying verse is that it also encodes the main message of the paper.


Note: I am on the editorial board of the Research Spotlights section of SIAM Review, but I had nothing whatsoever to do with the handling of this paper.

Tuesday, March 27, 2018

"Interval Signatures of Celebrities and Researchers on Twitter"

In case you ever wanted to compare my tweeting behavior to that of Lance Armstrong and other celebrities (and several network and data scientists), here is your chance.

It's always nice to be analyzed (for my tweeting behavior) alongside luminaries like Katy Perry, Shaquille O'Neal, and Steve Martin. Lots of my peeps from network and data science are also put under the data-analytic microscope (or perhaps I should write 'mesoscope') on this page.

Friday, March 02, 2018

Sunday, October 08, 2017

Visualizations of Music

These visualizations are really cool!

(Tip of the cap to Lucas Lacasa.

Wednesday, August 09, 2017

"A Roadmap for the Computation of Persistent Homology"

One of my papers just came out in published form. Here are the details.

Title: A Roadmap for the Computation of Persistent Homology

Authors: Nina Otter, Mason A. Porter, Ulrike Tillmann, Peter Grindrod, and Heather A. Harrington

Abstract: Persistent homology (PH) is a method used in topological data analysis (TDA) to study qualitative features of data that persist across multiple scales. It is robust to perturbations of input data, independent of dimensions and coordinates, and provides a compact representation of the qualitative features of the input. The computation of PH is an open area with numerous important and fascinating challenges. The field of PH computation is evolving rapidly, and new algorithms and software implementations are being updated and released at a rapid pace. The purposes of our article are to (1) introduce theory and computational methods for PH to a broad range of computational scientists and (2) provide benchmarks of state-of-the-art implementations for the computation of PH. We give a friendly introduction to PH, navigate the pipeline for the computation of PH with an eye towards applications, and use a range of synthetic and real-world data sets to evaluate currently available open-source implementations for the computation of PH. Based on our benchmarking, we indicate which algorithms and implementations are best suited to different types of data sets. In an accompanying tutorial, we provide guidelines for the computation of PH. We make publicly available all scripts that we wrote for the tutorial, and we make available the processed version of the data sets used in the benchmarking.

Friday, July 21, 2017

Math Profs are the Least Boring (And Other Observations from Data Analysis)

Here is some interesting data analysis using reviews from Rate My Professors.

You can see from this data analysis that math professors appear to be the least boring among all professors (yay!) and there are some expected (and perhaps more surprising) gender gaps (boo!) in the reviews.

(Tip of the cap to Nicholas Christakis.)

Tuesday, June 27, 2017

What Happens in London Stays in London

I'll be in London for a couple of days for a meeting with one of my industrial collaborators and the Royal Society Workshop on Mathematics for the Modern Economy. Come to the workshop!

Friday, June 16, 2017

How Do People from Different Cultures Draw Circles?

Here is a really interesting article about how people from different cultures draw circles.

I was once — well, at least once — told in elementary school that I was drawing circles the wrong way (because I was using the wrong sense around the clock). I think I responded with the definition of a circle and that the definition doesn't depend on the sense in which one draws it, and I think my teacher did not appreciate that.

On a similar note, in high school, I once lost 50% of the point total on an answer for misspelling ellipse (by using one 'l' instead of two), which was the correct answer. I called bullshit (on the grounds that it was my mathematics knowledge that was being tested), but unfortunately I lost.

More closely related to the article, one thing I noticed in the UK is that the most common way to write an 'x' there is with two arcs, so that they won't always cross if one writes quickly. In contrast, I write two attempts at lines that explicitly cross. (I haven't checked if this is US versus UK convention.)

And, indeed, most Americans drew their circles counterclockwise in this data set, and I draw mine clockwise.

(Tip of the cap to Improbable Research for their Facebook post.)

Update: Here is a lovely quote from the article: In a 1977 paper Theodore Blau, then-president of the American Psychological Association and creator of the torque test, argued that drawing clockwise circles was a sign of learning and behavioral aberrance.

Friday, May 26, 2017

Data Analysis of Gender in Film Dialog

The data set used in this analysis would be really cool to explore (perhaps in combination with movie networks).

Friday, May 12, 2017

Lyrical Repetition in Pop Music

Here is a cool article about lyrical repetition (and compression possibilities) in pop music and how it's changed over the last few decades.

I of course decided to look at how Depeche Mode stacks up, and I zoomed up on them as an individual artist.

From the song and artist library they used, Depeche Mode is listed as the band from the 1990s with the least repetitive lyrics, though it does only use a subset of their songs and it listed them in the 1990s instead of the 1980s. (Naturally, the employed songs span multiple decades.)

The most lyrically repetitive Depeche Mode song is very obvious.

(Tip of the cap to Taha Yasseri and I Fucking Love Data.)

Sunday, February 19, 2017

My Slides on Data Ethics for Mathematicians (and Others)

I have made some slides on data ethics for mathematicians (and others). I'll be presenting them later this term to my graduate course in network science.