Monday, November 25, 2019
Tales from the ArXiv: Mathematics Paper or Science-Fiction Novel?
Tropical PCA sounds very promising and fascinating. Additionally, I think the title of this paper would make a great title for a science-fiction novel.
Wednesday, July 31, 2019
UCLA's New Undergraduate Major in "Data Theory"
At @UCLA (joint between the Math & Stats departments), I helped design a brand new, very exciting Data Theory undergraduate major.
— Mason Porter (@masonporter) August 1, 2019
Courses: https://t.co/z27ZJybKQj
To look up course identities: https://t.co/cHcYrQzVcH
Catalog listing: https://t.co/MMOfGX4Cf4 pic.twitter.com/ZZA4E1G8tE
Friday, December 14, 2018
But Are These Data Cookies Reproducible?
(Tasty, tasty reproducibility.)
Last day of class tomorrow! I baked cookies for my Intro to Data Analytics students, with frosted data plots. Lucky them. #DAatDU pic.twitter.com/j3TCNeDOBe
— Dr. Sarah Supp (@srsupp) December 14, 2018
(Tip of the cap to Javier BuldĂș.)
Update: Now that I think of it, "Sweet, sweet reproducibility." would have been better phrasing, given its larger set of allusions.
Sunday, November 04, 2018
Tales from the ArXiv: The Power Law OF DEATH
And on the heels of The Power Law OF DOOM, we now have claims of The Power Law OF DEATH in a new arXiv paper called "Statistical study of time intervals between murders for serial killers".
In both cases (and as is common), the claims of a power law are unlikely to be justified statistically.
See also The Small-World Network OF LUST, this old blog entry, and this old blog entry.
Monday, September 24, 2018
Will the Authors of this Paper Look Back in Anger?
"Can high-density human collective motion be forecasted by spatiotemporal fluctuations?"
— Chris Danforth (@ChrisDanforth) September 24, 2018
Eigenmodes of crowd dynamics at an Oasis concert. #moshmathhttps://t.co/srSG8Biyui pic.twitter.com/XxnlBQuDFJ
Monday, July 30, 2018
A Cartoon Depiction of Deep Learning Versus Traditional Machine Learning
great meme pic.twitter.com/d5AW6WliTy
— Lex Flagel (@flagelbagel) July 29, 2018
(Tip of the cap to Michael Stumpf.)
Wednesday, May 09, 2018
The Pulse of Manhattan
The population of #Manhattan, hour-by-hour. #NYC #datavizhttps://t.co/KATB4uq9va pic.twitter.com/ov3xTQlVm0
— Randy Olson (@randal_olson) May 9, 2018
Tuesday, March 27, 2018
"Interval Signatures of Celebrities and Researchers on Twitter"
It's always nice to be analyzed (for my tweeting behavior) alongside luminaries like Katy Perry, Shaquille O'Neal, and Steve Martin. Lots of my peeps from network and data science are also put under the data-analytic microscope (or perhaps I should write 'mesoscope') on this page.
Friday, March 02, 2018
Proposal: Use Computational Topology to Study Aversion to Pictures of Holes
For science!
Sunday, October 08, 2017
Visualizations of Music
Data Visualisation of Music https://t.co/45mrRmaTW1 & https://t.co/xFiy9Gs5ay #music #dataviz pic.twitter.com/fLc071JyOG
— Benjamin Hennig (@geoviews) October 7, 2017
(Tip of the cap to Lucas Lacasa.
Sunday, September 03, 2017
Happiness is Not Communicating :)
I take that a step further and try not to spend too much time communicating with others (electronically or otherwise), and I am extremely happy!
Anybody chasing a User Active Minutes metric should spend time with this chart. It's a trap! https://t.co/aksBRKSRre pic.twitter.com/MksLi1JXNl
— William Pietri (@williampietri) September 3, 2017
(Tip of the cap to Hiroki Sayama.)
Wednesday, August 09, 2017
"A Roadmap for the Computation of Persistent Homology"
Title: A Roadmap for the Computation of Persistent Homology
Authors: Nina Otter, Mason A. Porter, Ulrike Tillmann, Peter Grindrod, and Heather A. Harrington
Abstract: Persistent homology (PH) is a method used in topological data analysis (TDA) to study qualitative features of data that persist across multiple scales. It is robust to perturbations of input data, independent of dimensions and coordinates, and provides a compact representation of the qualitative features of the input. The computation of PH is an open area with numerous important and fascinating challenges. The field of PH computation is evolving rapidly, and new algorithms and software implementations are being updated and released at a rapid pace. The purposes of our article are to (1) introduce theory and computational methods for PH to a broad range of computational scientists and (2) provide benchmarks of state-of-the-art implementations for the computation of PH. We give a friendly introduction to PH, navigate the pipeline for the computation of PH with an eye towards applications, and use a range of synthetic and real-world data sets to evaluate currently available open-source implementations for the computation of PH. Based on our benchmarking, we indicate which algorithms and implementations are best suited to different types of data sets. In an accompanying tutorial, we provide guidelines for the computation of PH. We make publicly available all scripts that we wrote for the tutorial, and we make available the processed version of the data sets used in the benchmarking.
Friday, July 21, 2017
Math Profs are the Least Boring (And Other Observations from Data Analysis)
You can see from this data analysis that math professors appear to be the least boring among all professors (yay!) and there are some expected (and perhaps more surprising) gender gaps (boo!) in the reviews.
(Tip of the cap to Nicholas Christakis.)
Tuesday, June 27, 2017
What Happens in London Stays in London
Friday, June 16, 2017
How Do People from Different Cultures Draw Circles?
I was once — well, at least once — told in elementary school that I was drawing circles the wrong way (because I was using the wrong sense around the clock). I think I responded with the definition of a circle and that the definition doesn't depend on the sense in which one draws it, and I think my teacher did not appreciate that.
On a similar note, in high school, I once lost 50% of the point total on an answer for misspelling ellipse (by using one 'l' instead of two), which was the correct answer. I called bullshit (on the grounds that it was my mathematics knowledge that was being tested), but unfortunately I lost.
More closely related to the article, one thing I noticed in the UK is that the most common way to write an 'x' there is with two arcs, so that they won't always cross if one writes quickly. In contrast, I write two attempts at lines that explicitly cross. (I haven't checked if this is US versus UK convention.)
And, indeed, most Americans drew their circles counterclockwise in this data set, and I draw mine clockwise.
(Tip of the cap to Improbable Research for their Facebook post.)
Update: Here is a lovely quote from the article: In a 1977 paper Theodore Blau, then-president of the American Psychological Association and creator of the torque test, argued that drawing clockwise circles was a sign of learning and behavioral aberrance.
Friday, May 26, 2017
Data Analysis of Gender in Film Dialog
The Largest Ever Analysis of Film Dialogue by Gender: 2,000 scripts, 25,000 actors, 4 million lines https://t.co/QV7NzluIpL pic.twitter.com/ampVQskeAu
— Sinan Aral (@sinanaral) May 26, 2017
Tuesday, May 16, 2017
"Quasi-Centralized Limit Order Books"
Title: Quasi-Centralized Limit Order Books
Authors: Martin D. Gould, Mason A. Porter, and Sam D. Howison
Abstract: A quasi-centralized limit order book (QCLOB) is a limit order book (LOB) in which financial institutions can only access the trading opportunities offered by counterpartieswithwhomthey possess sufficient bilateral credit. In this paper, we perform an empirical analysis of a recent, high-quality data set from a large electronic trading platform that utilizes QCLOBs to facilitate trade. We argue that the quote-relative framework often used to study other LOBs is not a sensible reference frame for QCLOBs, so we instead introduce an alternative, trade-relative framework, which we use to study the statistical properties of order flow and LOB state in our data. We also uncover an empirical universality: although the distributions that describe order flow and LOB state vary considerably across days, a simple, linear rescaling causes them to collapse onto a single curve. Motivated by this finding, we propose a semi-parametric model of order flow and LOB state for a single trading day. Our model provides similar performance to that of parametric curve-fitting techniques but is simpler to compute and faster to implement.
Friday, May 12, 2017
Lyrical Repetition in Pop Music
I of course decided to look at how Depeche Mode stacks up, and I zoomed up on them as an individual artist.
From the song and artist library they used, Depeche Mode is listed as the band from the 1990s with the least repetitive lyrics, though it does only use a subset of their songs and it listed them in the 1990s instead of the 1980s. (Naturally, the employed songs span multiple decades.)
The most lyrically repetitive Depeche Mode song is very obvious.
(Tip of the cap to Taha Yasseri and I Fucking Love Data.)
Sunday, February 19, 2017
My Slides on Data Ethics for Mathematicians (and Others)
My slides on "Data Ethics for Mathematicians" (n the era of Big Data) #appliedmathematics #bigdata https://t.co/JBqOcMyf6f via @SlideShare
— Mason Porter (@masonporter) February 19, 2017
Tuesday, January 31, 2017
A UCLA Math Undergrad's Data-Analysis Blog
It includes entries about community detection (in the Harry Potter universe), education issues, and UCLA salaries.