Showing posts with label data analytics. Show all posts
Showing posts with label data analytics. Show all posts

Monday, November 25, 2019

Tales from the ArXiv: Mathematics Paper or Science-Fiction Novel?

A new paper, which looks very fascinating, that just appeared on arXiv is called Tropical Principal Component Analysis on the Space of Ultrametrics.

Tropical PCA sounds very promising and fascinating. Additionally, I think the title of this paper would make a great title for a science-fiction novel.

Wednesday, July 31, 2019

UCLA's New Undergraduate Major in "Data Theory"

Friday, December 14, 2018

But Are These Data Cookies Reproducible?

I think it's very important to check the reproducibility of these data.

(Tasty, tasty reproducibility.)


(Tip of the cap to Javier BuldĂș.)

Update: Now that I think of it, "Sweet, sweet reproducibility." would have been better phrasing, given its larger set of allusions.

Sunday, November 04, 2018

Tales from the ArXiv: The Power Law OF DEATH

First I saw a researcher (Geoff West) claim that there is The Power Law of DOOM! (That's my name for it — not theirs.) At some point, I thought I had that in a blog post, but now I think it was only verbal snarky remarks and a Facebook entry. It started with a talk that he gave years at Oxford in which Geoff started off by saying that he was only going to have very modest conclusions in his seminar, but by the end of the talk, he was predicting the imminent decline of civilization.

And on the heels of The Power Law OF DOOM, we now have claims of The Power Law OF DEATH in a new arXiv paper called "Statistical study of time intervals between murders for serial killers".

In both cases (and as is common), the claims of a power law are unlikely to be justified statistically.

See also The Small-World Network OF LUST, this old blog entry, and this old blog entry.


Monday, September 24, 2018

Will the Authors of this Paper Look Back in Anger?

Key question: Will the authors look back in anger when they read their referee reports on this paper?

Tuesday, March 27, 2018

"Interval Signatures of Celebrities and Researchers on Twitter"

In case you ever wanted to compare my tweeting behavior to that of Lance Armstrong and other celebrities (and several network and data scientists), here is your chance.

It's always nice to be analyzed (for my tweeting behavior) alongside luminaries like Katy Perry, Shaquille O'Neal, and Steve Martin. Lots of my peeps from network and data science are also put under the data-analytic microscope (or perhaps I should write 'mesoscope') on this page.

Friday, March 02, 2018

Sunday, October 08, 2017

Visualizations of Music

These visualizations are really cool!

(Tip of the cap to Lucas Lacasa.

Sunday, September 03, 2017

Happiness is Not Communicating :)

Note that, on average, the happier people spend much less time communicating with other people electronically. (For example, notice the result for "Phone".)

I take that a step further and try not to spend too much time communicating with others (electronically or otherwise), and I am extremely happy!


(Tip of the cap to Hiroki Sayama.)

Wednesday, August 09, 2017

"A Roadmap for the Computation of Persistent Homology"

One of my papers just came out in published form. Here are the details.

Title: A Roadmap for the Computation of Persistent Homology

Authors: Nina Otter, Mason A. Porter, Ulrike Tillmann, Peter Grindrod, and Heather A. Harrington

Abstract: Persistent homology (PH) is a method used in topological data analysis (TDA) to study qualitative features of data that persist across multiple scales. It is robust to perturbations of input data, independent of dimensions and coordinates, and provides a compact representation of the qualitative features of the input. The computation of PH is an open area with numerous important and fascinating challenges. The field of PH computation is evolving rapidly, and new algorithms and software implementations are being updated and released at a rapid pace. The purposes of our article are to (1) introduce theory and computational methods for PH to a broad range of computational scientists and (2) provide benchmarks of state-of-the-art implementations for the computation of PH. We give a friendly introduction to PH, navigate the pipeline for the computation of PH with an eye towards applications, and use a range of synthetic and real-world data sets to evaluate currently available open-source implementations for the computation of PH. Based on our benchmarking, we indicate which algorithms and implementations are best suited to different types of data sets. In an accompanying tutorial, we provide guidelines for the computation of PH. We make publicly available all scripts that we wrote for the tutorial, and we make available the processed version of the data sets used in the benchmarking.

Friday, July 21, 2017

Math Profs are the Least Boring (And Other Observations from Data Analysis)

Here is some interesting data analysis using reviews from Rate My Professors.

You can see from this data analysis that math professors appear to be the least boring among all professors (yay!) and there are some expected (and perhaps more surprising) gender gaps (boo!) in the reviews.

(Tip of the cap to Nicholas Christakis.)

Tuesday, June 27, 2017

What Happens in London Stays in London

I'll be in London for a couple of days for a meeting with one of my industrial collaborators and the Royal Society Workshop on Mathematics for the Modern Economy. Come to the workshop!

Friday, June 16, 2017

How Do People from Different Cultures Draw Circles?

Here is a really interesting article about how people from different cultures draw circles.

I was once — well, at least once — told in elementary school that I was drawing circles the wrong way (because I was using the wrong sense around the clock). I think I responded with the definition of a circle and that the definition doesn't depend on the sense in which one draws it, and I think my teacher did not appreciate that.

On a similar note, in high school, I once lost 50% of the point total on an answer for misspelling ellipse (by using one 'l' instead of two), which was the correct answer. I called bullshit (on the grounds that it was my mathematics knowledge that was being tested), but unfortunately I lost.

More closely related to the article, one thing I noticed in the UK is that the most common way to write an 'x' there is with two arcs, so that they won't always cross if one writes quickly. In contrast, I write two attempts at lines that explicitly cross. (I haven't checked if this is US versus UK convention.)

And, indeed, most Americans drew their circles counterclockwise in this data set, and I draw mine clockwise.

(Tip of the cap to Improbable Research for their Facebook post.)

Update: Here is a lovely quote from the article: In a 1977 paper Theodore Blau, then-president of the American Psychological Association and creator of the torque test, argued that drawing clockwise circles was a sign of learning and behavioral aberrance.

Friday, May 26, 2017

Data Analysis of Gender in Film Dialog

The data set used in this analysis would be really cool to explore (perhaps in combination with movie networks).

Tuesday, May 16, 2017

"Quasi-Centralized Limit Order Books"

One of my papers got assigned its final journal coordinates today. (It came out a few months ago in advanced access.) Here are the details.

Title: Quasi-Centralized Limit Order Books

Authors: Martin D. Gould, Mason A. Porter, and Sam D. Howison

Abstract: A quasi-centralized limit order book (QCLOB) is a limit order book (LOB) in which financial institutions can only access the trading opportunities offered by counterpartieswithwhomthey possess sufficient bilateral credit. In this paper, we perform an empirical analysis of a recent, high-quality data set from a large electronic trading platform that utilizes QCLOBs to facilitate trade. We argue that the quote-relative framework often used to study other LOBs is not a sensible reference frame for QCLOBs, so we instead introduce an alternative, trade-relative framework, which we use to study the statistical properties of order flow and LOB state in our data. We also uncover an empirical universality: although the distributions that describe order flow and LOB state vary considerably across days, a simple, linear rescaling causes them to collapse onto a single curve. Motivated by this finding, we propose a semi-parametric model of order flow and LOB state for a single trading day. Our model provides similar performance to that of parametric curve-fitting techniques but is simpler to compute and faster to implement.

Friday, May 12, 2017

Lyrical Repetition in Pop Music

Here is a cool article about lyrical repetition (and compression possibilities) in pop music and how it's changed over the last few decades.

I of course decided to look at how Depeche Mode stacks up, and I zoomed up on them as an individual artist.

From the song and artist library they used, Depeche Mode is listed as the band from the 1990s with the least repetitive lyrics, though it does only use a subset of their songs and it listed them in the 1990s instead of the 1980s. (Naturally, the employed songs span multiple decades.)

The most lyrically repetitive Depeche Mode song is very obvious.

(Tip of the cap to Taha Yasseri and I Fucking Love Data.)

Sunday, February 19, 2017

My Slides on Data Ethics for Mathematicians (and Others)

I have made some slides on data ethics for mathematicians (and others). I'll be presenting them later this term to my graduate course in network science.

Tuesday, January 31, 2017

A UCLA Math Undergrad's Data-Analysis Blog

UCLA undergraduate student Ritvik Kharkar, who is doing some research with me, is doing some spiffy data analysis on his blog. Go take a look at it!

It includes entries about community detection (in the Harry Potter universe), education issues, and UCLA salaries.