Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Wednesday, July 31, 2019

UCLA's New Undergraduate Major in "Data Theory"

Tuesday, April 09, 2019

A Data Scientist is as Lucky as Lucky Can Be?

They have created an amazing opportunity for song lyrics. :)


Here is their link.

Sunday, September 03, 2017

Happiness is Not Communicating :)

Note that, on average, the happier people spend much less time communicating with other people electronically. (For example, notice the result for "Phone".)

I take that a step further and try not to spend too much time communicating with others (electronically or otherwise), and I am extremely happy!


(Tip of the cap to Hiroki Sayama.)

Wednesday, August 09, 2017

"A Roadmap for the Computation of Persistent Homology"

One of my papers just came out in published form. Here are the details.

Title: A Roadmap for the Computation of Persistent Homology

Authors: Nina Otter, Mason A. Porter, Ulrike Tillmann, Peter Grindrod, and Heather A. Harrington

Abstract: Persistent homology (PH) is a method used in topological data analysis (TDA) to study qualitative features of data that persist across multiple scales. It is robust to perturbations of input data, independent of dimensions and coordinates, and provides a compact representation of the qualitative features of the input. The computation of PH is an open area with numerous important and fascinating challenges. The field of PH computation is evolving rapidly, and new algorithms and software implementations are being updated and released at a rapid pace. The purposes of our article are to (1) introduce theory and computational methods for PH to a broad range of computational scientists and (2) provide benchmarks of state-of-the-art implementations for the computation of PH. We give a friendly introduction to PH, navigate the pipeline for the computation of PH with an eye towards applications, and use a range of synthetic and real-world data sets to evaluate currently available open-source implementations for the computation of PH. Based on our benchmarking, we indicate which algorithms and implementations are best suited to different types of data sets. In an accompanying tutorial, we provide guidelines for the computation of PH. We make publicly available all scripts that we wrote for the tutorial, and we make available the processed version of the data sets used in the benchmarking.

Friday, July 21, 2017

Math Profs are the Least Boring (And Other Observations from Data Analysis)

Here is some interesting data analysis using reviews from Rate My Professors.

You can see from this data analysis that math professors appear to be the least boring among all professors (yay!) and there are some expected (and perhaps more surprising) gender gaps (boo!) in the reviews.

(Tip of the cap to Nicholas Christakis.)

Tuesday, June 27, 2017

What Happens in London Stays in London

I'll be in London for a couple of days for a meeting with one of my industrial collaborators and the Royal Society Workshop on Mathematics for the Modern Economy. Come to the workshop!

Monday, June 26, 2017

Historical Naval Charts and Maps

The National Maritime Museum (one of the Royal Museums of Greenwich) has a really cool collection of historical sea charts and maps. It would be really cool to do some research involving them!

I found this out via the follow tweet.


(Tip of the cap to Sydney Padua.)

Friday, June 16, 2017

How Do People from Different Cultures Draw Circles?

Here is a really interesting article about how people from different cultures draw circles.

I was once — well, at least once — told in elementary school that I was drawing circles the wrong way (because I was using the wrong sense around the clock). I think I responded with the definition of a circle and that the definition doesn't depend on the sense in which one draws it, and I think my teacher did not appreciate that.

On a similar note, in high school, I once lost 50% of the point total on an answer for misspelling ellipse (by using one 'l' instead of two), which was the correct answer. I called bullshit (on the grounds that it was my mathematics knowledge that was being tested), but unfortunately I lost.

More closely related to the article, one thing I noticed in the UK is that the most common way to write an 'x' there is with two arcs, so that they won't always cross if one writes quickly. In contrast, I write two attempts at lines that explicitly cross. (I haven't checked if this is US versus UK convention.)

And, indeed, most Americans drew their circles counterclockwise in this data set, and I draw mine clockwise.

(Tip of the cap to Improbable Research for their Facebook post.)

Update: Here is a lovely quote from the article: In a 1977 paper Theodore Blau, then-president of the American Psychological Association and creator of the torque test, argued that drawing clockwise circles was a sign of learning and behavioral aberrance.

Friday, May 26, 2017

Data Analysis of Gender in Film Dialog

The data set used in this analysis would be really cool to explore (perhaps in combination with movie networks).

Friday, May 05, 2017

Replacing "Big Data" by "Batman" in Tweets

This Twitter account takes tweets and replaces "Big Data" with "Batman". Hilarity ensues.

Here is an example.



(Tip of the cap to Esteban Moro.)

Sunday, February 19, 2017

My Slides on Data Ethics for Mathematicians (and Others)

I have made some slides on data ethics for mathematicians (and others). I'll be presenting them later this term to my graduate course in network science.

Tuesday, January 31, 2017

A UCLA Math Undergrad's Data-Analysis Blog

UCLA undergraduate student Ritvik Kharkar, who is doing some research with me, is doing some spiffy data analysis on his blog. Go take a look at it!

It includes entries about community detection (in the Harry Potter universe), education issues, and UCLA salaries.

Tuesday, June 07, 2016

Quantifying the Smell of Urban Areas

My latest post in the Improbable Research blog is about using data analysis to quantify the smell of urban areas.

Monday, June 06, 2016

A Video of My Seminar on "A Simple Generative Model of Collective Online Behavior"

On Friday 27 May, I gave a talk at the Oxford Internet Institute on "A Simple Generative Model of Collective Online Behavior". The video was posted online today. My talk, which took place at the Oxford Internet Institute, was part of an event for the Computational Social Science Initiative London.

I spoke about this paper, which I like to call our "Small Data" paper.

You may also be interested in videos of some of my other presentations.

Thursday, August 06, 2015

Early-Career Research Positions Available at the New Alan Turing Institute for Data Science

The UK has a new "Alan Turing Institute" for data science, and there is a call for expressions of interest from early-career researchers.

One should click on "Research Positions" in the yellow panel on the right for the instructions regarding expressions of interest.

Quoting the website (see the site for more info): Prospective applicants are invited to submit a curriculum vitae and a one-page covering note explaining how their expertise is relevant to data science and the mission of the Institute via email to info@turing.ac.uk. Full details of the application procedure will be sent in the autumn to those who register their interest.

Monday, May 04, 2015

The Data Science Handbook

The Data Science Handbook looks pretty cool.

(Tip of the cap to Bradley Voytek.)

Sunday, February 22, 2015

The Big, the Rich, and the Good (Data)

Nate Silver has written a fascinating article about the stellar success and rapid progress of baseball analytics versus the less-than-rapid progress of data analytics in other areas (e.g., economic and earthquake forecasts).

Putting the baseball angle aside (and of course I like that angle very much), one thing I really like is the very concise comments about "big data" versus what Silver calls "rich data" and why sports analytics have genuinely improved answers to many important-for-it questions, whereas other situations still struggle immensely to use their data for genuinely large increases in understanding.

Note that I have sometimes used the term "good data" before as a contrast to "big data", though importantly good data can be either big or small, and I think that Silver is thinking of Rich Data strictly as a subset of Big Data. See this blog entry of mine as well as additional blog entries referenced therein.

One could also, I suppose, ask whether these other systems are "more complex" than sports competitions, but I'm not sure (a) whether that's actually true and (b) how to quantify it in a way that goes beyond "Ooh, it's complex." Of course, we have measures of information content, but I expect there would be a lot of assumptions involved in crunching such numbers in these cases. Anyway, it's a thought I had, so I figured that I should at least bring it up.

Tuesday, January 20, 2015

The Data Fairy

I greatly prefer terms like "Data Science" and "Data Analytics" to the ubiquitous "Big Data".

"Big Data" makes me feel like you put the data under your pillow in the hope that the Data Fairy comes during the night and leaves an answer there.

The point is that there is supposed to be actual science along with the data.

(See also an older blog entry of mine as well as one to which it links.)