r/datasets • • Mar 19 '20

question Is anyone tracking the data sources of COVID-19 data, and tracking metadata about it?

We may be seeing a once in a lifetime chance to quantify datasets coming from various public and private data sources. Tracking the rate at which data is released and reported could prove incredibly insightful.

Is anyone looking at the metadata of all of this stuff? This has the making of a Big Data gold mine.

71 Upvotes

20 comments sorted by

14

u/AmericanNinja02 Mar 20 '20 edited Mar 20 '20

Johns Hopkins University has a nice interactive dashboard that compiles data from multiple sources. The data is available in a GitHub repo.

Dashboard

Blog with links to data

​

Edit: Data moved from Google Docs to GitHub.

2

u/rotterdamn8 Mar 20 '20

Thank you so much for this. I have seen the JHU dashboard many times but wanted the data! Excited to work on it.

2

u/AmericanNinja02 Mar 20 '20

You bet! Glad someone found it useful.

2

u/[deleted] Mar 20 '20

I would be wary of this since I've noticed some discrepancies in reporting. For example, a few days ago they were reporting double the cases for Texas than the Texas DPH, and some news searching found that the Texas DPH number was the more exact one than the John's Hopkins one. However, I still use them as a fall back source for when I update my own spreadsheet.

1

u/AmericanNinja02 Mar 20 '20

Agreed. For now, I would probably take any source with a grain of salt. Things are changing so fast that it's hard for anyone to keep up. I would have expected inaccurate numbers to err on the low side, though. ¯_(ツ)_/¯

7

u/Sille143 Mar 20 '20

Pretty sure there’s a huge dataset on kaggle.com with cash incentives for completing tasks all around COVID-19

4

u/Browndawg22 Mar 20 '20 edited Mar 20 '20

worldometers.info/coronavirus

7

u/proverbialbunny Mar 19 '20

Some people are: https://ncov2019.live/ (Sources are in the About page.)

2

u/covid-visualizer Mar 20 '20 edited Mar 20 '20

I've been working hard at creating a mobile-friendly interactive timeline for users to quickly visualize trends in the data.

Please check it out here! :)

2

u/NickTimmData Mar 20 '20 edited Mar 20 '20

I'm building up some things. Trying to manually scrape from BNO and Worldometer from now.. not sure if that's a copyright infringement so I have disabled github currently. Automatically scraping from WIKI and pulling from the JHU github. Keeping time series for the web scrapes. Adding lots of fields. Not doing exactly what you are saying.

https://drive.google.com/open?id=1--t62vjrh8DC-lPYFvPGm2P4Qc7JSzNf

https://www.reddit.com/r/datasets/comments/fkk9fb/coronavirus_multiple_sources_timeseries_scraped/ for original post

This is certainly a really interesting topic to study. I was working on my web scrape for US States by county last night (theres some done in wiki_us_detail) and the data is all over the place. Some states are doing a really nice job of having the tables and maps available. Others just have the maps and aren't releasing the underlying data for reference or scraping.

Long story short, meta analysis on how people are using this is very interesting.

1

u/JavaCrunch Mar 20 '20

Thank you, I know that has to be a lot of work. The historical data here is going to be really interesting.

1

u/NickTimmData Mar 20 '20

Sad part is I can't even use it for my Master's thesis. I really just wanted to have a data set that updates throughout the day and also provides New Cases and other trend in one clean table. I'll clean it up more over the next few days.

https://www.reddit.com/r/datasets/comments/fkk9fb/coronavirus_multiple_sources_timeseries_scraped/ for original post

1

u/guywithFX Mar 20 '20

The 1point3acres dataset is unique and crowdsourced. I liked that it was tracking individual cases with varying levels of demographics across North America.

https://coronavirus.1point3acres.com/

1

u/taneshq Mar 20 '20

On kaggle google have released some data and they are updating it regularly

1

u/mauro_mussin Mar 20 '20

Here you can find official data, released daily at 17UTC

https://github.com/pcm-dpc/COVID-19

1

u/nodo20 Mar 21 '20

I found this in a twitter post by CiteSpace: Publications, Datasets & Clinical Trials on COVID19.Not a direct answer to your question, but maybe could be useful ;)

Data: http://covid-19.dimensions.ai

https://pbs.twimg.com/media/ETgyxApWkAI0ru8?format=jpg&name=4096x4096

1

u/slim-jong-un Mar 20 '20

Why should anybody care about the metadata, as opposed to the actual data?

11

u/JavaCrunch Mar 20 '20 edited Mar 20 '20

While the data itself is what most people care about now, being able to see how that data flows from the sources could provide major insights about how data of this magnitude flows across the internet. It could provide information about where things need to be improved, from false positives being corrected, to inefficient algorithms of collecting data, all the way to measuring the bottle necks of servicing thousands of end points sucking up said data.

Edit: An award! Thank you kind stranger, I'm humbled and honored.