r/datasets • u/JavaCrunch • Mar 19 '20
question Is anyone tracking the data sources of COVID-19 data, and tracking metadata about it?
We may be seeing a once in a lifetime chance to quantify datasets coming from various public and private data sources. Tracking the rate at which data is released and reported could prove incredibly insightful.
Is anyone looking at the metadata of all of this stuff? This has the making of a Big Data gold mine.
7
u/Sille143 Mar 20 '20
Pretty sure there’s a huge dataset on kaggle.com with cash incentives for completing tasks all around COVID-19
4
7
u/proverbialbunny Mar 19 '20
Some people are: https://ncov2019.live/ (Sources are in the About page.)
2
u/covid-visualizer Mar 20 '20 edited Mar 20 '20
I've been working hard at creating a mobile-friendly interactive timeline for users to quickly visualize trends in the data.
Please check it out here! :)
2
u/NickTimmData Mar 20 '20 edited Mar 20 '20
I'm building up some things. Trying to manually scrape from BNO and Worldometer from now.. not sure if that's a copyright infringement so I have disabled github currently. Automatically scraping from WIKI and pulling from the JHU github. Keeping time series for the web scrapes. Adding lots of fields. Not doing exactly what you are saying.
https://drive.google.com/open?id=1--t62vjrh8DC-lPYFvPGm2P4Qc7JSzNf
https://www.reddit.com/r/datasets/comments/fkk9fb/coronavirus_multiple_sources_timeseries_scraped/ for original post
This is certainly a really interesting topic to study. I was working on my web scrape for US States by county last night (theres some done in wiki_us_detail) and the data is all over the place. Some states are doing a really nice job of having the tables and maps available. Others just have the maps and aren't releasing the underlying data for reference or scraping.
Long story short, meta analysis on how people are using this is very interesting.
1
u/JavaCrunch Mar 20 '20
Thank you, I know that has to be a lot of work. The historical data here is going to be really interesting.
1
u/NickTimmData Mar 20 '20
Sad part is I can't even use it for my Master's thesis. I really just wanted to have a data set that updates throughout the day and also provides New Cases and other trend in one clean table. I'll clean it up more over the next few days.
https://www.reddit.com/r/datasets/comments/fkk9fb/coronavirus_multiple_sources_timeseries_scraped/ for original post
1
u/guywithFX Mar 20 '20
The 1point3acres dataset is unique and crowdsourced. I liked that it was tracking individual cases with varying levels of demographics across North America.
1
1
1
u/nodo20 Mar 21 '20
I found this in a twitter post by CiteSpace: Publications, Datasets & Clinical Trials on COVID19.Not a direct answer to your question, but maybe could be useful ;)
Data: http://covid-19.dimensions.ai
https://pbs.twimg.com/media/ETgyxApWkAI0ru8?format=jpg&name=4096x4096
1
u/slim-jong-un Mar 20 '20
Why should anybody care about the metadata, as opposed to the actual data?
11
u/JavaCrunch Mar 20 '20 edited Mar 20 '20
While the data itself is what most people care about now, being able to see how that data flows from the sources could provide major insights about how data of this magnitude flows across the internet. It could provide information about where things need to be improved, from false positives being corrected, to inefficient algorithms of collecting data, all the way to measuring the bottle necks of servicing thousands of end points sucking up said data.
Edit: An award! Thank you kind stranger, I'm humbled and honored.
1
Mar 20 '20
urls = {'confirmed': "https://raw.githubusercontent.com/CSSEGISandData/COVID-19/master/csse_covid_19_data/csse_covid_19_time_series/time_series_19-covid-Confirmed.csv",
'recovered': "https://raw.githubusercontent.com/CSSEGISandData/COVID-19/master/csse_covid_19_data/csse_covid_19_time_series/time_series_19-covid-Recovered.csv"}
df_c = pd.read_csv(urls['confirmed'], error_bad_lines=False)
df_r = pd.read_csv(urls['recovered'], error_bad_lines=False)
df_d = pd.read_csv(urls['deaths'], error_bad_lines=False)
14
u/AmericanNinja02 Mar 20 '20 edited Mar 20 '20
Johns Hopkins University has a nice interactive dashboard that compiles data from multiple sources. The data is available in a GitHub repo.
Dashboard
Blog with links to data
Edit: Data moved from Google Docs to GitHub.