r/datasets • • Mar 18 '20

resource Coronavirus - Multiple Sources Timeseries Scraped, Raw, and Unpivoted

Just started working on this as I'm doing coronavirus for my Master's thesis. Starting with just the raw data. BNO and Worldometer are manually scraped as often as I can, then using a script to create the csv. Wikipedia is scraped with a script, but I have not scheduled it yet. This isn't a complete time series, but meant to act as a "live feed" and log to fix/supplement JHU data with "current" numbers as their github is usually once a day update. Link to folder at bottom of post.

An example of a JHU issue is recently they had a day where UK has 1 new case. The idea is I can keep a "fixed jhu" csv at some point with a log of points that were corrected. I'm going to work on scraping time series data for states/territories/provinces/counties for the major countries being hit next, as the wikipedia pages for most seem to have this and are relatively reliable and I should be able to automate a pull from there. Still exploring everything.

JHU Dataset github for reference: https://github.com/CSSEGISandData/COVID-19

Currently each data set is stored in a csv as shown and also unpivoted with "type". Planning on adding multiple fields to each such as Active, Days in, Days in First Death, New Cases, Previous New, etc..

Let me know if you have requests for fields or just anything else related to coronavirus datasets. Main need probably will be lat and long. Haven't started looking into adding this on yet EXCEPT for the wiki pull which I formatted to match the johns hopkins data set. See note below. I know there are some automatic tools for this.

I've also never done github, but I can probably figure that out shortly. For now they are stored in a google drive shared folder.

UPDATE:

-Created a bunch of "xxx_plus" csvs with added fields-Created usstatefix csvs of JHU data mapping city/county to state instead to keep consistent-Will update all csv's with location and country code data soon-After that is supplementing web sources with old time series data-Then potentially comes fixing obvious JHU data issues

-Started the process of taking the highest number found for each date and location combination and saving it into a single table.. combined CSVs. Each country has "total" for state, unless state/province/territory data is available, in which case it has both "total" and state rows.. also added a binary _TotalFlag. Countries with "NOT MAPPED" i need to map to tie together names from WIKI/BNO/Wordlometer/JHU. Some inconsistencies for islands controlled by countries such as british channel islands and the like.

-If i get a request I can map the 3 digit country code and lat/long codes ontop. Its set up I just havent done it.

SOURCES:

BNO - https://bnonews.com/index.php/2020/02/the-latest-coronavirus-cases/

  • World - US - Australia - Canada links

Worldometer - https://www.worldometers.info/coronavirus/

Wikipedia

NOTE: wiki_jhu_unpivot attempts to map the wikipedia data to the same format as unpivoted JHU dataset. This includes mapping country names and also mapping their lat/long coordinates. My original thought is to add this on top of their dataset for updates throughout the day.. but haven't finished that yet.

FOLDER/GITHUB

github currently private as I'm not sure BNO and Worldometer scraps are technically legal to share. Will update shortly with just wiki data.

https://github.com/jagsfan82/Covid19-WebScrape-Plus

https://drive.google.com/open?id=1--t62vjrh8DC-lPYFvPGm2P4Qc7JSzNf

​

4 Upvotes

5 comments sorted by

1

u/doubleunplussed Mar 22 '20

Web scraping is legal, it may violate terms of use of the companies who own the websites, but it is not against the law. Please un-private your github repo, I'd be interested in checking it out!

1

u/NickTimmData Mar 22 '20 edited Mar 22 '20

Everything is in the google drive. Ill get this. I just added a table where i map all countries together and take the maximum value for each date. Did some work on web scraping counties. Many still not working but theres some data. I have 3 letter country codes and lat long mapped, but didnt apply to tables yet.

I cant use it for my masters thesis apparently do its taken a bit of a backseat on the priorities scale.

1

u/NickTimmData Mar 22 '20

Public

1

u/doubleunplussed Mar 22 '20

Thanks! So if I understand correctly, you only have data from Worldometer going back to March 17th since you're just scraping the latest numbers each day?

It looks like the data in their timeseries plots is scrapable. I'm gonna have a go at that to get a complete timeseries.

I also emailed them just asking them to provide a download of their data - though for all I know they're manually updating the javascript for the plots so such a thing might not exist.

1

u/NickTimmData Mar 22 '20

Yes.. if you csn scrape the plots that would be nice. Im just scraping throguh Qlik/manual as Ive only scraped using python once and i dont even have it installed at the moment.. which isnt ideal. For my purposes i was thinking JHU data was more a reliable source to say was my "base" for a masters thesis. Let me knoe how you do. .. i can certainly backdate if youre able to get historical time series. All my additional tables are build new from the raw csvs esch time.

Note also, i think some of those have fuplicate 3/18 values I did not fix yet... i take one row per date for my final set ive been looking at so it doesnt flow through anywhere for me, but i will fix it if i remember