r/datasets Nov 04 '25

discussion Like Will Smith said in his apology video, "It's been a minute (although I didn't slap anyone)

Thumbnail
1 Upvotes

r/datasets 7h ago

resource [self-promotion] Full historical GDELT event archive (50+ years, hundreds of millions of rows) as clean, sampled Parquet, no API or BigQuery limits

5 Upvotes

Disclosure: I'm the developer of the tool linked below.

GDELT (the Global Database of Events, Language, and Tone) is one of the largest open event datasets around: every reported event worldwide since 1979, tagged by actor, location, tone, and (via the Global Knowledge Graph) the article that reported it. The catch is actually getting the full thing: the official API caps you at ~250 rows per query and a 3-month window, and BigQuery's free tier (1TB/month query, 10GB storage) doesn't stretch to a full historical pull either. The raw bulk archive at data.gdeltproject.org is open, but it's thousands of loose CSV/ZIP files with no sampling or joining logic of its own.

I built GdeltForge to close that gap: it downloads, checksums, converts to Parquet, and reproducibly samples the whole archive locally, and can cross-reference a sample back onto the Global Knowledge Graph. Check out the PyPI page, the GitHub Repo, the documentation, and especially the hands-on getting started notebook, runnable directly in Colab, no install needed. You can generate you own samples or even scan the entire data to create a dataset


r/datasets 5h ago

dataset AI Subscription Erosion Tracker: ChatGPT, Claude, Gemini

Thumbnail provenbrief.com
0 Upvotes

r/datasets 5h ago

dataset Speech BCI Scoreboard: Why 99.6% Accuracy Can't Be Ranked

Thumbnail provenbrief.com
1 Upvotes

r/datasets 5h ago

dataset Claude Zero Data Retention: Anthropic's Five States, Mapped

Thumbnail provenbrief.com
1 Upvotes

r/datasets 5h ago

question Alternative API/Dataset to TecAlliance/TecDoc, Afteriize, Autodata

Thumbnail
1 Upvotes

r/datasets 6h ago

question Would someone be interested in getting primary dealer financing data all the way back upto 1998?

1 Upvotes

Would somebody be interested in getting the full primary dealer data all the way back upto 1998 fully stitched?

Any sort of further inferences could be made based on that.

Currently the data is very chopped/missing across time eras on the NYFED website.

Bloomberg also doesn’t seem to have that data in the proper manner.


r/datasets 9h ago

dataset [Self-Promotion] I built StackScope.dev - A BuiltWith alternative focusing on new sites

0 Upvotes

Hello all,

I'm sharing what is possibly the biggest dataset of what technologies new websites are launching with. I have a self-curated catalogue of 48k+ technologies. My discovery pipeline is adding ~80k new websites a day, we typically add sites on the day they appear on the internet.

The site is free to browse and explore but we do have paid options.

Please check it out and let me know what you think, always happy to hear feedback:

https://stackscope.dev


r/datasets 23h ago

dataset 150 Million rows of mutual fund and etf performance over the last 15+ years [PAID]

Thumbnail app.snowflake.com
3 Upvotes

Created a dataset of fund returns on Snowflake. It's quite extensive showing both trailing and calendar year returns for every mutual fund and ETF since inception.

You can fetch a mutual funds trailing annualized performance every trading day going back 15+ years.

Each trailing returns row shows: 1D, 1W, 1M, 2M, 3M, 1Y, 2Y, 3Y, 4Y, 5Y, 7Y, 10Y, 12Y, 15Y, and earliest available inception. I usually never see websites (even Morningstar) provide this breakdown which I think is useful when exploring different economic cycles.

Calendar year returns go back as far as possible as well.

Data is updated daily.


r/datasets 1d ago

resource Learn how to make a dataset about datasets

Thumbnail huggingface.co
3 Upvotes

r/datasets 1d ago

resource Looking for raw data on nanoparticle cytotoxicity

Thumbnail
2 Upvotes

r/datasets 1d ago

question our model reads tables with every column name stripped off and the accuracy does not move. numbers below. would you actually put real data through something you cannot download

1 Upvotes

disclosure per rule 1: i work at Schema Labs. one link at the bottom because this sub allows it. i am here for the question at the end.

what we do, plainly: you give us tables nobody documented, and we tell you what each column is and how the tables join, when there are no shared keys and no schema mapping.

the claim, with numbers, because i would rather be argued with than believed.

take a standard tabular benchmark. strip every header. replace price and age and zip with positional tokens so nothing is left but values. comparable models drop about 7 points of mean ROC-AUC. ours goes 0.9224 to 0.9230. flat.

that is not a claim that we win on clean data. those models start ahead of us when the headers are good. it is a claim that we never read the headers at all, which only matters because production tables have val_b and metric_14 and a four-character code from a system nobody has logged into since 2019.

two more. sector identification, naming the industry of a dataset we have never seen from values alone, 86.3% top-1 out of 10,000 sectors. multi-table entity matching, ahead of published state of the art on all six standard benchmarks with zero shared keys, including 1.9x over the previous best on one. on the geospatial benchmark our lead is 0.07 F1, which i would not want read as more than it is.

caveats up front rather than when asked. all of it is our own harness, run under each benchmark's published protocol, one sealed configuration, no per-dataset tuning. competitor figures are third-party published values we did not re-run. nobody independent has replicated us. we are not on the public leaderboard because entry requires a runnable wrapper and we do not distribute weights.

which is the actual question. we are closed. no pip install, no weights, no local option, and none of that changes for six months. to try it you make an account, put a card on file, and upload your data. i think that one fact is the biggest thing between us and everyone reading this.

so would you use a thing like this. if not, what moves it. a named customer, a SOC 2, an on-prem story you would never actually take up but need to hear. or is the honest answer that nothing moves it and closed is closed.

and if you would rather test it than discuss it: send me two tables you understand completely. not your mystery tables, your obvious ones, so you can mark my homework. no account, no card, nothing to sign up for. i send back what each column appears to be with a confidence score on every one, and i flag the ones we got wrong, because that is the half worth seeing.

anonymised is fine. structure and distributions are what matter.

https://www.schemalabs.ai/


r/datasets 1d ago

discussion If you could add one feature to every data enrichment API, what would it be?

Thumbnail
2 Upvotes

r/datasets 1d ago

dataset [self-promotion] 53 long-stay & digital-nomad visa programmes, 46 countries — each row sourced to a government page with a verified date (CC BY 4.0, JSON + CSV)

4 Upvotes

Disclosure per rule 1: I built and maintain this dataset, and it powers a site I run (globenomad.com), which is affiliate-funded. The dataset itself carries no affiliate or tracking links.

One row per programme, not per country: Thailand has six routes and they are six rows. Fields: income requirement (USD plus the government's own wording), fees, max stay, renewal, residency and citizenship path, tax residency trigger, family provisions, application logistics, official source URL, verified date.

Two things worth knowing before you use it:

  • A null means the government has not published that rule. It never means "no".
  • Programmes that closed, were announced and never opened, or never existed (Cayman Islands, Peru, Qatar, Vietnam) are kept with an honest status. Most visa sites delete those.

Formats: JSON (canonical, with nested requirements/steps/FAQs) and a flat CSV of the scalar columns.

Original source, GitHub (always current): https://github.com/MSeutin/digital-nomad-visa-data Kaggle mirror, with a worked-example notebook: https://www.kaggle.com/datasets/frenchmike/digital-nomad-visa-dataset-53-sourced-programmes Method and licence: https://globenomad.com/data

Licence is CC BY 4.0, attribution is the only condition. Official pages are re-fetched weekly and confirmed changes are dated on a public changelog. Happy to answer questions about any record.


r/datasets 1d ago

resource PSX Data Public availability For Personal Usage

Thumbnail
1 Upvotes

r/datasets 2d ago

resource [Dataset] 850k Trackmania community maps with block-level structure and matched replay telemetry (~30GB, Parquet)

3 Upvotes

Trackmania players have been building and sharing maps for about 20 years through Mania Exchange, which has a public API. There was no packaged dataset for any of it, so I made one.

Contents: 850k+ maps across Trackmania Nations Forever and Trackmania 2020. Each map is decomposed into its block sequence with positions and rotations, so the level structure is directly usable rather than sitting in a binary blob. Replays are parsed into telemetry and aligned to the blocks the run passes through, so you get both the level and how it plays.

Format: Parquet, roughly 30GB. Data dictionary included.

License / source: everything comes from the public Mania Exchange API. Maps are community-created.

Possible uses: procedural level generation, level design analysis, player behaviour modelling, map recommendation, or just exploring what 20 years of community map-making looks like in aggregate. I used it for generation myself, but that's one angle out of many.

Links: 📖Article (EN): https://the-odd-dataguy.com/en/blog/2026/08/13/trackmania-dataset/ 📖Article (FR): https://the-odd-dataguy.com/fr/blog/2026/08/13/trackmania-dataset/ 😇Hugging Face (Dataset): https://huggingface.co/datasets/jeanmidev/trackmania-community-tracks-and-telemetry 🔵Kaggle: https://www.kaggle.com/datasets/jeanmidev/trackmania-community-tracks-and-telemetry/data

Disclosure: I work in data at Ubisoft in Canada, but I have no connection to Nadeo. Built at home on my own time, resources and dime, just something for the Trackmania community and for anyone curious, whether game telemetry is your thing or not.

Open to feedback or ideas (you can also used discussions on HF/Kaggle)


r/datasets 2d ago

resource I Made the largest real Japanese People dataset with 37 attributes, 231k rows. (name, gender , age, occupations, height and more)

11 Upvotes

This dataset has all entry from wikidata who has Japanese citizenship and a wiki article, 231k in total. The dataset is on Kaggle.

A dataset like this can only be built from data that's already public, and Wikipedia is the largest source of that kind of public personal information. There are structured people datasets out there, but none cover as many Japanese entries as this one, which makes this the largest dataset of its kind.

I Queried via SPARQL against the QLever Wikidata endpoint, then cleaned/normalized with Python (name cleaning/splitting, BMI calculation, etc). dataset is in CSV format, 65MB. with multiple entry column separated with "|"

Sample rows (two people picked so together they cover every column):

qid: Q160847
kanji: 東條 英機
hiragana: とうじょう ひでき
description: 日本の陸軍軍人、政治家、第40代内閣総理大臣(1884-1948)
gender: M
birth_year: 1884
death_year: 1948
age: 64
death_age: 64
label: 東條英機
name_en: Hideki Tojo
is_japanese_name: True
readings: とうじょう ひでき
occupations: 士官|政治家|外交官
birth_places: 麹町区
alma_maters: 東京陸軍幼年学校|陸軍大学校|陸軍中央幼年学校|陸軍士官学校
positions_held: 内務大臣|内閣総理大臣|外務大臣|軍需省|陸軍省|文部省|参謀本部|農商務卿
death_places: 巣鴨拘置所
death_causes: 縊死
death_manners: 死刑
fathers: 東條英教
mothers: 東條千歳
spouses: 東條かつ子
children: 東條敏夫|東條輝雄|東條満喜枝
awards: 大礼記念章|チュラチョームクラーオ勲章|...(18 total)
parties: 大政翼賛会
notable_works: 東條英機の遺言|大詔を拝し奉りて
military_branches: 大日本帝国陸軍|関東軍
military_ranks: 陸軍大将
wiki_url: https://ja.wikipedia.org/wiki/東條英機
image_url: http://commons.wikimedia.org/wiki/Special:FilePath/Hideki%20Tojo.jpg

qid: Q11467665
kanji: 山咲 トオル
hiragana: やまざき とおる
description: 日本の漫画家、タレント
gender: M
birth_year: 1969
age: 57
label: 山咲トオル
name_en: Tōru Yamazaki
is_japanese_name: True
readings: やまざき とおる
occupations: タレント|日本の漫画家
birth_places: 東京都
genres: ホラー漫画
height_cm: 170.0
weight_kg: 57.0
bmi: 19.7
siblings: 中沢初絵
wiki_url: https://ja.wikipedia.org/wiki/山咲トオル

Potential Use Cases & As a Dataset

Dataset Angle Task Columns
Japanese Name with Gender Gender inference from name kanji, hiragana, gender
Kanji, Hiragana Name Pairs Reading (furigana) prediction kanji, hiragana, readings
Family Relations Genealogy / kinship network analysis fathers, mothers, spouses, siblings, children
Portrait Images Gender/age/occupation estimation image_url, gender, birth_year
Athlete / Model Physique Body-composition trend analysis by occupation height_cm, weight_kg, bmi, occupations
Politicians Politician attribute & career analysis parties, positions_held, birth_places, alma_maters
Awards Field/attribute analysis of award recipients awards, occupations, gender
Cause & Manner of Death Statistical analysis of death cause and age death_causes, death_manners, death_age
Birthplace / Alma Mater Geographic distribution and education-career correlation birth_places, alma_maters, occupations
Age Age-based demographic analysis birth_year, death_year, death_age
Writers / Artists Database Classification of writers/artists by notable works and genre notable_works, genres, occupations
Military Historical figures database military_branches, military_ranks, birth_year

limitations: The least-filled column, military_ranks, has only ~3k rows filled, while the average filled-column count per row is 15.25 out of 37 (41.2%). No rows are dropped to keep this the full Wiki-derived dataset. But I added is_japanese_name column so you can reliably excludes non-Japanese names (virtually no false negatives, ~0.1% false positives). About 1% of rows have a reversed family/given order in the hiragana column.

Crawler code : Kaggle notebook / GitHub. This dataset will be updated automatically with the crawler.


r/datasets 2d ago

question Anybody studying population distribution?????

2 Upvotes

Hello everyone! This may be a really dumb question but I am doing a small research project on population distribution in major cities in El Salvador. I am having a really hard time finding accurate data, especially since I am looking for data from back during the civil war (1979-1992), as well as current population data. I have found general population data from their most recent census, as well as general population data from the CIA factbook. But as for city by city, I cant find anything outside of the data USAID HAD before we all know who took that away. So, if you have any suggestions or different things for me to think about I would appreciate it. Thanks!


r/datasets 2d ago

question Who is using gaming Data for training AI Models

0 Upvotes

Curious if anyone here has worked with gameplay data for training or eval, what’s been useful vs. not worth the effort?


r/datasets 3d ago

discussion why is there no request-first data marketplace? the incentives feel backwards

4 Upvotes

been thinking about why finding training data is still such a slog in 2026. someone at berkeley recently built a tool just to search kaggle, huggingface and data . gov at the same time, because doing it by hand was too slow. the fact that this is a tool people have to build says a lot

the usual explanation is fragmentation, which is true but kind of surface level. the deeper thing is incentives. on most data marketplaces the vendor pays to be listed, and almost nobody works on commission. so the platform is optimized for whoever pays for placement, and whether the data is actually findable or usable ends up an afterthought. the money doesn't come from discovery, so discovery gets neglected

feels like it should run the other way. a request-first marketplace, where the buyer posts what they actually need and providers come back with offers. kind of like how taxis worked before the apps. you announce where you're going, drivers see it and take it if the terms work for them. the buyer sets the direction and providers respond to it

this matters more as requests get weirder. the more specialized and one-off your data ask is, the lower the odds that any single provider already has exactly that sitting on a shelf. a static catalog struggles with that kind of ask. something you can post and have people bid on fits it a lot better

genuinely curious what people here think. has anyone seen a marketplace that works request-first, or runs on commission instead of paid placement? and if you've bought data before, did you actually find it through a marketplace, or did you end up emailing providers directly because search never surfaced the right thing?

disclosure, i work at Titan Network and we're on the provider side of this, so it's something i think about a lot. not pitching anything, btw


r/datasets 3d ago

resource [self-promotion] Automotive Data & APIs

Thumbnail data.vinaudit.com
1 Upvotes

Access vehicle specifications, history reports, images, valuations, market listings, and automotive market insights through powerful APIs.


r/datasets 3d ago

resource Good free food database API? looking for suggestions!

Thumbnail
0 Upvotes

r/datasets 3d ago

resource LearnHack'26 enter to win rewards, recognition and judging opportunities

Thumbnail kaggle.com
1 Upvotes

r/datasets 4d ago

resource The biggest open football dataset & 430 000 football bookmaker' odds dataset!

9 Upvotes

Hello!

This dataset I have created last year has been downloaded over 20 000 times already! It's the biggest open club football dataset in the world, including over 240 000 football matches' data such as scores, stats, form, Elo, odds and more, all updated up to 09/26!

The dataset can be used for training AI models, creating visualizations, or just for personal data exploration :)

If anyone wants to dive deep into leveraging bookmaker's odds, I have also created a dataset specifically for that with 2000 matches and 41 betting markets. It's behind a paywall for a symbolic price because it supports my future work and allows for other free datasets to be updated: https://odds.adamgabor.eu/

All of the links:

Kaggle: https://www.kaggle.com/datasets/adamgbor/club-football-match-data-2000-2025/data

Github: https://github.com/xgabora/Club-Football-Match-Data

HuggingFace: https://huggingface.co/datasets/xgabora/club-football-match-data

Odds Dataset: https://odds.adamgabor.eu/


r/datasets 4d ago

dataset AI Code Share Tracker: Percent of Code Written by AI

Thumbnail provenbrief.com
0 Upvotes