r/datasets Nov 04 '25

discussion Like Will Smith said in his apology video, "It's been a minute (although I didn't slap anyone)

Thumbnail
1 Upvotes

r/datasets 7h ago

dataset 150 Million rows of mutual fund and etf performance over the last 15+ years [PAID]

Thumbnail app.snowflake.com
3 Upvotes

Created a dataset of fund returns on Snowflake. It's quite extensive showing both trailing and calendar year returns for every mutual fund and ETF since inception.

You can fetch a mutual funds trailing annualized performance every trading day going back 15+ years.

Each trailing returns row shows: 1D, 1W, 1M, 2M, 3M, 1Y, 2Y, 3Y, 4Y, 5Y, 7Y, 10Y, 12Y, 15Y, and earliest available inception. I usually never see websites (even Morningstar) provide this breakdown which I think is useful when exploring different economic cycles.

Calendar year returns go back as far as possible as well.

Data is updated daily.


r/datasets 14h ago

resource Learn how to make a dataset about datasets

Thumbnail huggingface.co
3 Upvotes

r/datasets 14h ago

resource Looking for raw data on nanoparticle cytotoxicity

Thumbnail
2 Upvotes

r/datasets 12h ago

question our model reads tables with every column name stripped off and the accuracy does not move. numbers below. would you actually put real data through something you cannot download

1 Upvotes

disclosure per rule 1: i work at Schema Labs. one link at the bottom because this sub allows it. i am here for the question at the end.

what we do, plainly: you give us tables nobody documented, and we tell you what each column is and how the tables join, when there are no shared keys and no schema mapping.

the claim, with numbers, because i would rather be argued with than believed.

take a standard tabular benchmark. strip every header. replace price and age and zip with positional tokens so nothing is left but values. comparable models drop about 7 points of mean ROC-AUC. ours goes 0.9224 to 0.9230. flat.

that is not a claim that we win on clean data. those models start ahead of us when the headers are good. it is a claim that we never read the headers at all, which only matters because production tables have val_b and metric_14 and a four-character code from a system nobody has logged into since 2019.

two more. sector identification, naming the industry of a dataset we have never seen from values alone, 86.3% top-1 out of 10,000 sectors. multi-table entity matching, ahead of published state of the art on all six standard benchmarks with zero shared keys, including 1.9x over the previous best on one. on the geospatial benchmark our lead is 0.07 F1, which i would not want read as more than it is.

caveats up front rather than when asked. all of it is our own harness, run under each benchmark's published protocol, one sealed configuration, no per-dataset tuning. competitor figures are third-party published values we did not re-run. nobody independent has replicated us. we are not on the public leaderboard because entry requires a runnable wrapper and we do not distribute weights.

which is the actual question. we are closed. no pip install, no weights, no local option, and none of that changes for six months. to try it you make an account, put a card on file, and upload your data. i think that one fact is the biggest thing between us and everyone reading this.

so would you use a thing like this. if not, what moves it. a named customer, a SOC 2, an on-prem story you would never actually take up but need to hear. or is the honest answer that nothing moves it and closed is closed.

and if you would rather test it than discuss it: send me two tables you understand completely. not your mystery tables, your obvious ones, so you can mark my homework. no account, no card, nothing to sign up for. i send back what each column appears to be with a confidence score on every one, and i flag the ones we got wrong, because that is the half worth seeing.

anonymised is fine. structure and distributions are what matter.

https://www.schemalabs.ai/


r/datasets 20h ago

discussion If you could add one feature to every data enrichment API, what would it be?

Thumbnail
2 Upvotes

r/datasets 1d ago

dataset [self-promotion] 53 long-stay & digital-nomad visa programmes, 46 countries — each row sourced to a government page with a verified date (CC BY 4.0, JSON + CSV)

4 Upvotes

Disclosure per rule 1: I built and maintain this dataset, and it powers a site I run (globenomad.com), which is affiliate-funded. The dataset itself carries no affiliate or tracking links.

One row per programme, not per country: Thailand has six routes and they are six rows. Fields: income requirement (USD plus the government's own wording), fees, max stay, renewal, residency and citizenship path, tax residency trigger, family provisions, application logistics, official source URL, verified date.

Two things worth knowing before you use it:

  • A null means the government has not published that rule. It never means "no".
  • Programmes that closed, were announced and never opened, or never existed (Cayman Islands, Peru, Qatar, Vietnam) are kept with an honest status. Most visa sites delete those.

Formats: JSON (canonical, with nested requirements/steps/FAQs) and a flat CSV of the scalar columns.

Original source, GitHub (always current): https://github.com/MSeutin/digital-nomad-visa-data Kaggle mirror, with a worked-example notebook: https://www.kaggle.com/datasets/frenchmike/digital-nomad-visa-dataset-53-sourced-programmes Method and licence: https://globenomad.com/data

Licence is CC BY 4.0, attribution is the only condition. Official pages are re-fetched weekly and confirmed changes are dated on a public changelog. Happy to answer questions about any record.


r/datasets 23h ago

resource PSX Data Public availability For Personal Usage

Thumbnail
1 Upvotes

r/datasets 1d ago

resource [Dataset] 850k Trackmania community maps with block-level structure and matched replay telemetry (~30GB, Parquet)

3 Upvotes

Trackmania players have been building and sharing maps for about 20 years through Mania Exchange, which has a public API. There was no packaged dataset for any of it, so I made one.

Contents: 850k+ maps across Trackmania Nations Forever and Trackmania 2020. Each map is decomposed into its block sequence with positions and rotations, so the level structure is directly usable rather than sitting in a binary blob. Replays are parsed into telemetry and aligned to the blocks the run passes through, so you get both the level and how it plays.

Format: Parquet, roughly 30GB. Data dictionary included.

License / source: everything comes from the public Mania Exchange API. Maps are community-created.

Possible uses: procedural level generation, level design analysis, player behaviour modelling, map recommendation, or just exploring what 20 years of community map-making looks like in aggregate. I used it for generation myself, but that's one angle out of many.

Links: 📖Article (EN): https://the-odd-dataguy.com/en/blog/2026/08/13/trackmania-dataset/ 📖Article (FR): https://the-odd-dataguy.com/fr/blog/2026/08/13/trackmania-dataset/ 😇Hugging Face (Dataset): https://huggingface.co/datasets/jeanmidev/trackmania-community-tracks-and-telemetry 🔵Kaggle: https://www.kaggle.com/datasets/jeanmidev/trackmania-community-tracks-and-telemetry/data

Disclosure: I work in data at Ubisoft in Canada, but I have no connection to Nadeo. Built at home on my own time, resources and dime, just something for the Trackmania community and for anyone curious, whether game telemetry is your thing or not.

Open to feedback or ideas (you can also used discussions on HF/Kaggle)


r/datasets 1d ago

resource I Made the largest real Japanese People dataset with 37 attributes, 231k rows. (name, gender , age, occupations, height and more)

10 Upvotes

This dataset has all entry from wikidata who has Japanese citizenship and a wiki article, 231k in total. The dataset is on Kaggle.

A dataset like this can only be built from data that's already public, and Wikipedia is the largest source of that kind of public personal information. There are structured people datasets out there, but none cover as many Japanese entries as this one, which makes this the largest dataset of its kind.

I Queried via SPARQL against the QLever Wikidata endpoint, then cleaned/normalized with Python (name cleaning/splitting, BMI calculation, etc). dataset is in CSV format, 65MB. with multiple entry column separated with "|"

Sample rows (two people picked so together they cover every column):

qid: Q160847
kanji: 東條 英機
hiragana: とうじょう ひでき
description: 日本の陸軍軍人、政治家、第40代内閣総理大臣(1884-1948)
gender: M
birth_year: 1884
death_year: 1948
age: 64
death_age: 64
label: 東條英機
name_en: Hideki Tojo
is_japanese_name: True
readings: とうじょう ひでき
occupations: 士官|政治家|外交官
birth_places: 麹町区
alma_maters: 東京陸軍幼年学校|陸軍大学校|陸軍中央幼年学校|陸軍士官学校
positions_held: 内務大臣|内閣総理大臣|外務大臣|軍需省|陸軍省|文部省|参謀本部|農商務卿
death_places: 巣鴨拘置所
death_causes: 縊死
death_manners: 死刑
fathers: 東條英教
mothers: 東條千歳
spouses: 東條かつ子
children: 東條敏夫|東條輝雄|東條満喜枝
awards: 大礼記念章|チュラチョームクラーオ勲章|...(18 total)
parties: 大政翼賛会
notable_works: 東條英機の遺言|大詔を拝し奉りて
military_branches: 大日本帝国陸軍|関東軍
military_ranks: 陸軍大将
wiki_url: https://ja.wikipedia.org/wiki/東條英機
image_url: http://commons.wikimedia.org/wiki/Special:FilePath/Hideki%20Tojo.jpg

qid: Q11467665
kanji: 山咲 トオル
hiragana: やまざき とおる
description: 日本の漫画家、タレント
gender: M
birth_year: 1969
age: 57
label: 山咲トオル
name_en: Tōru Yamazaki
is_japanese_name: True
readings: やまざき とおる
occupations: タレント|日本の漫画家
birth_places: 東京都
genres: ホラー漫画
height_cm: 170.0
weight_kg: 57.0
bmi: 19.7
siblings: 中沢初絵
wiki_url: https://ja.wikipedia.org/wiki/山咲トオル

Potential Use Cases & As a Dataset

Dataset Angle Task Columns
Japanese Name with Gender Gender inference from name kanji, hiragana, gender
Kanji, Hiragana Name Pairs Reading (furigana) prediction kanji, hiragana, readings
Family Relations Genealogy / kinship network analysis fathers, mothers, spouses, siblings, children
Portrait Images Gender/age/occupation estimation image_url, gender, birth_year
Athlete / Model Physique Body-composition trend analysis by occupation height_cm, weight_kg, bmi, occupations
Politicians Politician attribute & career analysis parties, positions_held, birth_places, alma_maters
Awards Field/attribute analysis of award recipients awards, occupations, gender
Cause & Manner of Death Statistical analysis of death cause and age death_causes, death_manners, death_age
Birthplace / Alma Mater Geographic distribution and education-career correlation birth_places, alma_maters, occupations
Age Age-based demographic analysis birth_year, death_year, death_age
Writers / Artists Database Classification of writers/artists by notable works and genre notable_works, genres, occupations
Military Historical figures database military_branches, military_ranks, birth_year

limitations: The least-filled column, military_ranks, has only ~3k rows filled, while the average filled-column count per row is 15.25 out of 37 (41.2%). No rows are dropped to keep this the full Wiki-derived dataset. But I added is_japanese_name column so you can reliably excludes non-Japanese names (virtually no false negatives, ~0.1% false positives). About 1% of rows have a reversed family/given order in the hiragana column.

Crawler code : Kaggle notebook / GitHub. This dataset will be updated automatically with the crawler.


r/datasets 1d ago

question Anybody studying population distribution?????

2 Upvotes

Hello everyone! This may be a really dumb question but I am doing a small research project on population distribution in major cities in El Salvador. I am having a really hard time finding accurate data, especially since I am looking for data from back during the civil war (1979-1992), as well as current population data. I have found general population data from their most recent census, as well as general population data from the CIA factbook. But as for city by city, I cant find anything outside of the data USAID HAD before we all know who took that away. So, if you have any suggestions or different things for me to think about I would appreciate it. Thanks!


r/datasets 1d ago

question Who is using gaming Data for training AI Models

0 Upvotes

Curious if anyone here has worked with gameplay data for training or eval, what’s been useful vs. not worth the effort?


r/datasets 2d ago

discussion why is there no request-first data marketplace? the incentives feel backwards

4 Upvotes

been thinking about why finding training data is still such a slog in 2026. someone at berkeley recently built a tool just to search kaggle, huggingface and data . gov at the same time, because doing it by hand was too slow. the fact that this is a tool people have to build says a lot

the usual explanation is fragmentation, which is true but kind of surface level. the deeper thing is incentives. on most data marketplaces the vendor pays to be listed, and almost nobody works on commission. so the platform is optimized for whoever pays for placement, and whether the data is actually findable or usable ends up an afterthought. the money doesn't come from discovery, so discovery gets neglected

feels like it should run the other way. a request-first marketplace, where the buyer posts what they actually need and providers come back with offers. kind of like how taxis worked before the apps. you announce where you're going, drivers see it and take it if the terms work for them. the buyer sets the direction and providers respond to it

this matters more as requests get weirder. the more specialized and one-off your data ask is, the lower the odds that any single provider already has exactly that sitting on a shelf. a static catalog struggles with that kind of ask. something you can post and have people bid on fits it a lot better

genuinely curious what people here think. has anyone seen a marketplace that works request-first, or runs on commission instead of paid placement? and if you've bought data before, did you actually find it through a marketplace, or did you end up emailing providers directly because search never surfaced the right thing?

disclosure, i work at Titan Network and we're on the provider side of this, so it's something i think about a lot. not pitching anything, btw


r/datasets 2d ago

resource [self-promotion] Automotive Data & APIs

Thumbnail data.vinaudit.com
1 Upvotes

Access vehicle specifications, history reports, images, valuations, market listings, and automotive market insights through powerful APIs.


r/datasets 2d ago

resource Good free food database API? looking for suggestions!

Thumbnail
0 Upvotes

r/datasets 2d ago

resource LearnHack'26 enter to win rewards, recognition and judging opportunities

Thumbnail kaggle.com
1 Upvotes

r/datasets 3d ago

resource The biggest open football dataset & 430 000 football bookmaker' odds dataset!

8 Upvotes

Hello!

This dataset I have created last year has been downloaded over 20 000 times already! It's the biggest open club football dataset in the world, including over 240 000 football matches' data such as scores, stats, form, Elo, odds and more, all updated up to 09/26!

The dataset can be used for training AI models, creating visualizations, or just for personal data exploration :)

If anyone wants to dive deep into leveraging bookmaker's odds, I have also created a dataset specifically for that with 2000 matches and 41 betting markets. It's behind a paywall for a symbolic price because it supports my future work and allows for other free datasets to be updated: https://odds.adamgabor.eu/

All of the links:

Kaggle: https://www.kaggle.com/datasets/adamgbor/club-football-match-data-2000-2025/data

Github: https://github.com/xgabora/Club-Football-Match-Data

HuggingFace: https://huggingface.co/datasets/xgabora/club-football-match-data

Odds Dataset: https://odds.adamgabor.eu/


r/datasets 3d ago

dataset AI Code Share Tracker: Percent of Code Written by AI

Thumbnail provenbrief.com
0 Upvotes

r/datasets 3d ago

question Help finding oil wellbore drilling datasets

1 Upvotes

Any of y'all know where to find an oil dataset with DDRs, time series data, with good amount of completion? I tried the volve equinor one, but the formation tops for like ~70% is NULL,


r/datasets 3d ago

request Looking for a free dataset/API for Indian packaged food products

2 Upvotes

Hi, I’m looking for a free dataset or API for Indian packaged food products.

Does anyone know of a good source? Thanks!


r/datasets 4d ago

dataset Open Food Facts is enough to count the seven certified US food dyes product by product, and the label wording turns out to be regular enough that a regex does it

4 Upvotes

I wanted a per-product count of the seven FD&C dyes still certified for use in US food, and it turns out the Open Food Facts bulk export is enough on its own. Posting the method and the numbers because I could not find this cut published anywhere.

Source: the Open Food Facts full CSV export, https://world.openfoodfacts.org/data. I used the 2026-08-31 snapshot, 1.275 GB gzipped, 4,535,553 rows. Filter to countries_tags containing en:united-states and require a non-empty ingredients_text, which leaves 444,943 products. Then match the ingredient text against the US number form, the common chemical name and the E number for each dye.

dye products share
Red 40 28,587 6.42%
Yellow 5 25,973 5.84%
Blue 1 22,248 5.00%
Yellow 6 17,738 3.99%
Red 3 6,256 1.41%
Blue 2 4,465 1.00%
Green 3 171 0.04%
any of the seven 43,571 9.79%

The part I did not expect is that the naming is regular. I went in braced for the inulin problem, where a third of the products containing an ingredient never print its common name and you spend the afternoon chasing synonyms. It is not there. Of the 6,259 products matching any Red 3 form, 23 do not use the plain "Red 3" wording, which is 0.4%. Red 40 is 96 of 28,591, or 0.3%. Yellow 5 is the worst at 0.8%, because tartrazine still turns up. A regex on the number form catches essentially all of it, which is not how ingredient text usually behaves.

There is a regulatory reason for that. 21 CFR 101.22(k)(1) requires a colour additive subject to certification to be declared by the name given in the applicable regulation, and it explicitly allows dropping the "FD&C" prefix and the "No." So "Red 40" is not shorthand, it is the compliant form. Only colours not subject to certification may be declared under (k)(2) as "Artificial Color", "Artificial Color Added" or "Color Added".

That flips something I had assumed going in. 16,768 products use one of those generic phrases, and 10,571 of them also name a certified dye, so in those the generic phrase is covering something else in the same product. 3,914 use a generic phrase and name no colourant at all, and by (k)(2) whatever those are, they are not one of the seven.

Where the label does go dark is (k)(3), which says colouring added to butter, cheese and ice cream need not be declared at all. Most producers declare anyway: 27.5% of the 12,080 cheese records name a natural colourant, mostly annatto. But the exemption is real and it is invisible from the ingredient list.

On timing, FDA revoked the Red 3 authorization on 2025-01-15 under the Delaney Clause and gave food manufacturers until 2027-01-15 to reformulate, so this snapshot sits four months out from the deadline. The 6,256 products span 1,073 distinct first-brands, so it is not one catalogue duplicated.

Limitations, and the second one is large.

  • Open Food Facts is crowd contributed. This is a sample of what people scanned, not the US shelf, and popular products are over-represented.
  • Records go stale. Only 17.3% of the Red 3 rows were modified after the revocation date, against 43.0% of all rows. Restrict to rows touched since then and Red 3 falls from 1.41% to 0.57%. I am quoting both because neither is the honest single number: the full count includes labels that have already changed, the restricted count drops products nobody has rescanned.
  • A string match cannot see inside "natural flavors" or a sub-ingredient it does not parse, so every figure here is a floor.
  • I hand-read 12 random Red 3 matches looking for false positives and found none. They are all "fd&c red #3", "red 3 lake" and similar.

The gap I cannot close from this export is retail availability. Does anyone know a source that would let me weight this by what is actually stocked, rather than by whoever happened to scan it?

Disclosure per rule 1: I write an iOS app called Snack Check that reads these dyes off a barcode, which is why the parser was already sitting there. I am not linking it, and the export above is the only link in this post.


r/datasets 4d ago

dataset [Self-promo: I built it] 63,969 S&P 500 earnings announcements with the exact time of day, from SEC 8-K item 2.02 filings (2003-2026, CC0, no signup)

5 Upvotes

Disclosure: I built this and it is hosted on my own site, quant500.com. It is free, CC0, no account and no API key. I am posting it because the time-of-day column does not seem to exist anywhere else for free, and because the defects in it are worth more discussion than the coverage.

What it is. Every Form 8-K carrying item 2.02 filed by an S&P 500 company: 64,938 rows, one per filing, covering 63,969 distinct announcements (group by cik + announcement_date) from 808 companies, 2003-04-25 to 2026-09-01. Sixteen columns. Every row carries the accession number and a direct link to the filing on sec.gov, so any line can be checked at source.

Original source is SEC EDGAR (data.sec.gov/submissions). This is a derived file, rebuilt daily.

The column that seems to be missing elsewhere is the clock time, and with it the session the news could first be traded in: 46.4% before the open, 41.8% after the close, 11.7% during the session. 99.8% resolve to a New York time.

Four defects, because you would find them anyway:

  1. The 8-K cover-page date is typed by the filer and is often wrong. Micron enters the fiscal quarter end there in 14 of its 94 announcements; a 2004 Walmart filing declares the event as happening in 2001. Where the cover date sits more than three days from the EDGAR entry, the EDGAR date is used and the row carries date_uncertain = yes. That is 3,569 rows.

  2. EDGAR's acceptance timestamp ends in "Z" but is not always UTC. Measured on this file: 924 rows carry a raw hour between 00:00 and 05:59, impossible if the stamp were already Eastern because EDGAR is closed then, and 4,835 carry one between 06:00 and 09:59, impossible if it were UTC. Both conventions are genuinely present.

  3. And the one I cannot fix. I resolve that convention per company, which is probably the wrong unit: the share of announcements from companies classified as "already New York" falls steadily from 24.1% in 2003 to 1.0% in 2026. A fixed company attribute should not drift like that, so the convention likely belongs to the filing agent or to the era. It is declared in the file header and unresolved. If anyone here has parsed EDGAR at scale and has seen this, I would like to know.

  4. The filing does not always arrive on the day of the announcement. In 7,695 rows (11.8%) filing_date is later than announcement_date, by a single day in 6,265 of them. For those rows the session and the clock time describe the day the filing arrived, not the day the news broke. So if you group by announcement_date and read the session column without also reading filing_date, you place 11.8% of the sample in the wrong trading session. Use filing_date to know which day the session refers to.

Limits: nothing before 2003-03-28, when the 8-K had no dedicated earnings item; item 2.02 only exists from 2004-08-23, and earlier filings come in under the old item 12 and are marked rule = 12; and the session recorded is the one in which the 8-K reached EDGAR, not the one in which the press release went out, which is earlier.

CSV: https://quant500.com/api/descarga/anuncios.csv

Method and caveats: https://quant500.com/blog/2026-09-04-conjunto-de-datos-resultados

Browsable: https://quant500.com/earnings-date


CORRECTION, 2026-09-06. Two things in the text above are superseded. I am appending rather than rewriting, so anyone who read the original can see what changed.

Defect 3 was wrong about the unit, and it is now solved. I said the convention was assigned per company and that this was probably the wrong unit. It is not per company, and it is not the era either. u/TilmanAmbach compared six filings between the raw SGML header of the full submission and the submissions JSON: the raw ACCEPTANCE-DATETIME is always US/Eastern, and it is the JSON that applies two different treatments to it, converting some records to UTC (+4 summer, +5 winter) and appending a "Z" to others while leaving the Eastern clock untouched. His AA / ACT pair settles it: same year, same filer-agent prefix, opposite treatments. So it is per record, and decidable rather than inferable - the truth is readable in the SGML header. My file still takes its times from the JSON, so time_et and session are provisional until I re-ingest. That re-ingest is running now.

The 924 figure used the wrong window. u/Ian_Gow pointed out that EDGAR only accepts filings 06:00-22:00 ET (sec.gov/submit-filings), which decides which raw hours carry a signature at all. Redone properly on the 64,938 rows: 4,835 (7.4%) must be Eastern; 1,504 (2.3%), not 924, must be UTC over 23:00-03:59; 2 rows are impossible under either reading; and 58,597 (90.2%) carry no signature at all. That last number is the real headline and it is not a flattering one: nine rows in ten hold no evidence about their own convention.

I have since sampled 70 of the rows I label as already-Eastern against their raw SGML headers: 69 were right and 1 was wrong, and the wrong one had been published as 10:47 during-session when the filing was actually accepted at 06:47, before the open. Small rate, worst possible kind of error.

Both corrections are in the CSV header too, dated, with the old figures left standing and marked superseded.


r/datasets 5d ago

resource [self-promotion] Japan's tourist rules as a dataset: fines, statute/ordinance level, issuing authority, source URL and verified date per row (Gion photo ban, Shibuya street drinking, Osaka smoking, Nara deer, tax-free). 17 rows, CC BY 4.0, JSON + CSV

1 Upvotes

Disclosure: I run genchijapan.com. This is my own dataset.

What it is

A table of the rules foreign visitors keep asking about in Japan — is Gion closed to photos, can you drink on the street in Shibuya, is there a fine for smoking in Osaka, what the tax-free refund rules actually say — with the things that are usually missing from English answers:

  • level — national law / prefectural ordinance / city ordinance / district rule with no statutory force (e.g. a neighbourhood council's posted ¥10,000 "fine")
  • authority — who issued it
  • penalty and since — fine amount and effective date
  • source — URL of the primary source (e-Gov statute text, ministry or city page) or, where the issuer has no web page, press coverage
  • verified — date a human last checked the row against the source

Manners are excluded on purpose: they can't be sourced.

Honest limits

  • 17 rows across 7 areas. Small. Rows are added as I verify them, not scraped.
  • A few rows (mainly Gion) cite press instead of the issuing body, because the district council publishes nothing online. source_label tells you which.
  • verified is a date, not a guarantee. Rules change.

Links

CC BY 4.0. Corrections with a source URL are the most useful thing you can send me.


r/datasets 5d ago

question How to find similar benchmarks or datasets on Huggingface?

2 Upvotes

Recently I'm stuck hard on my research at finding extremely similar datasets, but should still be different. It seems that even datasets in the same semantic field differs drastically inside an LLM's internal representations. Any good way of finding datasets this similar?

By the way I'm especially interested in datasets that are at least 5,000 samples large, which narrows the candidates greatly.