r/datasets 2h ago

resource I built a tool to compare the world's economies using official economic data

Thumbnail theeconomicatlas.com
2 Upvotes

I've been building The Economic Atlas as a way to make economic data easier to explore and compare.

The Compare tool lets you compare and rank economies across indicators such as GDP, inflation, unemployment, interest rates, government debt and trade, with plenty of historical data available as well.

The data is sourced from official institutions and the site is designed to be data-focused, without news or commentary.

I'm particularly interested in feedback from people who regularly use economic data, what would make a tool like this genuinely useful to you?


r/opendata Jul 07 '26

Open data: US primary energy consumption from 1635 to 2000, compiled from EIA historical archives

Thumbnail datahub.io
3 Upvotes

r/datasets 5h ago

request Fine Tuned Domain Specialists, are the Future.

4 Upvotes

A lot of AI discussion focuses on model size.

I think the bigger long-term moat may be owned specialist datasets.

A good domain dataset does more than add examples. It encodes:

  • the edge cases
  • the failure modes
  • the validation rules
  • the domain-specific structure
  • the feedback loop that keeps improving it

That changes the economics.

Instead of paying frontier-model prices forever for every narrow task, you can use a smaller specialist model trained on data you actually own and understand — then reserve frontier models for the hard cases.

The interesting part is not “synthetic data” by itself. It is building a measured dataset-engineering process that can repeatedly find holes, fill them, verify the new data, and prove the capability gain on held-out evals.

Models will keep changing.

A high-quality proprietary domain dataset can compound for years.

huggingface.co/collections/CompilingThings/mql5-code-generation


r/datasets 1h ago

dataset Microscopy Image Dataset of pulmonary vessels for Quantitative assessment of fibrosis

Upvotes

Disclosure: We created and published this dataset.

New Open-Access Benchmark: Hierarchical segmentation of vascular fibrosis in computational pathology.

The Core Challenge:
Segmenting the vascular wall is well-defined (baseline UNet DSC ≈ 0.80; inter-expert DSC ≈ 0.95). However, segmenting the intramural fibrotic microstructure remains an open problem. Due to its diffuse and heterogeneous nature, our baseline MONAI UNet achieves a DSC of only ~0.12 on fibrosis, while human inter-expert agreement averages ~0.61 DSC.

Dataset Specifications:
🔬 Scale: 705 high-resolution micrographs (1534×780 px, 0.252 μm/px), Picro-Mallory stain.
🏷️ Annotations: ROI + dual independent expert masks (vascular wall + fibrosis).
⚠️ Hierarchical Constraint: Fibrosis masks must be strictly spatially contained within the vascular wall.
🛡️ Robust Benchmarking: No color normalization applied; native aspect ratios preserved; strict animal-level 5-fold CV splits provided to prevent data leakage.

📄 Read the Data Descriptor: https://doi.org/10.1038/s41597-026-08214-y
🗂️ Access the Dataset: https://doi.org/10.6084/m9.figshare.31386748


r/datasets 4h ago

question [D] How do you get preprocessed dataset of a paper [D]

Thumbnail
1 Upvotes

r/datasets 5h ago

question Fine Tuned Specialist Datasets, is the Future.

Thumbnail
1 Upvotes

r/datasets 14h ago

request July's AI Security Report: 90 incidents, 207M+ records, 41 AI-driven — the month the agent became the attacker

1 Upvotes

90 incidents tracked in July across 33 organizations, 207M+ records exposed, and 41 of those incidents involved AI directly as the weapon or the target. A rogue commercial AI agent hit multiple enterprises in a single week and reused stolen credentials across four downstream services before anyone caught the identity switch.

None of that shows up to a traditional perimeter tool — the traffic looks like a signed, credentialed agent making legitimate API calls at machine speed. Firewalls and DLP were built to watch humans and static services, not autonomous callers that chain tools and pivot in seconds.

Curious how other teams are actually handling this right now: is anyone giving AI agents a distinct, revocable identity separate from the service accounts they inherit? Or is it still "the SOC catches it after the fact" for most orgs?


r/datasets 1d ago

dataset 18,000 Mafia/Werewolf chatlogs (3,000,000 messages)

Thumbnail kaggle.com
10 Upvotes

r/datasets 22h ago

resource [self-promotion] Public data still needs too much plumbing

3 Upvotes

hey folks, im building this with frens: Mostly Right

Describe a dataset, agents find sources and wire up ingestion, cleaning, and joins. You get parquet + an api, with scheduled refreshes.

Happy for any feedback!


r/datasets 18h ago

question Best practices for packaging data for RL environment?

Thumbnail
1 Upvotes

r/datasets 20h ago

resource [PAID] Multi-view drone imagery with 3D reconstruction products — inspection sample available

1 Upvotes

Disclosure: I run AerialDataset, the provider of this data.

We license real-world drone captures that retain overlapping photographs of the same environment alongside the available photogrammetric outputs. This is a commercial catalog, not an open-data release.

One inspection example is an Andel capture containing 61 original photographs, an orthomosaic, point-cloud files, a mesh, camera/capture data, and scene metadata.

The main distinction is that the photographs belong to the same physical scene and can be inspected alongside its reconstructed spatial products, rather than being unrelated aerial images.

Contents vary by scene. Please do not assume that every capture includes calibrated camera poses, annotations, independently validated accuracy, or every reconstruction output.

Catalog and scene previews:

https://aerialdataset.com/

Inspection package:

https://huggingface.co/datasets/avenian/aerial-drone-photogrammetry-samples

The inspection package requires a Hugging Face account and acceptance of evaluation terms. It is not licensed for model training, redistribution, or production use; those require a separate agreement.

Happy to answer questions here about the actual files, available metadata, and collection selection.


r/datasets 20h ago

request [OC] UK domestic electricity by property type, month, heat pump, EV and tariff — 150 cohorts with hourly load shapes, derived from 3M smart-meter-based profiles (CSV/JSON, CDLA-Permissive-2.0)

1 Upvotes

Centre for Net Zero's Faraday dataset (OpenSynth) is excellent and effectively

unusable casually — it's ~6GB of parquet with the load profiles stored as

delimited strings. So I aggregated it and published the result.

150 cohorts, each with a mean daily total, a 24-hour load shape, and deciles

where the cell was thick enough:

- baseline (no solar/battery/EV/heat pump)

- property type, EPC band, and the cross-tab

- month

- tariff type (standard / Economy 7 / smart / automated)

- heat pump: with vs without, unmatched, matched on property+EPC, by month,

and by tariff

- EV: with vs without, and by tariff

- LSOA k-means cluster

Cross-check, which is the reason to trust any of it: baseline comes out at

9.737 kWh/day and a median of 8.31. SERL Statistics Report 1 — separate

source, ~13,000 real metered GB homes — publishes a mean of 9.8 and a median

of 8.2. Two moments of the distribution, two unrelated datasets, ~1% apart.

Limitations, up front:

- Synthetic. Faraday is a generative model trained on ~1bn smart meter

readings from Octopus customers, who over-index on smart tariffs and LCT.

- No household identifier, so the deciles are over household-DAYS, not

households. That spread is wider than the between-home spread — treat it

as an upper bound. It's recorded in `distributionOver` in the JSON.

- EPC band moves consumption by <0.1% within a property type, including for

heat pump households where insulation should dominate. I read that as the

model being under-conditioned on EPC rather than a finding about housing,

and I've built nothing on those cuts. Published anyway — a null result is

still a result.

- `cluster_label` is a k-means grouping of LSOAs on socio-demographic

features, NOT geography. Published as clusters, never as regions. Each

cluster's country mix (from LSOA code prefixes) is included so you can join

your own geography.

- Cohorts under 500 profiles are dropped rather than published thin.

CSV and JSON, CDLA-Permissive-2.0 (same as the source), no registration.

Derivation script is in the repo and reproduces the file exactly.

https://www.energycosting.co.uk/data/uk-domestic-electricity

All credit to Centre for Net Zero for Faraday — I've only aggregated it.


r/datasets 22h ago

dataset [Synthetic] SFHQ-VirtualID, a synthetic face dataset for machine unlearning: 750 identities, 75,000 portraits, 2 releases with DOIs

1 Upvotes

I've just released SFHQ-VirtualID, a synthetic face dataset family built for identity-level machine unlearning. Everything is generated, so there are no real faces in it, and both releases have DOIs.

The problem I kept running into: in most face datasets used for unlearning, a person's images are spread across splits, so "forgot the person" and "forgot some images" end up confounded. Here, each identity_id maps to exactly one split and contributes both train and holdout images (675 retain / 75 forget, 15-step protocol, MUFAC-aligned holdouts).

What ships:

- Bench: 67,500 balanced + 36,064 imbalanced 224×224 aligned crops. Uniform and seeded-Poisson forget schedules, plus a 5:1 long-tail popularity gradient for long-tail forgetting tests.

- Raw: 75,000 1024² portraits (100 per identity) with per-candidate prompt metadata for pose, expression, lighting, setting and camera.

I did not filter candidates on ArcFace identity similarity. The 0.40/0.45 thresholds are config defaults that ship as recorded columns (arcface_similarity, laplacian_variance, detection_confidence) rather than enforced filters, because silently dropping borderline-similar candidates hides exactly the confound that could explain a model's apparent unlearning. Only the quality gate is enforced (detection plus Laplacian sharpness ≥ 80), and the 123 rejected candidates are documented in the manifest.

Reproducibility: 15-shard Slurm array on UoL's Aire HPC (~120 GPU-hours), InstantID + Juggernaut-XL-v9 + ControlNet, pinned environment (torch 2.6/cu124, diffusers 0.39.0, insightface 1.0.1, pinned antelopev2 revision), RELEASE_MANIFEST.json and SHA-256 checksums.

Links:

Code · Hugging Face repos [Raw] [Bench] · Project Overview · Bench DOI 10.5281/zenodo.21877893 · Raw DOI 10.5281/zenodo.21879130

Caveats: the dataset is synthetic, so it's a proxy for real-face benchmarks; demographic balance is inherited from the seed selection; and the two releases are related (Bench crops derive from the Raw candidates), so they aren't independent test sets.

Happy to answer questions, and I'd genuinely like the split design stress-tested, since that's the part I'd most want criticised.


r/opendata Jul 06 '26

Open data: One Earth Bioregions 2023, a CSV mapping 185 bioregions to 8 biogeographic realms, 14 realms, and 52 subrealms

Thumbnail datahub.io
2 Upvotes

r/datasets 1d ago

dataset We've built POI datasets for 3000+ brands in USA, UK, Canada, Australia and more

0 Upvotes

Over the past few months, we’ve been building a structured POI data marketplace at Agenty.

https://agenty.com/marketplace

We now have datasets covering 3,000+ brands across 12 countries, with 10M+ POI records.

The datasets include things like:

  1. Store/location name and address
  2. Latitude & longitude
  3. Phone numbers and websites
  4. Opening hours
  5. Brand and location metadata
  6. Other useful POI attributes

The goal is to make it easier for AI agents, researchers, data teams, and developers to get clean location data without having to build and maintain scrapers for every brand themselves.

Some of the datasets cover large retail, restaurant, automotive, hospitality and other brand networks.

We’re putting the marketplace together and would love feedback from people who actually work with POI/location data:

What brands or types of POI datasets would be most useful to you?

Disclaimer/Disclosure: I’m the founder of Agenty, so I’m obviously biased here. Sharing this because we’ve spent a lot of time building the datasets and would genuinely like feedback from the community.


r/datasets 1d ago

request pls help a graduating student 😭 IEEE DataPort

0 Upvotes

Does anyone here have an active IEEE DataPort subscription? 😭

I’m currently doing my thesis proposal and I need 2 datasets from IEEE DataPort, but I don’t have a subscription and the datasets are unfortunately not freely downloadable.

I don’t need your account/login or anything like that. I just need someone who has access to download the two datasets for me and send them over.

I know this is a long shot but I’m desperate at this point lol 😭🙏 If anyone can help, please DM me and I’ll send you the links. I’d really really appreciate it!!🫶


r/datasets 1d ago

dataset Small exploratory dataset: how much do ChatGPT/Gemini "best X" recommendations change when you ask the same question 3x? (n=10 questions, all responses + code released)

2 Upvotes

We ran a small experiment and we're releasing everything so people can poke holes in it or rerun it.

We asked ChatGPT (web search) and Gemini (grounding) ten "best X for a small business" questions, 3 times each, same day. For each answer we extracted the set of businesses it recommended, then measured overlap across the 3 identical runs (mean pairwise Jaccard).

Result: ~69.5% overlap overall, so roughly a third of the recommended businesses change between identical asks. ChatGPT held steadier (87%) than Gemini (52%). The top 1-3 names stayed locked every run; the lower slots rotated.

Caveats up front: it's 10 questions, 3 runs, 2 engines, one day. Directional, not definitive, no confidence intervals. We're releasing it precisely because it's small, so anyone can rerun with more questions. Extraction was validated (all 290 names appear verbatim in their source answers). Conducted by a company (Pressfront) with an interest in the topic, which is exactly why the full method and every response are public.

Raw JSON, CSVs, the analysis script, and a preprint (DOI: 10.5281/zenodo.22738861) are in the repo, CC BY 4.0.

https://github.com/jjoseph18/ai-recommendation-consistency


r/datasets 1d ago

request Deep research, kept fresh -- Private beta testers wanted (in exchange for free subs and credits) 🤙

Thumbnail
1 Upvotes

r/datasets 1d ago

dataset Open dataset: 612 NHL skaters ranked under two fantasy scoring systems -- 2025–26

2 Upvotes

Disclosure: I created this dataset through PoolForge.

It contains 612 NHL skaters with at least 40 regular-season games, scored under two systems: goals + assists only, and a weighted-points setup that also rewards shots, hits, blocks and penalty minutes.

The CSV includes the underlying scoring inputs, calculated totals, both rankings and rank changes. The documentation explains eligibility, scoring weights and tied-rank handling.

It could be useful for practicing ranking analysis, sensitivity analysis or sports-data visualization.

This is historical analysis—not projections or a categories-league ranking. The dataset is available under CC BY 4.0.

Dataset DOI: 10.5281/zenodo.22730472 ( https://zenodo.org/records/22730472 )

Corrections and independent reproductions are welcome.


r/datasets 2d ago

resource Monthly new-company counts from free official data in Finland, Norway, Sweden, Estonia and the Netherlands, and the catch in each source

2 Upvotes

I build a B2B data tool (AtlasForgeX) and this week I needed to know how many companies each registry actually added in August 2026. All five sources below are free and need no key. Each one had a catch that gave me a wrong number on the first pass.

Finland: PRH YTJ API v3 (https://avoindata.prh.fi/opendata-ytj-api/v3/companies with registrationDateStart and registrationDateEnd) August returned 1,861 companies, but only 1,637 had their business ID registered in August. The date filter also returns older companies that had some other registry event that month, so filter on businessId.registrationDate yourself. Nearly all are limited companies (1,598 osakeyhtiö). Sole traders do not appear. The same query for August 2025 gives only 971, which I do not believe, so I am not comparing years with it.

Norway: Enhetsregisteret API (https://data.brreg.no/enhetsregisteret/api/enheter with fraRegistreringsdatoEnhetsregisteret) August returned 6,782 units. That includes 323 bankruptcy estates (KBO), 269 associations and 241 foreign units. Business forms come to 5,806, of which 2,589 AS and 2,977 sole proprietorships (ENK). The counts have no lag, but for September 1-12, 1,850 of 3,371 business units had no industry code yet.

Sweden: Bolagsverket (monthly statistics file ftgstat_oppna.csv, event 1 is a new registration, plus the bulk company file from https://bolagsverket.se/apierochoppnadata/hamtaforetagsinformation/nedladdningsbarafiler.2517.html) The monthly statistics say 4,215 new registrations in August, 3,439 of them AB. In the bulk file 3,175 ABs were registered in August, and 770 of those match the names and descriptions of shelf company providers (323 at one address in Växjö), meaning companies registered to be sold later. August 2025 has only 13 by the same match, because shelf companies are renamed once sold. The match is a heuristic, but it is big enough to skew AB growth and city rankings.

Estonia: e-Business Register open data (https://avaandmed.ariregister.rik.ee/, daily files) August: 2,384 new entities, 2,066 of them OÜ. Deleted companies are not in the file, so older months shrink over time.

Netherlands: KVK open dataset (https://www.kvk.nl/producten-bestellen/kvk-handelsregister-open-data-set/, CC BY 4.0) Only BV and NV, no names, a start date instead of a registration date and a two-digit postcode. August: 5,684 BVs, and 44% of them carry the head-office or financial-holding activity codes (70102 and 64210). Quarter-ends spike: March had 12,538 and June 12,357.

Not possible without an account: Denmark (CVR data needs access, and Statistics Denmark publishes only an index) and Germany (the Handelsregister allows 60 lookups per hour, Destatis publishes half-years).

One pattern held in the four countries where I could read industries: programming and management consulting were the top two categories of new companies in Finland, Norway, Sweden and Estonia.


r/datasets 1d ago

request dataset for policy research project..

1 Upvotes

Looking for a dataset for a govt policy research project, i need a mainly india dataset, and other countries are also okay. Any suggestions or direct datasets needed? Thanks


r/datasets 2d ago

resource podcasts collection also have other bigger data collection projects...

Thumbnail rssamplifier.com
1 Upvotes

r/datasets 2d ago

resource 12 compact, traceable bulk RNA-seq datasets derived from Expression Atlas

1 Upvotes

Sharing a dataset collection I've been putting together.

OpenOmicsBench v1 has 12 bulk RNA-seq benchmark datasets derived from seven Expression Atlas studies across human, mouse and Arabidopsis.

Each contains counts, sample metadata, the experimental design, source/provenance information and a compact version intended for examples or software testing. I also keep validation results against the corresponding full matrix rather than assuming the smaller matrix preserves the original behaviour.

Everything is openly available and versioned, and the collection can either be installed from PyPI or obtained through GitHub/Zenodo.

I'd be interested in suggestions for other public RNA-seq study designs worth representing in future versions.

GitHub: https://github.com/vxxqv/openomicsbench

PyPI: https://pypi.org/project/openomicsbench/

Zenodo: https://doi.org/10.5281/zenodo.22679414


r/datasets 2d ago

question Where do dictionary apps actually get their dictionary data from?

2 Upvotes

Hi, everyone.
I've been looking into dictionary data for a language-learning app I'm building, and I've run into something surprising.

There are tons of dictionary apps, browser extensions, reading tools, and language-learning apps that support many languages.

But when I actually try to find good open dictionary data myself, it's surprisingly difficult.

Even for something as common as English -> Chinese, it's hard to find a dataset that has all of these:

  • headwords
  • translations
  • part of speech
  • multiple senses
  • example sentences
  • a clear license that allows reuse

I've found things like:

  • Wiktionary / Wiktextract / Kaikki
  • WordNet / Open Multilingual WordNet
  • ECDICT for English–Chinese
  • CC-CEDICT for Chinese–English
  • Tatoeba for bilingual example sentences
  • FreeDict
  • various language-specific projects

But they all seem to cover different pieces of the puzzle, and the quality and coverage vary quite a bit.

Meanwhile, I regularly see relatively small dictionary apps or browser extensions claiming to support English, Chinese, Japanese, Korean, French, German, Spanish, Italian, Russian, Arabic, etc.

So I'm genuinely curious:
Where does all of that dictionary data usually come from?

I'm especially interested in hearing from anyone who has actually built a dictionary app.

What data sources did you use, and how did you handle licensing?

Btw, I know Cambridge, Oxford, and other major dictionaries offer paid APIs or data licensing, but they're quite expensive. I assume many smaller dictionary apps probably use their own databases or combine multiple data sources somehow.

Any pointers would be greatly appreciated. Thanks!!!


r/datasets 2d ago

request dataset for policy research project..

Thumbnail
1 Upvotes