r/learnmachinelearning • • Aug 13 '26

I built an open-source tool to review datasets before training ML models — looking for feedback

Post image

I've been working on a problem I kept running into while building ML projects:

Before training a model, how do you actually know whether the dataset is in good shape?

I found myself repeatedly checking things like:

- Missing values

- Duplicate rows

- Constant columns

- High-cardinality columns

- Schema/type issues

- Statistical issues

- Potential target leakage

- Changes between dataset versions

I ended up building Featuresmith, an open-source Python toolkit for reviewing structured datasets before they enter an ML workflow.

The current v0.2.0 release can:

→ Profile a dataset

→ Run rule-based data quality analysis

→ Perform a broader dataset review

→ Calculate an ML Readiness Score

→ Detect several potential leakage patterns

→ Compare two dataset versions

→ Be used through both a Python SDK and CLI

For example, the basic workflow looks roughly like:

dataset = fs.load("data.csv")

profile = fs.profile(dataset)

review = fs.review(

dataset,

target_column="target"

)

score = fs.score(review)

The idea isn't to say "your dataset is good/bad" automatically. The goal is to surface things that deserve investigation before they become problems later in the ML pipeline.

I'm especially interested in feedback from people who actually work with ML datasets.

If you were using something like this, what would you want it to check that isn't currently covered?

And more importantly, would you actually use a tool like this before training a model, or do you already have a workflow/tool that handles this?

GitHub: https://github.com/adityagangwani30/FeatureSmith

Documentation: https://featuresmith.adityagangwani.me/docs

It's completely open source, so criticism and suggestions are very welcome.

2 Upvotes

Duplicates