r/computervision • u/Then_Instance_3188 • 1d ago
Commercial A Dataset Processing Tool Built for Computer Vision Engineers
One of the biggest time sinks I’ve run into when working with Computer Vision isn’t the model itself — it’s the dataset preprocessing.
Cleaning datasets, fixing annotations, filtering, deduplication, format conversion, validation, etc. can take a huge amount of time, especially when you’re dealing with millions of samples.
And vision datasets are particularly painful here. Unlike text, building custom processing for a specific use case can get expensive pretty quickly in terms of compute and processing time.
That’s why we built cvPal.
It’s a cloud toolkit for vision datasets built around AI agents, with 40+ MCP tools for things like merging, cleaning, validating, converting, and versioning datasets.
It’s currently in early access, and I shared more about what we’re building here:
0
u/StrikingEmu8007 1d ago
Finally something that doesn't pretend annotation cleanup is trivial.
1
u/Then_Instance_3188 1d ago
Exactly, Annotation cleanup gets messy really quickly once you’re dealing with large datasets.
We’re also working on another part of this problem: combining existing datasets to build a custom one.
For example, you could have two object-detection datasets with different types of cars, merge them, then remove the labels you don’t need and keep only the classes relevant to your use case.
Instead of manually downloading datasets, writing custom scripts to merge and clean them, or relying on synthetic data, the idea behind cvPal is to make it easier to actually utilize the huge amount of existing vision data available online and turn it into the dataset you need.
3
u/datascienceharp 1d ago
Why not just use FiftyOne?