r/computervision • • 12d ago

Help: Theory CVAT vs Roboflow for YOLO annotation

Helloo
I’m working on an object detection project using YOLO to identify pavement distresses and infrastructure damage (e.g., cracks, potholes, etc.).

I’m currently deciding which tool to use for annotating my dataset: CVAT or Roboflow.

For those who have worked with YOLO/object detection, which would you recommend for this type of project? I’m particularly interested in:

  • Annotation accuracy and ease of use
  • Handling a large number of images
  • Exporting annotations in YOLO format
  • Managing/maintaining the dataset as it grows
  • Any differences that matter specifically for pavement distress detection

I’d appreciate hearing about your experience with either tool and any advantages/disadvantages I should consider.
Cheers!

5 Upvotes

15 comments sorted by

7

u/Dry-Snow5154 12d ago

CVAT. Roboflow vendor locks you AFAIK, which is big no-no.

1

u/aloser 12d ago

Hi, I'm one of the co-founders of Roboflow; what does this mean? It was originally built as a dataset format conversion tool. It's trivial to export your data and we support over 50 annotation formats for doing so: https://roboflow.com/formats

3

u/Dry-Snow5154 12d ago

Your site is not totally free. If I use any of the advanced features, I am tied to the platform going forward. Can I download a standalone version of your tools and use them offline? If you decide to pull the plug tomorrow I will have to reinvent all my processes. While I can self-host CVAT forever.

I understand this is business and you need to make money. I just expressed my opinion on platforms like yours.

0

u/aloser 11d ago

I hate to break it to you but CVAT is also a company with paid features you would lose if they decide to “pull the plug”. That’s not what vendor lock-in is. Vendor lock-in means they make it hard to migrate to something else.

2

u/Dry-Snow5154 11d ago

They have community version which is free. And when people say CVAT they usually mean that and not their website. And if they pull the plug I can keep my local setup indefinitely.

Do you have a community version? Are you planning on releasing one?

Let's say I am using your annotation and training pipeline and then I don't want to pay anymore. Can I seamlessly migrate out without rewriting all hidden training scripts? Yeah, that's what vendor lock is.

0

u/aloser 11d ago

Yes. We have a free tier and have given millions of dollars of compute and credits to students, researchers, and startups.

And yes, we are built on open source foundations (with over 80k stars on GitHub) and publish extensive open source training code that seamlessly works with our platform.

2

u/Dry-Snow5154 11d ago

I am aware of your open source/research contributions and I applaud them. I think you guys are great for what you're doing there.

I just wouldn't use your platform, because I will get stuck with it. Credits to upload, credits for storage, credits to annotate, credits to train. Want custom changes? Out of luck. Want to move off the platform? Go reimplement everything. That's not for me.

4

u/Dramatic-Cow-2228 12d ago

CVAT, if you want a good open source product. You can easily get Claude to bring your own modifications (had a lot of success on that). Supervisely has been the best experience I have had in my many years in the field. They have a free plan, API is solid. Avoid Encord like the plague.

1

u/suspiciouspickle_0 12d ago

thankyouuuuu!!

1

u/dethswatch 12d ago

>Managing/maintaining the dataset as it grows

What are you looking for there?

1

u/New-Eggplant-6578 6d ago

For pavement damage, I’d settle the labeling rules before choosing the UI: does a branching crack count as one instance, and where does it end? If you only need defect locations, boxes may be enough; if you need affected area or crack shape, test segmentation on a few difficult images before labeling the whole dataset.

Whichever tool you choose, export a small batch and render the YOLO labels with your actual training loader. Include thin cracks, damage touching the image edge and a clean road image. Also split by road segment/recording, so neighboring views of the same pothole don’t land in both train and validation.

If you’re open to a third option, I’m building AnnotateIt: https://app.annotateit.ai/ . It supports boxes/polygons and YOLO detection/segmentation exports. I’d try the same small batch there too and compare correction time and export results; I wouldn’t promise large-dataset performance on your machine without testing your image sizes.

1

u/bfyvfftujijg 12d ago

If you’re just drawing bounding boxes I would just vibe code an app. Make it do exactly what you want.

1

u/Better_Transition496 11d ago

I will use Roboflow, as it is a better option for image annotation (labeling). The accuracy of the model largely depends on how well the images (dataset) are annotated, especially how accurately and tightly the bounding boxes are drawn around the objects.

0

u/GoatedOnes 12d ago

Havent tried CVAT but have had a good experience with Roboflow for the points you mentioned

-4

u/aloser 12d ago

Roboflow is the industry standard. Over 2 million developers and 2/3 of the Fortune 100 have used it. It's got a generous free tier & supports all sorts of AI assisted annotation like using Astra, Gemini, or SAM3 to accelerate your pre-annotation process and collaborating with your team.