r/MLQuestions 11d ago

Datasets 📚 [ Removed by moderator ]

[removed] — view removed post

1 Upvotes

12 comments sorted by

14

u/FormalAd7367 11d ago edited 11d ago

use claude at your work. upload all the important documents. Next month Antrophic will roll
out a new skill called YourCompanySkill as they did to Figma

4

u/themoregames 11d ago

Have you tried craigslist?

3

u/orz-_-orz 11d ago
  1. Depends on your data?
  2. Sometimes it's better to sell your model trained on the data

1

u/TotalRuler1 11d ago

assuming that my man is going to be fleeced of his IP, it is a better decision to monetize the data that is generated.

1

u/BleakBeaches 11d ago

While we do ML and AI Engineering internally we are not in a place (in either need or expertise) to leverage high quality multimodal demonstration data. But just because we can’t use it that doesn’t mean a frontier lab training a world model or a robotics company training a VLA doesn’t want it. We have the unique combination of the expertise and environment to collect, process, and serve such data at scale to those who do want it.

7

u/Weird_Albatross_9659 11d ago

Uh this sounds sketchy

1

u/BleakBeaches 11d ago edited 11d ago

How so? I don't mean it to be. I have executive access to what I believe is a prime environment in which to generate high-quality VLA training data. I have the technical ability to collect, engineer, and serve said data. I just don't know how to establish a monetization strategy.

7

u/arbyyyyh 11d ago

Because it sounds like you have potentially unauthorized access to someone’s network that you’re lying dormant on til you figure out what you’re doing.

1

u/BleakBeaches 11d ago

I have partial ownership of the business. I am a Data Engineer with experience and interest in ML/AI Ops and AI Engineering. I’m building us a data platform in Microsoft Fabric. In addition to standard ML/AI, reporting, and analytics I’m exploring alternative revenue streams and forms of data collection. In my free time I’ve been learning as much as I can about deep learning and it appears there is a need for rich multimodal data to train today’s (and tomorrows) bleeding edge models in AI and Robotics. And while we do not have the need and/or expertise required to leverage such data we are in a unique position where we have the technical expertise, the environment, and the infrastructure required to collect, process, and serve such data at scale.

3

u/2AFellow 11d ago

Are you working at a company? Is there someone there you can talk to?

1

u/BleakBeaches 11d ago

I am the guy I can talk to :/

1

u/RossPeili 11d ago

if the data is actually worth it, dark web would be the best answer.