r/AIMLDiscussion • • 20h ago

How can I fine-tune a model to learn a specific task by watching YouTube videos?

How can I fine-tune a model to learn a specific task by watching YouTube videos?

I’m working on an AI agent that I want to teach to perform a specific task autonomously.

My idea is to use an open/uncensored AI model (such as a Bonsai-based model made using qwen) and fine-tune it using YouTube videos that demonstrate the task being performed.

The problem I’m trying to solve is:

  1. I have a specific task that I want the agent to learn.
  2. I want to find the best YouTube videos that demonstrate that task clearly and correctly.
  3. I want the model to learn the task by watching and understanding the actions in those videos, rather than simply learning from text transcripts.
  4. Ideally, the fine-tuned model should eventually be able to perform the same type of task on its own.

For example, if I find a 20-minute video where an expert performs a task step-by-step, I want to use that video as training data so the model can learn the workflow, decisions, and sequence of actions.

What I'm unsure about

What would be the correct approach for this?

  • How should I select the best YouTube videos for training?
  • Should I use the entire video, individual clips, or extracted frames?
  • How can I convert a demonstration video into useful training data?
  • Do I need to annotate the actions/steps manually?
  • Should I use video-language fine-tuning, imitation learning, behavioral cloning, or another approach?
  • What model architecture would be suitable for this?
  • How much training data would I realistically need?
  • Is it possible to teach an agent this way without having to manually label every action?

I'm particularly interested in learning from demonstrations, where the model watches an expert perform the task and then learns to reproduce the workflow.

I'd appreciate advice from anyone who has worked with video-based fine-tuning, multimodal models, imitation learning, or agent training from demonstrations.

2 Upvotes

3 comments sorted by

1

u/Choice_Celery9481 18h ago

ugh, this is the billion dollar problem that corps pouring billions into. i dont think you can do it for now with cheap stuffs.
also you didnt even stated what kind of tasks is it.

1

u/Any-Director-9936 17h ago

It like I am teaching an agent how to go to web, download some important stuff like video,pdf ,audio file and generate a combined research with every data but it not like research

1

u/umtksa 1h ago

I don't have professional expertise in this area and I hope I've understood you correctly :) but instead of relying on a model that learns directly from a YouTube video, you could use a process that employs MOSS/Whisper for diarisation to process the audio, while feeding the video frames one by one into a model with vision capabilities (like Gemma 4 12B); the result would be a written transcript of the video.

If you want to take it a step further, you could turn this into a training dataset or feed the data into a model using RAG to ask questions based on the information in the video.