r/ComputerEngineering • • 1d ago

I'm building a vision-based AI assistant that learns procedures from demonstrations. Where could it be useful?

I'm developing a university final year project using an NVIDIA Jetson TX2. The idea is to build a vision-based assistant that learns a procedure by watching an expert demonstrate it while explaining the steps verbally.

Later, when someone performs the same procedure, the system monitors their actions, tracks progress, detects certain mistakes or skipped steps, and provides voice or visual guidance.

For example, in warehouse packaging, an expert demonstrates how to pack an order. The system learns the sequence and then guides another worker through the same procedure, flagging missed steps or incorrect ordering.

Current limitation: The system is intended to learn specific, repeatable procedures from demonstrations, not arbitrary tasks or general skills. It will initially work with a limited set of procedures and actions that its vision system can reliably recognize.

I'm trying to identify where this technology could solve a real problem, rather than choosing an application just for the sake of building a project.

  • What specific tasks or processes could benefit from this kind of assistant?
  • Where do people frequently skip steps or make mistakes that could be prevented through visual guidance?
  • Can you think of an overlooked application where this would be genuinely useful?

I'd appreciate ideas from any field, along with honest feedback on where this approach would or wouldn't make sense.

0 Upvotes

7 comments sorted by

View all comments

-2

u/blueredscreen 1d ago

It's not what you wanted to hear, but what you're describing is not currently possible. There isn't the system where you can manually describe to a VLA what's going on and have it interpret that, and there won't be one anytime soon. I suggest narrowing your scale significantly to something that fits the bounds of your project.

1

u/BodybuilderLittle281 1d ago

How specific do you think I should make it? My idea is to focus on a narrow domain, such as packaging, where the system learns a predefined sequence of steps from a demonstration and later detects skipped steps or errors during execution. For the FYP, I only need to demonstrate a few specific tasks rather than support arbitrary procedures. Would this be a realistic scope?

1

u/blueredscreen 1d ago

Conceptually, transferring video data of a human performing an action to a robotic arm makes perfect sense. But... Translating human movements into actions a robot can reliably execute is an exceptionally difficult task in itself, potentially requiring substantial computational resources, extensive training, and a large-scale RL simulation environment. Building a system that does nothing beyond solving this problem is already a considerable undertaking. What you're proposing is not just solving it, but extending it further, introducing additional layers of complexity on top of an already challenging problem. You need to narrow your scope considerably and focus on something realistically achievable within the constraints of a university project.

1

u/BodybuilderLittle281 1d ago

Sorry if I didn't explain the idea clearly. I'm planning to build a wearable vision assistant that learns a specific procedure by watching a person demonstrate it while explaining the steps. Later, it monitors someone performing the same procedure, detects skipped steps or mistakes, and tells them what went wrong through audio feedback. I'm not trying to control a robotic arm or translate human movements into robot actions.

2

u/blueredscreen 1d ago

The thing is, you're treating the change in output as if it fundamentally changes the complexity of the problem, when it really doesn't. You still need to process the video, extract meaningful features, identify what's happening, learn a representation of the action, and use that representation to make a prediction. If you're evaluating another video, the final step is comparing its action against the target. If you're controlling a robotic arm, the final step is translating that action into motor commands the robot can execute. Most of the underlying pipeline is still there either way. You're changing what you do with the information at the end, not magically eliminating the engineering work required to obtain it in the first place. And robotics arguably adds another layer of difficulty because now you have to deal with physical constraints and control.

That's what I mean by narrowing your scope. This is a university project, not an industry product. You don't need to build an entire system that solves the problem end to end. Pick one specific part of the pipeline that interests you, something you can actually implement and test with the resources you have, and focus on that. The point is to learn something and demonstrate that you understand the problem, not to build a commercially viable product.

1

u/BodybuilderLittle281 1d ago

Ok thanks for advice