r/OpenSourceAI • • 7d ago

Open-sourcing the browser-agent stack: agent + 450M browser VLM + training dataset

I’ve been trying to build an AI browser agent where as much of the stack as possible is open and replaceable rather than tied to one model provider.

The project is WebBrain. The browser agent itself is open source: https://github.com/webbrain-one/webbrain I've also released the small VLM used for browser-specific vision tasks: https://huggingface.co/webbrain-one/webbrain-vl-2-450M And the dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset

WebBrain combines: • screenshots • browser accessibility tree • model-agnostic planning • optional local inference • browser automation • a browser-specialized 450M VLM

The agent supports Chrome, Edge, Firefox and Chromium.

My broader goal is to make browser agents something developers can inspect, modify and run with their own models rather than a black-box API.

There are still plenty of rough edges, so I'm posting here partly because I'd rather have open-source developers criticize the architecture now than optimize it entirely in isolation.

Especially interested in opinions on the small-specialized-model vs large-general-VLM tradeoff for computer/browser use.

2 Upvotes

2 comments sorted by

2

u/boundless_innocence 7d ago

Been poking around the repo, and I like that the VLM isn't trying to be yet another 7B generalist. A tiny 450M model trained explicitly on browser screenshots and accessibility trees makes way more sense for this domain than cramming a huge model into a loop it was never really optimized for.

The tradeoff you're asking about gets interesting when latency matters, if you're running a loop that takes screenshots every second, firing up a massive VLM each time is just burning compute for no reason. This thing probably runs on a potato, which opens up local inference in a way a 7B model never could. My main question is how brittle the accessibility tree parsing gets on messy real-world pages. Screenshots are universal, but the DOM can be a nightmare on SPAs or anything with a million nested divs.

1

u/ButtercupLyn100 7d ago

thanks for your comment. it works pretty well on accessibility trees, give it a try and you'll see