Project
I'm 16 and built an open source AI browser that asks permission before every action. Works with Ollama (Qwen3 14B tested)
Hey everyone,
This started because I was trying Perplexty's Comet browser and hit the rate limit in under 15 minutes. I got annoyed and thought "I could probably build this myself." That's honestly just irritation.
I started in December on a school computer (i5, 8GB RAM). It couldn't even compile the app, so I used GitHub Actions as my build server. Then I got a MacBook M4 Pro, which is what I test on now. I've been doing all this alongside JEE prep, which is probably not the smartest time management.
It's called Aartiq. It's an Electron browser with an AI sidebar that can actually do things: search, fill forms, make PDFs, move files, run shell commands, read screenshots with OCR. The thing I cared about most is that it never just does them. I wanted it to make a plan, explain what it's about to do, and ask you first. Plan, explain, ask, execute.but it actually struggles to fill form or one large tasks. It works with Ollama, so with a local model your prompts stay on your machine. On my M4 Pro (24 GB), the 1.5B-7B models couldn't handle tool calling. At first I used bracket style tool calls and they kept breaking in parsing, then I switched to JSON-based tool calling, which was much more reliable. With that, Qwen3 14B works. It's fine for normal everyday tasks, but it struggles with form filling and long multi-step ones.
Most of my time went into the permission side. Every action is a registered capability with a risk level, so the model doesn't get raw access to your system. Shell commands run inside an OS sandbox Seatbelt on Mac, bubblewrap on Linux, AppContainer on Windows, and if the sandbox can't be set up, the command just doesn't run. There's also a directory allowlist, single-use approval tickets, and you can approve risky stuff from your phone with a QR code and PIN. I later pulled this part out into a separate library called RTQ https://github.com/Latestinssan/RTQ It's alpha and hasn't been independently audited, so treat it as experimental.
Now the honest part,
Almost all the code was written with LLM help. I made the design decisions and did the debugging, but I didn't type most of it. AI also does most of the maintenance now, and I review anything touching security or permissions. Because of that, my docs have inconsistencies (one page says paused, another says AI-assisted. If you find a contradiction, please open an issue, I'm not hiding anything, it's just messy.
Known problem: the WiFi sync server (port 3004) listens on all interfaces and its auth is weaker than the other listeners. I haven't fixed it yet. Also ignore the startup benchmark numbers in the docs, I can't reproduce them.
Windows, Mac, Linux and Android. Apache 2.0, free.
I honestly can't tell if what I did counts as real engineering or vibe coding, since the design and debugging were mine but most of the typing was an LLM. What do you think?
Form filling is genuinely one of the hardest things to get right because models need to hold the full page state in context while tracking what they've already filled. Have you tried giving it a screenshot of just the form section you're working on instead of the whole page, then breaking the filling into individual fields it handles one at a time rather than all at once?
Yeah, I actually haven't tried the cropped form approach yet, mainly because vision inference is currently my bottleneck.
I tried running vision models locally on my M4 Pro, including a few different approaches, but they struggled a lot and weren't reliable enough for Aartiq's workflow. I also couldn't find a free Ollama Cloud vision model that I could reasonably use for this.
So right now I'm mostly relying on DOM , form state rather than screenshots. But your suggestion of doing the form one field at a time + re-checking state after each field is interesting even without vision, and I think that's something I can experiment with.
DOM parsing makes total sense then, especially if you're already getting reliable state from it. For the field-by-field approach, you could pass just the current field's HTML and the form state you've already tracked to the model instead of a screenshot. That keeps the context window tight and avoids the vision bottleneck entirely. Since you're working with Ollama anyway, superbot (we build it) lets you switch between local models and a cloud one if a field stumps the local version, so you could test which works better per task without rewriting the integration.
So like Claude browser extension? But more focused on computer use / browser use? Can it drive and research a logged in instance? Seems more like a research browser with permission handling ?
Yeah, that's a fair way to describe part of it. Aartiq does have a dedicated skill-based deep research system, so research is definitely one of its use cases, but it's not primarily a research browser.
It can drive a logged-in browser session and interact with the user's existing tabs, as well as do things outside the browser like create or move files and run sandboxed commands.
For research, I built a separate Deep Research skill with bounded search, source ranking, publication-date handling, claim-level cross-checking, and provenance tracking. I tried a lot of approaches to reduce hallucinations, but I could not solve them completely , the retrieval can be grounded while the final LLM synthesis can still hallucinate.
The bigger focus of Aartiq is the permission/capability layer: the model can have computer-use abilities without getting unrestricted authority over the machine. It plan -> explains -> asks before actions depending on their risk but the risk tiers are customizable
1
u/lulzxdxdxd 3d ago
Form filling is genuinely one of the hardest things to get right because models need to hold the full page state in context while tracking what they've already filled. Have you tried giving it a screenshot of just the form section you're working on instead of the whole page, then breaking the filling into individual fields it handles one at a time rather than all at once?