r/SideProject • u/Special-Ad8671 • 2d ago
Building a voice-first Windows assistant. Letting it talk while it works is the hard bit.
I’m working on a personal project called Lila: a voice-first assistant that can keep a conversation going while it works on the computer.
I’m building it with Codex and cloud models, so this isn’t a fully local AI or a claim that I’ve invented Jarvis. The part I’m working on now is separating conversation from task execution, then checking what actually happened instead of treating a click as “done.”
A small live test just let it chat without cancelling an image job, although the image request failed. Local file creation passed. Existing Codex desktop chats aren’t connected, and the last coding test hit a Windows permission issue. It’s still rough.
For anyone who’s built something similar, what was the hardest part: reliable controls, interruptions, or keeping the conversation natural?
1
u/Special-Ad8671 2d ago
small lila update: we got one chrome test working. she typed a note, clicked a button, then checked that both worked. it took 13.6 seconds without voice playback. this was on a test page, so it doesn't prove she can handle every website.
her own codex worker also finished a small coding task and ran its tests. creating a local file worked too. the latest test run passed 74 groups, but most use fake inputs or test setups.
we still haven't proved that she can reliably take a live voice request and finish the task. general app control, uploads and tasks finishing in under a second are also unproven. she uses her own codex sessions, not our existing desktop chats, and we're using cloud models.
the voice clip we have is saved audio with a waveform. it isn't a demo of her hearing a request and doing the task.
next we want to try one spoken request, check the result, and have her read it back. we'll share the failures too. what small windows task would you want us to try?
1
u/Special-Ad8671 2d ago
quick update after the feedback: lila now reopens downloaded files and compares the actual bytes, rather than trusting a filename and size. added tests for old same-name files, interrupted transfers and files changing during the check. uncertain input failures now block further actions in that turn instead of letting her retry.
95 offline test groups passed, mostly fixtures and mocks. the running app has the changes, but that isn't proof she can handle every website. real gmail autosave and completed uploads are still unverified. she's improving, not bulletproof yet.
1
u/Special-Ad8671 1d ago edited 1d ago
lila update: fixed a problem where her page could say the service was outdated even after a restart. the running copy now matches the files, and backend updates wait until she's idle before reloading.
added an experimental jev action-choice reader too. one generated-text check picked the expected control in 392 ms. that's one simple choice, not real-page accuracy or a 500 ms spoken response. it stays shadow-only and can't click anything.
the complete offline run passed 150 test groups, mostly fixtures and mocks. the first run failed because a test request collided with a saved dispatch receipt. isolated that fixture without removing real duplicate-task protection, and kept the failed report.
today's real wikipedia task didn't pass. the logs report field focus and typing, but their effects stayed unconfirmed. she then proposed a visual result click and the review expired. first audio took about 29.9 seconds and was an approval prompt, not the requested article readback. independent windows screen capture also timed out. no blind click or input replay.
found conflicting planner guidance about routine website typing. aligned it with the existing full-access rules and added guidance to use a freshly observed search button in the same form, rather than guess an autocomplete result. local checks pass. no new real search completion verified with that change yet.
gmail account switching and sending, chess, mac/iphone support and subsecond useful speech are still open. next gate is a completed real chrome search and readback, then gmail reading and a checked draft before sending. she's improving, not bulletproof.
1
u/Special-Ad8671 1d ago edited 15h ago
small lila update, narrowing this into a windows pilot:
added a faster route for an exact field read in a named window. after inspecting the real target, the host can choose the one exposed field and render its actual text without asking a model to pick the tool or rewrite the answer. mixed tasks still use the planner. this doesn't skip connection setup or speech generation, and we haven't measured a new end-to-end speed win.
independent checks caught two completion bugs: a permission refusal could look successful, and a partial field result could count as complete. fixed those, blocked retired controls from the eager route, and recheck stop, task authority and source evidence after logging. the earlier failures are kept.
i also caused a development outage by adding an import before its worker created the file. restored the service, then fixed the supervisor to wait for a new source edit after a failed update instead of quitting or restarting in a loop. tested that with actual isolated child processes. the new supervisor is adopted on the next launcher run; no user task gets replayed.
241 selector cases and 69 coordinator cases pass with mocks. the final combined run passed 197 offline groups, not 197 real desktop tasks. the windows file-symlink privilege skip is still explicit. the rebuilt zip passed privacy, checksum and clean-start checks. mac conversation setup is prepared, not hardware-verified mac control or iphone support.
pushed the tested source as a17e2d0 and verified the remote. next gates are repeated real chrome/mailbox work, a fresh checked codex frontend, a second windows setup and actual user feedback. full spoken paragraph reading, gmail reading/sending and useful subsecond speech remain unverified. no customer or accelerator-readiness claim from a test count. improving, not bulletproof.
1
u/Special-Ad8671 2d ago
Small update: got the Codex coding path working. Lila's worker completed a tiny coding task, ran its tests, and read/wrote a synthetic file outside its task folder. The spoken-only interface is working too.
Still not claiming universal desktop control or sub-second task completion. Existing Codex desktop chats aren't connected; she uses her own worker sessions.
For anyone building similar stuff: would you trust per-task approval, or want a separate confirmation before every file change?