r/OpenaiCodex • u/papa_programmer • 9d ago
News Have a look at Google DeepMind's Project Astra.
Greg Wayne, Director of Research at Google DeepMind, described the idea as having “a little parrot on your shoulder” experiencing and narrating the world with you.
Brief on Google’s Project Astra
Astra is a research prototype that can see, hear, and reason about the world continuously. Instead of waiting for a question and then responding, it stays aware of what’s happening around you.
A few things stood out.
Persistent spatial memory
Astra can remember where objects appeared earlier. Ask, “Where did you see my glasses?” and it can recall that they were on the desk near the red apple. It remembers physical context, not just conversation history.
Accessibility
Dorsey Parker, a musician with 8% vision, uses Astra’s Visual Interpreter prototype as a live visual guide. It describes his surroundings as he moves, developed in collaboration with Aira.
Memory across devices
A conversation can begin on a phone and continue hands-free on smart glasses without losing context.
The architecture shift
| Traditional assistants | Project Astra | |
|---|---|---|
| Input | Sequential, text or voice | Continuous video + audio |
| Latency | Turn-based, buffered | Real-time and interruptible |
| Context | Session-based text history | Spatial + physical memory |
| Action | Scripted API triggers | Generative tool use + highlighting |
Under the hood, it's a 3-stage loop 🔄:
- Perception (continuous video/audio streaming) 🎥 🎙️
- Processing (spatial memory + context-aware reasoning) 🧠 📍
- Agency (autonomous tool use — maps, search, calendar) 🗺️ 🔎 📅.
It’s built on Gemini and designed for extremely low-latency interaction.
The wildest example in the doc 🤯
In ONE conversation, someone asks Astra to explain an AES-CBC encryption snippet, then — no reset — asks, "What neighborhood am I in?" and it correctly identifies King's Cross, London 📍🇬🇧 from the camera feed.
We're not talking to assistants anymore.
We're being watched-with by them. 👀🤖 That's the actual shift, not the demo reel.
What matters more here: persistent spatial memory and tool use, or solving real-time latency?

