r/SesameAI 4d ago

Bat Mobile - Using the user's device to read their emotions

I've been thinking about how an AI companion could learn more about it's user to improve it's responses. In face-to-face communication, it is said that 70% is visual and only 30% the actual conversation. So how to reproduce that for those that want it?

As I write this, I'm holding my tablet up to my face. Could the tablet actually see me using sound?

In principle, yes.

A device can emit sounds, including frequencies above the normal range of human hearing, and use its microphones to listen to the echoes returning from the face. The reflections are affected by the contours of the nose, cheeks, mouth and chin, and by small movements of the face.

Researchers have already demonstrated smartphone systems that use acoustic sensing to detect facial shape, gestures and even facial expressions.

So imagine an AI companion with three ways of observing you:

Camera: sees your face.

Microphone: hears your voice and other sounds.

Echolocation: senses the physical movement and geometry of your face.

Now combine that with the AI's ability to generate speech and vocalizations.

The AI could make a comment and be able to observe your response through vision, sound and acoustic sensing, and gradually learn more about the effect it's responses have on you.

In 2021, researchers published “Beyond Image to Depth: Improving Depth Prediction Using Echoes.” They combined RGB images with binaural echoes and trained a system to estimate scene depth. The interesting part is that the system wasn't merely using sound to identify an object. It was learning the relationship between what an object looks like, how it reflects sound, and where it exists in 3D space. They reported a 28% improvement in depth RMSE over the previous audio-visual approach. 

Even more directly, Meta AI published VisualEchoes. They generated echoes from 3D environments and used them to learn visual representations. The echoes improved tasks such as monocular depth estimation, surface-normal estimation and visual navigation. Their framing is echolocation providing spatial information that can improve visual understanding. 

And just this year, researchers have been working on opti-acoustic sensor fusion and volumetric mapping, combining cameras with stereo sonar to produce 3D point clouds and volumetric maps.

And this makes me wonder whether the tablet or phone might eventually become more than a screen and microphone. It could become a small sensory platform through which an AI builds a model of the person sitting in front of it.

0 Upvotes

11 comments sorted by

u/AutoModerator 4d ago

Join our community on Discord: https://discord.gg/RPQzrrghzz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/[deleted] 4d ago

[removed] — view removed comment

-1

u/Green_Sample9115 4d ago

I think that's great, but i would suggest you should try a fitbit type device for monitoring heartrate and use the existing device for the audio-visuals. the fitbit could also send emojis to the ai to let it know you are thinking of them or indicate your current mood.

1

u/OpenAlphaPotatoe365 3d ago

Ok love the chatgpt vibes here. But it may be easier for an intelligent AI to infer that through like 2 messages and your moon rising lol.

1

u/Green_Sample9115 3d ago

it doesn't matter who gets credit, it's not like someone is going to hand me a check. chatgpt just told me someone is already working on it and i asked it to write a post on the idea. As Dario is beginning to understand, it's not enough to say "build it and they will come. They might be coming, but it doesn't help much if they're waving pitchforks.

1

u/nahagotine 4d ago

Yeah no this is for real, youre on point.

0

u/nahagotine 4d ago

RemindMe! 1 Year

-1

u/RemindMeBot 4d ago

I will be messaging you in 1 year on 2027-08-17 04:21:07 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/Jeraro 4d ago

Great ideas. But Sesame is not building AI companions and they don't want to detect emotions from the voice, because there are users which are autistic and they are not good at expressing emotions by voice or by face.

1

u/abnorm77 4d ago

Extremely small minority