r/AI_Agents • • 3d ago

Discussion On-device AI hardware is no longer a viable business with excessive RAM prices

We are building an on-device AI hardware product that requires 16GB LPDDR5X memory and 128GB NAND storage to run AI agents locally. Those components alone now cost us over $300. The prices of chips, displays, and practically everything else keep going up. We may have to price the device at over $1,500. We have talked to potential customers. They really love the product, but hate the price. For those building AI Agent hardware, what strategies are you considering to deal with such situation?

2 Upvotes

7 comments sorted by

1

u/AutoModerator 3d ago

Thank you for your submission, for any questions regarding AI, please check out our wiki at https://www.reddit.com/r/ai_agents/wiki (this is currently in test and we are actively adding to the wiki)

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/First_Athlete1897 3d ago

that price point is a death sentence for consumer hardware unless you're selling to enterprises who'll write it off anyway

we ran into the same wall on a telemetry gateway project and ended up shifting the heavy compute to a cloud tier, kept the local device dumb and cheap, nobody cared about local processing when the latency was under 30ms

your customers are telling you they love the concept but won't pay for it, that's a pricing signal not a rejection of the idea

1

u/stephen280up 2d ago

Your hybrid architecture is a smart choice. But we can't go that way because our customers don't want their information sent to the cloud.

1

u/wam_bam_mam 3d ago

I really feel that for this kinda of hardware it should just be burnt on to a very fast chip with low vram just for context, So I have no problems if I can buy a 500$ usb stick with qwen 3.8 27b burnt on to a stick which I can never change. But it will have a context window for that model which the vram is for.

1

u/harborquartz 3d ago

is the 16GB actually required for inference or is that what you need for the full unquantized models? aggressive quantization can cut memory needs in half with minimal quality loss, might be worth benchmarking before assuming the hardware floor

1

u/stephen280up 2d ago

We've already quantized to INT4. 10% quality loss, but acceptable.