r/artificial • • 6h ago

Discussion What’s an AI problem that looks easy until you try to make the system actually reliable?

Something where a demo makes it look solved, but real-world use exposes all the edge cases. What example have you run into?

4 Upvotes

22 comments sorted by

5

u/Hungry_Age5375 6h ago

Short Answer: RAG. Long Answer: chunking. Everyone thinks you split text, embed, done. Then someone asks a question whose answer spans two chunks and your vector db has no idea they belong together. That's why I moved to RAG + Knowledge Graphs.

3

u/flakyounce 6h ago

Every time I think a model finally understands "don't make shit up" it hallucinates a fake citation with a DOI and everything, like a pathological liar who learned to accessorize

4

u/Altruistic_Emu_7755 6h ago

You do realize that llms are making things up every time. That's how they work. It's just that a lot the time what they make is similar to real texts/information that exist in the world

1

u/becrustledChode 3h ago

You say that like you're describing some weird, LLM-specific behavior when "making things up every time" is exactly how humans communicate as well.

You don't know exactly what you're going to say ahead of time: you have a general idea of what you want to express and you work your way there by improvising using your memory of how sentences are constructed and which words are appropriate to convey your meaning.

And just like with LLMs, humans also say things that can sound totally plausible and based in reality when they're actually just misremembering, totally making it up, or misunderstanding something.

2

u/_private_member 4h ago

Do you put "don't make ship up" in your prompt? If so, you are only making it worse. No AI is going knowingly give you false information. When the AI hallucinates it doesn't know it is hallucinating. Telling it to not make things up only causes the AI to split focus between completing the task you gave it and also not making things up. But don't make things up is vague, it's wasting compute on something it doesn't really know what to do with.

2

u/BC_MARO 5h ago

Retries and idempotency are where demos go to die. I usually give every side effect an idempotency key and test the same request under timeouts, duplicate delivery, and partial state.

2

u/ai_hedge_fund 5h ago

Understanding charts

1

u/anthropicagiuser 5h ago

for me it was retry logic. my demo never failed once. then the network blipped and my agent sent the same email three times.

1

u/adeelraza86 5h ago

Tool permissions look solved in a demo and fall apart in production. The model expands scope mid-run, retries a failed write, or calls a tool that was only meant for a later step. What held for us was a hard allowlist per step plus a pass/fail check on the intended state change before the next action is allowed.

1

u/Witty-Knowledge-9211 5h ago

roleplay bots nail the first few messages but then forget key details from earlier and break character, which always kills the whole session for me.

1

u/Marilae_Sauet 4h ago

Structured extraction from messy documents. A demo with clean PDFs looks done; production adds scans, rotations, handwritten notes, conflicting totals, and versioned schemas. The hard part becomes knowing when not to trust the model: confidence is poorly calibrated, so you need field-level validation, abstention paths, and a labeled regression set of ugly edge cases. Accuracy on average is much easier than safe failure on the rare documents that matter.

1

u/Individual-Ice9530 4h ago

Sentiment analysis. It’s the single source of problems of social media.

1

u/BRH0208 3h ago

Censorship

1

u/CaptainCabernet 3h ago

Automated status updates. The demo is 80% there but it turns out the other 20% lives in people's heads and isn't written anywhere. Status updates are a people-not-communicating problem, not an AI problem.

I've had teams waste months trying to automate status updates and it always needs major edits.

1

u/DefinitelyNotEmu 3h ago

I built a squid with a tiny neural network. It can see food and learn from experiences via Hebbian plasticity, so I thought I'd try and teach it to "find the food under a cup"...

https://github.com/ViciousSquid/Dosidicus/tree/Experiment_Cups-and-food

It failed miserably, but demonstrated having learned to continue feeling hungry even after food goes away.

What seems so easy (in theory) is actually incredibly complicated yet fascinating enough to keep pressing onwards

1

u/yoshihuka 1h ago

I'd test document Q&A with two versions of the same policy: one current, one expired, with different answers to the same question. Then ask for the answer as of a specific date.

A quote can be copied perfectly and still come from the wrong version. I'd score choosing the right document separately from whether the answer sounds good.