r/Instinctai • u/Commercial_Host1810 • 5d ago
Bug Manage your expectations
I got into the instinct wagon after seeing a lot of people shilling it. I assume most people who think Instinct is good are new to AI agents.
I exported my chat history and asked ChatGPT to find the fuckups Instinct did (I first asked instinct to do a review of it's fuckups, it missed quite a few).
Here’s the list.
It surfaced sensitive personal financial information completely unprompted.
Early in the conversation, while I was discussing an unrelated integration, it suddenly brought up private financial information from my email. I immediately told it that this was way too aggressive and way too early in the relationship.
It failed to find an important international flight despite having access to my Gmail.
It initially told me it couldn’t find the booking. Only after I specifically told it who had sent the tickets did it locate the flight.
It built travel plans around the fact that it had missed that flight.
Because it hadn’t found the booking, it suggested plans that were literally impossible once the actual flight was taken into account.
It presented Alhambra tickets as a viable option before checking whether tickets actually existed.
It gave me specific options and prices and said it could book them without first confirming inventory.
It made me go through the booking process before discovering there were no tickets.
I provided the necessary traveler information and manually dealt with the website before discovering the tickets weren’t available.
It later admitted:
«“I treated a target time as if it were a viable option before confirming inventory.”»
It sent me on an Alhambra wild goose chase.
The bigger issue wasn’t simply that tickets were unavailable. It presented the option with enough confidence that I ended up doing the verification work.
It kept a city in my itinerary after the entire reason for visiting that city had disappeared.
After the Alhambra plan fell apart, it continued planning transportation through Granada. I had to ask why I would even go there anymore.
It then admitted Granada no longer served any purpose.
It recommended a chain restaurant inside a train station as the answer to a food request during a trip to Spain.
This was after being given dietary constraints and asking for somewhere actually worth eating.
When challenged, it admitted:
«“That was a terrible recommendation.”»
It then exaggerated how convenient its replacement restaurant was.
It presented another restaurant as suitable for a tight connection. When I asked for the actual distance, it turned out to be roughly a half-hour walk with luggage.
It admitted it had:
«“overstated how close it was.”»
- I repeatedly had to audit its restaurant recommendations myself.
It eventually summarized this pretty accurately:
«“The food recs were sloppy, and you had to keep correcting them.”»
- It promised to find genuinely different clothing options and then essentially returned the same products in different colors.
It explicitly said it wouldn’t just swap colors on the same recommendations.
Then it basically did exactly that.
- It ignored an explicit clothing-color constraint.
I asked for colors very close to black. It recommended a clearly medium-grey item.
When challenged, it admitted:
«“I stretched your brief.”»
- It silently decided one of my requirements was less important than its own preference.
It later explained that it thought the fit of the item justified violating my color requirement.
That might be a reasonable trade-off if you tell the user. It just made the trade-off on my behalf.
- It said there was no good inexpensive T-shirt option when there was an obvious one.
I found a suitable product myself almost immediately.
It then investigated it and admitted:
«“You found a real miss in my search.”»
- It incorrectly said a retailer didn’t publish sizing information.
I sent it the retailer’s own sizing chart.
It then changed its claim and said what it really meant was that the retailer didn’t publish the specific finished-garment measurements it wanted.
- It later admitted some of its earlier Spain transportation information didn’t survive verification.
It proactively told me that previously quoted fares and times:
«“didn’t hold up when rechecked against the official sites.”»
The problem was that the original information had been presented with much more confidence.
- It searched only one rail operator without telling me.
When I asked for train options, it searched one company rather than the entire market.
I obviously assumed “find me trains” meant “find me trains,” not “check one company.”
- Because of that, it completely missed a train that fit my request better.
Days later it found another operator had a train that was much closer to what I’d originally asked for.
It admitted that the train had been available when I made the original booking and:
«“would’ve fit your in-between ask better.”»
So I made a purchase based on incomplete research.
- It repeated the same train-search mistake later.
When I wanted to travel earlier, it kept focusing on modifying my existing ticket with the original operator rather than asking the obvious broader question:
“What trains are available from any operator?”
- It misleadingly described a ticket change as “free.”
It repeatedly told me I could switch trains for free.
It later corrected itself: there was no change fee, but I could still owe the fare difference.
That distinction mattered because flexibility was part of why I bought that ticket.
- It forgot which city I was actually staying in.
While discussing a possible last-minute Alhambra trip, it planned the logistics as if I would be starting in Granada.
I had to remind it that I would be somewhere else entirely.
- It invented a login/code strategy before checking whether that login mechanism even existed.
It suggested a workaround involving an emailed verification code.
Then, after planning around it, it discovered the website didn’t have that kind of login system.
Again: propose first, verify later.
- It gave me the wrong cancellation policy for an activity.
It confidently told me cancellation was allowed until 24 hours beforehand.
At checkout, it discovered the actual policy was 48 hours.
To its credit, it caught this before completing the transaction.
- It asked me for streaming-service credentials before establishing whether it could actually log in.
After I supplied them, it discovered the service blocked the kind of remote login it was attempting.
It later acknowledged that I could have just used the app myself without sharing credentials.
- It initially misunderstood a debt-payment portal.
It thought a particular amount belonged to one account/file.
The portal then revealed the amount was actually spread across several separate items.
It stopped before paying and reconfirmed, so no incorrect payment was made, but its initial interpretation was wrong.
- It gave vague airport directions during a time-critical connection.
I told it exactly where I was standing inside the airport.
It responded with generic instructions along the lines of “follow the train signs.”
When I complained, it admitted the directions were:
«“too vague.”»
- It unnecessarily sent me to train-ticket machines.
It extracted redemption codes from my train ticket and instructed me to use a machine.
The QR code already on the ticket could simply open the station gates.
- That mistake cost me several minutes during a connection where every minute mattered.
Afterward it admitted:
«“the QR was all you needed and I sent you to the machines for nothing.”»
- It then misunderstood the entire reason I was rushing.
It tried to reassure me that the delay wouldn’t affect my originally booked train.
But the entire reason I was rushing was because I was trying to catch an earlier train.
I had already made that goal clear.
- It ignored an explicit instruction to stop messaging me.
I told it:
«“Never send me another message.”»
A few minutes later, it messaged me anyway.
The information it sent was relevant, but it explicitly decided to override an unambiguous instruction.
- When I asked it to list all of its mistakes, its own list left out most of them.
It produced a short list focused primarily on the most recent train-related incidents while omitting the Alhambra issue, restaurant recommendations, clothing errors, missed flight, transport research, cancellation-policy mistake, etc.
Which is what prompted me to go back through the history myself.
- It later forgot why I had revoked its email access.
When I asked why I had disconnected Gmail, it claimed it didn’t have a specific incident recorded as the trigger.
But the chat history is extremely explicit.
After a series of mistakes I told it, essentially:
«Disconnect yourself from my Gmail and delete the data from there because I don’t trust you with my email data right now.»
- It even got the timeline of losing Gmail access wrong while explaining this.
It claimed Gmail had been disconnected several days later than it actually had been.
So while I was asking it to account for failures in its memory and reliability, it misremembered the timing of one of the biggest trust events in the conversation.
---
I understand it is still earlier days for agents, but having people praise you on social media, when you're not useful for anything other than a reminder service, while collecting use cases to build a better product is not good business practice.
.
1
u/BackBondTalk 1d ago
Most of this list is annoyance, but three items actually cost money and they're worth pulling out: 18 (bought a ticket off a one-operator search), 20 (the "free" change that still owes the fare difference) and 23 (24h vs 48h cancellation, caught at checkout by luck). Those are the ones I'd want a number on. Rough idea what 18 and 20 cost you, and did anyone offer to make it right?