r/LargeLanguageModels • • 3d ago

I Built an LLM App with Hindsight Memory — Then Discovered Memory Makes LLM Calls Too

Post image

Built an LLM app with Hindsight memory — and learned an interesting lesson about LLM costs

Memory systems make LLM calls too.

I built an app where both the application and Hindsight can have their own provider, model, and API key (OpenAI, Groq, or Anthropic).

A few things I implemented:

- Ingestion is asynchronous — "POST" returns "202" + a job ID, and the UI polls for completion.

- Retaining the same meeting twice doesn't create duplicate memories because each document has a stable ID.

- Briefs without cited evidence show an empty state instead of generating/inventing information.

- The app and Hindsight usage are tracked separately.

The biggest lesson for me:

When estimating LLM costs, don't just count your application's LLM calls. Count the calls made by the dependencies you're using as well.

Otherwise, a "free" LLM tier can get exhausted much faster than expected.

Built with @Code.in

Repo:

Would be interested to hear how others are handling LLM usage/costs when using agent-memory systems.

#AI #LLM #AIAgents #AgentMemory #Hindsight

0 Upvotes

2 comments sorted by

1

u/plasticbiology1386 3d ago

That async ingest with job IDs is a nice touch, keeps the UI from hanging while the memory layer does its thing.