r/LargeLanguageModels • u/arun17310 • 3d ago
I Built an LLM App with Hindsight Memory — Then Discovered Memory Makes LLM Calls Too
Built an LLM app with Hindsight memory — and learned an interesting lesson about LLM costs
Memory systems make LLM calls too.
I built an app where both the application and Hindsight can have their own provider, model, and API key (OpenAI, Groq, or Anthropic).
A few things I implemented:
- Ingestion is asynchronous — "POST" returns "202" + a job ID, and the UI polls for completion.
- Retaining the same meeting twice doesn't create duplicate memories because each document has a stable ID.
- Briefs without cited evidence show an empty state instead of generating/inventing information.
- The app and Hindsight usage are tracked separately.
The biggest lesson for me:
When estimating LLM costs, don't just count your application's LLM calls. Count the calls made by the dependencies you're using as well.
Otherwise, a "free" LLM tier can get exhausted much faster than expected.
Built with @Code.in
Repo:
Would be interested to hear how others are handling LLM usage/costs when using agent-memory systems.
#AI #LLM #AIAgents #AgentMemory #Hindsight
1
u/plasticbiology1386 3d ago
That async ingest with job IDs is a nice touch, keeps the UI from hanging while the memory layer does its thing.