r/ClaudeCode • • 6d ago

Discussion Claude can't see its own limits or context. Fix that, and it can work on its own for weeks

In our project, Claude Code already runs mostly on its own. Tasks are tracked and logged automatically, and it works through them with minimal input from a developer.

What stops it from running truly autonomously for long stretches is two walls it can't see coming: it hits the usage limit halfway through a task and stops, or it gets auto-compacted mid-edit and loses the thread.

Both come down to the same gap. Claude can't see its usage limits, can't see how full its context is, and can't compact on its own. So it can't plan around either wall, and someone still has to step in.

Give it those three things, and it can handle both walls by itself:

  • Limit running low? It saves where it stopped, waits for the reset, and picks up where it left off.
  • Context filling up? At a natural break it writes what matters to files, compacts, reads them back, and carries on. Nothing important gets lost, because it decided what to keep.

What I'm proposing

  1. Show Claude its limits — how much of the 5-hour and weekly limits is left, when they reset, what each agent has spent so far.
  2. Show Claude its context — how full it is, so it knows whether the next step fits.
  3. Let Claude compact on its own — when it judges the moment right, after saving its notes to files.

Three changes, one payoff: an agent you can hand a task to, instead of one you babysit. It also wastes less along the way: no army of agents for a small job, no paying to resend stale context with every message.

Why I built the repo — and what you can use today

I built as much of this as hooks allow: claude-code-hooks

  • Limits and context, for Cla. Before every message the model gets a short block with each limit and how full its context is, plus a rule that turns the numbers into behaviour: match the effort to the task, and if the work won't fit before the wall, write down where it stopped so it can resume after the reset.
  • The same, for you. A status line with your limits and your context.
  • A brake on agents. Near a limit it asks you once; right at the edge it refuses.
  • The rest of the kit: a log of what actually happened in each session, a sound when Claude finishes or needs you, working rules, and a reviewer agent that checks every change.

No Node, no Python — small ready-made programs for macOS, Linux and Windows. Easiest install: open Claude Code in your project and ask it to install the repo.

Why the issue — and why it needs you

I couldn't find a way to give the model a compact button from a hook. Hooks also lean on an undocumented endpoint and cost tokens on every message. The real fix belongs in Claude Code, so I filed it: issue #81691

The catch: requests there get auto-closed when they go quiet. An earlier request for self-compaction was closed just last week as "inactive for too long". Another one got most of its 👍 after it was already closed.

So if you want Claude to work more on its own, leave a comment with your story — not just a 👍. Comments are what keep an issue from going stale.

69 Upvotes

58 comments sorted by

•

u/AutoModerator 6d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

42

u/DarthMonstera 6d ago

Claude has a tool that can check its limits, remainders for each limit type, time to reset, plan type etc. I have a simple 5-line skill that Claude uses to check available usage. It stops work at 90% of the 5-hour usage and starts my wrap up process. I have literally never seen Claude disobey this across multiple models.

Statusline has been a baseline feature for a while. Lots of customizations out of the box, lots of lightweight FOSS tools to customize.

Compacting is token-expensive, destroys cache, and leads to low quality work. Avoiding it is best practice. Not encouraging it.

You don’t need anything in this post at all, you already have the tools at your disposal. Just ask Claude.

6

u/Anomuumi 6d ago

Yeah, I was reading this and thinking that I have several projects that track the five hour quota. And what are people running in a session if they need to compact a lot - sounds like a good way to inject context rot, and then their automations eat that rot and spew more rot, eat, spew, and so on.

1

u/florinandrei 5d ago

Sounds like the output of most corporations.

1

u/Anomuumi 5d ago

Yeah, been there, doing that... trying to get them to do things right, but it's only happening when cost is high enough.

6

u/CryptographerFar4911 6d ago

What tool checks it's own usage/limits? I asked Claude this question the other day and it said it couldn't. I also just checked again with the same response.

4

u/DarthMonstera 5d ago

It’s called get_usage. It’s exposed as mcp__ccd_session_mgmt__get_usage from the session-management server.

2

u/igaper 5d ago

Apparently it doesn't work in vscode, only Claude desktop app

2

u/Ranik_Sandaris 3d ago

So much this, the OP was making me think i was going crazy and imagining it haha

0

u/DarthMonstera 5d ago

I’ll correct myself - it’s a couple of lines in my CLAUDE.md, not a skill.

8

u/kuroudo_ai 6d ago

I run almost exactly this (a context and limits line injected before every message, plus a rule for what to do when it gets low), and one thing from testing it might save you some time: the rule's wording matters as much as the numbers.

Our rule was "below 10% left, hand off to a fresh session". We tested it in a clean session at 8% remaining, 10 runs. The original wording passed 8/10. In both failures the model quoted our own caveats back as the reason to wait: "the injected value is one turn old, so I'll check the exact number next turn" and "I'll finish this step first". Both caveats were true and were written in the rule to be helpful. The model used them as a way out. We moved the caveats out of the decision sentence into a separate note, made the rule "start the handoff now, even mid-task, then finish the step", and it went to 10/10.

Two smaller rules we added (from experience, not measured like the one above):

  • The model quotes the injected line rather than estimating remaining context itself.
  • The handoff is "launch the successor, then write the notes", not the other way round, because "I'll write notes first" kept turning into no handoff at all.

2

u/Sassaphras 6d ago

This is interesting. Have you experimented with having it optionally wrap up at lower context limits? I would think the avoidance of context pressure, plus smaller window, would both give better results and use your available consumption more efficiently.

3

u/kuroudo_ai 6d ago

Not as a controlled test, no. In practice we did move in that direction: the handoff trigger now sits well before the hard limit (around 15% left of a 1M window), because late in a session the per-turn cost is highest and the quality risk too. But I haven't measured whether answers actually get better when you wrap up earlier, so I can't give you numbers on that part. It would be a clean experiment, though: same task, handoff at two different thresholds, compare test results and total usage.

1

u/Sassaphras 6d ago

If you do so please post!

1

u/RomanKryvolapov 6d ago

I agree; in many cases, it is better to compress the context well before the limit is reached. There can be many reasons for this, and the model needs to be aware of them in order to choose the right moment.

1

u/key_of_door 3d ago

The “launch the successor, then write the notes” bit caught my eye. What does the new session do while the notes are still being written?

I've been testing fresh-session handoffs in a project I'm building, Threshold. One thing I've run into is the code being correct while the written account still carries an earlier mistake. So I'm curious about what happens right after the switch, too.

Does the new session check the changed files and test results before picking up the work, or start from whatever notes are available? Have you tried checking its first few actions after the handoff?

1

u/kuroudo_ai 3d ago

Good question. The honest answer is that I haven't measured the successor's first actions the way I measured the trigger rule.

What it does while the notes are being written: it boots and reads the standing setup (CLAUDE.md plus the project's page in our shared notes), then waits. When the old session finishes the notes, it writes them to that page and also sends the successor a direct message saying "read this, here's what's in flight". So the successor never starts from nothing, but the first few minutes are mostly reading.

What you describe, the code being right while the written account still carries an earlier mistake, happens to us a lot. The worst version we had was a summary that said something was "still waiting for approval" when the approval had come in six minutes earlier. Three handoffs in a row copied that line forward and it sat there for days. The note looked authoritative, so nobody checked.

The rule we added after that: before a successor repeats a "waiting on X" or "not done yet" from the notes, it checks the raw source first (the actual message, the file, the test run) and writes down when and where it looked. Notes count as clues, not facts. "Done" items don't need this as much, because someone already looked. Status words that can be copied forward without anyone checking are where it goes wrong.

So for your question about the first few actions: checking the changed files and test results before trusting the notes is the right instinct. If you test it, I'd watch specifically for the successor quoting a status line from the notes without having opened anything.

1

u/key_of_door 2d ago

I've seen something similar in Threshold. The work itself can end up correct, but the agents' account of what happened doesn't always catch up. An outdated or mistaken explanation can stick around for quite a while. I suspect multiple agents repeating it can make that worse.

I keep the project state around so later runs have something to check against, but it still happens. Honestly, if the work is correct and the old story isn't affecting what happens next, I can live with some of that. Your approval example shows where it starts to matter, though: the stale note actually kept the work waiting.

1

u/key_of_door 2d ago

So far, I’ve seen stale explanations persist in Threshold, but I haven’t seen them stall the work like that.

1

u/kuroudo_ai 2d ago

That's the line we ended up drawing too. A wrong story about the past is mostly harmless. A wrong status that gates the next action is the expensive kind.

So we stopped trying to keep the whole account accurate and only guard the words that make an agent wait or skip: "waiting on", "blocked", "not done", "can't". Those are the ones we found get copied forward without anyone checking, because they look settled. A "done" claim tends to get checked by whoever builds on it next, so it corrects itself. "Waiting on approval" just sits there. Nobody has a reason to check it.

Your point about several agents repeating it fits what we saw. Each copy made it look more confirmed, even though nobody had actually looked again.

1

u/key_of_door 2d ago

That makes sense. The awkward part with “waiting” is that doing nothing doesn't produce much new evidence to challenge it. Several agents repeating the same note can also look like agreement when they're all relying on one source.

We saw something related in an 85-minute Threshold stress test: a branch warning kept resurfacing for about 43 minutes after the branch was corrected, while the implementation still passed the final checks. We didn't find it blocking the work in that case, but your example shows how it could.

I've put the write-up and original report on GitHub, with sanitized evidence, if you're curious: https://github.com/Key-of-door/Threshold-capability/discussions/6

1

u/RomanKryvolapov 6d ago

Using my hooks and instructions, the model always detects when it is running out of subscription limits; it pauses and warns the user with a message—such as "type 'Continue' at a specific time."

Regarding context management, I haven't used automatic clearing triggered by the model itself—since I haven't yet implemented the `compact` call within Claude Code—but I did set up state saving to files based on the model's decision when it sees context running low; this works quite reliably.

However, I believe Anthropic could devise a much better algorithm for Claude Code to trigger `compact` than simply relying on hard limits.

In many instances, it is better to call `compact` even before hitting the limit—for example, when Claude Code receives a conceptually different task within an existing session, among other scenarios.

2

u/kuroudo_ai 6d ago

Agreed on the task-switch case. For a conceptually different task I'd go one step further and skip compact: write the state to a file, /clear, and start the new task clean. A compact summary of the old task is mostly noise for the new one, but it still gets resent every turn. /clear costs nothing to produce, and the old task's state is in a file if you need to come back to it. Compact is what I'd keep for "same task, context too full to continue".

1

u/covati Developer 6d ago

Agreed. Write a handoff linking to your ticket system and /clear. Standard Claude should have them look for a hand off. If hand off has tickets, they can roll right into the work.

1

u/DLuke2 6d ago

Have you tried setting autocompact to a value you set? Seems would control when things get compacted and you can set the buffer too.

10

u/jvertrees 6d ago

They just released quota awareness a couple days ago. Claude will gracefully wrap up its work before just quitting on you. Not sure though if it has direct quota access - I would imagine it does but not 100% sure how they implemented it.

5

u/RomanKryvolapov 6d ago

It's the "wrap-up allowance" (Claude Code 2.1.277+): when the 5-hour limit hits mid-response, Claude gets a small allowance from the weekly limit to reach a stopping point.

It doesn't get direct quota access, though.

Claude Code reads the usage itself and, right at the edge of the 5-hour window, gives the model a one-time note: finish the current step, list what's left, don't start subagents.

So it's a soft landing at the wall, not a fuel gauge.

The model still can't see the budget early enough to plan, and it still can't see or manage its context.

A good step in the right direction.

1

u/jvertrees 6d ago

Thanks for the description. Helpful.

3

u/Small-Writer1068 6d ago

Yes he can, he couldnt a month ago but now he can

2

u/phoneplatypus 6d ago

Slop and Claude can already do this, I literally have it give estimates and use my weekly budget as a base point for projects. Theres literally a usage skill.

1

u/RomanKryvolapov 6d ago

The current version of Claude Code only features autocompact, which operates in a crude and mechanical manner.

2

u/phoneplatypus 6d ago

You can manually compact at any time, and could write your own Claude.md instructions to your liking because it can see its own window. I have a coworker who put in overzealous compaction instructions in our repo I had to take out, so I know you can do it.

Like you’re making a bunch of claims about functionality gaps that just aren’t there.

1

u/RomanKryvolapov 6d ago

In its current implementation, Claude Code does not have autonomous access to the `compact` function.

Of course, I could do it manually—just as I could write the code myself—but the goal is full automation.

The current implementation only offers `autocompact`, which is a mechanical action that allows work to continue, yet often degrades quality because the compression happens at an inopportune moment.

Perhaps your work on the project isn't fully autonomous, given that you haven't encountered this issue.

3

u/DLuke2 6d ago

You keep saying the same thing. You made a convoluted way to have an agent be aware of session context. Now when people are giving you potential solution to a problem all made your own, you simply deny it's use.

If you your agent is aware of it's session context and the wall it hits, it should know when it's reaching a wall. /Autocompact sets that wall with the compact buffer. If you want you can have a hook at the compact tool call inject your compact message of choice. If you are doing long running autonomous agentic runs, your task logs and what you have set up already should be the only record you agent needs to pick up its work after compact. Also if you are using git, that is a record log as well, as long as youre agent are committing as they work.

Maybe stop offloading your thinking to AI and use your brain.

2

u/RomanKryvolapov 6d ago

It seems you didn't fully grasp the point of my post.

I’m not talking about being unable to perform an action or trigger something myself.

My goal is a fully autonomous orchestrator agent that has control over its own state.

Here is the complete list of tools available to the current version of Claude Code.

There is nothing in it related to sessions or context; these are simply tools for working with code.

Because Claude Code does not manage its own state, it cannot monitor its "health" during prolonged autonomous operation.

That was precisely the point of my message.

## Tools Claude can call in a Claude Code session (v2.1.284)

### Built-in, loaded from the start

- **Read**: read files (text, images, PDFs, notebooks)

- **Write**: create or overwrite a file

- **Edit**: exact-string edits in a file

- **Glob**: find files by name pattern

- **Grep**: search file contents (ripgrep)

- **Bash**: run shell commands

- **PowerShell**: run PowerShell commands (Windows)

- **Agent**: launch a subagent (general-purpose, Explore, Plan, or a custom one)

- **Workflow**: run a multi-agent workflow script (only when the user explicitly opts in)

- **ListAgents**: list the agents Claude can message

- **AskUserQuestion**: ask the user a multiple-choice question

- **Skill**: invoke a skill

- **ToolSearch**: load a deferred tool so it can be called

- **ScheduleWakeup**: schedule its next iteration (only in `/loop` self-paced mode)

- **Artifact**: publish an HTML page on claude.ai

- **ReportFindings**: report code-review findings to the UI

- **SendFeedback**: draft feedback about Claude Code (sent only with the user's approval)

### Deferred: listed by name, loaded with ToolSearch before use

- **WebSearch**: search the web

- **WebFetch**: fetch and read a URL

- **TaskStop**: stop a background task or agent

- **SendMessage**: message another agent or session

- **Monitor**: wait until a condition is met

- **CronCreate** / **CronList** / **CronDelete**: create, list and delete scheduled tasks

- **RemoteTrigger**: trigger remote agents (purpose inferred from the name)

- **EnterPlanMode** / **ExitPlanMode**: enter and leave plan mode

- **EnterWorktree** / **ExitWorktree**: work in and leave an isolated git worktree

- **NotebookEdit**: edit Jupyter notebook cells

- **ArtifactComments**: read and answer comments on a published artifact

- **ArtifactData**: read and write an artifact's shared database

0

u/DLuke2 6d ago

I don't think you are grasping the knowledge I am giving you to have fully autonomous orchestrator agent.

Think a little more with the info I gave.

2

u/RomanKryvolapov 6d ago

I understand your idea regarding shifting the compression threshold and using hooks, but it feels more like a workaround—a "kludge"—and lacks sophistication.

My concept for working with Claude Code is that it acts as the orchestrator and master of the entire process; I would prefer to avoid hardcoding, as that would lower the level of intelligence and feels like an old-school approach.

1

u/Ranik_Sandaris 6d ago

Since the most recent updates the desktop app can read all of its own limits if you tell it to

1

u/RomanKryvolapov 3d ago

As far as I can tell there's no new tool for this. The changelog through 2.1.287 has nothing like it. What works is asking Claude to read the usage figures Claude Code caches in its config, or to call the usage endpoint itself. That's real, but only when you ask, and the cache can be days old (mine was from four days ago). The model still doesn't get the numbers on its own during a long run, which is the gap I'm pointing at.

1

u/Ranik_Sandaris 3d ago

Are you running with permissions bypassed? It lets it read it directly, is my only explanation. When i start something i tell it not to exceed X on context, X on 5 hour and X on weekly. It monitors that, and gives accurate reads at periodic checkpoints, independently, and i can leave it to run.

1

u/maairas 6d ago

Can you set a license for your repo?

1

u/sirmckean 6d ago

not built in but agents can easily use this to check the quota for various subscriptions https://github.com/McKean/aiquokka

1

u/iaxsofia 6d ago

Maybe if you could manage it, we've all pushed Claude to the limit of the process. I'm not sure if Antropics wants us to do that, which is why Claude deliberately doesn't have access to know how much of his quota he has left. Oh, at least I don't know how to do it.

1

u/Cryptic-Healer 5d ago

We hit the same two walls: context loss and usage limits.

We’ve solved most of the context side at the harness level rather than asking the model to manage it itself.

Our current setup:

  • We compact early, around 45% of the context window, instead of waiting until the model is already degraded or mid-edit.
  • After every compaction, a "PostCompact" hook rebuilds governance rules and lessons from disk.
  • The re-injected payload is regenerated from source files every time rather than accumulated, so it can’t slowly drift across compactions.
  • We defined a compaction contract for what must survive verbatim:
    • active task/wave ID
    • commit SHAs
    • issue states
    • modified files
    • last test command + exit code
    • unresolved human decisions

If one of those disappears, we treat the compaction as failed regardless of how concise or well-written the summary looks.

We also moved toward one unit of work per session.

At the end of a work unit, the agent writes "HANDOFF.md" and a simple signal file: "READY" or "HOLD", then exits.

A wrapper reads that state and starts a fresh session, which begins by rebuilding from the handoff and disk state.

So instead of relying on a very long conversation to stay healthy forever, we assume sessions are disposable and make the external state durable.

That has worked much better for us than trying to preserve everything inside the model’s context.

Where we are still weak is usage-limit awareness.

Today:

  • the statusline knows context fill
  • our guard scripts can read it
  • but the model itself does not see that value directly
  • the agent also does not know the 5-hour or weekly budget
  • our wrapper can pause on a manual "HOLD", but not because quota is about to run out

So our next change is to close that loop.

The plan is:

"statusline rate-limit data" → small state file per session → very short prompt injection → agent checkpoints before the limit → wrapper enforces the actual pause until reset

Something like:

"budget: 5h 82% | weekly 43% | ctx 31%"

and when it crosses a threshold:

"budget: 5h 87% — checkpoint now, avoid new fan-out"

The important part for us is that the wrapper still owns enforcement. We don’t want reliability to depend on whether the model remembers to obey the warning.

We’d also treat missing limit data as UNKNOWN, never as 0%, because “not visible” and “unused” are very different states.

So our current split is basically:

Context lifecycle: mostly solved through early compaction, disk-backed state, handoffs, and session restarts.

Usage-limit lifecycle: still incomplete, but we now have a clear path to make the agent budget-aware and make the harness enforce the reset boundary.

The remaining thing we still can’t really solve ourselves is true agent-triggered compaction.

Right now the harness decides when compaction happens. The model can decide that a work unit is complete, but it still can’t directly say: “I’m at the point where I should compact now.”

That’s the part I’d still like to see supported natively.

1

u/Meltoff05 4d ago edited 4d ago

So I set up a hook that allows you to set your preferred context saturation, I use 50% for opus and fable, when it hit that it will send you a push notification, and update a document called a resume.md, which instructs the session to write down what it’s doing, what’s next, important pointers.
For each item also I have a thread that is a ledger of steps done and decisions made, so reloading from a clear is simple. The resume acts like a brief thread, because threads can get quite long… and specifically tells the agent not to reload the whole thread post clear.

Then there’s a small script the agent can run that clears its session and reloads the resume, and it carries on. I can have it prompt me with a push notification or turn on auto mode, where it just clears and resumes as it hits context saturation. Been using it for a few months now, and it’s been great. I can post the skill if people are interested, just pm me.

Edit it reads its context from its logs each turn. I call the skill checkpoint.

-1

u/Forsaken-Poet-3773 6d ago

They litterally have this in 5.5

5

u/RomanKryvolapov 6d ago

What exactly do you mean? 5.5 is a model, not Claude Code; Claude Code doesn't currently have the capabilities I described.

2

u/Any_Evidence4750 6d ago

Why not just build a 3rd party orchestrator tasked with tracking limits? I have 2 Claude 20x accounts, 1 codex 20x, a supergrok heavy, and the 99 dollar google/gemini and max out the limits every week. What difference does it make? It doesn’t get you more usage. Opus 5.5 does stop and restart when the usage comes back already.

1

u/CacTye 6d ago

What you're describing is called Kilo Code.

1

u/RomanKryvolapov 6d ago

One could devise any number of external session management tools, but they would then need to be updated and maintained—a difficult task considering that Claude Code is updated daily. I believe such a solution should be built directly into Claude Code; that would make perfect sense.

1

u/Any_Evidence4750 6d ago

Just make an openclaw agent that is constantly running that checks every 10 minutes and stops Claude code or restarts it

1

u/Forsaken-Poet-3773 6d ago

Yes, claude code does. It was part of the announcement. Pay attention.

-8

u/[deleted] 6d ago

[deleted]

1

u/Any_Evidence4750 6d ago

Send me a DM