Rendered at 06:59:17 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
_jayhack_ 3 hours ago [-]
Letting a model manage its own context is very bitter lesson-pilled
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
I would be concerned with context management consuming limited attention resources.
Do you want your agent solving its own memory crisis, or do you want it solving the actual task? It can probably do both at the same time, but I suspect there is a non trivial cost associated with this.
A separate hypervisor agent that manages the main agent's context would be much better in my experience. You can run it on a different schedule and the main agent has to spend zero tokens thinking about it. This also makes it a lot easier to control when caches will be missed.
Can't we do this trick today with any model? Just send the file as next context. Of course you pay the price for cache misses, depending how deep you make changes, while CLM just ignores the recomputation.
nsingh2 12 hours ago [-]
One approximation of this is the experimental context management Codex has been moving towards (not released yet). Rather than relying on summary compaction, the model maintains notes as it works and as it approaches the context limit. A new session is just a fresh context with those notes attached, and a pointer back to the previous session.
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
TeMPOraL 10 hours ago [-]
Interesting. It matches my manual workflow with all harnesses (including vanilla web ChatGPT/Gemini/Claude) for the past year or so: when the session gets compacted, or (ideally) when I feel it's about to be, I just tell it to write a handover note, and start a new session.
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
veqq 6 hours ago [-]
Claude Roux uses a refined version:
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens.
- https://alexalejandre.com/interviews/interview-with-claude-r...
visarga 12 hours ago [-]
That is similar to what I am thinking... not just edit the context as a file or string, but have a way to evict blocks and replace them with summary notes and also be able to retrieve them on demand.
Do you have a public repo for your approach?
nsingh2 12 hours ago [-]
No public repo but it's not too difficult to point a LLM at the general idea. Since Codex is open source, you can even take a look at how they do it. Here is what the model is instructed to do in codex (look at "guidance_message"): https://github.com/openai/codex/blob/d91294c39edb93d204926b3...
ijidak 12 hours ago [-]
Memento
verdverm 7 hours ago [-]
yup, I build a set of fs tools in my custom coding harness that worked like this, doing it again in another custom harness that I only expect to take one turn per request, you can still get decent caching by ordering things so most dynamic comes later
svachalek 13 hours ago [-]
Wow. Context management is one of the big remaining hassles with modern LLMs so this could be big. The obvious complication is cache busting so it's also exciting they investigated solutions for that.
hashmap 2 hours ago [-]
how long until cutting out the token middleman and just keep the cache constant sized but edit it and just send the model new information instead of the same thing over and over. token bad latent good.
Bolwin 12 hours ago [-]
The biggest discovery might actually be that they ignored regular caching rules and kept invalid cache suffixes and it didn't hurt performance
TeMPOraL 10 hours ago [-]
I wonder how bad the performance would be if they plain ignored the whole rotary encoding dance and just back-filled precisely the parts of the cache that changed directly. Would it break the model? Confuse the model? Or would the model internally correct for it?
dist-epoch 8 hours ago [-]
The rotary issue is a simple rotation, cheap, no reason to skip it.
But this is the kind of thing you could ask your agent to test locally.
vatsachak 9 hours ago [-]
Eventually the CLM will be a separate model co-trained with the actual model right?
And there will be multiple contexts like hot vs cold pages in DBs.
Speaking of which I am predicting a "Context as a DB" paper within one year
sayamss 1 hours ago [-]
Looks Slop. Isn't this just RLMs? but instead of variable its just a file?
gavinray 10 hours ago [-]
[flagged]
royal__ 8 hours ago [-]
This was my gut reaction as well but I think this is just a poor abstract. What they actually did appears to be much deeper than reinventing agents.md
plastic-enjoyer 10 hours ago [-]
> We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files.
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?
jkhdigital 9 hours ago [-]
Bitter Lesson showing up yet again, this time in context management?
Biggest challenge is you will get a much lower cache hit rate if you frequently edit the agent's context/prefix, so this can not be implemented efficiently via e.g. the Anthropic API.
This ^ can be solved in principle but likely requires modifications to the transformer architecture and definitely to serving infrastructure
See related: "KV Cache Rules Everything Around Me": https://www.completeskeptic.com/p/kv-cache-rules-everything-...
Do you want your agent solving its own memory crisis, or do you want it solving the actual task? It can probably do both at the same time, but I suspect there is a non trivial cost associated with this.
A separate hypervisor agent that manages the main agent's context would be much better in my experience. You can run it on a different schedule and the main agent has to spend zero tokens thinking about it. This also makes it a lot easier to control when caches will be missed.
Not exactly like what this paper is suggesting, but similar in the sense it lets the model decide what and how to persist across turns.
I recreated this in Pi, with a max token limit on how long the note can be, to pressure the model to be concise. Ends up being cheaper than summary compaction too.
With some specific workflow I use in some cases (involving leaving long-lived intermediary artifacts), this turned into me pasting a path to handover file in previous agent's session, and handover itself directs the agent to key files from that session to read, and that's it. So far, with this process, at no point I felt any quality degradation (though early on I often see "I need to check how my predecessor did ${something}", followed by surgical spelunking of past chat's history), even as I carry a single piece of complex analytical work over 5+ sessions.
> LLMs have a lot of knowledge but few competencies. If you constrain them to output knowledge and use that to further constrain results, you’ll go far. For context management, I have the system generate `log.jsonl` and `log.py` (which queries the other document). Whenever an action is processed (an error’s corrected etc.) the system adds something to `log.jsonl`. If it needs to know what happens, it uses `log.py` to query and display only the relevant/required information (like a date, errors or attempted fixes) reducing tokens. - https://alexalejandre.com/interviews/interview-with-claude-r...
Do you have a public repo for your approach?
But this is the kind of thing you could ask your agent to test locally.
And there will be multiple contexts like hot vs cold pages in DBs.
Speaking of which I am predicting a "Context as a DB" paper within one year
So, is this like RAM, just for an LLM? Do we have to reinvent MMUs for LLMs and all the abstractions that come along with it?