After maxing out sub and burning $700/day per person on api, we fine tuned a compression model to trim codex's tool call output to reduce input token + cache. It cut down tokens by 29.6% and now I just leave it on by default in Codex.
To avoid messing up w/ cache, we use proxy + fine tuned qwen model trained on preserving agent trajectory to remove tool call results before they go back to the model, leaving kv cache untouched.
The cli is free for everyone to use (https://github.com/spenmcke/compress). Just lmk ur feedback and hacks to shave even more costs on astra! If you want to integrate it into your product to offer the best models at low cost, I can set you up with an sdk and api keys
PS: It’s built for coding agents, not conversational agents. I optimized it for file retrieval accuracy, trajectory preservation, and quality to get up to 30% cost reduction depending on how context-heavy the task is.
On security and privacy side, it's a proxy wrapping your local codex and ZDR so it doesn't retain any queries. It’s on by default in codex and when you don’t want compression, you can use `codex --uncompress` to disable it.
Give it a try: code is in https://github.com/spenmcke/compress
You can install the cli using
`curl -fsSL https://install.everestagi.com/install.sh | sh && source ~/.config/everest/shell.sh`
Love to hear any feedback and learn your hacky ways to save token costs too!
2 comments