Measured, not claimed
A compressor only counts if the task still passes. We score both, and publish the runs.
Every turn, a coding agent re-sends the whole conversation — files it already read, tool output it already used, its own reasoning. You pay for all of it. SOMA compresses the context and the chain-of-thought on the way to the model, and keeps what the task needs.
This session has sent 658k tokens so far. 39k of the last request was context the model no longer needed — 45% of it.
Illustration — a typical 14-turn coding session. Real figures land with the benchmark release.
SOMA
SOMA is a research company working on one problem: how little context a model needs to do the job well. The app is where that work ships first — it isn’t the only place.
Point your coding agent at SOMA instead of the model. We compress the conversation and the reasoning on the way through, and charge the model’s real cost — no markup.
Paste text or drop a PDF, pick how hard to compress, and get a shorter version back. The same compression, without writing any code.
A context compressor that plugs into the OpenClaw agent loop. Public repository, run it yourself.
Private deployment inside your cloud, per-team budgets, and a savings report your finance team can read.
A compressor only counts if the task still passes. We score both, and publish the runs.
You pay what the model costs us. We earn when compression earns its keep, not on the spread.
Our compressor comes from an open subnet where anyone can beat the current one.
Benchmarks
We run the same agent on the same task twice — with compression off and on — and compare cost and whether the task still passed. Every figure says whether it is measured or estimated.
List price against the SOMA price for the same input. Measured runs replace estimates as they land.
ESTIMATED SAVINGS — TODAY
Savings depend on the task, context size, and how much of it the provider already caches. Long sessions with heavy tool output save the most; short ones can save nothing.
Agents & models
SOMA speaks the OpenAI API, so the agent doesn’t have to change. Point it at our endpoint and it starts sending less.
YOUR AGENT
THE MODEL
Setup
Sign up, pick your agent and a model, copy your sk-soma key, run one command. The agent you already use keeps working, just with less going out.
# point Copilot CLI at SOMA with your key and model curl -s https://app.thesoma.ai/copilot \ | bash -s sk-soma-YOUR-KEY deepseek/deepseek-v4-pro # that’s it — start coding copilot
macOS / Linux. On Windows: run it in WSL or Git Bash, not PowerShell. Compression runs before the request leaves us; if unavailable, the request passes through uncompressed rather than failing.
Research & development · Bittensor subnet 114
SOMA runs subnet 114 on Bittensor. Miners compete to build the best compressor, validators score every submission on real work, and after each competition our compression team takes the winning code and adapts it to the app.
Independent teams train and submit compressors to the subnet.
Every submission runs on real tasks. Cost and task success, both.
The winning code becomes production-ready for real traffic.
Usage funds the next round of open competition.
Get started