01Problem02Products03Benchmarks04Agents & models05Setup06Bittensor — Subnet 114

Your agents pay for tokens they never needed.

Every turn, a coding agent re-sends the whole conversation — files it already read, tool output it already used, its own reasoning. You pay for all of it. SOMA compresses the context and the chain-of-thought on the way to the model, and keeps what the task needs.

Context grows every turn. So does the bill.

  • AAgents re-send history on every request — the same files, the same tool output.
  • BCost per task grows faster than the task itself: every turn pays for all the turns before it.
  • CMost of that context is no longer needed to get the next step right.
AT TURN 14

This session has sent 658k tokens so far. 39k of the last request was context the model no longer needed — 45% of it.

TOKENS SENT PER AGENT TURN
Needed Removable
100k75k50k25k0k
1
3
5
7
9
11
13

Illustration — a typical 14-turn coding session. Real figures land with the benchmark release.

SOMA

We build the layer that makes AI cost what it should.

SOMA is a research company working on one problem: how little context a model needs to do the job well. The app is where that work ships first — it isn’t the only place.

01

Measured, not claimed

A compressor only counts if the task still passes. We score both, and publish the runs.

02

Zero margin

You pay what the model costs us. We earn when compression earns its keep, not on the spread.

03

Written by competition

Our compressor comes from an open subnet where anyone can beat the current one.

Benchmarks

Same agent. Same model. Fewer tokens.

We run the same agent on the same task twice — with compression off and on — and compare cost and whether the task still passed. Every figure says whether it is measured or estimated.

MODELSAVED
DeepSeek V4 Pro$0.92 → list, $0.78 with SOMA−15% est.
DeepSeek V4 Flash$0.03 → list, $0.03 with SOMA−15% est.
DeepSeek V4.1 Flash$0.15 → list, $0.13 with SOMA−15% est.

List price against the SOMA price for the same input. Measured runs replace estimates as they land.

ESTIMATED SAVINGS — TODAY

0%15%
GitHub Copilot CLI + DeepSeek V4 Pro, for this pair.

What we test on

  1. 1Real issues from real GitHub repositories
  2. 2Navigating large codebases and changing many files
  3. 3Debugging, running and fixing tests
  4. 4Long tasks with many tool calls and reasoning rounds

Savings depend on the task, context size, and how much of it the provider already caches. Long sessions with heavy tool output save the most; short ones can save nothing.

Agents & models

Keep your agent. Keep your model.

SOMA speaks the OpenAI API, so the agent doesn’t have to change. Point it at our endpoint and it starts sending less.

YOUR AGENT

GitHub Copilot CLIone-line installerLive
Codexby OpenAILive
Claude Codeby AnthropicSoon
Zedmanual setupPlanned
Hermesby Nous ResearchPlanned
Cursorcustom model providerPlanned
Compress

THE MODEL

DeepSeek V4 ProDeepSeekLive
DeepSeek V4 FlashDeepSeekLive
DeepSeek V4.1 FlashDeepSeekLive
GLM 5.3Z.aiSoon
GPT modelsOpenAISoon
Claude modelsAnthropicSoon

Setup

One line, then keep working.

Sign up, pick your agent and a model, copy your sk-soma key, run one command. The agent you already use keeps working, just with less going out.

GitHub Copilot CLIshell
# point Copilot CLI at SOMA with your key and model

curl -s https://app.thesoma.ai/copilot \
  | bash -s sk-soma-YOUR-KEY deepseek/deepseek-v4-pro

# that’s it — start coding

copilot
NEW ACCOUNTS $5 in creditsMARKUP nonePAY WITH card or TAO

macOS / Linux. On Windows: run it in WSL or Git Bash, not PowerShell. Compression runs before the request leaves us; if unavailable, the request passes through uncompressed rather than failing.

Research & development · Bittensor subnet 114

Our compressors come from an open competition.

SOMA runs subnet 114 on Bittensor. Miners compete to build the best compressor, validators score every submission on real work, and after each competition our compression team takes the winning code and adapts it to the app.

  1. 01

    Miners submit

    Independent teams train and submit compressors to the subnet.

  2. 02

    Validators score

    Every submission runs on real tasks. Cost and task success, both.

  3. 03

    Our team adapts it

    The winning code becomes production-ready for real traffic.

  4. 04

    It ships to the app

    Usage funds the next round of open competition.

SUBNET
SN 114
COMPRESSES
Context + CoT
MINERS LAST ROUND
120
LEADERBOARD
public

Get started

Pay for the context that matters.