
22 Sep 2026
Subconscious raises $5.1M for inference built for long-running agents
Cambridge startup Subconscious announced $5.1 million across a pre-seed and a seed, led by MassVentures, for an inference platform that compresses and caches agent context so long-running coding agents finish faster and cost less.
SOFTWARE desk — agent bills are the new tax. Inference that compresses long traces is how teams keep agents running without blowing the budget. Inference means running the model, and a trace is the record of every step.
What the company says it is shipping. Subconscious calls the product an inference platform designed for long-running agents. Inference, here, means running a model so it can take the next step. It is not the training that builds the model. An agent, here, is software that takes a series of steps with tools, not only a chat reply. Long-running means that job can go on for a long trace. A trace is the record of those steps. The post says the platform is available to engineering teams today. Enterprises can use a cloud deployment or an on-premises one. On-premises, often shortened to on-prem, means the software runs on the customer’s own computers. The same post says a team can start in about 30 seconds with a command-line tool, shortened to CLI, a short typed command, and names the coding agents Claude Code, Codex, Pi, Copilot, and OpenCode. Claude Code is Anthropic’s coding agent. Codex is OpenAI’s. The page does not define Pi or OpenCode, and it does not say which company’s Copilot. Do not add a vendor the page left off. A team can also run the system on its own graphics chips, which the page calls GPUs, for what it says is the largest cost saving, with fully air-gapped data protection. Air-gapped means those computers are not connected to the public internet. File that availability list as the company’s. This desk did not install the command.
Where the company says the method comes from. The post is bylined by Jack O’Brien, co-founder and chief executive, and Hongyin Luo, co-founder and chief technology officer. It says MIT researchers found a way to compress up to 80 percent of an agent’s context, and that the company is turning that into this platform. Context is the text the model can see for one job, including the conversation and the results of tools. Compression means shrinking that text so the model processes less of it. The post also says the system was born from MIT research on inference optimizations, and that it uses dynamic context compression plus highly efficient caching. Caching means reusing work the system already computed, instead of doing it again. Up to 80 percent is the company’s ceiling for how much of the context it says it can shrink. The post does not name a paper. Do not add one. The post says this does not require a change to the hardware, the models, or the apps that call them. File the MIT line, the 80 percent context line, and the two founder titles as the post’s. This desk did not read a lab notebook.
The performance claims, and whose numbers they are. For agents that use more than 200,000 tokens, the post says Subconscious can finish tasks twice as fast, extend a model’s effective context to 5 million tokens or more, improve results on key agent benchmarks such as coding and workflows by 1 to 10 percent, and cut cost by up to 80 percent. A token is a small chunk of text, roughly a word piece. More than 200,000 tokens is a long job. Five million is more than twenty-five times that. Twice as fast means the work takes about half as long, on this claim. Up to 80 percent lower cost means the bill, at the top of the claim, is about one fifth. Up to is a ceiling. It is not a promise that every job lands there. One to 10 percent is a small accuracy range, not a single score. These figures are the company’s. This desk did not time a job or open a bill.
Two benchmark lines, still the company’s, and not a test this desk reran. On TriE, which the post calls a benchmark built to measure system performance, Subconscious’s runtime finished tasks twice as fast as SGLang and supported 2.3 times as many requests at the same time. SGLang is the other way of serving a model that the page uses as the comparison. It does not define the letters. Concurrent means those requests run together, not one after another. The post says handling more of them at once squeezes more work out of the graphics chips, and that this effectively doubles or triples the size of a cluster. That doubling is the company’s reading of the 2.3 times line. It is not a machine count this desk made. On DeepSWE, which the post calls a benchmark of long coding tasks, GLM 5.2 hosted on Subconscious solved 46 percent of problems at an average cost of $2.79. The same model on what the page calls standard inference infrastructure scored 44 percent and cost $3.92. Forty-six percent is two points above forty-four. $2.79 is about seventy-one cents of each dollar of the $3.92, roughly a 29 percent lower bill on that one comparison. That is not the same number as the up-to-80-percent cost line. The post says the compression and the caching improve both efficiency and accuracy. File both benches as Subconscious’s. This desk did not rerun the tasks.
The closed-beta story, and what it is not. The post says the platform is live in production today and has been powering coding agents and other agent products in a closed beta for months. A closed beta is a limited trial. Production, here, means the company says real work is already running on it. In July, the post says, a 20-person engineering team switched from Claude to the GLM 5.2 model hosted on Subconscious to power their coding agents. The team is not named. Do not name one. In the two months since, the post says they cut monthly AI spend from $40,000 to $6,000, engineers report faster token throughput, and they have not hit their rate limits. $6,000 is 15 percent of $40,000, a cut of $34,000 a month, on the company’s account of that team. Throughput means how fast the tokens come back. A rate limit is a cap on how much the team can use. One engineer on that team, the post says, ran a trace of 4,571 turns and 9,556 tool calls. A turn is one step in the back-and-forth. A tool call is the agent using a piece of software. Subconscious recorded 449 million tokens where a conventional runtime would have billed 2.6 billion, which the post calls an 82 percent reduction. 449 million is about 17 percent of 2.6 billion, which is that 82 percent cut. The post says that despite aggressive compression, the team reported no loss in model capability. Reported is the company’s word for what that team said. This desk did not see the invoice or the trace.
Two quotes, as color only. Jack O’Brien said open models got good enough this summer that their quality against closed models stopped being a compromise, that every engineering team he talks to has put a ceiling on spend per developer, that teams need to spend less while engineers keep wanting more, and that Subconscious runs open models in a way built for coding agents so teams can have both. Stacy Swider, vice president of investments at MassVentures, said the team is the best on the planet for one of AI’s toughest challenges, and that MassVentures is excited to be part of growth toward thousands of companies and billions or even trillions of agents. Best on the planet, and trillions of agents, are her words. File the names, the titles, and those sentences as the post’s. A quote is not a customer count, and it is not a price for the company.
The about box, and the forecast that is not a result. Subconscious says it is an inference platform for long-horizon agents, based in Cambridge, Massachusetts, and that it has raised $5.1 million led by MassVentures. Long-horizon means the same thing as long-running on this page: a job that keeps going. The about box says the aim is infrastructure for trillions of agents. The mission line is to let everyone do deep work and automate the rest. Earlier, the post says agents do the most valuable work and are the most expensive way to use the models, and that on recent trends agents will consume virtually all inference soon. That last sentence is a forecast. It is not a measurement. The page does not split the $5.1 million into a pre-seed check and a seed check. Do not split it. It does not name a lead other than MassVentures, and it says MassVentures led the rounds, plural.
Plain English for the rest of the card: pre-seed = the earliest private check. seed = the next early round. $5.1 million is the total the page prints across both. It is not a valuation. The page does not print one. inference = running the model, not training it. agent = software that takes steps with tools. trace = the record of those steps. context = the text the model can see for one job. compression = shrinking that text. caching = reusing work already computed. token = a small chunk of text. 200,000 tokens = the long-job line the claims start from. 5 million tokens = the context the company says it can stretch to. CLI = a short typed command. on-prem = the customer’s own computers. air-gapped = not connected to the public internet. GPU = the graphics chip that runs the model. closed beta = a limited trial. $40,000 to $6,000 = the unnamed team’s monthly spend, the company’s account. 449 million versus 2.6 billion = the one trace, an 82 percent cut on the page’s arithmetic. SGLang = the comparison runtime named on the TriE line. GLM 5.2 = the model name on the DeepSWE line and in the July story. This filing is the 22 Sep announcement. It is not a bill this desk audited.
PRIMARY here: Subconscious’s 22 Sep 2026 post, “Subconscious Raises $5.1 Million to Build the Inference Platform for Long-Running Agents,” bylined by Jack O’Brien and Hongyin Luo, at subconscious.dev — Tier A PRIMARY, the company’s own record. The page says September 22, 2026, and Cambridge, MA. It does not print an hour. The $5.1 million across a pre-seed and a seed, the MassVentures lead on the rounds, Foothill Ventures, Underscore VC, E14 Fund, Oakseed Ventures, the Agent Fund, the among-others line, the available-today line, the cloud or on-prem choice, the about-30-second CLI, the Claude Code, Codex, Pi, Copilot, and OpenCode names, the own-GPU and air-gapped line, the MIT research and up-to-80-percent context compression, the caching line, the beyond-200,000-token claims of twice as fast, 5 million tokens, 1 to 10 percent, and up to 80 percent lower cost, the TriE comparison with SGLang, the DeepSWE 46 percent at $2.79 versus 44 percent at $3.92, the closed beta, the unnamed 20-person team, the July switch from Claude to GLM 5.2, the $40,000 to $6,000 line, the 4,571 turns and 9,556 tool calls, the 449 million versus 2.6 billion tokens, the O’Brien and Swider quotes, the Cambridge headquarters, and the trillions-of-agents aim are the post’s. NOT claimed: a valuation, a split of the $5.1 million, an hour stamp, a paper title, a vendor for Pi, OpenCode, or Copilot, that this desk installed the CLI, reran TriE or DeepSWE, saw the invoice, or named the 20-person team, a stock tip, or investment advice. Distinct from the already-filed firecrawl-alexandria-75m, snorkel-350m, and outerlimit-16m.
RELATED
On 22 Sep 2026, Subconscious announced $5.1 million raised across a pre-seed and a seed. The record is the company’s own post, “Subconscious Raises $5.1 Million to Build the Inference Platform for Long-Running Agents.” The page says September 22, 2026. It does not print an hour. The dateline in the announcement line is Cambridge, MA. A pre-seed is the earliest private check. A seed is the next early round. The page prints one total, $5.1 million, across both, and says MassVentures led the rounds. Participants named on the page: Foothill Ventures, Underscore VC, E14 Fund, Oakseed Ventures, and the Agent Fund, among others. Among others means the page does not claim that list is complete. Do not add a firm it did not name. The page does not print a valuation. A valuation would be a price on the whole company. Do not invent one. These lines are the company’s. This desk did not see a term sheet.