
28 Sep 2026
Gimlet Labs partners with Cerebras as CS-4 launch partner for ultrafast inference cloud
Gimlet Labs and Cerebras Systems said Monday they are collaborating to deliver ultrafast AI inference through Gimlet Cloud, combining Cerebras wafer-scale compute with GPUs in a disaggregated architecture and naming Gimlet a launch partner for Cerebras CS-4.
Safety teams are still arguing about how tightly to box in AI agents. Cerebras is putting CS-4 into a cloud built to answer them at up to 3,000 tokens a second, and if that speed holds in production, serving AI stops being only an NVIDIA story.
On Monday, 28 September 2026, Gimlet Labs and Cerebras Systems said they will work together to deliver very fast AI inference at large scale through Gimlet Cloud. Inference is the step where a trained model produces an answer. Cerebras trades on Nasdaq as CBRS. The dateline is San Francisco and Sunnyvale, California. The same announcement went out on GlobeNewswire. Those lines are the companies’, on Cerebras’s press page.
The companies say they plan to deliver speeds of up to 3,000 tokens per second for demanding agent and real-time applications. A token is a small piece of text the model reads or writes. Tokens per second is how fast the answer comes back. An agent, here, is software that takes a series of steps, not only one reply. The first Gimlet Cloud datacenter that runs Cerebras hardware is expected to come online later this year. The speed and the date are plans in the release. The page does not say a public cloud is already serving at that speed.
Gimlet Cloud puts the Cerebras Wafer Scale Engine and GPUs into one inference product. A GPU is the graphics chip most AI clouds use for that math. The Wafer Scale Engine is Cerebras’s very large chip, built from a whole round of silicon instead of a small square cut from it. The cloud splits the job so each phase runs on the chip best suited to it. Cerebras calls that inference disaggregation. The release does not say which phase sits on which chip. Those lines are Cerebras’s.
Zain Asgar, co-founder and chief executive of Gimlet Labs, said inference speed decides how productive AI can be, and that fast inference opens new markets. He said Gimlet’s software, written for more than one kind of chip, plus the Wafer Scale Engine, lets them run each phase on the hardware best suited to it. He said they plan to deliver up to 3,000 tokens per second at production scale. Production scale means a real service customers use, not a lab demo. That quotation is Asgar’s, as Cerebras prints it.
Sean Lie, co-founder and chief technology officer of Cerebras, said pairing Cerebras’s fastest tokens with high-throughput GPUs is the best economics for a datacenter. Throughput means how much work the chips finish, not only how quickly the first word arrives. He said making Cerebras a native part of Gimlet’s inference cloud will bring that speed to more developers. He also said Cerebras is building with Gimlet as a launch partner for CS-4, “giving customers a direct path to our latest technology.” CS-4 is the name he uses for that latest system. The release does not give a day when CS-4 is on sale for everyone, and it does not print a price. The quotation is Lie’s, as Cerebras prints it.
The work builds on customer projects the two companies have had going since last year, and on a combined setup that is already serving tokens in private deployments. Private means those customers are already getting answers. It does not mean the new public datacenter is open. Gimlet said it will widen the partnership to software, how the datacenter is designed, the developer APIs, tools, tuning, testing, and day-to-day operations, so the fast inference is available through Gimlet Cloud. An API is the door a developer’s program uses to call the service. Those lines are Cerebras’s.
Gimlet is backed by Andreessen Horowitz and Menlo Ventures and is based in San Francisco. The release says Gimlet’s aim is a large jump in AI performance, and that its inference cloud grew out of the company’s own research. Cerebras, in Sunnyvale, says customers can run its systems in their own buildings or in the cloud. Those lines are the about sections of the release.
In plain terms, Gimlet is putting Cerebras’s giant chip and ordinary GPUs in the same cloud, and the companies say that mix is meant to reach up to 3,000 tokens a second. Gimlet is a named launch partner for CS-4. The first datacenter in the plan is later this year. The release does not name a power capacity or a dollar value for the deal.
The picture is Cerebras’s product graphic for its Nexus rack-scale platform. Three backpack modules sit in a rack, and the side of the machine reads Wafer-Scale. It is the company’s image.