
1 Oct 2026
Nebius acquires Inferize to cut idle-GPU cold starts in Token Factory
Nebius said it acquired Inferize so Token Factory can spin up and scale large-model inference with less idle GPU time when models load, demand spikes, or weights update mid-run.
Production inference is where GPU bills hide — idle capacity waiting for cold starts is pure waste. A Nasdaq AI cloud buying a cold-start killer into Token Factory is a concrete bet that elastic inference, not just more chips, is the next margin fight.
On Thursday, 1 October 2026, Nebius said it acquired Inferize. The newsroom page is titled “Nebius acquires Inferize to strengthen Nebius Token Factory's production inference stack.” The page dates the item October 1, 2026. It does not print an hour. The dateline is Amsterdam. Nebius (Nasdaq: NBIS), which calls itself the AI cloud company, said it acquired Inferize, an inference-optimization company. Inference is the step where a trained model answers a request. Optimization, here, means making that step faster and less wasteful. Nebius says Inferize's technology shortens the time needed to launch and scale large AI models, and that it makes inference workloads elastic. Elastic means the computers can grow and shrink with the work, instead of sitting at one fixed size. Inferize's technology and team have joined Nebius Token Factory, Nebius's managed inference platform for production AI. Managed means Nebius runs the service and the customer sends the work. Production means the models are already serving real use, not a lab demo. Those lines are Nebius's.
What a cold start costs, in Nebius's account. Cold starts are the time models need to load before they can serve a single request. Nebius says that wait adds time and cost for teams running models in production at scale. It leaves assigned GPUs idle when a model first loads, when demand spikes and new instances spin up, and when weights are updated in the middle of a run, for example during reinforcement learning. A GPU is the chip that does the model's math. Idle means the chip is reserved and is not answering anyone. An instance is one running copy of the model. Weights are the numbers that make the model what it is. Reinforcement learning, in that example, is a training loop that changes those numbers while the model is already running. Nebius says the pattern forces a platform to hold spare capacity just to hit its service-level targets. A service-level target is the promise of how fast, or how reliably, the service answers. Those lines are Nebius's.
The phrase Nebius uses for that waste. Inferize's technology cuts this “idle GPU tax,” so capacity can scale much more closely with actual usage. The page says that drives higher capacity utilization and better token economics. Utilization is how much of the reserved hardware is doing useful work. A token is a small piece of the text a model reads or writes. Token economics is Nebius's phrase for what it costs to produce those pieces. The bullets under the headline say the deal brings technology that cuts the time to launch and scale large AI models, improving token economics and capacity utilization; adds deeply experienced systems engineering talent to Nebius's inference team; and continues the production-inference build-out, with optimizations across the stack. A stack, here, is the layers of software and machines that serve the model. Those lines are Nebius's. The page does not print a measured speedup, a utilization percentage, or a dollar saved per token.
What the chief technology officer said. Danila Shtan, chief technology officer of Nebius, is quoted on the page. “Running inference well takes more than fast GPUs and optimized models. The whole system needs to respond when demand changes, including how quickly additional capacity is ready to serve customers. Inferize brings technology that accelerates that process and a team with deep expertise in GPU systems. We’re bringing both into Token Factory to make it more responsive to customer demand and get more useful work out of our infrastructure. The team’s contribution will extend well beyond this first integration.” That quotation is his.
Where Nebius says this layer sits. Inferize adds another layer to how Nebius runs production inference. Eigen AI brought optimization at the model, kernel, and system levels into Nebius Token Factory. Clarifai's core team and licensed technology added system-level inference and compute orchestration. A kernel, in that sentence, is the low-level code that runs the math on the chip. Compute orchestration is the work of deciding which machines run which jobs. Licensed technology means the right to use that software under a license. Those lines are Nebius's account of the earlier layers. The page does not restate when those steps happened, and it does not print what Nebius paid for them.
What Inferize's chief executive said, and where the engineers go. Guy Bortnikov, co-founder and chief executive of Inferize, is quoted on the page. “Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do, and Nebius is where it can go straight into the platform. Our team will work across the stack with one objective: serving more customer demand from every GPU.” That quotation is his. The page also says Inferize's engineers will work across Nebius Token Factory, starting with the integration of their technology. Integration means fitting that technology into Token Factory so it runs there.
How new the company is, and what the page says about the money. Nebius says Inferize was founded in January 2026 and that the team had a working prototype within three months. A prototype is a working first version. The terms of the transaction were not disclosed. The page does not print a purchase price. Those lines are Nebius's.
What the about box says Nebius is. Nebius calls itself the AI cloud company, building the full-stack platform for developers and companies to take charge of their AI future, from data and model training to production deployment. It says Nebius operates at scale with a rapidly expanding global footprint and serves startups and enterprises building AI products, agents, and services. Full-stack, in that line, means the layers from the data through training and into the running product sit with one company. Nebius is listed on Nasdaq as NBIS and is headquartered in Amsterdam. Those lines are Nebius's. The page does not print a revenue figure or a customer count.
The picture is the official newsroom graphic. A lime frame surrounds a white card. A lime pill at the upper left reads NEBIUS. The right side shows a pale grid, like a sheet of paper, with white bars and a white play button. The headline on the card reads “Nebius acquires Inferize to strengthen Nebius Token Factory's production inference stack.” The frame does not print a calendar date. It is the newsroom graphic. It is not a photograph of a data center or a person.
In plain terms, Nebius said on Thursday, from Amsterdam, that it bought Inferize and folded the technology and the team into Token Factory, the managed service it uses to run models in production. Nebius says the point is to cut idle time on GPUs while a large model loads, while new copies spin up for a spike in demand, and while the model's weights change in the middle of a run. Chief technology officer Danila Shtan said the team's work will go past the first integration. Inferize co-founder and chief executive Guy Bortnikov said the company was built to remove the cost of keeping spare GPUs ready. Nebius says Inferize was founded in January 2026, had a prototype within three months, and that this is another layer after Eigen AI and Clarifai's technology in Token Factory. The page does not disclose the terms.