← News

CoreWeave Forge product diagram of The AI Loop: Run, Observe, Curate, Improve, Evaluate

30 Sep 2026

CoreWeave

CoreWeave launches Forge to close the model-and-agent improvement loop

CoreWeave said Wednesday it launched CoreWeave Forge, a development layer that runs the full AI loop — run, observe, curate, improve, evaluate, and repeat — in one connected environment open across models, frameworks, and other clouds, with MasterClass and Canva already building on it.

GPU clouds used to sell the rack. Forge is CoreWeave trying to own the loop after the rack: the place where a production trace feeds the next training run, so a model or an agent gets better without a scavenger hunt across five vendors.

On Wednesday, 30 September 2026, CoreWeave Inc. (Nasdaq: CRWV) announced CoreWeave Forge. The newsroom and the Business Wire release, as the company’s investor page carries it, are datelined San Francisco. CoreWeave calls itself The Essential Cloud for AI. Forge, it says, is a development layer for teams building and improving models and agents. It runs the entire AI loop — run, observe, curate, improve, evaluate, and repeat — in one connected environment, and it stays open across the models, frameworks, and other clouds a team already uses. MasterClass and Canva are already building on it. Those are the only customers the announcement names. The news was shared at Fully Connected, CoreWeave’s AI cloud conference. CoreWeave says the conference brings together more than 4,500 customers, partners, developers, and AI leaders. That headcount is the conference, not a Forge customer count. Those lines are CoreWeave’s.

The problem CoreWeave names is a split toolkit. Model and agent development, the release says, has been split across tools from separate vendors, none built to work together. For the people carrying a model or an agent through its whole life, that split has a cost. A production trace does not feed the next training run. A finished experiment does not inform the next evaluation. Every handoff loses a signal or costs more time. A trace, here, is the record of what a model or an agent actually did in use. An agent, here, is software that takes a next step, such as calling a tool, rather than only answering a question. Those sentences are CoreWeave’s.

What Forge pulls into one place. CoreWeave says it unifies Weights & Biases Models, post-training expertise from OpenPipe, and the open-source marimo notebook project with CoreWeave’s own services. The aim is one connected environment built for continuous improvement. Weights & Biases is the experiment tracker in that set. A notebook is a page where code, notes, and charts sit together. marimo is the open-source project those charts are built on. Post-training is the work after a model’s first big training run: teaching it further for a narrower job. The release says this layer closes the AI loop and helps a customer improve a model or an agent, whatever their level of expertise, while the team stays free to build with any model, framework, or cloud. Training runs, experiment tracking, evaluations, and agent traces live where the model or the agent actually runs. CoreWeave Registry keeps every checkpoint and agent configuration versioned. A checkpoint is a saved copy of a model at one moment in the work. Those lines are CoreWeave’s.

Chen Goldberg, executive vice president of product and engineering at CoreWeave, said that as more people build with AI, they put models to work with their own data and workflows, and that is where the gaps between what a model can do and how the whole system performs become clear. He said engineering teams need to understand those gaps, pick out the signals that matter, improve the next version, and measure whether the change worked under real operating conditions. He said all of that has to function as a single, connected system. He said Forge brings these capabilities onto one platform, so what a business learns from running AI becomes part of how engineers improve it. That quotation is in the CoreWeave release.

CoreWeave says Forge is available today, with new and expanded pieces across the loop. CoreWeave ARIA is now generally available. ARIA is a coding agent. It analyzes large-scale experiment data and agent observability data, surfaces what drove a change, and proposes the next experiments worth running and the right agent worth running, with the evidence to support the proposal. Observability, here, means being able to see what the system did. Weights & Biases Models, CoreWeave says, helps a team track and compare tens of thousands of experiments and millions of metrics, with visualizations and automated workflows. CoreWeave Notebooks lives in the environment the work runs in, so a prototype carries into training, evaluation, and production instead of dying in a rewrite. Custom visualizations are built on marimo. Those lines are CoreWeave’s.

CoreWeave Agent Lens is a new observability tool for agents in production, meant to keep improving them. CoreWeave says it analyzes tens of millions of traces and turns them into insights that drive fixes. On a benchmark against a general-purpose frontier LLM, CoreWeave says Agent Lens detects 20 percent more critical failures and fixes issues at one-tenth the cost. A frontier LLM, in that sentence, is a large general model at the top of the market. Twenty percent more means, on that test, it caught a fifth more of the serious failures than the general model did. One-tenth the cost means the fix cost about 10 cents for each dollar the comparison spent. CoreWeave does not name the model, and the announcement does not publish the benchmark’s method. Those figures are CoreWeave’s. CoreWeave Sandboxes is now generally available. It is a fresh, isolated environment for every run, whether an agent is using a tool, a team is doing reinforcement learning, or a team is evaluating a model. CoreWeave says it is already connected to the rest of the loop, so a team does not stand up a separate sandbox to use it. Reinforcement learning, shortened to RL, is training by reward: the system tries something, gets a score, and adjusts. CoreWeave Registry versions every model checkpoint and agent configuration in open, portable formats, with a clear lineage, so the next attempt can start from any point in a team’s history, on any stack. Those lines are CoreWeave’s.

CoreWeave Post-Training, the release says, covers serverless reinforcement learning, serverless supervised fine-tuning, and model distillation, so the improvement can go as deep as the weights. Serverless, here, means the team does not stand up its own training machines. Supervised fine-tuning is further training on examples a person labeled. Distillation carries what a larger model learned into a smaller, faster one. CoreWeave says serverless RL trains 1.4 times faster at 40 percent lower cost than a self-managed setup. One point four times faster means the same work finishes in a bit under three-quarters of the time, on that comparison. Forty percent lower cost means the bill is six-tenths of the self-managed bill. The announcement does not name the self-managed setup. Those figures are CoreWeave’s. The product page adds that this serverless RL service trains agents against reward signals with nothing to provision, and that the automatic scoring does not require hand-labeled data or a reward function written by hand. CoreWeave Inference is now part of Forge. It offers a playground and deployment access to the latest open-weight models. Open-weight means the model’s weights are published so others can run them. Dedicated Inference is adding RL Rollouts in preview. CoreWeave says that feature hot-loads a checkpoint into a live deployment, so a training loop can keep running without a redeploy. Preview means it is not the general release. Those lines are CoreWeave’s.

The release says Forge comes in Free, Pro, and Enterprise editions, so a team can start without a procurement cycle and grow across the loop. It points readers to a signup for Forge Pro and a free 30-day trial. On the product page, Forge Free is listed at $0 a month, for personal development of AI applications and models. Forge Pro starts at $60 a month, billed monthly, with a 30-day free trial. The page says Pro is for early-stage teams of fewer than 50 employees, and that a customer who passes that line has to move to Forge Enterprise. Enterprise is a custom plan. Those prices are the product page’s. The page also says the traces a team flags become the training data, and the evaluations a team uses to gate a release, and that the assets from a production run stay portable and stay the team’s.

The release also quotes Nick Patience, vice president and practice lead for AI platforms at the Futurum Group. He said every platform decision in this market has involved weighing how much choice a team is willing to give up for a connected toolchain, and that Forge is built on the premise that teams should not have to make that calculation. He said closing the loop from production back into training, in one environment, while staying open across models, frameworks, and clouds, is a harder engineering problem than either half alone. That quotation is in CoreWeave’s release. It is Patience’s view of the product, as CoreWeave prints it. It is not a test of the 20 percent figure or the 1.4 times figure.

The wire’s about box, as the investor page prints it, says CoreWeave was established in 2017 and completed its Nasdaq listing in March 2025. It says the company is trusted by 9 of the 10 leading foundation model providers. That count is CoreWeave’s description of the company. The same release says CoreWeave has posted record MLPerf results in inference and training, and that it is the only AI cloud to earn SemiAnalysis’s top Platinum ClusterMAX ranking three times in a row. MLPerf is a shared speed test for training and serving models. ClusterMAX is SemiAnalysis’s rating of AI clouds. Those lines are CoreWeave’s account of its infrastructure. They are not a score for the Agent Lens or serverless RL comparisons above. The media contact on the release is press@coreweave.com.

The picture is CoreWeave’s product diagram of The AI Loop. Five words sit outside the rings: Run at the top, Observe on the right, Curate toward the lower right, Improve toward the lower left, and Evaluate on the left. The black center reads The AI Loop, over an infinity mark. Hexagons, triangles, squares, and circles mark points on the rings. Blue tick marks cluster on the Evaluate side and the Observe side. It is the product diagram. It is not a photograph of a machine room, and it is not a wire dateline.

In plain terms, CoreWeave said on Wednesday that Forge is the layer where a team runs a model or an agent, watches what it did, picks the traces worth keeping, trains the next version, and checks whether the change worked, without splitting that loop across separate logins. The environment is meant to stay open to other models, frameworks, and clouds. MasterClass and Canva are the customers named at launch. ARIA and Sandboxes are generally available. RL Rollouts is a preview. Free, Pro, and Enterprise are the editions, with Pro’s trial and the $60 starting price on the product page. The 20 percent, the one-tenth cost, the 1.4 times speed, and the 40 percent cost cut are CoreWeave’s comparisons.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources