← News

Corvex launches Token Factory for open-weight inference with zero retention

Corvex said Wednesday it launched Token Factory, a serverless inference platform that serves open-weight models through an API with zero data retention by default and no customer-managed GPU clusters.

Teams want DeepSeek-class open weights without standing up GPUs or shipping prompts to someone else’s trainer. Token Factory’s pitch is API-shaped inference on Corvex-run hardware with zero retention by default — a security-first on-ramp, not another idle-cluster bill.

On Wednesday, 30 September 2026, Corvex, Inc. announced Corvex Token Factory. The company trades on Nasdaq as MOVE. It calls itself an engineering-led AI computing company. Token Factory is its serverless inference platform for open-weight AI models, reached through an API. Serverless, here, means the customer sends a request and Corvex runs the computers. An API is the address a program uses to send that request. Open-weight means the model’s weights, the numbers that make it answer, are published so others can run them, rather than locked inside one vendor’s closed product. The PR Newswire page carries the release at 9:40 a.m. Eastern. The dateline is Arlington, Virginia. The source line is Corvex. Those lines are Corvex’s.

What a customer gets. Corvex says Token Factory gives access to open-weight models through an API, and the customer does not have to deploy or operate GPU clusters. A GPU is the chip that does the heavy math for a model. A cluster is a group of those machines. The first models the release names are GLM 5.3 from Z.ai and DeepSeek V4 Flash 0731. Corvex says the launch is for developers and enterprise technology teams that want to use AI as a service. Those lines are Corvex’s.

Seth Demsey, co-chief executive and co-founder, said companies putting AI to work on source code, customer records, and internal documents need to know where that information goes and what happens to it. He said that with Token Factory, by default, Corvex does not store end-user prompts or responses and does not train on any user data. A prompt is the text a person or a program sends in. He said those controls sit in a service developers can use with the coding tools and API calls they already have, “so that they can safely enjoy the power of leading open-weights models at a fraction of the cost of closed-source alternatives.” That quotation is his, on the wire. The release does not print that fraction.

Where a request runs. Inference is the step that turns a request into the model’s reply. Corvex says it runs every request on hardware it manages, and that it operates the models itself. It says requests are never routed to third-party inference providers and are never forwarded to the companies that built the models. A third-party provider, in that sentence, is another company running the model for a fee. Those lines are Corvex’s.

Four principles the release lists at launch. Assurance: zero data retention is the default. Prompts and responses are processed in memory, never logged or stored, and never used to train models. Corvex keeps limited operational metadata for security, service operations, and billing, and that metadata does not include the prompt or the response. Operational metadata, here, is the record that a request happened and what it takes to run and bill it, not the words inside it. Reliability: Corvex says the service is built to hold up under sustained workloads, so it can sit inside a production workflow. Performance: Corvex says it tunes the stack, balancing how fast the first words come back and how much work it can carry under load, for interactive applications and for agent workflows. An agent workflow is a program that takes several steps, not one chat reply. Simplicity: OpenAI-compatible and Anthropic-compatible APIs let a developer point an existing client at Token Factory by changing the base URL, the API key, and the model name. The base URL is the server address the client calls. Corvex manages the machines underneath. The release says which features work depends on the client and on the model. Those lines are Corvex’s. The reliability and performance lines are the company’s description of the service. They are not a published uptime figure or a speed table.

Security paperwork, and the bill, as the release states them. Corvex says it is SOC 2 Type II certified. SOC 2 Type II is an outside audit of how a company protects customer data across a stretch of time, not a one-day checklist. The company also says it signs a business associate agreement with customers that are subject to HIPAA. HIPAA is the U.S. health-privacy law. That agreement is the contract a health organization uses when a vendor handles patient information. Corvex says the security controls are documented in the Corvex Trust Center. Customers pay for the input tokens and the output tokens they use, and they do not pay an idle GPU charge. A token, here, is a chunk of text, often a word or part of a word. Input is the text sent in. Output is the reply. Idle means a fee for leaving a GPU switched on when no request is running. The press release does not print a per-token price. Those lines are Corvex’s.

What the company says comes next. Token Factory, Corvex says, is a starting point for a deeper customer relationship. The roadmap names dedicated enterprise deployments and secure private inference environments, a path to more services as usage grows. A roadmap is the plan the company is stating. Those lines are forward-looking. They are not a switch a customer can flip from this announcement. The release says the launch follows a closed alpha, a private test with a limited set of users. A customer can register at no cost at tokenfactory.corvex.cloud/app and get an API key. The product documentation is at docs.corvex.cloud. Those lines are Corvex’s.

What the documentation says on the same day. The Get Started page calls Token Factory a managed inference service for open-weight models. A developer uses an OpenAI-compatible or Anthropic-compatible client, chooses a model from the live catalog, and sends requests without deploying or operating a GPU cluster. The page says Token Factory publishes a separate input rate and output rate for each model. The dashboard is where a customer manages API keys, tries a prompt in a playground, and reads usage. A notice on the page says Token Factory is currently in a free Alpha, and that inference is not charged during the Alpha. The same notice says the rates on the pricing page are what Corvex uses to calculate the equivalent value of that Alpha usage. Those lines are on the docs. The press release describes a launch after a closed alpha, with registration at no cost. The docs, read the same day, still call the service a free Alpha.

How the docs describe retention, beside the press release. The data-security page says zero data retention means prompt and generation data are not logged or stored permanently by default, unless the user explicitly opts in. Generation data is the model’s reply. That data sits in volatile memory for the length of the request. Volatile memory is temporary. It does not stay as a file on disk after the request ends. If prompt caching is turned on, the page says some of the prompt, and the matching key-value cache, can remain in that volatile memory for several minutes. A cache, here, is a short-term copy kept so a repeated request can start faster. The page says Corvex retains operational metadata, such as token counts, in order to deliver the service. It points readers to the Corvex Trust Center for security documentation and audit reports. Those lines are on the docs.

The two models, as the model page lists them. Both support the OpenAI-compatible chat-completions API and the Anthropic-compatible messages API. GLM 5.3 has a context window of 393,216 tokens. The page also writes that as 384K. A context window is how much text one request can hold. DeepSeek V4 Flash 0731 has a context window of 1,048,576 tokens, which the page writes as 1M. The page says both models take text, and both support tool calls, a JSON mode, and structured output. A tool call lets the model ask a program to run a step. JSON is a common format for structured data. Structured output is a reply forced into a set shape, such as fields in a form. The page says neither model takes an image as input. Those lines are on the model page.

The about box describes the company. It says Corvex is an AI cloud computing company for GPU-accelerated infrastructure, and a publicly traded AI compute platform. The products it names are AI Factories and GPU Clusters, the Assured AI confidential-computing platform, and Corvex Token Factory. Confidential computing, in that product name, is Corvex’s label for a platform meant to keep a workload private while it runs. Those are product names. The about box does not name a customer for Token Factory. The media contact on the release is Chris Donahoe at Stillpoint, corvex.media@stillpointglobaladvisors.com. The release says statements about product capabilities, customer deployment, and growth plans are forward-looking and can turn out differently. Those lines are Corvex’s.

In plain terms, Corvex said on Wednesday that a developer can call GLM 5.3 or DeepSeek V4 Flash 0731 through an API Corvex runs, without standing up a GPU cluster and without a charge for leaving those chips idle. By default, Corvex says it does not keep the prompt or the reply and does not train on them. The request stays on hardware Corvex operates. An existing OpenAI or Anthropic client reconnects by changing the address, the key, and the model name. The same day’s docs still mark the service as a free Alpha and say inference is not charged while that Alpha lasts. A key is available at no cost after the closed alpha the release describes.

The picture is the Get Started page in the Corvex Token Factory documentation. The sidebar has Corvex Token Factory selected under Get Started, then Quickstart and Authentication. Under Models & Pricing it lists Models, Pricing, GLM 5.3, and DeepSeek V4 Flash 0731. The page headline is Corvex Token Factory. The line under it says the service offers open-weight model APIs for production inference, with usage-based pricing and no GPU infrastructure to manage. A notice says Token Factory is currently in a free Alpha and that inference is not charged during the Alpha. The data-security section points to zero data retention for inference and says Token Factory retains operational metadata. It is a screenshot of that docs page. It does not print a date.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources