Skip to main content
Corvex Token Factory is a managed inference service for open-weight models. Use OpenAI- or Anthropic-compatible clients, choose a model from the live catalog, and send requests without deploying or operating GPU infrastructure. Token Factory publishes separate input and output rates for each model. The dashboard provides API-key management, a playground, and usage reporting.
Token Factory is currently in a free Alpha. Inference is not charged during the Alpha; see Pricing for the published rates used to calculate the equivalent value of Alpha usage.

Data security

Read about zero data retention for inference. Token Factory retains operational metadata; the Data Security page explains what is retained and how volatile prompt caching works.

API compatibility

OpenAI-compatible API

Use Chat Completions, the Responses API, model listing, and supported OpenAI SDK features.

Anthropic-compatible API

Use the Messages API, token counting, the Anthropic SDK, and Claude Code.

Models and pricing

Compare the models currently served by Token Factory, including context windows, capabilities, and per-token rates.

Usage API

Query account-scoped usage from GET /api/v1/usage with your API key.
Compatibility applies to the documented endpoints and fields. Model features such as image input, tool use, context length, and structured output vary by model; check the model catalog before relying on them.

Start here

Send your first request

Create an API key and call POST /v1/chat/completions.

Protect your API key

Choose the correct authentication header and store keys safely.

Connect a client or tool

Configure an SDK or coding agent.

Read the API reference

Review endpoint schemas generated from the public OpenAPI 3.1 specification.
Last modified on September 22, 2026