API
OpenAI-compatible, fully documented.
Standard chat completions with streaming, authenticated with a Cloud key from the console. Any client that takes a base URL works unchanged.
https://inference.runanywhere.ai/v1Copy this into your agent to get started
API
Standard chat completions with streaming, authenticated with a Cloud key from the console. Any client that takes a base URL works unchanged.
https://inference.runanywhere.ai/v1The rule
Cloud and local are explicit choices. Nothing moves a request between them silently.
Questions
It's how you reach RunAnywhere's biggest models without leaving your own workflow. Sign in once with Wally, tell it to run in the cloud, and you're talking to an OpenAI-compatible endpoint. Your local SDKs and accelerators still handle everything that fits on your device, and Wally never moves a request between the two on its own.
The live lineup and its prices sit in the Console rather than here, since they change as we add capacity. Every model bills per million tokens and speaks the same streaming chat-completions format, so switching between them is a config change, not a rewrite.
Sign in for the current catalogYes. Point your existing OpenAI client, curl command, or SDK at Wally's base URL and model ID, and it works. Nothing new to install or learn.
Most people never touch a key. Run wally account login, approve it in your browser with Google or GitHub, and the credential comes back to Wally on its own. A raw key only matters if you're calling the endpoint directly with curl or another SDK outside Wally.
Production Wally uses zero data retention for prompt and completion bodies: we do not store the text you send or the text the model returns. Usage and billing records are separate. Read the privacy page for the full policy.