Wally

Wally by RunAnywhere

Open models for work that outgrows a laptop.

Copy this into your agent to get started

API

OpenAI-compatible, fully documented.

Standard chat completions with streaming, authenticated with a Cloud key from the console. Any client that takes a base URL works unchanged.

https://inference.runanywhere.ai/v1
Read the docs

The rule

Hosted is a flag, never a fallback.

Cloud and local are explicit choices. Nothing moves a request between them silently.

Questions

Frequently asked.

What is Wally?

It's how you reach RunAnywhere's biggest models without leaving your own workflow. Sign in once with Wally, tell it to run in the cloud, and you're talking to an OpenAI-compatible endpoint. Your local SDKs and accelerators still handle everything that fits on your device, and Wally never moves a request between the two on its own.

Which models are served?

The live lineup and its prices sit in the Console rather than here, since they change as we add capacity. Every model bills per million tokens and speaks the same streaming chat-completions format, so switching between them is a config change, not a rewrite.

Sign in for the current catalog
Is the API OpenAI-compatible?

Yes. Point your existing OpenAI client, curl command, or SDK at Wally's base URL and model ID, and it works. Nothing new to install or learn.

Do I need an API key, or does wally account login handle it?

Most people never touch a key. Run wally account login, approve it in your browser with Google or GitHub, and the credential comes back to Wally on its own. A raw key only matters if you're calling the endpoint directly with curl or another SDK outside Wally.

What happens to my prompts?

Production Wally uses zero data retention for prompt and completion bodies: we do not store the text you send or the text the model returns. Usage and billing records are separate. Read the privacy page for the full policy.