ANY REDIS CLIENT
redis-cli, redis-py, redis-rb — the client you already have is the SDK. Nothing to install but the server.
Any Redis client works unchanged. Pre-1.0, MIT licensed. Turn text — or any tensor a script builds — into vectors.
INSTALL
Fast, portable inference.
Download models by repo id.
High throughput, low latency.
Serve hot embeddings fast.
Load and switch models on the fly.
Your code around the model call.
Raw text, or any tensor
a script builds.
ONNX Runtime,
batched runs.
High-dimensional
vectors (e.g. 384D).
Speak Redis.
RESP2 or RESP3.
A real emb process answers every command below, over RESP, through a bridge that may ask it questions and may not change it: the same binary you would run locally, the sandbox's own models, no configuration, no scripts from this page, nothing written to shared state.
Run a row from the ledger, or type your own command and press Enter; ↑ and ↓ walk what you have already run. Each reply ends with how long the server took to answer it. Nothing below is replayed: the bytes are what this server sent, this time.
redis-cli, redis-py, redis-rb — the client you already have is the SDK. Nothing to install but the server.
redis-cli EMB minilm "hello world" # → \x7c\x8e\x80\xbd… 384 float32s × 4 bytes redis-cli -3 EMB minilm VALUES "hello world" # → dtype FLOAT shape [1 384] values [-0.1974, 0.1776, …]
EMB.MULTI queries several models in one command and answers MGET-style: a failed pair returns null in its slot, not the call.
The plain embed command is a fixed pipeline: model(input) → output. A script replaces the edges around that same call with your own code — preprocessing, postprocessing, and the reply itself.
model(input) output EMB
model(fn(input)) output EMB.EVAL
local enc = emb.tokenize.encode(KEYS[1], 512) local out = emb.run({ input_ids = { shape = {1, #enc.ids}, data = enc.ids }, attention_mask = { shape = {1, #enc.ids}, data = enc.mask }, }) local probs = emb.math.softmax(out.logits.data) local idx, score = emb.math.argmax(probs) -- labels arrive in the model's training order return { label = labels[idx], confidence = score }
Redis-shaped INFO and CONFIG, including live cache resizing — no restart, no new config format.
EMB.READY answers +OK or says why not — loading, draining, no models.
At the concurrency cap the server answers ERR busy instead of stacking unbounded work. Control commands still answer.
RSS, CPU, goroutines, connections and per-model latency, in the same EMB.STATS reply.
Req/s, p50/p95/p99, cache hit ratio and RSS — polled in one round trip, over the same protocol you already speak.
emb-top v0.4.0.pre5 · 127.0.0.1:16379 · uptime 29s · 4 models · poll 1s · 85.6 r/s · p95 12.2ms · ● connected ╭──────────────────────────────────────────────────────────────────────────────────────────────────────────────╮ │ req/s · models × recent polls ████████ │ │ e5-small ██████████████████████████████████████████████████████████████████████████████████████████████ │ │ minilm ██████████████████████████████████████████████████████████████████████████████████████████████ │ │ jina-small ██████████████████████████████████████████████████████████████████████████████████████████████ │ │ bge-small ██████████████████████████████████████████████████████████████████████████████████████████████ │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────╯ ╭────────────────────────────────────────────────────────╮╭────────────────────────────────────────────────────────╮ │ │ ╭─╮ ╭──╮│ ╰─╮ ││ │ ╭────────╯ │ │ 136│ │ ╰─╯ ╰╯ ╰╮ ││ 8.1ms│ ╭──╯ │ │ │ ╭───╯ ╰── ││ │ ╭───╯ │ │ 67.8│ │ ││ 4.1ms│ ╭╯ │ │ │ ╭╯ ││ │ │ │ │ 0.0│ ──╯ ││ 0µs │ ──╯ │ ╰────────────────────────────────────────────────────────╯╰────────────────────────────────────────────────────────╯ e5-small ▂██▆▆▅▃▄▃ 34.8 r/s 419 t/s p50 4.9ms p95 14.2ms err 0 dim 384 · mean · fp32 · batch 32/16384 workers 1 minilm ▂▇██▆▇▇▅▅▆▄▅▃▄▄▃▃▃▃▂▂ 26.9 r/s 245 t/s p50 2µs p95 5.5ms err 0 dim 384 · mean · fp32 · batch 32/16384 workers 1 jina-small ▃▆█▆▅▅▃▃▃▂▃▂▂ 15.9 r/s 145 t/s p50 2µs p95 5.7ms err 0 dim 512 · mean · fp32 · batch 32/16384 workers 1 bge-small ▁▇█▆▅▅▄▄▃▂▃▁▃▁▁ ▁ 8.0 r/s 72.7 t/s p50 2µs p95 9.5ms err 0 dim 384 · cls · fp32 · batch 32/16384 workers 1 cache 54.1% cpu 90.9% mem 1.1kMB █████████████████████████████████████ █████████████████████████████████████ █████████████████████████████████████ █████████████████████████████████████ █████████████████████████████████████ █████████████████████████████████████ conns 2 · active 0 · goroutines 25 · truncated texts/pairs/images 0/0/0 event e5-small · 1 texts · 5.9ms ✓ q quit · p pause · r reset · j/k scroll · ? help
Measured run · 2026-09-15 · emb-top v0.4.0.pre5 · 4 models · 127.0.0.1:16379
Busiest poll of the recorded run: 224 requests/s, 2.4k tokens/s, p95 latency 8.6ms, cache hit rate 37%, CPU 174%, 1050MB resident. Per model: minilm 47/s (avg 1.3ms, dim 384, mean, fp32); bge-small 18/s (avg 2.4ms, dim 384, cls, fp32); jina-small 44/s (avg 1.8ms, dim 512, mean, fp32); e5-small 115/s (avg 2.4ms, dim 384, mean, fp32).
EMBED
EVERYTHING
FURTHER
HIGHER
DIMENSIONS
BRIGHTER
APPLICATIONS