emb
Menu

BYTES IN.
VECTORS OUT.

OPEN SOURCE
EMBEDDING SERVER
FOR THE STACK
YOU ALREADY RUN.

RESP3
MULTI-MODEL
ONNX

A fast embedding server that speaks the Redis protocol.

Any Redis client works unchanged. Pre-1.0, MIT licensed. Turn text — or any tensor a script builds — into vectors.

EVERYTHING SERVER-SIDE

  1. ONNX Runtime

    Fast, portable inference.

  2. Hugging Face

    Download models by repo id.

  3. Smart Batching

    High throughput, low latency.

  4. LRU Cache

    Serve hot embeddings fast.

  5. Multi-Model

    Load and switch models on the fly.

  6. Lua Scripting

    Your code around the model call.

Architecture

  1. 01

    INPUT

    Raw text, or any tensor
    a script builds.

  2. 02

    INFERENCE

    ONNX Runtime,
    batched runs.

  3. 03

    EMBEDDINGS

    High-dimensional
    vectors (e.g. 384D).

  4. 04

    SERVE

    Speak Redis.
    RESP2 or RESP3.

THE SERVER IS THE DEMO

A real emb process answers every command below, over RESP, through a bridge that may ask it questions and may not change it: the same binary you would run locally, the sandbox's own models, no configuration, no scripts from this page, nothing written to shared state.

Run a row from the ledger, or type your own command and press Enter; ↑ and ↓ walk what you have already run. Each reply ends with how long the server took to answer it. Nothing below is replayed: the bytes are what this server sent, this time.

ANSWERS
emb v0.4.0 · minilm · sst2
RUNS
EMB · EMB.MULTI · the read commands
REFUSES
config · raw Lua · images · writes
BOUNDS
2 commands/s · 8 texts · 2 KB per text
MEASURED
the server's time, under each reply
RESETS
the sandbox may restart between visits

Console

IDLE

SANDBOX · MAY RESET

THE PROTOCOL IS THE INTEGRATION

ANY REDIS CLIENT

redis-cli, redis-py, redis-rb — the client you already have is the SDK. Nothing to install but the server.

redis-cli EMB minilm "hello world"
# → \x7c\x8e\x80\xbd…           384 float32s × 4 bytes

redis-cli -3 EMB minilm VALUES "hello world"
# → dtype FLOAT  shape [1 384]  values [-0.1974, 0.1776, …]
bytes by default; VALUES for clients that cannot unpack float32

MANY MODELS, ONE ROUND TRIP

EMB.MULTI queries several models in one command and answers MGET-style: a failed pair returns null in its slot, not the call.

WIRE
RESP2 · HELLO 3 for RESP3
REPLY
BLOB bytes · VALUES envelope
ACCESS
AUTH · TLS
CLIENTS
any Redis client, unchanged

THE MODEL, AS A FUNCTION

The plain embed command is a fixed pipeline: model(input) → output. A script replaces the edges around that same call with your own code — preprocessing, postprocessing, and the reply itself.

model(input) output EMB

model(fn(input)) output EMB.EVAL

local enc = emb.tokenize.encode(KEYS[1], 512)
local out = emb.run({
  input_ids      = { shape = {1, #enc.ids}, data = enc.ids },
  attention_mask = { shape = {1, #enc.ids}, data = enc.mask },
})
local probs = emb.math.softmax(out.logits.data)
local idx, score = emb.math.argmax(probs)
-- labels arrive in the model's training order
return { label = labels[idx], confidence = score }
examples/scripts/snippets/sst2.lua — a raw-logits text model becomes a labelled classifier
  1. sst2 classify
  2. qa extractive QA
  3. rerank sigmoid scores
  4. gliner2 span extraction
  5. siglip2 image + text
CALLS
EMB.EVAL · EMB.EVSHA
CACHE
by SHA1, per model
BOOT
preload from config
SAFETY
pure compute · no os, io, require

OPERATIONS, REDIS-SHAPED

READ AND TUNE, LIVE

Redis-shaped INFO and CONFIG, including live cache resizing — no restart, no new config format.

A REAL READINESS PROBE

EMB.READY answers +OK or says why not — loading, draining, no models.

BACKPRESSURE, NOT QUEUING

At the concurrency cap the server answers ERR busy instead of stacking unbounded work. Control commands still answer.

WHAT IT DOES TO THE MACHINE

RSS, CPU, goroutines, connections and per-model latency, in the same EMB.STATS reply.

INFO
server · cache · stats · memory · cpu
CONFIG
GET / SET, live
MONITOR
recent requests, no text payloads
EMB.READY
health probe

emb-top

Req/s, p50/p95/p99, cache hit ratio and RSS — polled in one round trip, over the same protocol you already speak.

emb-top v0.4.0.pre5 · 127.0.0.1:16379 · uptime 29s · 4 models · poll 1s · 85.6 r/s · p95 12.2ms · ● connected
╭──────────────────────────────────────────────────────────────────────────────────────────────────────────────╮
│  req/s · models × recent polls  ████████                                                                     │
│ e5-small      ██████████████████████████████████████████████████████████████████████████████████████████████ │
│ minilm        ██████████████████████████████████████████████████████████████████████████████████████████████ │
│ jina-small    ██████████████████████████████████████████████████████████████████████████████████████████████ │
│ bge-small     ██████████████████████████████████████████████████████████████████████████████████████████████ │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
╭────────────────────────────────────────────────────────╮╭────────────────────────────────────────────────────────╮
│     │                                 ╭─╮ ╭──╮│ ╰─╮    ││       │                                  ╭────────╯    │
│  136│                                 │ ╰─╯  ╰╯   ╰╮   ││  8.1ms│                               ╭──╯             │
│     │                             ╭───╯            ╰── ││       │                           ╭───╯                │
│ 67.8│                             │                    ││  4.1ms│                          ╭╯                    │
│     │                            ╭╯                    ││       │                          │                     │
│  0.0│                          ──╯                     ││   0µs │                        ──╯                     │
╰────────────────────────────────────────────────────────╯╰────────────────────────────────────────────────────────╯
e5-small                       ▂██▆▆▅▃▄▃ 34.8 r/s 419 t/s p50 4.9ms p95 14.2ms err 0
   dim 384 · mean · fp32 · batch 32/16384 workers 1
minilm             ▂▇██▆▇▇▅▅▆▄▅▃▄▄▃▃▃▃▂▂ 26.9 r/s 245 t/s p50 2µs p95 5.5ms err 0
   dim 384 · mean · fp32 · batch 32/16384 workers 1
jina-small                 ▃▆█▆▅▅▃▃▃▂▃▂▂ 15.9 r/s 145 t/s p50 2µs p95 5.7ms err 0
   dim 512 · mean · fp32 · batch 32/16384 workers 1
bge-small              ▁▇█▆▅▅▄▄▃▂▃▁▃▁▁ ▁ 8.0 r/s 72.7 t/s p50 2µs p95 9.5ms err 0
   dim 384 · cls · fp32 · batch 32/16384 workers 1
cache   54.1%                          cpu     90.9%                          mem     1.1kMB
█████████████████████████████████████  █████████████████████████████████████  █████████████████████████████████████
█████████████████████████████████████  █████████████████████████████████████  █████████████████████████████████████
  conns 2 · active 0 · goroutines 25 · truncated texts/pairs/images 0/0/0
  event e5-small · 1 texts · 5.9ms ✓
q quit · p pause · r reset · j/k scroll · ? help

Measured run · 2026-09-15 · emb-top v0.4.0.pre5 · 4 models · 127.0.0.1:16379

Busiest poll of the recorded run: 224 requests/s, 2.4k tokens/s, p95 latency 8.6ms, cache hit rate 37%, CPU 174%, 1050MB resident. Per model: minilm 47/s (avg 1.3ms, dim 384, mean, fp32); bge-small 18/s (avg 2.4ms, dim 384, cls, fp32); jina-small 44/s (avg 1.8ms, dim 512, mean, fp32); e5-small 115/s (avg 2.4ms, dim 384, mean, fp32).

Scale

EMBED
EVERYTHING
FURTHER

HIGHER
DIMENSIONS
BRIGHTER
APPLICATIONS

Image