B4.run

Menu

Site

An agent framework, the way I'd build it.

Ridiculous speed.
Readable code.

Write the agent in TypeScript, give it tools, and set its limits.
You ship code you can actually read.

Runs on LangGraph.js. You keep the graph.

Get started →
What npm create b4-app scaffolds

npm create b4-app@latest my-agent

Created my-agent (basic template)

my-agent/

  • AGENTS.md, guide for your coding agent
  • b4.config.ts, runtime
  • src/app/hello/, the agent
    • index.ts, model + prompt
    • tools/greet.ts, a typed tool
    • evals/smoke.eval.ts, behavior check
  • test/agent.test.ts, passes, no API key

cd my-agent && npm install && npm test

Your first agent

An agent is a folder.

Every file in src/app/hello/ is a feature. The scaffold starts you with three. Add a file and the agent gets the tools that come with it.

my-agent/src/app/hello/

src/app/hello/index.tsscaffolded

One export is the agent.

index.ts picks the model and writes the system prompt. Every TypeScript file in tools/ becomes one of its tools.

Agents →
src/app/hello/index.ts
import { agent } from "@b4run/sdk" export default agent({  model: "gpt-5-mini",  systemPrompt: "You are a friendly assistant. Use greet to greet people by name.",})

src/app/hello/tools/greet.tsscaffolded

Your types are the tool schema.

The file name is the tool's name, and the JSDoc above it is the description. B4 reads the input type and writes the JSON schema the model sees.

Tool descriptions →

src/app/hello/tools/greet.ts

/** Greet someone by name. */export default async (input: { readonly name: string }) => {  return { message: `Hello, ${input.name}!` }}

What the model sees

{  "name": "greet",  "description": "Greet someone by name.",  "parameters": {    "type": "object",    "properties": {      "name": { "type": "string" }    },    "required": ["name"],    "additionalProperties": false  }}

src/app/hello/plan.md+ added

Add plan.md and it plans.

A markdown checklist next to index.ts turns planning on and seeds the todo list. The agent gets a writeTodos tool to keep the list current.

Planning →
src/app/hello/plan.md
- [ ] Understand the customer request- [ ] Check account context- [ ] Decide whether to answer or escalate- [ ] Write the final response

src/app/hello/memory.ts+ added

Add memory.ts and it remembers.

defineMemory declares the shape of the records and where they are scoped. The agent gets remember and recall tools typed from your schema.

Long-term memory →
src/app/hello/memory.ts
import { defineMemory } from "@b4run/sdk"import { z } from "zod" export default defineMemory({  kind: "semantic",  scope: ["workspace", "route"],  schema: z.object({    subject: z.string(),    predicate: z.string(),    value: z.string(),  }),})

src/app/hello/skills/greetings/SKILL.md+ added

Add a skill for long instructions.

The model sees each skill's description, not its body. It loads the full text only when it needs it, with readSkill.

Skills →
src/app/hello/skills/greetings/SKILL.md
---description: How to greet someone formally or for the time of day.--- Use this skill when the user asks for a formal greeting or one that fits the time of day. 1. Say "Good morning", "Good afternoon" or "Good evening" for the local time.2. For a formal greeting, use the person's title and family name.3. Keep the greeting to one sentence.

src/app/hello/subagents/translator/index.ts+ added

Add a subagent to hand work off.

A folder under subagents/ is a full agent with its own prompt and tools. The parent gets a task tool and uses each description to decide when to delegate.

Subagents →
src/app/hello/subagents/translator/index.ts
import { agent } from "@b4run/sdk" export default agent({  model: "gpt-5-mini",  description: "Translate a short greeting into the language the user asks for.",  systemPrompt: "You translate greetings. Reply with the translation only.",})

src/app/hello/evals/smoke.eval.tsscaffolded

The scaffold ships an eval.

smoke.eval.ts replays a scripted reply and scores it, so npm run eval needs no API key. Run npm run eval -- --live to try the real model.

Evals →
src/app/hello/evals/smoke.eval.ts
import { contains, defineEval } from "@b4run/evals"import { script } from "@b4run/testing" export default defineEval({  name: "greets by name",  dataset: [    {      name: "ada",      input: "Say hello to Ada",      fixtures: script().user("Say hello to Ada").replies("Hello, Ada!"),    },  ],  scorers: [contains("Hello", { threshold: 1 })],  threshold: 1,})

Guardrails

Four checks decide what a call can do.

Tool scope, permissions, the sandbox and delegation each answer one question. Pick a call this support agent might make and see which checks it meets.

Pick a call the support agent makes
  1. 1 · Tool scope passedreadFile comes with the workspace, and deny doesn't name it.
  2. 2 · Permission passedA path inside the workspace needs no approval.
  3. 3 · Sandbox containedIt reads the file inside the sandbox container.
  4. 4 · Delegation not involvedOnly a task call to a subagent reaches this check.

readFile runs inside the sandbox, and no one is asked.

src/app/support/index.ts

import { agent, type DelegationConstraintPredicate } from "@b4run/sdk"import translator from "./subagents/translator/index.js" // The translator gets one reply at a time, never a whole thread.const oneReply: DelegationConstraintPredicate = ({ input }) =>  input.length <= 2_000 || "Send the translator one reply at a time." export default agent({  model: "gpt-5-mini",  systemPrompt: "You answer support questions. Refund an order only when the policy allows it.",  tools: {    approve: ["refund"],    deny: ["deleteUser"],  },  subagents: { translator },  delegation: {    rules: { translator: { action: "constrain", predicate: oneReply } },  },})

Neither tools list names readFile, and a path inside the workspace needs no approval.Workspace permissions →

You keep the graph

Pick the shape per route.

A route's index.ts exports one of four shapes. Let the model drive with an agent, or write the steps yourself as a workflow, a LangGraph graph or a chain.

Route shape

src/app/hello/index.ts

import { agent } from "@b4run/sdk" export default agent({  model: "gpt-5-mini",  systemPrompt: "You are a friendly assistant. Use greet to greet people by name.",})

export default agent({ … })The model decides which tools to call, and when.Agents →

Test and ship

Test it offline, then pick where it runs.

The scaffold's test and eval answer from script() fixtures instead of a model, so they need no API key. One line of b4.config.ts picks what b4 build emits.

npm test and b4 eval

Recorded in a fresh scaffold, with no OPENAI_API_KEY set.

my-agent

npm test -- --reporter=verbose> test> vitest run --reporter=verbose  RUN  v4.1.11 my-agent  ✓ test/agent.test.ts > greets by name  Test Files  1 passed (1)      Tests  1 passed (1)npx b4 evalPASS greets by name › ada mean=1.00 [contains(Hello)=1.00]PASS greets by name mean=1.00

b4 build

node and langsmith are the defaults. Setting build.targets replaces them.

Deploy target

b4.config.ts

import { config } from "@b4run/cli" export default config({  build: { targets: ["node"] },})
npx b4 buildBuild complete: .b4/build  1 route(s) compiled  targets: node  wrote .b4/build/workspace.json  wrote .b4/build/modules.mjs  wrote .b4/build/server.mjs  wrote Dockerfile

The full B4 HTTP runtime as a Node server, with a Dockerfile.Node and Docker →

The last mile

The parts you'd write next are already here.

Twelve things an agent needs before real users reach it. Turn a tile over to see the file, config or command that handles it.

12 of 12 openedHandled.

  1. b4 typegen reads the tool's TypeScript types and its doc comment.

    src/app/hello/tools/greet.ts

    /** Greet someone by name. */
    export default async (input: { readonly name: string }) => {
    Tools →
  2. Every route streams over Agent Protocol and AG-UI.

    POST /threads/:thread_id/runs/stream
    POST /agui/{routeId}
    Agent Protocol →
  3. One middleware file runs before every route request.

    src/middleware.ts

    export default defineMiddleware(async (req) => {
    Middleware →
  4. One policy file decides who may create, read, change or delete each thread.

    src/thread-access.ts

    export default defineThreadAccess({
    Thread access →
  5. The run pauses, and a person answers once, always or deny.

    src/app/support/index.ts

    tools: { approve: ["refund"] },
    Permissions →
  6. The file tools and runBash run in a Docker container, or in a Kubernetes Pod with kubernetesSandbox.

    b4.config.ts

    provider: dockerSandbox({ scope: "my-agent", image: "node:24-slim" }),
    Sandbox →
  7. A memory.ts file gives the agent remember and recall tools.

    src/app/hello/memory.ts

    export default defineMemory({
      kind: "semantic",
    Long-term memory →
  8. A model call that fails with a rate limit, a server error or a network error, before anything has streamed, retries with backoff.

    src/app/support/index.ts

    retry: { maxAttempts: 5 },
    Retry →
  9. @b4run/postgres-storage gives you the checkpointer, the thread store and the permission store.

    b4.config.ts

    checkpointer: postgresCheckpointer({ pool }),
    threadsStore: createPostgresThreadsStore({ pool }),
    Postgres backend →
  10. script() fixtures stand in for the model, so npm test needs no API key.

    test/agent.test.ts

    fixtures: script().user("Say hello to Ada").replies("Hello, Ada!"),
    Testing agents →
  11. b4 eval replays each case offline, and b4 eval --record captures new fixtures from the model.

    src/app/hello/evals/smoke.eval.ts

    export default defineEval({
      name: "greets by name",
    Evals →
  12. b4 inspect opens the Inspector, a browser view of the app's long-term memory.

    b4 inspect
    Inspector →

Get started

Build your own agent.

Scaffold a new B4 app with one command:


Getting Started →