System One Skills for Codex, Claude Code, and Devin

Skills that save your agent’s context for the real work.

Coding agents burn minutes and context on noisy logs and questions with short answers. System One Skills is our skill pack for that. Its first skill, system-one-verify, runs a noisy test or build once, hands your agent a short result, and keeps the full log on disk. Across 563 real validation runs it returned 35% less text, with no model and no Sys1 required. When your agent also needs fast decisions, Sys1 adds review and completion-check skills backed by Jev, TypeSafe’s fast hosted decision model, or a local model.

v0.17.0 Sys1 · v0.4.1 System One Skills · MIT · Open source. Review and verify are experimental and advisory. Hosted Jev is opt-in; an experimental local model runs on your machine.

A small decision, made explicitIllustrative example
Your context

The check failed because config.json is missing.

Image
Sys1

You choose Jev, local Qwen, or your server.

Which category fits: configuration, code, or network?

Answer excerpt"choice": "configuration"

One of your named options, with probabilities in the full response.

How Sys1 answers the small questions behind its skills: three question forms, shown with authored examples. This demo makes no model calls.

Three skills your agent runs itself.

Install a project skill for Codex, Claude Code, or Devin. The first works on its own. The other two ask Jev or a local model the small questions, so your agent saves its own reasoning for the findings worth a closer look.

Keep the full log out of contextsystem-one-verify · System One Skills

Run a noisy test or build once, hand your agent a short result, and keep the full log on disk. Across 563 real runs it returned 35% less text. It needs neither Sys1 nor a model.

Use compact check output →
Review the changesys1-review

Check a batch of Git changes against your repository’s rules. The preview is free, an unchanged batch costs no new requests for 24 hours, and findings you have already seen stay out of the output. Turn a real mistake into a new rule. Two bundled rules cover empty catch blocks and removed test assertions in JavaScript and TypeScript.

Set up agent review →
Check the completion claimsys1-verify

Before your agent says “done,” compare its message with Git, pull requests, and live pages. A claimed push with nothing pushed shows up as a contradiction. Missing evidence stays unverifiable. Works with a message from any coding agent.

Verify a completion claim →

Review scores rank candidates; they are not calibrated defect probabilities. Keep your repository’s tests and review.

Benchmarks against other skill packs are in progress.

People reasonably ask how System One Skills compares with gstack, pstack, Superpowers, and Anthropic’s skills. We have not run that comparison yet, so we publish no head-to-head numbers. Those packs cover planning, review, shipping, and document workflows. Ours starts with check output and completion claims, so the fair test is the same agent tasks with and without each pack.

QuestionStatusWhere to look
Compact check outputsystem-one-verify

Measured. 563 replayed validation outputs from one developer’s Codex and Devin sessions returned 35% less text.

Read the study and its limits

Cost per decisionSys1 with Jev

Cost per decision and output size are measured. Whole-task token savings are not yet.

See our evaluations

Skill packs head to headgstack, pstack, Superpowers, Anthropic

Planned, not run. Fixed tasks, pinned pack versions, and the same agent and model, scored on task success, tokens, time, and false “done” claims.

Read the benchmark plan

We will publish results, including the ones where another pack wins, with the task set and raw logs. Until then, treat any comparison you read here as a question, not a result.

Meet Sys1.

Why a fast model beside a smart one saves tokens and time, how ALGAL fits in, and what our first trial did and did not show.

Read the launch story
Introducing Sys1 · 52 secondsRead the film transcriptOriginal motion graphics and instrumental score, rendered for Sys1 with Slopcamera.

Three parts, one system.

People think at two speeds: slow and deliberate for hard problems, fast and instinctive for everything else. Coding agents only have the slow speed. Sys1 adds the fast one, and ALGAL gives the system a way to keep what works.

PartWhat it doesWhat you get
Your coding agentThe large language model

Plans, writes code, and reasons through the problems that need it.

Its context and budget stay on real work instead of routine checks.

JevThe fast decision model

Answers yes/no, choice, and score questions with probabilities. Sys1 routes each question and checks every answer.

About $0.0004 and 0.4 seconds a decision, from TypeSafe’s published workflow figures. Output tokens are free.

ALGALThe language that remembers

A programming language for agent programs that wait for approval, resume after a crash, and replay from receipts.

A procedure that proves itself on tested cases becomes a reusable program, so the harness improves with use.

Sys1 already learns on a small scale: draft a rule from a mistake your agent made, and every later review checks for it. Sys1 and ALGAL are separate open-source projects today that call the same Jev decision API. We have measured cost per decision and output size, not whole-task token savings yet.

Put fast decisions in your own code.

System One models answer small, structured questions with probabilities. Jev is TypeSafe’s hosted model for this approach. Ask a yes/no, choice, or score question and get numbers your code can act on, not text to parse.

Call POST /v1/systemone through a loopback gateway or use the Node and Bun client. Choose hosted Jev, experimental local Qwen, or an OpenJev server you run. Local answers approximate Jev; they are not calibrated Jev output. Understand System One and Jev.

  1. Supply contextSend state and the questions your application needs answered.
  2. Choose a backendPin Jev, an experimental local model, or a compatible server.
  3. Validate the responseSys1 checks the answer format and reports malformed responses as errors.
  4. Act in your codeYou decide the thresholds, permissions, and next action.

Test the decision on your task.

A September 28, 2026 trial submitted 48 synthetic examples once each to Jev 1.13.0. Claim support is the strongest next candidate for a larger evaluation; failure triage tied its baseline.

Read the results and limitations →
WorkflowJev · 16 cases eachBaseline
Claim support16 / 16No errors6 / 16
Failure triage16 / 16No errors16 / 16
Excerpt relevance13 / 163 errors6 / 16
Eight development and eight separately authored screening cases per workflow. The baselines are simple error matching, word overlap, and a weak literal-match heuristic. These small tests do not establish production accuracy. Method and all results.

Three ways to call Sys1.

All three take the same request and return the same answer format.

Node & Bun client
@hraness/sys1/client

Typed requests, validated responses, and cancellation. Installs without the optional native runtime.

Embedded Bun router
createRouter()

Routing and local models run inside your Bun app. Call dispose() when you finish.

Local HTTP daemon
POST /v1/systemone

Listens on loopback only (127.0.0.1 by default), for Rust, Python, and other languages. sys1 up, sys1 status, and sys1 down manage it.

Inside your applicationTypeScript
import { createClient } from "@hraness/sys1/client";

const sys1 = createClient();
const { response, metadata } = await sys1.evaluate({
  state: "The build failed after a dependency update.",
  questions: {
    next: {
      type: "choice",
      criteria: {
        repair: "Fix the build",
        continue: "Continue work"
      }
    }
  }
});

console.log(response.answers.next);
console.log(metadata.backend);
Read the integration guide

The client also works in Node without the optional native runtime. The CLI and embedded router require Bun. Choose an integration or inspect the evaluation records.

What Sys1 connects to.

Every backend answers in the same format, but not equally well. Accuracy and calibration differ by model, so test each one on your task before you reuse thresholds or change defaults. Compare the models.

BackendWhat it isHow Sys1 uses it
Qwen3 1.7BLocal model, experimental

The default local model, run on your hardware.

sys1 setup downloads verified weights and turns it on. It stays selected when you install another model.

JevHosted, opt-in

TypeSafe’s hosted System One decision model.

Off until you run sys1 jev enable with your own TypeSafe API key. Enabling it selects hosted-only routing, so requests do not fall back to a local model.

OpenJevSelf-hosted server

An open implementation of a Jev-compatible decision API.

Register a server you run and test it with sys1 backend check. Image, chat, and advanced sampling extensions are not supported.

llama.cppLocal runtime

Native local inference, through node-llama-cpp.

Sys1 loads the model, cancels requests, and maps first-token probabilities to answers.

Install Sys1.

Install the release from GitHub, inspect your setup, then add the project skill for your agent. Choose hosted Jev or a local model in the setup guide.

Install Sys1
npm install --global --allow-scripts=node-llama-cpp \
  https://github.com/hraness/sys1/releases/download/v0.17.0/hraness-sys1-0.17.0.tgz
sys1 doctor
sys1 review setup codex

The CLI requires Bun 1.3.14 or newer on macOS, Linux, or Windows. --allow-scripts=node-llama-cpp permits the native inference package’s install script. Installation downloads no model weights and activates no hosted backend. Use setup claude-code or setup devin for those agents.

Before you begin.

How is Sys1 different from CodeRabbit, Cursor Bugbot, or Claude Code Review?

Those are hosted AI reviewers that look for many kinds of bugs. Bugbot and Claude Code Review comment on pull requests, and the CodeRabbit CLI also reviews local changes. For broad bug-finding they cover more than Sys1, whose bundled rules cover two JavaScript and TypeScript checks. Sys1 checks only the rules you select, can keep your source on your machine with a local model, and lets you record whether each candidate was useful. For a pure syntax pattern such as an empty catch block, an ESLint or Semgrep rule is exact and free; use Sys1 for rules that need judgment.

Can everything stay local?

Yes, for decision requests using local models and the local-only routing policy. Installing Sys1 and running model setup download packages and weights. Hosted Jev is off by default.

Can I replace my existing Jev client?

Yes. Run sys1 jev enable, which turns on hosted Jev and sets routing to hosted-only so no request falls back to a local model. Pin jev-1.13.0 so your app keeps the same model while you switch to the Sys1 client. Test your request sizes, question limits, cancellation, and error handling. Read the adoption guide.

Does Sys1 pick the most accurate model?

No. You choose the model. Fresh setup uses Qwen3 1.7B; enabling Jev selects hosted-only operation, and a hosted outage returns an error. Local fallback needs an explicit policy change after you evaluate it. Installing a smaller model does not change the selection.

What does it cost?

Sys1 is free, open-source software under the MIT license. Local inference uses your hardware. Hosted providers may charge separately under their own plans.

How do I keep noisy test logs out of my agent’s context?

System One Skills provides the separate system-one-verify skill for Devin, Claude Code, and Codex. It runs a known noisy check once, returns a short result, and keeps the full log on disk. It works without Sys1 or any model backend.