v0.5.1, Jev the dominant exemplar

Place judgment. Keep proof exact.

Design-judgment skill for the decision-model class. TypeSafe Jev Choice/Score/Noul is the dominant exemplar most users will call. Classical decision methods, composition algebra, and a validation gate.

Jev-class

With vs without Augustus

Call the model and act on the score, or place the judgment. Jev is the exemplar. Encoders, open heads, gameplay specialists, and conversion on-ramps sit in the same class. Not TypeSafe-only.

  • Jev
  • kev
  • Laya
  • OpenJev
  • GLiNER
  • SemIf
  • NanoJev
  • Jeff-1
  • TypeLLM

Without

Call, then act

  1. Call the model
  2. Act on the score

Quiet failure modes

  • Soft Noul treated as hard gate
  • GPT bakeoff framing
  • No falsifier
  • Polarity unchosen

With Augustus

Place, then judge

  1. State
  2. Pillar / family map
  3. Question design
  4. Fail-open vs fail-closed
  5. Typed judgment Choice Score Noul
  6. Code owns effects
  7. Named falsifying experiment

Shared class: Jev, kev, Laya, OpenJev, GLiNER, SemIf, NanoJev, Jeff-1, localjev. Augustus owns placement. The model owns narrow judgment. Code owns the effect. Class recipes (problem → without → with → measure).

Recipes

Same split, whole class

Not Jev-only. Problem → without → with → what to measure. Full cards in the v0.5.1 notes (GEPA domain-adapt) and the v0.5.0 notes. No invented scores.

Inference type

Decide vs generate

Problem
A typed judgment treated as another token stream.
Without
Call tryDecide, then parse the chat. Ship the speedup.
With
decide is not generate. tryDecide returns typed calibrated judgments, not a token stream.
Measure
Type of the return value. Third-party benches stay *theirs*.

Encoder

GLiNER / GLiClass

Problem
Locate or categorize treated as a decision head.
Without
Swap, hard-gate spans, bake off a chat LLM.
With
Species map. Extractive remainder. Soft scores ≠ hard gates.
Measure
Span quality separately from ECE. Softmax ≠ Noul.

Open heads

Laya, SemIf, kev, Jeff-1

Problem
Wire-compat or argmax agree treated as a replica.
Without
Drop-in swap. Ship the speedup. Skip OOD.
With
Softmax ≠ calibrated Noul. Systems timing ≠ semantic equivalence.
Measure
Held-out ECE/Brier and accuracy. In-distribution vs OOD.

Gameplay

NanoJev

Problem
Game success treated as a calibrated Noul.
Without
Quote a win rate as a production gate.
With
Specialist S1. Local boolean ≠ TypeSafe noul.
Measure
Held-out game metrics on one ledger. ECE on another.

Domain adapt

GEPA on Jev

Problem
A schema-valid Choice treated as a correct label, or API confidence treated as P(correct).
Without
Ship F1 as the review-queue policy. Treat GEPA as a new scoring-table species.
With
schema-valid is not the same as correct. API confidence is not P(correct). GEPA revises Choice instructions/criteria with weights fixed. review-queue policy is not F1. Soft is not gate. Not an 18th scoring-table species.
Measure
Brier and F1 stay *theirs*. Track FN tradeoff separately from review-queue retention.

Conversion

llm-to-jev

Problem
A chat prompt assumed equivalent to Choice/Score/Noul.
Without
Paste, convert, ship.
With
Heuristic on-ramp. Review the Score rubric. Prose stays with the LLM. heuristic conversion ≠ calibrated Noul.
Measure
Suitability labels and human review. Not equivalent behavior.

Lookup

jcr

Problem
Finding a documented command treated as permission to run it.
Without
The agent executes whatever the tree returned.
With
Returns context. Does not execute. Routing ≠ permission.
Measure
Lookup+explain only. Not Harbor task-execution.

Readout

localjev / prompted JSON

Problem
Parsed JSON from a generator treated as a Noul.
Without
Schema-valid taken as picked-right.
With
Prompted JSON ≠ structured logit read. Schema-valid ≠ picked-right.
Measure
Schema pass separately from calibration and from the effect.

Client / contract

sdk 0.7 / MLX 400

Problem
A Pydantic client model or a 400 SchemaError treated as the head.
Without
Ship 0.7 as logit-equiv. Treat 400 as a Noul.
With
Pydantic response models ≠ logit-equiv. Error contract is not a Noul. Locate ≠ decide.
Measure
Server logits separately from client decode. coverage-at-error-budget stays *theirs*.

Serving

Dual surface / hosted port

Problem
A second HTTP path or a free hosted URL treated as the same species as decide.
Without
Point chat at /v1/chat/completions and call it System One. Treat Codiv as TypeSafe.
With
dual serving is not generate. chat 501 on MLX. Hosted Codiv ≠ TypeSafe.
Measure
Which path returned the answer. Third-party benches stay *theirs*.

Scores / tools

Relative p, advisory MCP, LoRA

Problem
Candidate p, an MCP recommendation, or a LoRA adapter treated as calibrated truth.
Without
Hard-gate 0.8. Block on the server. Quote 94.71% as Harbor.
With
candidate probabilities are relative not correctness. recommendation is advisory. LoRA ≠ RLCD replica. pass-min 0.8 still soft.
Measure
Relative ranking separately from ECE. The server never blocks on its own.

Constrained AR

TypeLLM vs Noul

Problem
Typed constrained decode treated as a calibrated Noul.
Without
Quote 5.8x. Ship thinking=True as a probability. Treat 0.8B thinking On 0/18 as a Noul.
With
Constrained AR ≠ calibrated Noul. truncated thinking then constrained decode. type safety does not guarantee factual accuracy. type-valid ≠ exact.
Measure
Batch 5.8x *theirs*. 0.8B thinking On 0/18 *theirs*. forced closure 20/20 type-valid *theirs*. Thinking budget separately from ECE.

Open heads

kev family isolation

Problem
Sibling-blind questions treated as option-order immunity.
Without
Quote 0.812 as Harbor. Hard-gate 0.9. Treat Kev-0.8B as Archer.
With
Questions share the input text but cannot read each other. option order can change an answer. Kev-0.8B completes family. Qwen3.5 ≠ Archer.
Measure
4B new-source 0.794/0.832 *theirs*. 9B new-source 0.812/0.837 *theirs*. transfer-v9 Kev-9B 5% Jev 9% Kev-8B 26% *theirs*. SemIf Kev-9B 0.917 Jev 0.965 *theirs*.

Fail polarity

Closed routing vs open RUN

Problem
One cutoff treated as both a router and a permission.
Without
0.65 auto-allows. Uncertainty skips the suite.
With
JEV_THRESHOLD 0.65 still soft. routing ≠ permission. fail-open uncertainty means RUN. fail closed never auto-allows.
Measure
Name the polarity per act. Thresholds stay application policy.

Eval / catalog

classifier ≠ authorizer

Problem
An estimate row or a curated list treated as a grant.
Without
Ship 76/81 as Harbor. Install because it is listed.
With
classifier ≠ authorizer. estimates not Harbor. catalog ≠ endorsement.
Measure
76/81 vs 77/81 *theirs*. 118 entries are an index, not a proof.

Serving pin

Restructured head vs replica

Problem
A package bump or a vLLM subclass treated as logit-equivalent truth.
Without
Ship 0.3.0 as TypeSafe. Treat Codiv as the hosted product.
With
restructured vLLM head ≠ logit-equiv. MODEL_VERSION stays openjev-0.1. dual serving is not generate. Hosted Codiv ≠ TypeSafe.
Measure
Package version separately from wire id. Third-party benches stay *theirs*.

Split jobs

Route, write, render

Problem
One model asked to choose, write, and draw.
Without
A swarm. Luna as the judgment. Diffusion as structure.
With
Models participate. Real tools execute. documentation is read not judged. json-render is the only renderer. Thresholds are policy not model.
Measure
Who chose, who wrote, who rendered. Soft scores ≠ hard gates.

Done-check

Facts block, Jev advises

Problem
A probability treated as the thing that stops an agent from claiming done.
Without
Refuse done because Jev said 0.8. Skip the ledger.
With
Facts go to code. Judgments go to Jev. Only facts can block. Jev never blocks.
Measure
Replay the ledger. A missing passing check is the block, not a Noul.

Serving port

ANE argmax vs replica

Problem
155/155 argmax treated as logit-equivalent Kev.
Without
Ship ANE fp16 as the 0.8B family. Quote 7.16 ms as Harbor.
With
155/155 argmax *theirs*. serving substrate ≠ calibrated replica. MidasMulli/kev-ane ≠ jaredpalmer/kev.
Measure
Argmax separately from probability spread. *theirs* not Harbor.

Gameplay / sim

Win rate vs Noul

Problem
A seeded traffic win treated as a calibrated probability.
Without
Quote delay as ECE. Hard-gate the city.
With
game success ≠ calibrated Noul. Fixed vs Adaptive vs the decision head on one ledger.
Measure
Sim metrics separately from Brier/ECE. *theirs* not Harbor.

Open-head training

From-scratch vs warm-start

Problem
A from-scratch adapter treated as equal to a released head plus your labels.
Without
Train from the base. Quote 0.88 as Harbor. Skip the held-out file.
With
from-scratch ≠ warm-start. JSONL labels ≠ Harbor. --init_from keeps what the head already knows.
Measure
0.33 vs 0.84 vs 0.83/0.88 *theirs*. Same split for any Choice/Score/Noul head.

Independent assay

Split verdict vs global certificate

Problem
One corpus ECE treated as the model's calibration everywhere.
Without
Ship CLINC150 as proof. Ignore Banking77 overconfidence.
With
assay-001 split verdict. A result applies to the artifacts examined.
Measure
CLINC150 ECE 0.0204 *theirs*. Banking77 ECE 0.0936 *theirs*. Not Harbor.

Reconstruction

Random weights vs replica

Problem
A from-first-principles sketch treated as TypeSafe logits.
Without
Swap the serving URL. Quote the catalog as a grant.
With
reconstruction ≠ replica. unofficial research implementation with random weights. catalog ≠ endorsement.
Measure
Namesake lock first. Third-party benches stay *theirs*.

Open LoRA specialist

Latency vs meaning

Problem
A local P50 treated as proof the head matches hosted Jev.
Without
Quote 85 ms as parity. Skip the 1024/32 slowdown. Treat 94.71% as a Noul.
With
systems latency ≠ semantic equivalence. hard acc ≠ calibrated Noul. LoRA ≠ RLCD replica. not merged base models.
Measure
customer-service P50 85.03 vs 295.26 *theirs*. 1024/32 1015.90 vs 301.37 *theirs*. Open-Jev TREC pending.

Decision support

Dashboard vs fill

Problem
A Buy/Sell/Hold card treated as an executed order.
Without
Wire the Choice to the exchange. Skip the human.
With
platform does not execute trades. does not execute. Code owns the fill.
Measure
Count fills separately from cards. *theirs* not Harbor.

Independent assay

AI review vs gold / one trial vs Harbor

Problem
62.69% or one seed-0 robot run treated as a certificate.
Without
Ship AI labels as gold. Quote $0.018825 as Harbor.
With
AI-reviewed labels ≠ gold. one-trial robot ≠ Harbor. 10.59× systems ≠ ECE. agreement ≠ accuracy.
Measure
62.69% vs 67.26% *theirs* not gold. Spanish −6.4 pp XNLI *theirs*. Not Harbor.

On-ramp / catalog

Desc rewrite vs replica

Problem
A GitHub description rewrite treated as a new compiler, or a catalog as a grant.
Without
Mint a sibling card. Quote 485 as endorsement.
With
desc rewrite ≠ SHA/behavior change. catalog ≠ endorsement. heuristic conversion ≠ calibrated Noul.
Measure
SHA unchanged 234058ab372d. 485 is an index. *theirs* not Harbor.

Serve calibration

Temperature scaling vs ECE

Problem
A serve temperature treated as proof the head is calibrated.
Without
Ship T≈2.0 as Harbor. Fold grouped T into the raw row.
With
temperature scaling ≠ ECE unless measured. grouped T rejected. Report a calibrated row separately.
Measure
Brier 0.291→0.267 ECE 0.105→0.039 *theirs*. 7.5%→3.2% *theirs*. Not Harbor.

Hub pin

Revision vs replica

Problem
A Hub --revision treated as hosted Jev logits.
Without
Swap the serving URL to night2-du. Skip the locked-test costs.
With
Hub --revision is a pin not a replica. SHA move is not a replica. main untouched awaiting sign-off.
Measure
0.852 OOD *theirs*. coverage@5% 0.62 from 0.66 *theirs*. *theirs* not Harbor.

Trained encoder runtime

OpenJev runtime vs TypeSafe

Problem
A trained OpenJev loader treated as the hosted System One API.
Without
Point the .NET SDK at DeBERTa. Treat generated_text: False as TypeSafe.
With
trained runtime ≠ TypeSafe. OpenJev.from_pretrained. encoder class member not Jev replica.
Measure
kind typed-decisions/open-jev-v1. generated_text: False. DeBERTa 0.855 / 42 ms *theirs*.

Class members

Stock Qwen, CoreML, SQL, locate, legal LoRA

Problem
A letter readout, on-device bundle, SQL predicate, span locator, or legal LoRA treated as RLCD Jev.
Without
Quote 0.812 as Archer. Treat DuckDB as a Noul. Collapse GLiNER into Choice.
With
Qwen3.6 ≠ Archer. serving substrate ≠ calibrated replica. Locate ≠ decide. legal LoRA ≠ RLCD replica. catalog ≠ endorsement.
Measure
jqv stock Qwen3. swev CoreML. duckdb-jev SQL predicates. gliner2-skill locate. *theirs* not Harbor.

Provider eval

CPU tests vs GPU quality

Problem
A CPU-passing provider suite treated as completed GPU quality.
Without
Ship 65/76 as Open-Jev TREC. Quote 48 CPU tests as Harbor.
With
provider pipeline ≠ completed Open-Jev quality. CPU tests ≠ GPU scores. Open-Jev TREC pending.
Measure
808 requests 1841 labelled *theirs*. 65/76 72/76 66/76 60/76 71/76 *theirs*. Not Harbor.

Fine-tune vs hosted

Kev CartPole vs TypeSafe

Problem
A fine-tuned Kev lab treated as hosted Jev, or 52/64 as a majority proof.
Without
Swap the CartPole URL for TypeSafe. Treat softmax as a Noul.
With
fine-tuned Kev ≠ TypeSafe Jev. one record of 64. softmax ≠ calibrated Noul.
Measure
81.25% 52/64 *theirs*. majority 79.69% 51/64. *theirs* not Harbor.

Broker / cutoff / catalog

Mock orders, 95% gates, 80.1% gold

Problem
A mock/dry QMT sidecar treated as live fills, 95% as a hard gate, or 80.1% as gold.
Without
Fire orders from AUC 0.532. Hard-gate Luna fallback. Quote 3.69ms as Harbor.
With
QMT mock/dry default no orders. does not execute. cutoff 95% still soft. catalog ≠ endorsement. option order can change an answer.
Measure
AUC 0.532 *theirs*. 79.6% / 80.1% *theirs*. 3.69ms *theirs* not Harbor. 87 of 144 order-unstable *theirs*.

Measurement recipe (hysteresis, equal-width vs quantile ECE, hop-ECE, Harbor schema-pass ≠ joint, DecisionOps FALLBACK): v0.5.0 notes. GEPA domain-adapt: v0.5.1 notes. Ranking ≠ calibration. Soft Noul ≠ hard gate.

What it is

A gate for where judgment belongs

Augustus is named for Augustus De Morgan, mentor of William Stanley Jevons. TypeSafe Jev is the dominant exemplar most users will call. Exact work stays in code or policy. The model owns narrow judgment. A soft Noul is not a proof. The atlas lives in .agents/skills/augustus/ and research/notes.md. This page is a gate, not a rewrite of the repository README.

Install

Two paths

Claude Code via the marketplace, or any skills-compatible agent.

Claude Code

claude plugin marketplace add 24601/Augustus
claude plugin install augustus@augustus

skills.sh / npx

npx skills add 24601/Augustus --skill augustus

Pillars

Four placements, then a family

Pick the pillar from the hole, then the family, then the vendor. Ranking is not calibration. A soft Noul is not a hard gate.

Placement

Pillar, family, fail polarity

Name where judgment sits, which family matches the action, fail-open vs fail-closed, and the experiment that could prove the design wrong.

Classical methods

Mental models, not a vendor how-to

Expected utility, abstention, VOI, MCDA, signal detection, search and control, Leveson-style org and safety. Across AI, software, business, knowledge work, and life.

Formal methods

Proof stays proof

Alloy, TLA+, contracts, DST (Antithesis, Resonate, PufferLib). A Noul is a sensor. Never launder it as a proof.

Validation

A gate, not a scoreboard

Harbor and jevals practice: Score is 0..n-1. Noul has no confidence field. 0.85 / minProbability is not a hard Harbor gate. VERIFY needs discriminating evidence.

Companions

Contracts, integrity, atlas

Not a TypeSafe product. Augustus owns placement. Neighbors own their jobs.

Last updated 2026-09-21 (v0.5.1).