Design-judgment skill for the decision-model class. TypeSafe Jev Choice/Score/Noul
is the dominant exemplar most users will call. Classical decision methods,
composition algebra, and a validation gate.
Call the model and act on the score, or place the judgment. Jev is the
exemplar. Encoders, open heads, gameplay specialists, and conversion
on-ramps sit in the same class. Not TypeSafe-only.
Jev
kev
Laya
OpenJev
GLiNER
SemIf
NanoJev
Jeff-1
TypeLLM
Without
Call, then act
1Call the model
2Act on the score
Quiet failure modes
Soft Noul treated as hard gate
GPT bakeoff framing
No falsifier
Polarity unchosen
vs
With Augustus
Place, then judge
1State
2Pillar / family map
3Question design
4Fail-open vs fail-closed
5Typed judgment
ChoiceScoreNoul
6Code owns effects
7Named falsifying experiment
Shared class: Jev, kev, Laya, OpenJev, GLiNER, SemIf, NanoJev, Jeff-1,
localjev. Augustus owns placement. The model owns narrow judgment. Code
owns the effect.
Class recipes
(problem → without → with → measure).
Recipes
Same split, whole class
Not Jev-only. Problem → without → with → what to measure. Full cards
in the
v0.5.1 notes
(GEPA domain-adapt) and the
v0.5.0 notes.
No invented scores.
Inference type
Decide vs generate
Problem
A typed judgment treated as another token stream.
Without
Call tryDecide, then parse the chat. Ship the speedup.
With
decide is not generate. tryDecide returns typed calibrated judgments, not a token stream.
Measure
Type of the return value. Third-party benches stay *theirs*.
Encoder
GLiNER / GLiClass
Problem
Locate or categorize treated as a decision head.
Without
Swap, hard-gate spans, bake off a chat LLM.
With
Species map. Extractive remainder. Soft scores ≠ hard gates.
Measure
Span quality separately from ECE. Softmax ≠ Noul.
Open heads
Laya, SemIf, kev, Jeff-1
Problem
Wire-compat or argmax agree treated as a replica.
Without
Drop-in swap. Ship the speedup. Skip OOD.
With
Softmax ≠ calibrated Noul. Systems timing ≠ semantic equivalence.
Measure
Held-out ECE/Brier and accuracy. In-distribution vs OOD.
Gameplay
NanoJev
Problem
Game success treated as a calibrated Noul.
Without
Quote a win rate as a production gate.
With
Specialist S1. Local boolean ≠ TypeSafe noul.
Measure
Held-out game metrics on one ledger. ECE on another.
Domain adapt
GEPA on Jev
Problem
A schema-valid Choice treated as a correct label, or API confidence treated as P(correct).
Without
Ship F1 as the review-queue policy. Treat GEPA as a new scoring-table species.
With
schema-valid is not the same as correct. API confidence is not P(correct). GEPA revises Choice instructions/criteria with weights fixed. review-queue policy is not F1. Soft is not gate. Not an 18th scoring-table species.
Measure
Brier and F1 stay *theirs*. Track FN tradeoff separately from review-queue retention.
Conversion
llm-to-jev
Problem
A chat prompt assumed equivalent to Choice/Score/Noul.
Without
Paste, convert, ship.
With
Heuristic on-ramp. Review the Score rubric. Prose stays with the LLM. heuristic conversion ≠ calibrated Noul.
Measure
Suitability labels and human review. Not equivalent behavior.
Lookup
jcr
Problem
Finding a documented command treated as permission to run it.
Without
The agent executes whatever the tree returned.
With
Returns context. Does not execute. Routing ≠ permission.
A fine-tuned Kev lab treated as hosted Jev, or 52/64 as a majority proof.
Without
Swap the CartPole URL for TypeSafe. Treat softmax as a Noul.
With
fine-tuned Kev ≠ TypeSafe Jev. one record of 64. softmax ≠ calibrated Noul.
Measure
81.25% 52/64 *theirs*. majority 79.69% 51/64. *theirs* not Harbor.
Broker / cutoff / catalog
Mock orders, 95% gates, 80.1% gold
Problem
A mock/dry QMT sidecar treated as live fills, 95% as a hard gate, or 80.1% as gold.
Without
Fire orders from AUC 0.532. Hard-gate Luna fallback. Quote 3.69ms as Harbor.
With
QMT mock/dry default no orders. does not execute. cutoff 95% still soft. catalog ≠ endorsement. option order can change an answer.
Measure
AUC 0.532 *theirs*. 79.6% / 80.1% *theirs*. 3.69ms *theirs* not Harbor. 87 of 144 order-unstable *theirs*.
Measurement recipe (hysteresis, equal-width vs quantile ECE, hop-ECE,
Harbor schema-pass ≠ joint, DecisionOps FALLBACK):
v0.5.0 notes.
GEPA domain-adapt:
v0.5.1 notes.
Ranking ≠ calibration. Soft Noul ≠ hard gate.
What it is
A gate for where judgment belongs
Augustus is named for Augustus De Morgan, mentor of William Stanley
Jevons. TypeSafe Jev is the dominant exemplar most users will call.
Exact work stays in code or policy. The model owns narrow judgment.
A soft Noul is not a proof. The atlas lives in
.agents/skills/augustus/
and research/notes.md.
This page is a gate, not a rewrite of the
repository README.
Install
Two paths
Claude Code via the marketplace, or any skills-compatible agent.
Claude Code
claude plugin marketplace add 24601/Augustus
claude plugin install augustus@augustus
skills.sh / npx
npx skills add 24601/Augustus --skill augustus
Pillars
Four placements, then a family
Pick the pillar from the hole, then the family, then the vendor.
Ranking is not calibration. A soft Noul is not a hard gate.
Placement
Pillar, family, fail polarity
Name where judgment sits, which family matches the action, fail-open
vs fail-closed, and the experiment that could prove the design wrong.
Classical methods
Mental models, not a vendor how-to
Expected utility, abstention, VOI, MCDA, signal detection,
search and control, Leveson-style org and safety. Across AI,
software, business, knowledge work, and life.
Formal methods
Proof stays proof
Alloy, TLA+, contracts, DST (Antithesis, Resonate, PufferLib).
A Noul is a sensor. Never launder it as a proof.
Validation
A gate, not a scoreboard
Harbor and jevals practice: Score is 0..n-1. Noul has no confidence
field. 0.85 / minProbability is not a hard Harbor gate. VERIFY needs
discriminating evidence.
Companions
Contracts, integrity, atlas
Not a TypeSafe product. Augustus owns placement. Neighbors own their jobs.