Skip to content

Repository files navigation

Tailor

Tailor is a job-application copilot with a personal memory: it researches the company, tailors a LaTeX resume from a verified experience library, and drafts application answers where every factual claim is cited to a library entry, while a human approves everything before it is sent.

Demo video: not published yet. The shot list the recording follows is docs/demo-script.md.

The Tailor application pipeline, a seven-column board from Saved to Replied holding one synthetic application per column, with a card mid-run labelled Tailoring resume

Synthetic applications. Real cards carry real companies and real application material and are never screenshotted.

Why it is different

Auto-apply tools optimize for volume. Tailor optimizes for answers Basil would actually send, and it gives up throughput to do it.

Verified grounding. Resume bullets and answer claims may only come from entries in an experience library that a human has marked verified. Retrieval is verified-only and category-supplemented, so a run that finds no verified experience, project or skill entry stops rather than guessing. Answer citations resolve to a specific source_ref; anything the model asserted without one is surfaced as an unsupported_claims flag rather than quietly kept. Numeric grounding preserves qualifiers and local claim context so a count cannot be reassigned to an unrelated subject.

An edit-memory loop. Saving an edited answer records the exact character edit distance and asynchronously extracts at most one typed memory: a voice rule, a preference, or a fact. Active voice rules and preferences are retrieved into later resume and draft prompts and their ids are stored on the resulting immutable version, so the review page can show which learned rules shaped a given artifact.

A human approval boundary in software, not in policy. Nothing submits an application. Runs start explicitly. Every resume and answer is an immutable numbered version, saved edits reject stale browser state, and memories can shape wording but can never become citations or evidence.

Architecture

flowchart TB
  B["Browser<br/>single /32 allowlist"] --> ALB["ALB<br/>HTTPS, ACM certificate"]
  MC["MCP client<br/>Claude Code or Claude chat"] --> A
  ALB --> W["Next.js 16<br/>standalone production server"]
  W -->|"/api rewrite over Cloud Map"| A["FastAPI<br/>REST plus MCP endpoint"]
  A -->|"enqueue run"| R[("ElastiCache Redis<br/>TLS queue")]
  R -->|"poll"| K["Worker<br/>ingest, research, tailor, draft"]
  A --> P[("RDS PostgreSQL 16<br/>with pgvector")]
  K --> P
  A --> E[("EFS<br/>compiled PDFs")]
  K --> E
  K --> AN["Anthropic claude-opus-5<br/>generation and web search"]
  K --> V["Voyage voyage-4<br/>embeddings"]
  K -.->|"optional"| LF["Langfuse tracing"]
Loading

Web, API and worker are three ARM64 ECS Fargate services; Postgres, Redis and EFS are managed, private and encrypted; the load balancer is the only thing anything outside the VPC can reach. The worker owns the fixed four-stage graph: ingest parses the posting and its application questions, research produces a structured company brief, tailor retrieves verified entries and compiles a .tex resume to PDF, and draft writes cited answers against every stated word and character limit. Local development runs the same five services under Docker Compose with Postgres, Redis, and a bind-mounted artifact directory in place of RDS, ElastiCache, and EFS.

Measured results

Two paid evals ran in CI on this branch, on identical code and the same 30 draft-stage cases, with claude-opus-5 generating and claude-sonnet-5 judging. Both are shown, because a single run of a sampled pipeline is not a measurement:

Measure Run 1 Run 2 Basis
Grounding, LLM judge 4.86 / 5 4.90 / 5 30 cases, judged per answer
Voice match, LLM judge 4.17 / 5 4.17 / 5 the first two runs ever to score this axis, against a 3-sample corpus
Job-description coverage, LLM judge 3.52 / 5 3.50 / 5 30 cases, judged per answer
Cases that failed outright 1 0 a failed case is scored on neither axis
Answers with flagged unsupported claims 4 10 deterministic count, not a judge score
Answers over a stated limit 0 0 deterministic count
Citations per answer 7.43 7.93 deterministic count
Judge-human agreement not measured not measured 0 human labels exist, so the judge is not calibrated
Retrieval recall@5, hybrid 0.958 0.958 12 labelled queries, real Voyage embeddings
Retrieval recall@5, vector only 0.958 0.958 measured delta 0.000

Both runs passed the CI gate against the committed baseline.

Measure Value Basis
Normalized edit distance 0.634, 0.338, 0.119 3 accepted edits, in order
Complete four-stage run, wall clock 233.8 s 1 uninterrupted run on the deployed AWS stack
Claude-only generation cost per run $0.80 same run, published rates, not total cost

Voice reproduced; the unsupported-claim count did not. Voice landed at 4.172 and 4.167, a gap of 0.005, which is the most stable of the three axes across the pair. That is agreement between two run means, not the per-case exact-agreement and mean-absolute-error that the grounding and coverage tolerances were derived from, so voice still has no gate tolerance of its own. Read it as two consistent readings from an uncalibrated judge against three accepted answers, which is a thin definition of anybody's voice.

The flagged-unsupported-claim count is the number that moved: 4 answers in one run, 10 in the next, from identical code. That count is deterministic given an answer, so the variance is in what the model generated, not in how it was judged. It is a property of one sample, and quoting a single run's flag count as a quality metric would be wrong.

The first run's worst cases were noise, not a regression. Run 1 had one outright failure and a grounding worst case of 2; run 2 had zero failures and a worst case of 4. Neither recurred, so nothing here is evidence of a per-case regression. The pair is a useful reminder that the gate compares means, and that a single red-looking run needs a repeat before anyone acts on it.

The judge is not calibrated. Basil has read judge outputs but no human label rows exist, so no judge-human agreement number can be reported. What is measured instead is judge self-consistency across two complete runs of the same 30 cases: grounding agreed exactly on 83% of cases with a mean absolute error of 0.17, and coverage agreed exactly on 73% with a mean absolute error of 0.27, with every score inside one point on both axes. Those two errors are the regression tolerances the CI gate uses. Treat every judge score above as a repeatable automated signal, not as a human-quality measurement.

The committed baseline is a different measurement. api/evals/baselines/current.json holds grounding 4.90 and coverage 3.37 with voice skipped, recorded on an earlier 20-entry library; both runs above used the current 19-entry one, so they are not a controlled comparison against it and the baseline was deliberately not rewritten here. Coverage is the axis that moved, from 3.37 to 3.50 and 3.52, and a library change is at least as plausible an explanation as anything in the prompts.

Hybrid retrieval shows no advantage here. Hybrid and vector-only search return the same recall@5 on this corpus, and both miss the same one of two expected entries on the same retrieval-augmented-generation case. Twelve queries against a small library is not enough to conclude hybrid is useless, but it is enough to say the measured delta is 0.000 and nothing in this repository demonstrates a hybrid win.

Cost is Claude only. The $0.80 is input, output, cache-read and cache-write tokens for one complete run priced at Anthropic's published claude-opus-5 rates effective 2026-08-03 ($5.00 / $25.00 per million input / output tokens; see Anthropic pricing). It excludes Voyage embeddings and all AWS infrastructure. The run spent its time as ingest 7.6 s, research 64.7 s, tailor 67.6 s, draft 93.5 s.

Accepted-answer edit distance

Three accepted edits plotted in order at 63.4%, 33.8% and 11.9% normalized edit distance, falling from left to right, annotated that no learned memory was applied to any of the three drafts

Three accepted edits is three points. They fall, and that is the honest description of the data, but no learned memory had been applied to any of those three drafts, so the direction is not evidence that the memory loop caused it. No trend line is fitted below five points. The figure is rendered deterministically from an aggregate-only export that contains no job id, company, question, answer, or evidence text:

cd api
uv run python -m app.metrics_cli export --out evals/reports/product-metrics.json
uv run python scripts/render_metrics.py --metrics evals/reports/product-metrics.json --out ../docs/assets/edit-distance.svg

How grounding and memory boundaries work

The experience library is chunked from a master content file and from public GitHub repositories, embedded with voyage-4, and stored in Postgres with pgvector. Only entries a human has explicitly verified are citable; GitHub-sourced rows stay unverified by design and can inform nothing.

Tailoring retrieves verified entries, then prints the library's own titles, dates and numbers rather than the model's wording of them. Drafting accepts a citation only when it resolves to an entry that was actually retrieved for that question. Every remaining assertion is returned as a flagged unsupported claim, visible in the review page and in the MCP get_drafts response, so a reader always knows which sentences are backed and which are not.

Memories govern how to write, never what is true. Retrieval injects only active voice_rule and preference memories whose scope matches the current company and role, into a separately labelled instruction block, and records the injected ids on the new artifact version. fact memories are stored and visible in the ledger but are deliberately not injected into generative prompts, because prompt wording alone cannot prove a free-text factual claim; that stays blocked until claim-level validation exists. Deactivating a memory preserves the ledger and historical attribution while excluding it from future runs.

Evals and CI

The eval harness is separate from the application board and never writes into it. Two corpora:

  • Draft quality: 30 representative application questions, scored by an LLM judge on grounding, voice and job-description coverage, plus deterministic counts of flagged unsupported claims, over-limit answers and citations per answer. It evaluates the draft stage only. Ingest, research, tailoring and PDF compilation are covered by deterministic tests, not by this suite.
  • Retrieval: 12 labelled queries with expected source_ref values, scored as recall@5 in hybrid and vector-only modes.

Calibration is implemented but unused: label walks a report interactively so Basil can score cases by hand, and calibrate compares those labels to the judge. The label file holds 0 rows. Agents are forbidden from writing labels, so this number only moves when Basil types the scores himself.

cd api
FAKE_EMBEDDINGS=0 uv run python -m app.evals.cli retrieval
FAKE_EMBEDDINGS=0 uv run python -m app.evals.cli run
uv run python -m app.evals.cli gate --report evals/reports/latest.json
uv run python -m app.evals.cli label --report evals/reports/latest.json --count 20
uv run python -m app.evals.cli calibrate --report evals/reports/latest.json

.github/workflows/ci.yml runs deterministic API tests and lint, web type checking, lint and build, the fully mocked Playwright suite, CloudFormation lint, and a production image build on every pull request. .github/workflows/evals.yml runs the paid real-model eval daily, by manual dispatch, and on pull requests touching generation, retrieval, corpus, baseline or eval code; it gates against the committed baseline and uploads only an allowlisted aggregate summary even when the gate fails. Pull requests skip the paid job with a visible notice when the corpus secrets are unavailable, while scheduled and manual runs fail instead, so missing production configuration cannot pass silently.

Every pipeline run, node and model generation is wrapped in optional Langfuse tracing. Leave the keys empty and the wrapper is a verified no-op: the worker reports tracing disabled, reaches its polling loop, and the API stays healthy.

MCP server

Tailor exposes a streamable HTTP MCP endpoint so a job link can be fired at it from any Claude session:

claude mcp add --transport http tailor http://localhost:8000/mcp
  • ingest_job saves pasted text or a posting URL and can optionally queue the pipeline.
  • get_pipeline returns every board card and counts by application status.
  • get_drafts returns the newest accepted-or-generated answer, keeps generated and accepted text separate, names the current source, recomputes over_limit from the returned text, and states that citations and unsupported claims describe the generated answer.
  • add_memory stores one reusable voice_rule, fact or preference.
  • get_profile_facts returns curated standard-form values and never guesses a missing one.

MCP_TOKEN is optional locally. When it is non-empty every request must carry Authorization: Bearer <token>. DNS-rebinding protection stays enabled in every environment, so each Host and Origin the endpoint should answer is listed explicitly in MCP_ALLOWED_HOSTS and MCP_ALLOWED_ORIGINS; setting a token does not widen that allowlist. MCP is reachable locally but not through the deployed URL - see limitations.

AWS deployment

Three ARM64 ECS Fargate services in ca-central-1: a public Next.js service behind an HTTPS Application Load Balancer, an internal FastAPI service found through AWS Cloud Map, and one worker. RDS PostgreSQL 16 with pgvector holds durable state, a single-node TLS ElastiCache Redis replication group owns the queue, and encrypted EFS keeps compiled resumes across task replacement. Two CloudFormation templates and an exact runbook live in infra/.

This is single-user hosting, not public hosting. The load balancer security group takes a mandatory /32 at deploy time and has no 0.0.0.0/0 default. RDS and ElastiCache are private, encrypted at rest and in transit. EFS requires IAM authorization and transit encryption. ECS tasks carry public IPs only for outbound provider calls, which is how the deployment avoids a NAT Gateway, and no security group admits public ingress to a task. Provider credentials exist only in Secrets Manager, and the library is seeded from a private encrypted S3 object that is deleted immediately after import.

Verified against the live stack:

  • HTTPS with a certificate that verifies for tailor.basilliu.dev, and the board loading through the web service's /api rewrite
  • all 9 Alembic migrations against RDS, including CREATE EXTENSION vector
  • rediss:// TLS queueing, with the worker reaching its polling loop
  • the library seeded from S3, 19 entries imported, and the seed object deleted
  • real Voyage retrieval on the deployed API
  • one complete four-stage run, then a PDF download of 103,958 bytes
  • EFS persistence: after the API task was killed, its replacement served a byte-identical PDF with the same sha256
  • direct connections to task public IPs, to RDS and to Redis all refused

Two things are not true of the deployment yet. tailor.basilliu.dev does not resolve: the ACM certificate is issued and the TLS chain verifies for that name, but the second DNS record pointing the host at the load balancer has not been added, so the stack is reachable only through the load balancer's own hostname. And the MCP endpoint is not reachable through the deployed URL; the MCP server itself passed its checks in a live container, returning 401 unauthenticated, 421 for an unlisted Host and 403 for an unlisted Origin, but the public route to it is broken. See limitations.

Cost while tailor-app is up is roughly $0.13 per hour, dominated by three Fargate tasks, RDS, ElastiCache and the ALB. That is one coarse console reading over the first live day, consistent with the $0.12 per hour published-rate estimate in the runbook; the Cost Explorer API still returns near-zero for this account, so do not read it as a precise figure. Stopping the ECS services saves only the Fargate share, because RDS, ElastiCache, EFS and the ALB bill regardless, so teardown means deleting the stack. Section 11 of the runbook is the exact procedure.

Local setup

cp .env.example .env     # add ANTHROPIC_API_KEY and VOYAGE_API_KEY
docker compose up --build
open http://localhost:3000

Add an application, then start its run explicitly. The API and worker images bundle TeX Live so they can compile the vendored pdfTeX template, which is why those images are large.

Library management runs inside the worker:

docker compose exec -T worker python -m app.cli import-master --path /path/to/master.md
docker compose exec -T worker python -m app.cli import-github --repo owner/repo
docker compose exec -T worker python -m app.cli list-unverified
docker compose exec -T worker python -m app.cli verify --id <entry-id>
docker compose exec -T worker python -m app.cli search --verified-only "agent orchestration"

GET /jobs/{id}/review returns the parsed posting, the research brief, the newest resume version with resolved evidence, and the newest answer per question with any accepted edit. GET /jobs/{id}/resume.pdf serves the latest compiled resume and GET /jobs/{id}/resume.tex is the fallback after a failed compile. Artifacts land in data/artifacts/<job_id>/; only paths relative to that root are stored in Postgres.

Tests:

cd api && uv run ruff check && uv run pytest -q
cd web && npx tsc --noEmit && npx eslint . && npm run test:e2e

The Playwright suite runs against deterministic API mocks and makes no provider calls.

Limitations and future work

  • The judge is not calibrated. Zero human labels. Every LLM-judge number here is an automated signal with measured self-consistency and no measured agreement with a human.
  • Voice scores 4.17 and Basil disagrees with it. His consistent criticism of generated answers is that the openings are too generic, which is a voice problem, and the axis scored 4.17 twice against a corpus of three accepted answers with no human calibration behind the judge. Two consistent readings mean the measurement is stable, not that it is right. A comfortable score sitting next to an unhappy human is a reason to doubt the metric, and the cheapest way to settle it is human labels.
  • The gate compares means, so it cannot see a single bad case. One run had an outright failure and a grounding score of 2; the repeat had neither. Both passed. That is the correct outcome for noise, but it also means a genuine per-case regression would pass the same way, and today only a human reading the report would catch it.
  • The learned rule that addresses that criticism is switched off. The first extracted voice rule says to avoid opening with scene-setting narrative, and it is active=false, so retrieval never injects it. Why it is inactive has not been established; if overlapping extractions silently deactivate earlier rules, that is a memory-lifecycle bug worth finding before flipping the row back on.
  • MCP is unreachable through the deployed URL. Next strips the trailing slash and answers 308, FastAPI's mount then answers 307 built from the Host it saw, and the client is sent to an internal Cloud Map hostname it cannot resolve. Every other /api/* path works. The likely fix is skipTrailingSlashRedirect in next.config.ts.
  • The board is the wrong shape for how applications actually move. Seven columns, most of them usually empty, and horizontal scroll to hold a handful of cards. Most applications end up submitted, and most postings need only a tailored resume.
  • A submitted application is still an editor. There is no read-only view, so a finished application still shows Save buttons.
  • No first-run guidance. An empty board says what to do in one line and nothing walks a new user through it.
  • fact memories cannot reach a prompt until claim-level validation can prove free-text claims against verified entries.
  • Named reusable resume variants and a skip-tailor run mode are specified but unbuilt, so every run pays for tailoring even when a stored resume would do.
  • Deferred findings with their reasoning live in docs/notes/backlog.md.

Privacy posture. This repository is private and is not safe to publish as is. The experience library, accepted answers, eval reports and compiled PDFs are gitignored, and metrics exports are aggregate-only by construction, but the vendored resume template still carries real contact details and the eval corpus secret holds real application material. A sanitized public mirror is future work, not something this repository is one flag away from.

About

A job-application copilot: a four-stage LangGraph agent that researches a posting, tailors a resume, and drafts answers where every factual claim is cited to a human-verified experience library.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages