We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

The U.S. government is legally required to route hundreds of billions in contracts to small businesses every year. Most small firms never bid — not because they would lose, but because government business development is a full-time job that an eleven-person company cannot staff.

Someone has to check SAM.gov every day, read a 90-page solicitation, extract every "shall" into a compliance matrix, track four different deadlines, and notice the moment an amendment silently moves the due date and adds a cybersecurity requirement to an HVAC contract. That is three days of work per solicitation, repeating forever.

I built BidBeacon around a single question: what if that work ran in the background, all the time, without being asked?

What it does

BidBeacon is a five-agent system that handles the continuous, asynchronous work of government contract pursuit for small businesses:

  1. The Watcher polls the live SAM.gov Get Opportunities API on a schedule. Before a notice reaches the expensive model, a cheaper Gemma triage pass filters the structurally unbiddable ones — roughly four in five — at ~$0.000025 each. Only plausible notices reach Gemini Flash.

  2. The Qualifier scores each opportunity against a structured company profile. A deterministic rubric in code computes the composite score and the bid/no-bid badge. Gemini writes the rationale a busy owner reads; it may nudge the fit score by ±10. It cannot move a hard disqualifier. A test asserts this.

  3. The RFP Analyst turns a solicitation into a compliance matrix: every "shall", "must" and "will" extracted by a parser (not a model call), mapped against company evidence, cited to its source page. The analyst holds no network capability at all — a solicitation is a document a stranger uploaded, and it must have no channel to send the company profile anywhere.

  4. The Drafting Agent writes proposal sections grounded in the structured profile and runs a claim linter on its own output before a human sees it. Critically, the linter also runs on human edits — the dangerous direction is the proposal writer at 11pm reaching for "industry-leading" and rounding three contracts up to forty. A paragraph carrying a blocking flag cannot be accepted at all; the button is disabled, not just warned.

  5. The Deadline Sentinel manages timers (durable records, not setTimeouts), diffs every amendment, recomputes compliance, re-scores the bid, and drafts the clarification question — before the Q&A window closes, without being asked. Missed escalations are recovered on the next scheduler tick.

How we built it

Stack: Next.js 15 App Router with TypeScript, Tailwind CSS v4, Vitest.

Models:

  • Gemini 3.6 Flash (Vertex AI) — long-context analysis: solicitation enrichment, qualification rationale, draft generation, clarification questions. Falls back to Gemini 3.5 Flash if 3.6 is unavailable on the project, and only then to the deterministic offline provider
  • Gemma 3 — triage classifier: runs on every incoming notice before Flash touches it

Google Cloud:

  • Vertex AI — Gemini Flash + Gemma, with per-call fallback to an offline deterministic provider (so the demo runs with zero credentials)
  • Cloud Run — standalone Next.js build, containerized
  • Firestore — single-document working set for atomic multi-collection writes; the whole amendment apply (register + score + timers + activity) lands as one transaction
  • Cloud Pub/Sub — event bus; agents publish, the bus fans out; splitting into separate Cloud Run services is a matter of subscribing each service to its topics
  • Cloud Tasks — per-escalation timers at their exact scheduled moment; Cloud Scheduler is the safety net
  • Cloud Trace — every span carries a real W3C trace id; the activity feed in the console is the same data, so "it worked unattended" is something you can go and check
  • Secret Manager — all credentials

Architecture: Five agents, each with an explicit capability scope checked on every repository call. Agents never call each other — they publish events, and the bus fans out. The three critical boundaries that carry real weight: the Analyst holds no net:external, the Sentinel cannot write drafts, and no agent holds profile:write.

Paywall: Polar for billing, Standard Webhooks signature verification (with 5-minute replay protection) so entitlements change only when a signed webhook says so — the browser is never trusted to report that it paid. Admin accounts bypass everything without a subscription.

Demo film: A Remotion composition with real application captures (Puppeteer at 3200×1800), Microsoft neural voice narration (edge-tts), and an animated cursor placed using recorded button coordinates — so the pointer lands on the actual control, not an artist's guess.

Challenges we ran into

1. The compliance-mapping bug that would have killed the demo.

The centrepiece is the amendment beat: Amendment 0002 adds CMMC Level 2 to an HVAC contract, and the decision card should show four newly non-compliant critical requirements. First run: "Newly non-compliant: none."

The bug was in evidence matching. A past-performance record shared one common word with a CMMC requirement, which was enough to promote the row from gap to partial — suppressing it from the non-compliance list entirely. The system was telling a company it partially satisfied a cybersecurity certification because a previous contract used the word "contract."

Fix: named credentials (CMMC, ISO, clearances, licences) are resolved first against the certifications list, authoritatively. Adjacent experience is never partial credit. Past performance now requires two overlapping key terms, not one. gap is reserved for security and certification categories — a procedural prohibition is critical but not evidenceable.

2. The linter that never fired.

Building the demo's linter beat, the linter would not fire: the drafting agent is constrained to the profile, so it only ever writes supportable claims. The guarantee was real and completely invisible.

The obvious move was to stage a fake screenshot. The better move was to notice the threat model was backwards: the dangerous direction is not the model, it is the human editing the draft at 11pm. So draft paragraphs became editable, and edits are linted identically — including withdrawing any existing acceptance when text changes.

3. Least privilege is a claim, not a description.

Every agent project says "each agent has limited permissions." Most of the time that sentence is doing no work. Making it executable caught a real design flaw: the Watcher was calling getOpportunity to deduplicate notices — which required opportunity:read, enough to read the entire pipeline from a network-facing agent. The fix was a narrower right: opportunity:discover, which permits insert-if-absent only and returns {created: false} for a known notice without revealing anything about it.

A comment in a diagram cannot fail. A capability check can.

4. \b byte corruption in production code.

A Python heredoc wrote literal backspace bytes (0x08) into the regex GOVERNMENT_OBLIGATION_RE in requirements.ts where \b word-boundary bytes were intended. The filter never matched. Found by scanning every source file for control characters after several tests unexpectedly failed.

5. The server build and zero-credential demo.

output: 'standalone' requires running from the bundle's own directory (the standalone server resolves .next relative to its cwd). npm start used ${PORT:-3000} shell syntax that PowerShell cannot expand. Both found during demo capture.

The zero-credential property — the whole demo works with an empty .env, every test passes, no API keys needed — was not a nicety. A demo that depends on a rate-limited government API is a demo that fails on stage.

Accomplishments that we're proud of

  • The amendment beat is real. The decision card — due date −9 days, score 79→63, four newly non-compliant critical requirements, clarification question drafted — is the Sentinel's real handler running on an unattended trace. Nothing is scripted or simulated at the application layer.

  • 141 tests, nothing mocked. The differ, rubric, claim linter, injection screen, timer recovery after a cold start, every capability boundary, and the full billing signature verification (including forged webhooks, tampered payloads, and replayed requests). The logic under test is never mocked; only the store is swapped for an in-memory driver.

  • The linter beat turned a staging problem into a product improvement. Finding that the linter never fires on agent output (because the agent is constrained to the profile) led to the more important feature: linting human edits. That is now the most useful twenty seconds of the demo.

  • The injection screen passes the test of the thing it is defending against. Fixture arc 2 (the USACE solicitation) carries a real prompt-injection attempt in an attachment. The screen neutralises it in place — wrapping with [[SCREENED:rule]] rather than deleting, so surrounding requirement text is preserved. The analyst then processes the sanitised document. The screen is pattern matching, not a model call, because it has to be deterministic and impossible to talk out of.

  • The film is honest. Every frame is a real capture of the running application. The cursor is the one drawn element, and even that is placed using recorded button coordinates from the capture script.

What we learned

  • The model boundary is the most important architectural decision in an agentic system. Not which model, not which framework — what the model is and is not allowed to decide. The rubric computes the badge; the model writes the rationale. That split is what lets the system degrade gracefully when the model is wrong or unavailable.

  • Deterministic extraction before model enrichment. A model that silently drops one requirement in ninety produces a matrix that looks complete and is not. A parser that over-extracts produces visible noise. The asymmetry decides the design.

  • Enforcement beats documentation for capability boundaries. The code that catches a scope violation mid-build is worth more than any architecture diagram.

  • The threat model for fabricated claims is backwards. The model is constrained; the human under deadline pressure is not. Lint both.

  • Rejected signals need to be stored with their reason. A filter you cannot inspect is a filter you cannot trust. Every notice Gemma triaged out is kept with the reason, so the owner can audit what the agent decided not to show them.

  • Timers need to be durable records. A setTimeout on a Cloud Run instance that scales to zero is a missed deadline. Escalations with firedAt stamps, a Cloud Tasks queue per escalation, and a Cloud Scheduler sweep as the safety net ensure that a missed escalation is delivered late rather than lost.

What's next for it

  • Real PDF pipeline. Current fixtures are page-delimited text. Production needs Document AI or pdf-parse feeding the same paginate() contract.
  • Firebase Auth or IAP to replace the demo-grade signed-cookie sessions.
  • Multi-tenant Firestore split — one document per company, each agent service with IAM rules matching its capability scope.
  • Calibrated rubric weights against real win/loss data; currently reasonable priors.
  • Live escalation channels — email and SMS beyond the console Countdown chip.
  • The one constraint that will not change: BidBeacon will never submit a bid. Submission portals stay human.

Built With

  • cloudrun
  • gemini3.6
  • vertexai
Share this project:

Updates

Submission history