Inspiration
WebMCP is a powerful emerging standard for AI tool integration, but developers currently have no way to test whether agents will actually select the right tools, use correct parameters, or follow valid multi-step journeys before shipping. We built ContractLab to fill this gap—providing a deterministic lab where you can refine contracts, catch ambiguity, and prove improvements against live browser agents.
What it does
ContractLab is a focused workbench for designing, linting, simulating, and live-evaluating WebMCP tool contracts. It features:
- Design mode: Author tools with structured editors, validate schemas, annotations, state conditions, and safe effects
- Lint engine: Detects ambiguous names, vague descriptions, missing enums, annotation mismatches, and overlapping tools
- Live eval mode: Switches from authoring tools to your compiled draft contracts, exposing them only to browser agents
- Deterministic mock domain: A ticket support system with clean state, version guards, and visible UI updates
- Call recording: Captures immutable arguments, results, failures, timings, state diffs, and UI effects
- Seven-dimension grader: Scores tool choice, parameters, order, prohibited calls, executor success, final state, and UI effects
- Version comparison: Compare repeatable evidence across contract versions to prove improvements
The app itself uses WebMCP in design mode for collaborative authoring, then compiles your safe contracts into real WebMCP tools for live agent evaluation. No LLM is embedded and no user code is executed.
How we built it
- Frontend: React 19, TypeScript, Vite, and an Apple-inspired design system with translucent chrome and dense developer surfaces
- WebMCP integration: Centralized adapter in
src/lib/webmcp.tsfollowing Chrome's imperative API withAbortControllerlifecycle management - Safe contract model: Typed JSON Schema subset with forms-based editing, finite mock-effect DSL, and deterministic ticket domain reducer
- State management: Event-sourced contract versions and eval runs stored in localStorage with reset and export/import
- Testing: Vitest, Testing Library, and Playwright for unit, component, and end-to-end coverage including mode isolation
- Backend companion: Small Node server for signed admin sessions and Polar checkout verification with signed webhooks
- Security: No embedded LLM, no arbitrary code execution, no cross-origin exposure, bounded descriptions, and strict annotation validation
Challenges we ran into
- Registry lifecycle isolation: Ensuring design-mode tools never leak into eval mode required careful
AbortControllermanagement and explicit human confirmation for mode switches - State-dependent tool registration: The
close_support_tickettool must appear only after a resolution note is added, requiring dynamic registry rebuilding on state changes - Annotation validation: Preventing
readOnlyHintfrom binding to mutating effects and ensuringuntrustedContentHintmatches actual user-authored content - Deterministic grading: Building a flexible grader that handles exact ordered calls, unordered groups, optional calls, and various argument matchers while remaining deterministic
- Browser compatibility: Adjusting for current WebMCP draft status and OpenAI's built-in browser limitations (declarative tools not discovered, iframe registration unsupported)
Accomplishments that we're proud of
- Complete two-mode architecture: Design and eval modes with full registry isolation and visible mode switching
- Live agent integration: Successfully tested with Codex in-app Browser WebMCP, achieving 100/100 grade on urgent triage case
- Comprehensive lint engine: 12+ deterministic checks covering contract quality, ambiguity, and security signals
- Seven-dimension grader: Separate scoring for tool selection, parameters, order, prohibited calls, executor success, final state, and UI effects
- Seeded project with deliberate flaws: Starter contract demonstrates real problems (overlapping tools, vague descriptions, missing enums, annotation issues) that users can fix and test
- Full test coverage: 5 test suites, 11 unit tests, and 1 Playwright journey all passing
- Security boundaries: No embedded LLM, no arbitrary code execution, same-origin only, bounded outputs, and strict validation
- Apple-inspired UI: Restrained typography, translucent chrome, grouped inspectors, and platform-style controls using only local fonts
What we learned
- WebMCP is still a Community Group draft, not a W3C Standard—documentation can change and browser support varies
- Annotations are hints for agents, not authorization systems—normal server and application authorization must remain authoritative
- Deterministic linting cannot predict every model behavior, but it can dramatically reduce ambiguity and common failure modes
- Mode isolation is critical: never expose authoring tools and draft tools simultaneously
- State-dependent tool registration requires careful lifecycle management and explicit UI feedback
- Browser agents need clear, bounded descriptions and schemas—vague tools lead to wrong selections
- Version guards on mutating commands prevent race conditions in multi-agent scenarios
- Untrusted content must be explicitly labeled throughout the stack (schemas, results, UI)
What's next for it
- Additional seeded projects: Expand beyond the support-ticket domain to demonstrate WebMCP contracts for file systems, databases, and external APIs
- Enhanced lint rules: Add more sophisticated ambiguity detection, semantic overlap analysis, and security-focused checks
- Collaborative features: Multi-user project editing with conflict resolution and shared eval runs
- Advanced grading: Add machine learning-based pattern recognition for common agent failure modes
- Export formats: Generate OpenAPI specs, TypeScript types, and other contract representations from WebMCP definitions
- Browser agent marketplace: Curated library of tested WebMCP tool contracts for common domains
- Real backend integration: Option to connect contracts to actual APIs while maintaining safety boundaries
Built With
- webmcp
Log in or sign up for Devpost to join the conversation.