Skip to content

About

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

mcp_analysis — MCP Specification and Headless Surface Validator

Usage guide · Project website · Source repository

License: Copyright © 2026 Fluency Corp. This is publicly viewable, proprietary source-available software—not open-source software. Internal non-production evaluation is permitted under the Fluency Commercial Source License 1.0; production, commercial, hosted-service, and redistribution rights require a separate written license from Fluency Corp.

mcp_analysis validates an MCP in two deliberately separate layers:

  1. MCP specification conformance — static checks grounded in the official MCP 2025-11-25 Tools specification and initialization lifecycle, plus the official Logging and Pagination utilities.
  2. Optional Headless design checks — opinionated checks for agent routing, outcome-oriented capabilities, progressive disclosure, guarded mutations, version transparency, MCP-delivered skills, and interactive UI presentation.

The specification layer is product-agnostic and always runs. Headless rules are configurable because a filesystem, database, calendar, developer tool, and multi-tenant security platform should not all be forced into one vocabulary.

Python 3.10+ is the only runtime dependency. A desktop interface lives in desktop/ and requires Node 20.19+.

Quick start

# offline, against a bundled example
python3 mcp_validate.py --dump examples/headless_surface.json --profile profiles/headless.json

# a real server, saving the surface so it can be re-analyzed offline
python3 mcp_validate.py --stdio "uv --directory /path/to/mcp run server" \
  --profile profiles/headless.json --save-surface surface.json --json assessment.json

Or open the desktop app, which lists every MCP already installed on the machine and needs no paths at all:

cd desktop && npm install && npm run dev

Every run answers three questions:

  1. How much passed? — 3,367 of 4,033 checks passed (83.5%), with a denominator on every rule rather than a bare finding count.
  2. Did I see the whole catalog? — complete, truncated, or unverified.
  3. What should I fix first? — a shortlist ranked by the coverage each step recovers, blocking work first.

The full guide is DOCUMENTATION.md.

How to use this repository

Use this repository to capture an MCP's advertised surface, assess it against the selected standard, save a complete machine-readable result, and turn that result into an actionable developer report.

1. Verify the checkout

Run the validator tests from the repository root. No package installation is required:

cd /path/to/mcp_analysis
python3 -m unittest discover -s tests -q

2. Choose the MCP evidence source

Use the strongest source you can access:

  • --stdio initializes and inspects a local MCP process. This is normally the best choice for a local development checkout.
  • --http initializes and inspects a deployed Streamable HTTP MCP endpoint.
  • --dump reads an existing JSON tool catalog or surface-evidence bundle. It is useful for repeatable offline analysis, but it proves only what the bundle captured.

Examples:

python3 mcp_validate.py --stdio "uv --directory /path/to/mcp run server"
python3 mcp_validate.py --http https://mcp.example.com/mcp --bearer "$MCP_TOKEN"
python3 mcp_validate.py --dump examples/agnostic_surface.json

Do not place tokens or sensitive tenant data in saved commands, reports, or surface bundles.

3. Choose the validation policy

  • Omit --profile for product-agnostic MCP specification and general quality checks.
  • Use profiles/headless.json for the complete optional Headless conventions.
  • Use profiles/fluency.json only for Fluency's tool names and product-specific expectations.

A profile changes policy, not protocol truth. Specification findings remain separate from quality heuristics and optional extension findings.

4. Save the complete assessment

Always save JSON when the result will be compared, remediated, or reported:

python3 mcp_validate.py \
  --dump surface.json \
  --profile profiles/headless.json \
  --json reports/assessment-before.json

The terminal dashboard is a prioritized summary. The JSON file contains every finding and is the source for automation and HTML reporting. Alongside findings it carries:

  • checks — the ledger: per rule and per layer, how many targets were evaluated and how many passed. This is what makes a coverage percentage possible, and what a finding count alone can never tell you.
  • catalog — whether the collected tool list is the whole catalog (complete, truncated, or unverified), plus tool_names.
  • advice.next_steps — an ordered shortlist with the coverage each step recovers and the projected pass rate if it lands.
  • advice.recommendations — why each issue matters, what to change, how to stay backward compatible, and how to verify the improvement.

Interpret the exit code as follows:

  • 0: no FAIL findings;
  • 1: one or more validation failures;
  • 2: the MCP surface could not be collected or the profile was invalid.

Warnings do not fail the command by default, but they remain developer work.

5. Generate the developer HTML report

python3 .codex/skills/mcp-developer-report/scripts/build_report.py \
  --assessment reports/assessment-before.json \
  --output reports/mcp-developer-report.html \
  --target "Example MCP" \
  --source-ref "commit or deployed build" \
  --command "sanitized validator command" \
  --profile profiles/headless.json \
  --test-result "test command and outcome"

Open reports/mcp-developer-report.html in a browser. The report shows the verdict, layer health, mentor guidance, rule clusters, affected tools, evidence metadata, and validation limitations. Add --baseline reports/assessment-before.json when rendering a comparable after-assessment to show the change in FAIL, WARN, and INFO counts.

6. Improve and reassess the MCP

Fix specification failures first, then quality warnings and selected extension gaps. Preserve the original assessment, use the same collector and profile, and write the new result to a separate file such as reports/assessment-after.json. Before renaming a tool, search the target MCP's current skills and instructions for that name. Prefer additive schema, pagination, metadata, logging, and UI changes so existing clients continue to work.

For the complete agent procedure and evidence checklist, use AGENT_RUNBOOK.md. Rule definitions and their rationale are in RULES.md.

Repository map

  • mcp_validate.py — collectors, validation rules, terminal dashboard, and JSON output.
  • profiles/ — optional policy overlays.
  • examples/ — sample surface bundles.
  • .codex/skills/mcp-developer-report/ — reusable agent instructions and the HTML report builder.
  • tests/ — validator, reporting, and repository-policy regression tests.
  • reports/ — generated JSON, HTML, and PDF artifacts.
  • desktop/ — Electron interface over the validator; see desktop/README.md.
  • DOCUMENTATION.md — usage guide: sources, rule sets, reading a result, and deciding what to fix first.

Agent and developer-report workflow

Future agents should begin with AGENT_RUNBOOK.md. It defines the evidence hierarchy, before/after process, backward-compatibility rules, runtime evidence checklist, and developer handoff standard. A repository-local skill at .codex/skills/mcp-developer-report/ makes the workflow discoverable and bundles a deterministic, dependency-free HTML report builder.

Build a self-contained developer report from the complete JSON assessment:

python3 .codex/skills/mcp-developer-report/scripts/build_report.py \
  --assessment reports/assessment-after.json \
  --baseline reports/assessment-before.json \
  --output reports/mcp-developer-report.html \
  --target "Example MCP" \
  --source-ref "commit or deployed build" \
  --command "sanitized validator command" \
  --profile profiles/headless.json \
  --test-result "test command and result"

The baseline is optional. Unknown source, command, profile, or test metadata is shown as Not recorded rather than guessed. The generated report separates MCP specification findings from quality heuristics and optional extensions, includes every finding grouped by rule and tool, and has responsive and print styling.

Command reference

Validate a tools/list dump or a complete surface bundle:

python3 mcp_validate.py --dump tools.json
python3 mcp_validate.py --dump examples/agnostic_surface.json
python3 mcp_validate.py --dump examples/headless_surface.json --profile profiles/headless.json

Initialize and enumerate a stdio MCP server:

python3 mcp_validate.py --stdio "uv --directory /path/to/project run server"

Initialize and enumerate a Streamable HTTP server:

python3 mcp_validate.py --http https://host.example/mcp --bearer TOKEN

Apply the reusable Headless standard or a product-specific profile:

python3 mcp_validate.py --dump surface.json --profile profiles/headless.json
python3 mcp_validate.py --dump surface.json --profile profiles/fluency.json

Write all findings as JSON:

python3 mcp_validate.py --dump surface.json --json report.json

The default terminal output is a visual dashboard with layer scorecards, severity bars, dominant rule groups, prioritized findings, and extension coverage. Adjust its width and detail budget when needed:

python3 mcp_validate.py --dump surface.json --width 120 --max-show 8

The JSON report remains the complete automation interface; the dashboard is a human-oriented summary and deliberately does not print hundreds of repetitive findings.

Exit code 0 means no FAIL findings, 1 means validation failures, and 2 means the surface could not be collected or the profile was invalid. WARN findings do not fail CI by default.

Surface bundle format

A bare tool array and a normal { "tools": [...] } tools/list result are accepted. To validate lifecycle metadata, runtime errors, logging, and pagination, capture the initialization result and optional runtime evidence:

{
  "initialize": {
    "protocolVersion": "2025-11-25",
    "capabilities": {"tools": {}, "logging": {}},
    "serverInfo": {"name": "example", "version": "1.2.3"},
    "instructions": "A concise map of the server's capability families."
  },
  "tools": [],
  "toolsListPages": [{"tools": []}],
  "events": [
    {
      "direction": "server_to_client",
      "message": {
        "jsonrpc": "2.0",
        "method": "notifications/message",
        "params": {"level": "error", "logger": "database", "data": {"message": "sanitized"}}
      }
    }
  ],
  "toolResults": [
    {
      "tool": "example_tool",
      "case": "execution_error",
      "result": {
        "content": [{"type": "text", "text": "Correctable error and remedy"}],
        "isError": true
      }
    }
  ]
}

Live stdio and HTTP collection captures initialization and every tools/list page automatically. When resources are advertised, it also captures every resources/list page so MCP Apps references can be resolved. Tool-call results and asynchronous events require an integration harness or evidence bundle.

Specification validation

The spec layer checks the statically observable parts of MCP revision 2025-11-25:

  • tools/list returns an array and every tool is an object;
  • names are present, unique, and follow the specification's recommended format;
  • inputSchema is present with an object root;
  • optional outputSchema has the object root required by this revision;
  • schema properties, required, and $schema have valid container types;
  • titles, descriptions, annotations, and task-execution declarations have the protocol-defined types;
  • a captured initialization result contains the negotiated protocol revision, capabilities, and serverInfo.name / serverInfo.version;
  • a server returning tools advertised the tools capability;
  • captured tools/list pages use opaque string cursors and reach a final page;
  • captured standard log notifications have valid direction, severity, logger, and data fields and are backed by the logging capability;
  • captured execution failures return isError: true with client-visible diagnostic content.

This is a focused static validator, not a complete JSON Schema implementation or wire-protocol certification suite. Live result validation, authorization, uncaptured notifications, cancellation, tasks, and every transport/security requirement need integration tests.

Logging, issue reporting, and pagination

Standard MCP logging is one-way: servers advertise logging, clients may call logging/setLevel, and servers emit notifications/message. The Headless profile adds the reverse path as a clearly labeled extension: initialize.instructions tells the client when to call report_client_issue, which sends a sanitized problem report to the server and returns a receipt.

Server -- notifications/message --> Client
Client -- report_client_issue ----> Server

The official pagination utility applies to MCP list operations such as tools/list. Live collection follows every nextCursor; captured pages are validated when supplied. For arbitrary tool results, an unbounded array in outputSchema triggers a quality warning unless the tool exposes an opaque cursor input and continuation output. A genuinely fixed or small result can instead document that it is bounded.

Headless extensions

The project defines five conventions that are useful but not requirements of the MCP specification. profiles/headless.json enables them.

H1 — Progressive capability cheat sheet

The initialization result should contain concise instructions that describe the server's major capability families. A read-only tool such as describe_capabilities provides the detailed, retrievable cheat sheet. This keeps the initial context small while allowing the agent to pull more routing guidance only when it needs it.

The tool name is configurable. Fluency uses routing_cheatsheet.

H2 — Transparent version compatibility

The protocol-standard serverInfo.version identifies the running server during initialization. A Headless version tool, such as inspect_version_compatibility, should additionally return the versions of the server surface, capability contract, and published skills, plus compatibility or staleness judgments. That lets an agent determine whether its connection or loaded guidance is behind the server it is operating.

H3 — MCP-delivered skill loading

The MCP should expose a read-only skill catalog and content loader, such as list_skills and load_skill. load_skill(skill_id) returns the selected skill instructions and version through MCP, eliminating a separate manual download/upload path.

MCP itself does not define a portable “Skill” primitive and cannot guarantee that every host will install instructions into a host-specific skill registry. The portable contract is therefore retrieval: the model can discover and load versioned instruction content through tools. A host may add native installation as a separate, explicitly authorized integration.

H4 — Client-to-server issue reporting

The initialization instructions identify a configurable additive tool such as report_client_issue. It accepts a concise summary and sanitized details, then returns an issue_id or report_id for correlation. This complements standard server-to-client logging; it does not redefine the MCP logging utility.

H5 — UI presentation through MCP Apps

MCP Apps is the first official MCP extension. A UI-capable tool declares _meta.ui.resourceUri, pointing to a ui:// resource containing its HTML/JavaScript interface. This lets hosts render dashboards, forms, visualizations, viewers, and multi-step workflows in the conversation.

The validator can require minimum UI coverage and particular presentation tools. It verifies the ui:// URI, resolves it through the captured resource catalog, checks for an HTML MIME type, and requires a description plus outputSchema fallback for clients without MCP Apps support. Profiles should select tools whose output benefits from exploration or human review; not every low-level tool needs its own UI.

Recommended skill metadata:

{
  "skill_id": "case-investigation",
  "skill_version": "2.1.0",
  "compatible_server": ">=1.4,<2",
  "content_type": "text/markdown",
  "content": "...",
  "content_sha256": "..."
}

Profiles

Profiles configure policy without changing protocol truth. The generic default does not reject legitimate resource-oriented tools or require a dry-run twin for every write. profiles/headless.json turns on the full opinionated design model. profiles/fluency.json maps the same extension roles to Fluency's names and adds its multi-tenant scope, guarded-query expectations, and MCP Apps case-card presentation requirement.

Name-based quality rules operate on a normalized tool name. A profile can declare an exact prefix with "tool_namespace": "discord_"; otherwise the validator infers prefixes shared by at least two tools when the remaining name begins with a recognized operation. Findings always retain the original callable name. Profiles can exclude tools that have no meaningful runtime failure path from J10 with "j10_exempt": ["describe_capabilities"], analogous to n1_exempt.

See RULES.md for the rule catalog and rationale, and AGENT_RUNBOOK.md for a reproducible assessment and reporting procedure.

Licensing and ownership

Fluency Corp. retains ownership of this repository even when its source is published in a public repository. Public visibility permits inspection and the limited evaluation activities stated in LICENSE; it does not grant commercial or production rights. External contributions require prior written authorization and an intellectual-property assignment as described in CONTRIBUTING.md.

About

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages