We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Phone agents are becoming good at completing real-world tasks, but placing a call is only half of the problem. After the call ends, an operator still needs reliable answers to harder questions: Did the call achieve the exact objective? Did the agent stay within the authorized limits? Can the result be traced back to what the recipient actually said? And what should happen when the evidence is incomplete or contradictory?

We built CallProof because a provider's structured result should be treated as a claim, not as proof. A phone agent may report that a delivery was moved or that a surcharge was accepted, while the transcript tells a different story. In a system with real-world side effects, silently accepting that mismatch is unsafe.

Our goal was to close that trust gap. CALL-E gives an agent the ability to act over the phone; CallProof adds an auditable control loop around that action so an operator can define boundaries before the call, verify evidence afterward, and route uncertain outcomes to a human instead of guessing.

What it does

CallProof is a closed-loop verification system for goal-driven outbound calls.

Before a live call, the operator enters the recipient, objective, and permitted commitment limit. CallProof creates an immutable CallContract, masks the phone number in the interface, displays a no-side-effect preview, and requires a request-specific typed confirmation. A real CALL-E call is possible only when the provider, live-call switch, credentials, authenticated operator, and HTTPS requirements all pass server-side checks.

After CALL-E returns a terminal result, CallProof normalizes the provider response and verifies the recipient status, phone binding, task-completion evidence, confidence, and transcript. It then submits the transcript, structured result, and original contract to Call Analyzer.

Call Analyzer does not use keywords or free-form language as automatic proof. Its deterministic evaluator checks a finite protocol generated from typed contract claims: the provider value must equal the expected value, the agent must make the exact canonical confirmation statement, and the recipient must give an exact adjacent response. If the recipient later says anything outside that closed protocol, the earlier confirmation is no longer considered final and the result is routed to human review. Rules that require open-ended semantic interpretation are also reported as unevaluated rather than assumed safe.

AgentKit Rails then persists the workflow, evidence, audit trail, telemetry, and human-in-the-loop decision. The default demo uses fictional data and deterministic adapters, so anyone can run both a compliant scenario and a policy-violation scenario without credentials or real phone calls.

How we built it

CallProof is a monorepo with three independently owned planes connected through versioned contracts:

  • Execution plane: CALL-E performs the outbound phone task and returns status, transcript, and structured results. A Ruby provider adapter sends one idempotent create request and uses the documented retrieval endpoint for terminal recovery.
  • Control plane: A Rails application uses AgentKit Rails for resumable orchestration, persisted human-review gates, audit records, memory boundaries, and telemetry. PostgreSQL with pgvector stores product and workflow state.
  • Evaluation plane: A Python FastAPI service accepts analysis requests asynchronously, persists received → analyzing → completed states, runs inline for deterministic development or through Redis/RQ workers, and returns HMAC-signed completion webhooks with timestamp and replay protection.

Rails and Python do not share database tables. They communicate through strict JSON Schema 2.0 envelopes containing stable identifiers, correlation metadata, immutable typed claims, transcript turn references, and explicit confidence and review fields. Missed analyzer webhooks can be recovered through an idempotent status endpoint.

The live workflow was designed as a two-step operation. Creating a preview has no external side effect. Confirming it uses one stable idempotency key for one intended call. Timeouts, server errors, and idempotency conflicts remain unresolved until the original request can be reconciled; CallProof never converts an ambiguous create attempt into a replacement call.

We also built normalizers for the different terminal shapes returned by CALL-E's REST and MCP surfaces, a complete fake provider and fake analyzer for offline demonstrations, shared fixtures based on a redacted controlled call, and adversarial regression tests for unsafe transcript interpretations.

Challenges we ran into

The hardest challenge was defining what “verified” should mean. Early deterministic approaches relied on topical words, affirmation lists, stemming, or nearby turns. Adversarial examples repeatedly showed why that was unsafe: an agent's own question could look like supporting evidence, words such as “correct” could appear inside a denial, and a canonical “yes” could be retracted later in the conversation. We ultimately replaced open-ended lexical judgment with a closed proof protocol and a fail-closed human-review path.

Real call creation introduced a different class of ambiguity. A timeout does not prove that no call was placed, and a 409 idempotency response may mean the original request exists. Likewise, a later rejected retry says nothing conclusive about whether the first attempt was accepted. We had to model unresolved as a real operational state and reconcile only through the stable idempotency key instead of reporting a reassuring but unsupported failure.

Securing the live path also required defense in depth. Preview URLs expose objectives and recipient metadata, so operator-created and live records must remain authenticated even across migrations and legacy data. We added operator ownership, live-record backfills, authorization based on both live and operator-initiated state, masked recipient displays, enforced HTTPS in production, and explicit readiness gates that cannot be bypassed by changing the browser UI.

CALL-E's REST and MCP interfaces return different result shapes, which made normalization and evidence binding more involved than expected. We also chose not to invent a CALL-E webhook signature verifier when the accessible public reference did not specify enough details; the current durable path uses the documented retrieval endpoint instead.

Finally, integrating AgentKit Rails exposed framework-level details around engine load order, Ruby compatibility, and serialization of persisted human-review suggestions. Solving those issues was essential for making the workflow genuinely resumable rather than merely demonstrating a happy path.

Accomplishments that we're proud of

We are proud that CallProof demonstrates an end-to-end safety loop rather than a single API call. An operator can preview a contract, execute or simulate a call, normalize the terminal result, evaluate typed claims, inspect transcript-linked evidence, and resume a persisted human-review decision.

The project is safe to evaluate by default: automated tests, the local stack, previews, and demo scenarios place no calls and require no external credentials. Live execution is isolated behind explicit, auditable opt-in controls and stable idempotency.

We also completed one controlled live test with a consenting recipient through the official CALL-E surface. The call completed in 28 seconds, returned task_completed=true with 0.93 confidence, included the recording disclosure, made no unauthorized commitment, and produced a redacted fixture that now drives our normalizer tests. Earlier carrier-declined attempts were correctly treated as failed and unverifiable instead of being sent to the analyzer as successful calls.

The exact-protocol evaluator is another major accomplishment. It now rejects provider-only success claims, unsupported paraphrases, agent-authored “evidence,” ambiguous policy assertions, explicit denials, conflicting amounts, and later retractions. Unknown meaning remains unknown and requires a human.

Finally, we packaged CallProof as a runnable reference application for the awesome-phone-call-agents repository. The current verification suite includes 41 Python tests and 47 Rails tests, together with repository validation, RuboCop, Zeitwerk, and Brakeman checks.

What we learned

We learned that structured extraction and evidence are fundamentally different things. A provider result is useful for locating a claim, but the original contract and recipient-side transcript evidence must determine whether that claim is accepted.

We also learned that deterministic natural-language shortcuts are dangerous precisely because they look convincing on ordinary examples. Negation, morphology, speaker attribution, conversational context, and later retractions make open language impossible to verify safely with a growing collection of word lists. A small protocol that can prove a narrow claim is more trustworthy than a broad evaluator that sometimes invents certainty.

Side effects need a richer state model than success or failure. In distributed systems, “we do not know whether the provider accepted the call” is a valid and important outcome. Preserving that ambiguity prevents accidental duplicate calls and misleading operator messages.

Human review works best as an explicit boundary, not as a patch applied after the fact. The contract should define in advance which claims can be proven mechanically, which policies require semantic judgment, and which outcomes can become accepted evidence or future learning signals.

Most importantly, safety cannot live only in prompts or interface controls. It has to be enforced across authentication, persistence, migrations, transport security, idempotency, schemas, provider reconciliation, evaluation logic, and the human decision ledger.

What's next for Call Proof

The next step is to complete the learning loop: approved human corrections will become scoped, provenance-rich AgentKit preferences that influence later comparable contracts without silently expanding the agent's authority.

We also plan to build a live call timeline, a plan-versus-actual evidence view, and dedicated screens for learned policies and pending human decisions. These interfaces will make the verification trail understandable without requiring operators to inspect raw payloads.

For production deployment, we will move analyzer persistence from local SQLite to PostgreSQL, add multi-host worker hardening, and integrate an official signed CALL-E webhook flow when its verification contract is available. We also want provider-supported cancellation and stronger reconciliation by stable idempotency key.

On the evaluation side, we will preserve the deterministic protocol as the trusted baseline while exploring a constrained semantic evaluator for individual claims. Any semantic result must cite recipient-side evidence, expose uncertainty, and fall back to human review rather than overriding an unresolved deterministic verdict. Audio ingestion can later provide an independent fallback when transcript quality is insufficient.

Finally, we plan to expand the adversarial golden set, add more bounded use cases beyond delivery changes, deploy a free-to-test build, and measure how often CallProof prevents unsupported outcomes while reducing unnecessary human interruptions over time.

Share this project:

Updates

Submission history