Inspiration

Insurance carriers receive more submissions than their underwriters can carefully investigate. Each submission may require information from policies, claims, insureds, locations, buildings, and external risk sources before it can be compared with the carrier's appetite.

We built UnderwriteIQ to answer one question: Which submissions deserve an underwriter's attention first, and why? and how can we do this in a time and token efficient way?

What it does

UnderwriteIQ is an AI underwriting agent that:

  1. Dynamically discovers Federato’s schema and retrieves relevant policy, claims, location, and building data through querying the given API endpoint in an agentic loop until enough evidence is found.
  2. Enriches submissions with real-world risk signals.
  3. Evaluates risks against carrier-specific appetite rules.
  4. Classifies and ranks submissions as "Target", "Acceptable", "Needs Review", or "Out of Appetite".
  5. Explains every recommendation with the determining rules and supporting evidence. Its custom Qwen3-8B model was distilled from 10,900 GPT-5.6 Sol-reviewed decisions and LoRA fine-tuned on Baseten. It achieved 96.1% disposition accuracy, 92.6% exact match, 95.9% rule-attribution F1, and 100% valid JSON on held-out tests.

How we built it

UnderwriteIQ begins by calling Federato’s schema-discovery endpoint to map available resources, fields, and relationships. It then uses validated tools to retrieve connected submission, policy, insured, claims, location, and building data. Every query is checked against the discovered schema before execution, and available external risk signals can be incorporated without inventing missing evidence. We use a hybrid decision architecture:

  • The AI agent selects tools, gathers evidence, and identifies missing or conflicting information.
  • A deterministic rule engine applies eligibility requirements, numerical thresholds, and appetite preferences to produce the final classification.
  • Evidence provenance is preserved, unsupported conclusions are rejected, and unresolved cases are escalated for human review.

For underwriting classification, we distilled GPT-5.6 Sol’s decision quality into Qwen3-8B. GPT audited 1,900 benchmark decisions, producing 89 corrections that were propagated across 241 equivalent cases. We then created leakage-safe splits containing 692 training, 217 validation, and 231 held-out test examples. Qwen3-8B was fine-tuned with LoRA on a single Baseten H100 using BF16 precision, assistant-only loss, gradient checkpointing, and a strict JSON output contract. On the held-out test set, it achieved:

  • 96.10% disposition accuracy
  • 92.64% exact-match accuracy
  • 95.87% rule-attribution micro-F1
  • 100% valid JSON output The trained adapter is served through Baseten’s OpenAI-compatible vLLM endpoint. Results appear in a ranked underwriting queue where users can inspect matched preferences, failed requirements, missing information, recommended actions, supporting evidence, and the complete agent tool trace.

InsuranceBench

We also created InsuranceBench, a human-verified evaluation suite for underwriting agents.

We manually curated representative cases and defined the correct classification, ranking, applicable rules, required evidence, and expected actions. InsuranceBench then runs the same cases through different models and measures:

  • Classification accuracy
  • Ranking quality
  • Guideline adherence
  • Tool and query correctness
  • Evidence retrieval
  • Explanation faithfulness
  • Structured-output reliability
  • Latency
  • Inference cost

Objective results are scored deterministically. A language-model judge is used only for subjective qualities such as explanation clarity and usefulness.

Results

Metric UnderwriteIQ Qwen3-8B GPT-5.6 Sol
Model-call latency 581 ms 18,966 ms
Input tokens 480 44,976
Output tokens 18 1,621
Total tokens 498 46,597
Final rule result out_of_appetite needs_review
Material evidence Found renewal and ineligible state Retrieval query failed
Model Disposition accuracy Exact match Rule micro-F1 Valid JSON
GPT-5.6 Sol 94.30% 91.34% 93.67% 100%
UnderwriteIQ Qwen3-8B 96.10% 92.64% 95.87% 100%

The model also achieved 99.57% schema compliance. Base-Qwen and frontier-model results should only be added after running them on the same held-out split; valid tool calls, p95 latency, and relative cost were not measured by this classification benchmark.

Challenges we faced

The hardest challenge was proving that each underwriting decision was correct, not simply generating a convincing response. We built an evaluation harness with schema validation, deterministic rule checks, evidence tracking, and explicit handling of missing or conflicting data.

Creating reliable training data was equally difficult. GPT-5.6 Sol audited 1,900 generated labels, while deduplication and leakage-safe splits ensured we measured generalization rather than memorization.

One problem we encountered was the model weights not completely fitting within memory all at once. To fit Qwen3-8B on one H100, we used LoRA, BF16, gradient checkpointing, gradient accumulation, and most importantly: chunked evaluation. A shared test harness then compared models using structured-output validity, disposition accuracy, exact match, rule-attribution F1, latency, and token usage.

What we learned

We learned that building the agent harness was as difficult as building the model. The harness had to manage dynamic schema discovery, validate and repair generated queries, enforce tool and turn budgets, track evidence across related records, preserve provenance, and recover safely from incomplete data or failed tool calls. We also had to clearly separate responsibilities: the AI decides what evidence to retrieve, while deterministic code owns the final underwriting rules and ranking.

Agent evaluation also required more than checking the final answer. We built structured benchmark cases and execution traces to measure output validity, decision accuracy, rule attribution, evidence coverage, latency, token usage, and tool-call efficiency. Effective agent evals must test both the outcome and the trajectory, whether the agent selected the correct tools, queried valid fields, improved the evidence ledger, handled uncertainty properly, and reached a supported conclusion without unnecessary calls. The biggest lesson was that trustworthy domain agents depend on the entire system: high-quality data, constrained tools, deterministic safeguards, traceable evidence, and evaluations designed to reveal failures that fluent model responses might otherwise hide.

What's next

Next, we would expand InsuranceBench across more lines of business, carrier guidelines, external risk sources, and adversarial cases involving incomplete or contradictory evidence. We would also add portfolio-level constraints, underwriter feedback loops, agent-trajectory evaluations, and production monitoring for accuracy, schema drift, latency, and cost.

We also plan to distill the underwriting workflow into multiple smaller open models and compare different model sizes, architectures, quantization levels, and LoRA configurations. This would let us route simple classifications to lightweight, inexpensive models while reserving larger models for complex or ambiguous cases reducing inference cost and latency without sacrificing reliability.

Our long-term goal is a dependable underwriting AI that can evaluate every submission affordably at scale while keeping each recommendation transparent, traceable, and under human control.

Built With

Share this project:

Updates

Submission history