Tokens are the unit of cost, and agents spend them like a fleet, not a person. Every call pays for context in and text out; workflows make many calls per task, and fan-out multiplies them. Engineer cost with the same discipline as quality.
The figures below are costed for specialty work, where a single account carries far more document handling than a standardized risk. For why that is, and the carrier results behind it, see Specialty & E&S.
1 · Token economics: where the money actually goes
Four drivers set the cost of an agentic workflow, and none of them is the model's sticker price:
- Context per call. An agent that reads a 40-page submission packet per step spends those input tokens on every step. Context discipline (send what the step needs, not the file room) is the biggest lever most teams miss.
- Calls per task. A five-stage pipeline makes five-plus calls; a retry loop makes more. Orchestrated agent teams have been measured at roughly 15x the token cost of a single agent (Anthropic's multi-agent research write-up); fan-out multiplies spend roughly linearly.
- Model tier per step. Frontier models cost several times more than mid-tier ones. Extraction and classification rarely need the frontier; judgment steps sometimes do. Route by step, not by habit.
- Rework loops. Failed output is regenerated, and regeneration is spend. Evaluation gates stop repeated low-quality attempts before they reach a person.
Budget anchors. Mid-size enterprise LLM agreements run roughly $250k to $1M+ a year; volume discounts are material but must be verified in procurement (benchmark guide). Insurer IT spend averages ~4.5% of GWP (Datos Insights), and two-thirds of insurance CEOs plan to allocate 10–20% of budget to AI (KPMG CEO Outlook, PDF). Vendor purchases succeed about 67% of the time versus 33% for internal builds (MIT, PDF), supporting a buy-commodity, build-differentiation strategy.
Measure cost per completed task and dollars per underwriter hour returned, not price per token. A use case that cannot clear both bars is a demo. The winning comparison is machine cost per submission versus loaded human cost per submission reviewed, and it usually is not close.
The four drivers multiply, which is why sticker price misleads
None of these is the model's price per token. They compound: halving context and routing two steps down a tier does not add up, it multiplies down.
Worked anchor: orchestrated agent teams have been measured at roughly 15x the token cost of a single agent, because fan-out raises calls per task and each agent pays its own context. Measure cost per completed task, not price per token.
2 · The operational realities the demo never shows
| Consideration | What it looks like in production | The practical rule |
|---|---|---|
| Latency and broker SLAs | Multi-step agents take seconds to minutes; brokers notice turnaround, not your architecture | Put speed where the broker sees it (acknowledgment, triage, appetite answer) and depth behind it |
| Enterprise terms | Zero data retention and no-training clauses, audit logging, regional processing | No enterprise agreement, no company data. This is the first control in the governance section |
| Model deprecation and drift | Vendors retire and upgrade models; behavior shifts silently under a pipeline that passed its evals in March | Re-run the golden dataset on every version change, and pin versions where the provider allows it |
| Rate limits and surge | A CAT event is a volume spike; provider rate limits are a hard ceiling | Capacity-plan for surge events, and keep a queue-and-degrade path that keeps humans working when the cap is hit |
| PII and residency | Claimant and policyholder data in prompts and logs; state and partner rules on where it may be processed | Minimize and de-identify by default; log retention is a compliance surface, not an IT detail |
| Vendor viability | 95% of H1 2026 insurtech funding went to AI startups (funding data), which means consolidation is coming | Due diligence on funding and runway; an exit plan for any vendor whose output feeds a regulated decision |
| Lock-in and portability | A harness written to one provider's quirks is a migration project later; both core vendors shipped agentic frameworks in 2026 (Guidewire, Duck Creek) | Keep context specs, evals, and hooks model-agnostic; they are your portable assets |
3 · The most impactful use cases for insurance companies, ranked
Ranked by measured value divided by implementation risk, from the carrier and vendor results behind the integration phases (verbatim source passages for the top entries are on the evidence page).
| # | Use case | Why it pays (measured) | Where it sits |
|---|---|---|---|
| 1 | Submission intake and triage | 2–5x underwriting speed and 370k+ submissions a year at AIG; Markel's 113% productivity uplift; 50–97% faster processing and +15% hit ratios at Sixfold customers (evidence) | Underwriting; the proven first move |
| 2 | Document extraction and summarization | Loss runs, SOVs, claims files: routine 50–80% time cuts on document-heavy work; the foundation every other use case reads from | Everywhere; assistive, low scrutiny |
| 3 | Claims triage, severity and litigation prediction | Attorney-involved claims cost ~4.9x more (CLARA data); early triage moves both cycle time and indemnity | Claims; decision support with human authority |
| 4 | Bordereaux processing | 85–94% processing time savings (Verodat); the unglamorous pain point of program business | Program/delegated-authority operations |
| 5 | Fraud detection | 5x more fraud detected at Tokio Marine (case study, vendor-reported) | Claims; scoring with human review |
| 6 | Knowledge access (RAG over guidelines) | Appetite and procedure answers in seconds; multiplies every other use case by keeping context current | Enterprise-wide; internal only |
| 7 | Actuarial filing research and pricing workbench | Filing research from weeks to hours (Akur8); 13 pricing tools in 13 weeks at Allianz Commercial (hx) | Actuarial/pricing; assistive |
| 8 | Leakage and subrogation | AI pre-payment controls prevent 90–95% of detectable leakage (analysis); $15–20B of subrogation goes uncollected annually (industry estimate) | Claims finance; Phase 3 material |
| 9 | Bounded agentic quoting | 3 days to ~3 minutes at Hiscox London Market; CFC's agentic pilot (evidence). Real, but gated: human authority, governance gates, bounded segments only | Underwriting; last, under Phase 3 gates |
The same nine, arranged by how much scrutiny they demand
Grouped using the sequencing rule stated below the table and each entry's own "where it sits" column. Read left to right: the highest-ranked moves are also the least governed, and the one that needs the most governance is ranked last. Numbers are the value-over-risk rank from the table.
Start here · assistive
AI drafts, a person decides. Low scrutiny, highest confidence.
1Submission intake and triage 2Document extraction and summarization 6Knowledge access over guidelines 7Actuarial filing research and pricing workbenchThen · decision support
The model scores or recommends; a named human holds authority.
3Claims triage, severity and litigation prediction 4Bordereaux processing 5Fraud detectionLast · bounded automation
Gated work: bounded segments only, behind Phase 3 controls.
8Leakage and subrogation 9Bounded agentic quotingRanking answers what to do first; this arrangement answers what has to be in place before you can. A high rank never licenses skipping the governance a stage demands.
Sequencing rule: start assistive (ranks 1–2), move to decision support with human authority (3–5), and reach bounded automation last (9). That is the integration phases's phase logic applied to a single function; the harness section of the practice ladder is the build manual for each step.