Plan an insurer's first gen AI initiative — decision exercise
Exercise: plan an insurer's first gen AI initiative
Original fictional scenario. No cloud account, API calls or paid services are required. Difficulty: intermediate · Estimated duration: 20 minutes
Harbourline Insurance has four proposals on the table and budget for one pilot. All names and figures are invented.
| Proposal | What it would produce | Data available | Sponsor's stated goal |
|---|---|---|---|
| A — Claim-note digests | Short summaries of long adjuster notes | 5 years of notes, unlabeled | Cut adjuster reading time |
| B — Fraud flags | A risk flag on each new claim | 10 years of labeled claims in tables | Catch more fraud |
| C — Policy-letter rewrite | Plain-language versions of policy letters | Current templates | Fewer customer calls about letters |
| D — "An AI strategy" | Not specified | Not specified | "Be seen to use AI" |
Your decision
- For each proposal, name the kind of solution it needs — generative (text, image, code, personalization), traditional, pre-trained, or none — and why.
- Pick one proposal for a generative pilot and state its measurable business goal.
- List three integration steps beyond building the model that the pilot needs.
- Choose two success measures from different ROI families and say what baseline you would take before launch.
Rubric (10 house points)
- 3 points: classify A–D correctly, with the reason for each.
- 2 points: choose A or C for the generative pilot and state a measurable goal.
- 3 points: name users, process change and a release that can be rolled back (or stakeholder involvement).
- 2 points: two measures from different ROI families, each with a pre-launch baseline.
Reference solution
A needs a generative text solution (summaries). B is classification on structured, labeled history — a traditional model fits better than a generative one. C is generative text (rewriting). D has no measurable goal, so it is not ready for any approach yet.
Choose A or C. For A, a measurable goal could be "reduce average reading time per claim file by a third within two quarters". Integration needs the adjusters (end users) to know how to use and check the summaries, the claim-handling workflow changed so summaries arrive with each file, domain experts reviewing summary quality, and a release that can be switched off quickly if summaries mislead.
Measures: average handling time per claim (operational efficiency) and complaint or satisfaction scores on claim updates (customer experience), each measured for a period before launch.
Rejected alternatives. Running B as a generative pilot ignores that predictive work on structured data suits traditional AI. Starting with D builds before any goal exists. Counting "summaries generated" as success measures activity, not impact.
Sources and scope
All names, data and figures above are house-authored. The general principles are grounded in: