
How we taught an AI to pass the insurance agent’s exam

Methodology
General-purpose AI is astonishing and, for high-stakes insurance questions, not enough. GPT-class models write beautiful sentences about insurance. But ask what a specific waiting-period clause means for your mother’s knee surgery and a wrong answer isn’t a hallucination — it’s a family making a decision on fiction. Insurance advice has zero tolerance for confident nonsense.
So we built Hibiscus, an engine specialized for exactly one domain: Indian insurance.
Hibiscus is not one model. It is fourteen specialist agents working as one intelligence on a nine-node orchestration pipeline — a policy analyzer, a hospital-bill auditor, a grievance navigator, a surrender calculator, a risk detector, a regulation engine, and others — each doing one job well, coordinated like a team.
Underneath sits the knowledge that generic models lack: a vector corpus of over five and a half lakh passages built from IRDAI-filed product documents, and a knowledge graph of 7,516 insurance products with their riders, exclusions, and regulatory lineage. When Hibiscus answers, it answers from documents that exist — and cites the clause.
Then we made rules that most AI products won’t make. Every answer must carry a label: from-your-document, general-rule, or needs-confirmation. If Hibiscus cannot ground a claim, the claim is blocked — not softened, blocked. In adversarial testing across 53 attack categories, 100% of jailbreak and data-exfiltration attempts were stopped.
And then, the exam. India licenses its insurance advisors through the IC-38 syllabus. We benchmarked Hibiscus on the 280-question life-insurance agent bank — the same material every human agent studies — inside a 700-question suite spanning life, health, and general insurance. It scored 88.9%.
We publish that number for a simple reason: trust in AI should be earned the way trust in people is earned — by taking the test, in public, and showing your work. The full technical report on Hibiscus’s architecture, benchmarks, and guardrails is available to anyone who wants to check ours.
Every Eazr Protection Score is computed by this engine. That’s what it means when we say the score is auditable: not “trust us” — “check us.”
The Hibiscus Technical Report (v1.0, June 2026) details the architecture, methodology, and full benchmark results.
Share this article
Relevans posts
Your family’s score is waiting.
Three minutes. One policy. The truth about the people you love.

DPDP-aligned

Zero commissions
Welcome to Eazr
What would you like to know about your family’s protection?
Explain my policy like I’m not an expert
Score
Policies
Claims
Renewals
Summarize our product in simple terms for new users
Draft a friendly support reply using our help docs


