AI in Insurance Testing: How Swiss Re Built an AI-Supported Testing Ecosystem

Recap of the talk by Stephany Daneri, Ananya Misra, and Amos Njoroge at Swiss Testing Day 2026

 

From 30 minutes per test case to 60 seconds, and pipeline failures analyzed in 5 instead of 60 minutes. The Quality Engineering team at iptiQ (Swiss Re) showed at Swiss Testing Day 2026 how they transformed their testing ecosystem with three AI-powered pillars. Without external tools, entirely in-house.

 

Key Takeaways

  • AI amplifies what’s already there: with weak foundations, it amplifies the weaknesses
  • Three pillars: Unified Visibility (E2E coverage + business risk), Smart Test Data (60s instead of 30 min), Pipeline Analyzer (5 min instead of 60 min)
  • The key was a Business Process Framework as a shared language across all stakeholders
  • MCP (Model Context Protocol) as an architectural decision for decoupling and extensibility
  • Human in the loop remains critical. AI hallucinations occurred especially with payment methods and European jurisdiction differences

 

The Starting Point: “Are Our Foundations Strong Enough?”

Stephany Daneri asked the decisive question: while most AI discussions in testing start with tools — faster automation, better pipelines — her team started differently: “Are our foundations strong enough? Because AI doesn’t fix a weak system. It only amplifies whatever already exists.”

On paper, everything looked mature: end-to-end tests, CI/CD, observability, incident management. But three friction points blocked real progress:

  1. Silo thinking: Stakeholders with different backgrounds had different views of the platform. Result: increasing incidents.
  2. Test data bottleneck: Insurance test data must replicate a policy’s lifecycle — up to 70 years. And only QE could create it.
  3. Pipeline analysis overhead: Every failure analysis took 30–60 minutes, multiple times daily.

 

Pillar 1: Unified Visibility — A Shared Language

Before: No correlation between business processes, tests, code, and incidents. After: A Business Process Framework as a shared language

The team asked stakeholders about their E2E test coverage. Everyone said 100%. The problem: theoretically high coverage doesn’t translate into understanding business risks. There was no meaningful correlation between business processes, test cases, code, incidents, and test reports.

The solution: A Business Process Framework as a YAML file; simple, versionable, jointly maintained with product management. AI scans all test code, understands the intent behind tests, and maps them to the business taxonomy. Key: explainability (why was a test classified this way?) and an experienced reviewer in the loop catching hallucinations.

 

Pillar 2: Smart Test Data — From 30 Minutes to 60 Seconds

The architecture: Slack as user interface, multi-agent orchestration with MCP (Model Context Protocol), full observability

Ananya Misra presented the “Policy Assistant”: anyone — QA, analyst, product engineer — can request an insurance policy via Slack in natural language. In 60 seconds instead of 30–45 minutes.

The architecture uses a multi-agent workflow: a Planner Agent (the “brain”) orchestrates a Creator Agent that communicates with APIs via MCP (Model Context Protocol). Deliberate decoupling: when APIs change, only the tools change, not the agents. Complete with OpenTelemetry tracing, Grafana monitoring, and cost tracking.

“This solution is not only for insurance: any domain with complex multi-step workflows where requirements keep changing.”

 

Pillar 3: Pipeline Analyzer — AI as a Diagnostic Tool

Before: A pipeline failure triggered a 30-60 minute odyssey. Slack ping, cryptic errors, wrong contact person, waiting

Amos Njoroge presented the third piece: a sidecar agent running alongside the pipeline. On failure:

  • Fingerprinting: identical errors aren’t reanalyzed (saves tokens/cost)
  • Parallel data collection: Docker logs, Jaeger traces, test results
  • Agent skills: schema analysis, log parsing, trace analysis
  • Report: Markdown with summary, affected service, probable root cause, and reproduction steps

 

Result: 15 failures → 1 root cause. Analysis time: 5 minutes instead of 30–60. Amos emphasized: “AI is not gatekeeping — it’s helping you as a tool.”

 

From the Audience: Questions and Answers

“How does the Pipeline Analyzer work in detail? Do you just dump log files to the AI?” Amos: The AI agent starts as a sidecar alongside the tests. On failure, it first analyzes the failed steps from the test report. Then it starts parallel data collection: Docker logs, trace logs from test results, Jaeger traces from the microservice architecture. All processed via agent skills. Everything runs in parallel within the same pipeline.

 

What This Means for Your Team

The message is clear: before introducing AI tools, check your foundations. A Business Process Framework as a shared language for all stakeholders is the basis. On top of that, you can build targeted AI solutions for visibility, test data democratization, and automated diagnosis. All in-house, all with internal talent.

 

About the Speakers

Stephany Daneri: Quality Engineering Manager at iptiQ / Swiss Re. Ananya Misra: Senior Software Engineer in Test at iptiQ / Swiss Re. Amos Njoroge: Software Engineer in Test at iptiQ / Swiss Re. Together, they support the iptiQ insurance platform and built the AI-supported testing ecosystem entirely with internal talent, from first learning steps to production-ready solutions.

Based on the talk “Building an AI-supported Testing Ecosystem in Insurance” at Swiss Testing Day 2026 in Zurich (Main Stage, 10:45–11:15). Conference theme: “Defining Quality in a Dangerous Decade”.

Swiss Testing Day — the leading conference for software quality in Switzerland.

 


 

KI im Versicherungstesting: Wie Swiss Re ein AI-gestütztes Test-Ökosystem aufgebaut hat

Rückblick auf den Vortrag von Stephany Daneri, Ananya Misra und Amos Njoroge am Swiss Testing Day 2026

Von 30 Minuten pro Testfall auf 60 Sekunden und Pipeline-Failures in 5 statt 60 Minuten analysiert.Das Quality Engineering Team von iptiQ (Swiss Re) zeigte am Swiss Testing Day 2026, wie sie mit drei KI-gestützten Säulen ihr Testing-Ökosystem transformiert haben. Ohne externe Tools, komplett in-house.

 

Key Takeaways

  • KI verstärkt, was bereits da ist. Bei schwachen Fundamenten verstärkt sie die Schwächen
  • Drei Säulen: Unified Visibility (E2E Coverage + Business Risk), Smart Test Data (60s statt 30 Min), Pipeline Analyzer (5 Min statt 60 Min)
  • Der Schlüssel war ein Business Process Framework als gemeinsame Sprache aller Stakeholder
  • MCP (Model Context Protocol) als Architekturentscheidung für Entkopplung und Erweiterbarkeit
  • Human in the Loop bleibt kritisch. KI-Halluzinationen traten besonders bei Zahlungsmethoden und europäischen Jurisdiktionsunterschieden auf

 

Der Ausgangspunkt: “Sind unsere Fundamente stark genug?”

Stephany Daneri stellte die entscheidende Frage: Während die meisten KI-Diskussionen im Testing mit Tools beginnen — schnellere Automation, bessere Pipelines — startete ihr Team anders: “Are our foundations strong enough? Because AI doesn’t fix a weak system. It only amplifies whatever already exists.”

Auf dem Papier sah alles gut aus: End-to-End-Tests, CI/CD, Observability, Incident Management. Aber drei Reibungspunkte blockierten echten Fortschritt:

  1. Silo-Denken: Stakeholder mit verschiedenen Hintergründen hatten unterschiedliche Sichten auf die Plattform. Ergebnis: steigende Incidents.
  2. Testdaten-Engpass: Testdaten für Versicherungen müssen den Lebenszyklus einer Police abbilden, bis zu 70 Jahre. Und nur QE konnte sie erstellen.
  3. Pipeline-Analyse-Overhead: Jede Failure-Analyse dauerte 30–60 Minuten, mehrmals täglich.

 

Säule 1: Unified Visibility — Eine gemeinsame Sprache

Vorher: Keine Korrelation zwischen Geschäftsprozessen, Tests, Code und Incidents. Nachher: Ein Business Process Framework als gemeinsame Sprache

Das Team fragte Stakeholder nach ihrer E2E-Testabdeckung. Alle sagten: 100 %. Das Problem: Theoretisch hohe Coverage übersetzt sich nicht in Verständnis der Geschäftsrisiken. Es gab keine sinnvolle Korrelation zwischen Business-Prozessen, Testfällen, Code, Incidents und Testreports.

Die Lösung: Ein Business Process Framework als YAML-Datei; einfach, versionierbar, gemeinsam gepflegt mit Product Management. KI scannt den gesamten Test-Code, versteht den Intent hinter den Tests und mappt sie auf die Business-Taxonomie. Wichtig: Explainability (warum wurde ein Test so klassifiziert?) und ein erfahrener Reviewer im Loop, der Halluzinationen erkennt.

 

Säule 2: Smart Test Data — Von 30 Minuten auf 60 Sekunden

Die Architektur: Slack als User-Interface, Multi-Agent-Orchestrierung mit MCP (Model Context Protocol), vollständige Observability

Ananya Misra präsentierte den “Policy Assistant”: Jeder — QA, Analyst, Product Engineer — kann via Slack in natürlicher Sprache eine Versicherungspolice anfordern. In 60 Sekunden statt 30–45 Minuten.

Die Architektur nutzt einen Multi-Agent-Workflow: Ein Planner Agent (das “Gehirn”) orchestriert einen Creator Agent, der über MCP (Model Context Protocol) entkoppelt mit den APIs kommuniziert. Bewusste Entkopplung: Wenn sich APIs ändern, ändern sich nur die Tools, nicht die Agents. Komplett mit OpenTelemetry Tracing, Grafana-Monitoring und Kostentracking.

“This solution is not only for insurance — any domain with complex multi-step workflows where requirements keep changing.”

 

Säule 3: Pipeline Analyzer — KI als Diagnose-Werkzeug

Vorher: Eine Pipeline-Failure löste eine 30-60-minütige Odyssee aus Slack-Ping, kryptische Fehler, falscher Ansprechpartner, Warten

Amos Njoroge zeigte das dritte Puzzlestück: Ein Sidecar-Agent, der parallel zur Pipeline läuft. Bei einer Failure:

  • Fingerprinting: identische Fehler werden nicht erneut analysiert (spart Tokens/Kosten)
  • Parallele Datensammlung: Docker Logs, Jaeger Traces, Test Results
  • Agent Skills: Schema-Analyse, Log-Parsing, Trace-Analyse
  • Report: Markdown mit Zusammenfassung, betroffenem Service, wahrscheinlicher Root Cause und Reproduktionsschritten

 

Ergebnis: 15 Failures → 1 Root Cause. Analysezeit: 5 Minuten statt 30–60. Amos betonte: “AI is not gatekeeping — it’s helping you as a tool.”

 

Aus dem Publikum: Fragen und Antworten

“Wie funktioniert der Pipeline Analyzer im Detail? Werden einfach Log-Files an die KI geschickt?” Amos: Der KI-Agent startet als Sidecar parallel zu den Tests. Bei einer Failure analysiert er zuerst die fehlgeschlagenen Schritte aus dem Test-Report. Dann startet parallele Datensammlung: Docker Logs, Trace Logs aus Testergebnissen, Jaeger Traces aus der Microservice-Architektur. All das wird über Agent Skills verarbeitet. Alles läuft parallel innerhalb derselben Pipeline.

 

Was bedeutet das für Ihr Team?

Die Botschaft ist klar: Bevor Sie KI-Tools einführen, prüfen Sie Ihre Fundamente. Ein Business Process Framework als gemeinsame Sprache aller Stakeholder ist die Basis. Darauf lassen sich dann gezielt KI-Lösungen aufbauen für Visibility, Testdaten-Demokratisierung und automatisierte Diagnose. Alles in-house, alles mit internem Talent.

 

Über die Speaker

Stephany Daneri: Quality Engineering Manager bei iptiQ / Swiss Re. Ananya Misra: Senior Software Engineer in Test bei iptiQ / Swiss Re. Amos Njoroge: Software Engineer in Test bei iptiQ / Swiss Re. Gemeinsam betreuen sie die Versicherungsplattform von iptiQ und haben das KI-gestützte Testing-Ökosystem komplett mit internem Talent aufgebaut. Vom ersten Lernschritt bis zur produktionsreifen Lösung.

 


 

Basierend auf dem Vortrag “Building an AI-supported Testing Ecosystem in Insurance” am Swiss Testing Day 2026 in Zürich (Main Stage, 10:45–11:15). Konferenzmotto: “Defining Quality in a Dangerous Decade”.

Swiss Testing Day — die führende Konferenz für Software-Qualität in der Schweiz.

AI in Insurance Testing: How Swiss Re Built an AI-Supported Testing Ecosystem

Swiss Testing Day 2027: Proof, Not Promises

From Test Data to Trust: How AI Governance Becomes Reality in the Software Lifecycle