Scaling Voice Biometrics Evaluation Across 20,000+ Vishing Attack Profiles for a Multinational Bank
With Polyagentic Autonomous Voice Agents.
Step 1:
Asymmetry Audit
We worked with the bank to identify the highest-leverage challenge: continuously strengthening its voice biometrics defenses against evolving vishing attacks. The objective was to evaluate four leading voice biometrics platforms against the incumbent system using 20,000+ attack profiles spanning landline, mobile, VoIP, and varying background noise conditions. To scale the evaluation, autonomous voice agents captured authentic caller voice profiles, generated AI voice clones, and autonomously placed follow-up calls to challenge each biometrics platform under realistic attack scenarios.
The outcome: A scalable evaluation framework that enabled the bank to identify the most resilient voice biometrics solution and continuously strengthen its defenses against evolving vishing threats.
Step 2:
Architecture Fit
To align with the bank's existing technology stack, we designed the solution around its Genesys contact center and real-time streaming architecture. The architecture leveraged bi-directional audio streaming through Confluent to orchestrate live transcription, large language model reasoning, AI voice generation, and voice biometrics evaluation as a unified streaming pipeline. It also incorporated replay capabilities through Twilio to repeatedly exercise each voice biometrics platform against realistic attack scenarios, ensuring the architecture could support continuous evaluation, scalability, and future evolution without disrupting existing operations.
The outcome: A scalable, streaming-native architecture that seamlessly integrated with the bank's environment and enabled continuous, real-world evaluation of voice biometrics systems.
Step 3:
Instrumented Pilot
We deployed the solution into a production-like environment to validate the end-to-end evaluation pipeline under real-world conditions. Autonomous voice agents conducted live conversations through the bank's Genesys contact center, while every interaction was instrumented to capture transcripts, voiceprints, biometrics decisions, and evaluation metrics. The platform automatically measured true positives, false positives, true negatives, and false negatives across each voice biometrics system, providing continuous, data-driven feedback to refine both the evaluation process and the AI agents.
The outcome: A production-grade pilot that continuously measured voice biometrics performance at scale, enabling objective comparison across platforms under realistic vishing attack scenarios.
Step 4:
Expertise Extraction
We established a human-within-the-loop expertise extraction process by performing deep analysis on every evaluation outcome, including authentic calls (true acceptance and false rejection) and deepfake voice attacks (true rejection and false acceptance). Each result was further classified to distinguish human vishing attempts from other causes of failure, creating a structured feedback loop for the voice biometrics platform. This continuously captured operational expertise enabled the bank to fine-tune detection models, improve decision accuracy, and adapt to emerging attack patterns.
The outcome: Every evaluation continuously strengthened the voice biometrics system, transforming operational insights into a compounding enterprise asset that improved fraud detection over time.
Step 5:
AI Compression & Funneling
As a separate pipeline, we enhanced the voice biometrics system by funneling higher-level speech reasoning into the decision process. While the voice biometrics platform remained focused on voiceprint analysis, a reasoning model analyzed the caller's speech patterns and utterances from the historical conversations to provide additional behavioral intelligence. By combining acoustic biometrics with AI-driven conversational analysis, the overall detection capability became more accurate and resilient against sophisticated vishing attacks.
The outcome: A layered AI approach that significantly improved fraud detection by augmenting existing voice biometrics with contextual speech reasoning, without replacing the underlying platform.
Step 6:
Govern and scale
We transformed the vishing penetration testing platform into a governed, repeatable capability that can evaluate every voice biometrics system with a single trigger across 20,000+ voice attack profiles. This enables the bank to continuously validate detection performance, support compliance and model governance requirements, and rapidly assess new biometrics technologies as attack techniques evolve.
The outcome: A scalable, enterprise-wide vishing evaluation capability that keeps the bank's voice biometrics defenses continuously validated, compliant, and at the forefront of emerging threats.
SPEAK TO OUR POLYAGENTIC AI ARCHITECTS
Discover how GoodLabs Polyagentic AI Engineering helps you identify the biggest asymmetries, scale human expertise, earn autonomy, and turn each successful implementation into a blueprint for enterprise-wide impact.