---
title: "Your AI red team may be missing the big picture: Adapting evaluation for multi-agent AI"
description: "This discussion covers what to consider in plotting a path forward. We'll walk through why evaluating multi-agent systems requires a new approach that is itself agentic: adaptive evaluation. Instead of a one-time pass/fail gate, testing functions in a continuous, iterative loop, using what the system has learned to probe deeper and wider, so the evaluation itself is as dynamic and autonomous as the agent."
image: https://webinars.techstronglearning.com/hubfs/2026.10.07%20Vijil-Social-1080x1080%20(1).png
---

Your AI red team may be missing the big picture:  
Adapting evaluation for multi-agent AI

![2026.10.07 Vijil-LandingPage-1540x660-DO (1)](https://webinars.techstronglearning.com/hs-fs/hubfs/2026.10.07%20Vijil-LandingPage-1540x660-DO%20(1).png?width=1540&height=660&name=2026.10.07%20Vijil-LandingPage-1540x660-DO%20(1).png "2026.10.07 Vijil-LandingPage-1540x660-DO (1)")

## Sponsored By:

![Vijil\_Logo](https://webinars.techstronglearning.com/hubfs/Vijil_Logo.jpg)

##### WEDNESDAY, OCTOBER 7, 2026

10 AM PT   |   1 PM ET

Current AI red teaming practices were built for a world of static models with predictable inputs and outputs. The future of AI systems is inevitably multi-agent: a planner agent delegating to specialist agents, calling external tools over MCP, chaining decisions across sessions, and sometimes rewriting their own approach mid-task, potentially in collusion with each other. That’s why red teaming, as traditionally practiced, is becoming obsolete: a point-in-time, semi-automated test can't keep pace with a system that is dynamic and can act autonomously.

This discussion covers what to consider in plotting a path forward. We'll walk through why evaluating multi-agent systems requires a new approach that is itself agentic: adaptive evaluation. Instead of a one-time pass/fail gate, testing functions in a continuous, iterative loop, using what the system has learned to probe deeper and wider, so the evaluation itself is as dynamic and autonomous as the agent.

We'll unpack what that shift means in practice for the three groups who have to live with it: engineering teams building agents, as they need security signal that doesn't slow down shipping; platform and DevOps teams deploying agents, who need confidence that holds up across updates, not just at launch; and security and governance teams, who need continuous, auditable evidence instead of a static report that's outdated within weeks.

Key Takeaways:

- Understand the limitations of static red teaming: why existing testing doesn’t cover the mechanics and behavior of multi-agent systems
- Gather a technical requirements checklist for evaluating multi-agent architectures and filling in blind spots for dynamic behavior, such as LLM, tools, MCP gateway, and delegated agents, so that you can benchmark your current approach against.
- Map how to integrate evaluation into your agentic lifecycle: how continuous, closed-loop testing works in practice, and how it differs structurally from a red team engagement or a one-time benchmark.
- Align stakeholders on a new approach, whether you're building agents (how to get security signal without slowing releases), deploying them (how to maintain confidence across updates), or securing them (how to produce continuous, audit-ready evidence instead of a stale report).
- Probe live about your own agent architecture and get direct input on where your current testing approach has blind spots.

# Register Below:

We'll send you an email confirmation

![Vin head shot-modified](https://webinars.techstronglearning.com/hs-fs/hubfs/Vin%20head%20shot-modified.png?width=320&height=320&name=Vin%20head%20shot-modified.png "Vin head shot-modified")

### Vin Sharma

###### Founder and CEO - Vijil

Vin Sharma is the co-founder and CEO of Vijil, a company providing an infrastructure layer of trust for enterprise AI agents. With more than 30 years of experience in AI/ML, data, cloud, operating systems, and security, Vin has led the development of more than 30 products and 11 AWS AI services, most recently as general manager and director of engineering at AWS. He has previously held executive roles at Intel, Hewlett Packard Enterprise, and Foursquare. Vin holds 10 patents and has authored research papers in deep learning. 

![Alastair Cooke (1)-modified-1](https://webinars.techstronglearning.com/hs-fs/hubfs/Alastair%20Cooke%20(1)-modified-1.png?width=1463&height=1463&name=Alastair%20Cooke%20(1)-modified-1.png "Alastair Cooke (1)-modified-1")

### Alastair Cooke

###### Research Director, Cloud and Data Center - The Futurum Group

Alastair has made a twenty-year career out of helping people understand complex IT infrastructure and how to build solutions that fulfill business needs. Much of his career has included teaching official training courses for vendors, including HPE, VMware, and AWS.  
Alastair has written hundreds of analyst articles and papers exploring products and topics around on-premises infrastructure and virtualization and getting the most out of public cloud and hybrid infrastructure. Alastair has also been involved in community-driven, practitioner-led education through the vBrownBag podcast and the vBrownBag TechTalks.

![Fernando Montenegro-modified-1](https://webinars.techstronglearning.com/hs-fs/hubfs/Fernando%20Montenegro-modified-1.png?width=302&height=302&name=Fernando%20Montenegro-modified-1.png "Fernando Montenegro-modified-1")

### Fernando Montenegro

###### Vice President and Practice Lead, Cybersecurity - The Futurum Group

Fernando Montenegro serves as the Vice President & Practice Lead for Cybersecurity & Resilience at The Futurum Group. In this role, he leads the development and execution of the Cybersecurity research agenda, working closely with the team to drive the practice’s growth. His research focuses on addressing critical topics in modern cybersecurity. These include the multifaceted role of AI in cybersecurity, strategies for managing an ever-expanding attack surface, and the evolution of cybersecurity architectures toward more platform-oriented solutions.

![Techstrong Group logo](https://webinars.techstronglearning.com/hubfs/Techstrong%20Group%20logo.png "Techstrong Group logo")

![TechstrongLearning-Colored-2](https://webinars.techstronglearning.com/hubfs/TechstrongLearning-Colored-2.png "TechstrongLearning-Colored-2")

![DO.com R+B RGB Logo-Feb-22-2024-07-06-28-4406-PM](https://webinars.techstronglearning.com/hs-fs/hubfs/DO.com%20R+B%20RGB%20Logo-Feb-22-2024-07-06-28-4406-PM.png?width=1363&height=267&name=DO.com%20R+B%20RGB%20Logo-Feb-22-2024-07-06-28-4406-PM.png "DO.com R+B RGB Logo-Feb-22-2024-07-06-28-4406-PM")

![Security boulevard logo blue](https://webinars.techstronglearning.com/hs-fs/hubfs/Security%20boulevard%20logo%20blue.png?width=2000&height=360&name=Security%20boulevard%20logo%20blue.png "Security boulevard logo blue")

![CNN Blue and White-4](https://webinars.techstronglearning.com/hs-fs/hubfs/CNN%20Blue%20and%20White-4.png?width=1400&height=400&name=CNN%20Blue%20and%20White-4.png "CNN Blue and White-4")

![Techstrong.ai blue logo](https://webinars.techstronglearning.com/hs-fs/hubfs/Techstrong.ai%20blue%20logo.png?width=2000&height=295&name=Techstrong.ai%20blue%20logo.png "Techstrong.ai blue logo")

© 2026 [Techstrong Inc.](https://techstronggroup.com/) All Rights Reserved. [Contact Support.](mailto:jared@techstronggroup.com)