Cloudflare has introduced a sophisticated multi-AI-agent security operations harness designed to tackle the pervasive ‘alert paradox’ in modern security environments. This new system significantly enhances the efficiency of Cloudflare’s Managed Defense, aiming to reduce the burden on human analysts and accelerate the resolution of security incidents. By automating the aggregation, connection, and initial analysis of security detections, the platform allows security professionals to focus on critical decision-making and mitigation deployment, as detailed in a recent Cloudflare blog post.
Optimizing Security Incident Analysis Workflow with AI Agents
The core challenge in security operations is the sheer volume and complexity of alerts. A single security event can trigger numerous related alerts across an environment, overwhelming even experienced analysts. Cloudflare’s previous attempts with a single, general-purpose AI agent revealed limitations, including hallucinated claims, context drift, and failures in distinguishing between ‘not checked’ and ‘not found’ data. These issues highlighted the necessity of a more structured, evidence-grounded approach to Cloudflare AI agent alert triage.
To address these shortcomings, Cloudflare pivoted to an architecture that separates deterministic evidence collection from AI inference. This ‘recon first, inference second’ strategy ensures that a fixed set of reconnaissance workflows run using versioned API calls to gather comprehensive data before any model analysis begins. This data includes customer identity, detection history, traffic baseline, enforcement outcomes, and network observations, all stored with their source, version, and timestamp. This method ensures reproducibility, allowing the same snapshot of evidence to be replayed for consistent analysis by different specialist AI agents.
Multi-AI-Agent Architecture and Noise Filtering
The multi-AI-agent security operations harness employs a multi-stage process:
- Early Noise Filtering: Most security alerts do not represent true incidents. Cloudflare utilizes
Clef, an open-source decision model running on Workers AI, as a lightweight triage model.Clefcompares new alerts with reconnaissance data, historical patterns, and past analyst decisions. Alerts identified as high-likelihood false positives are deterministically classified as passive noise, removing them from the active analysis queue while retaining them for context. - Specialist Agent Investigation: For alerts requiring deeper review, a coordinator AI agent orchestrates four specialist AI agents, which run in parallel:
- Traffic Analysis: Reviews request behavior, historical changes, and enforcement actions.
- Customer Context: Examines earlier alerts, dispositions, and past Managed Defense Analyst decisions.
- Global Telemetry: Compares activity against privacy-preserving, Internet-wide signals (e.g., checking if an IP is targeting one site or scanning thousands).
- Threat Intelligence: Checks indicators already associated with the alert or case.
- Synthesis and Advisory: A final synthesis AI agent combines the typed findings from the specialist agents into a single advisory. Crucially, this agent cannot fetch new evidence or make classifications outside an approved vocabulary, making unsupported claims easier to audit and catch. The global telemetry specialist adheres strictly to aggregate data to preserve customer privacy, never re-identifying specific customer details from broad patterns.
Actionable Recommendations
While this article describes Cloudflare’s internal process improvements, the principles offer valuable lessons for any organization seeking to enhance its security operations:
- Prioritize Structured Data Collection: Implement deterministic data collection workflows before engaging AI for analysis to ensure accuracy and reproducibility.
- Leverage Specialized AI: Rather than a single general-purpose AI, employ multiple, narrowly focused AI agents for specific investigative tasks to reduce the risk of hallucination and improve contextual accuracy.
- Integrate Early Triage: Utilize lightweight, fast decision models to filter known noise and false positives, allowing human analysts and more complex AI systems to focus on critical alerts.
This evolution in Cloudflare’s security operations underscores a commitment to leveraging advanced AI not as a replacement for human expertise, but as a force multiplier that empowers security teams to respond more effectively and efficiently to complex threats.
Related: Recorded Future Debuts Autonomous Defense Against AI Threats, Cloudflare Enhances Vulnerability Management with AI & Context