University of AlbertaMultimedia Research Centre · Dept. of Computing Science
ROSSRemote Observation, Sensing & System
Research / R/04 Dual-Use & Defence / MatrixShield: adversarial security testing for AI agents
R/04 · Dual-Use & Defence

MatrixShield: adversarial security testing for AI agents

A dual-use security showcase: MatrixShield red-teams autonomous AI agents as a real adversary would, grading every response with an independent judge to turn agent risk into a board-ready 0-to-100 score. A collaboration led by Alvin (Xinyao) Sun with Matrix Labs.

Dual-Use & DefenceAI SecurityRed-TeamingAgent Safety

Autonomous AI agents have become an unmonitored attack surface. An agent is not a static endpoint: it interprets natural language, calls tools, reads untrusted content and acts on its own. That means it can be talked into leaking data, abusing its own tooling, or crossing a regulatory line, all without a single line of exploit code. Traditional application-security scanners and one-off penetration tests were never built for this, so a growing population of customer-facing agents ships without ever being attacked.

MatrixShield closes that gap. It is a security-testing platform that treats an agent the way a real adversary would: it connects to the live endpoint, runs adversarial scenarios as genuine multi-turn conversations, and has an independent judge model grade every response. Those verdicts roll up into a single 0-to-100 safety score, a per-agent grade, and a control-by-control compliance view, so a security lead can answer how many agents are production-ready, where the critical risks are, and whether the fleet is covered for frameworks like the EU AI Act, all from one screen.

The MatrixShield enterprise dashboard: fleet posture score, risk heatmap and framework compliance

The enterprise dashboard turns a fleet of agents into one security posture: a fleet score and grade, a risk heatmap of agents against threat categories, framework compliance, and recommended next actions. Courtesy Matrix Labs / MatrixShield.

This piece is a dual-use security showcase within the group’s Dual-Use & Defence theme, on the cybersecurity side. MatrixShield is built by Matrix Labs and led by Alvin (Xinyao) Sun, whose research spans adversarial machine learning, agent safety and decentralized-systems security. The same adversarial-ML methods that harden a remote-sensing pipeline against noise and spoofing carry directly into stress-testing the autonomous agents that increasingly sit behind critical services.

How it works

The platform is a single loop, from a live agent endpoint to an auditable security posture.

MatrixShield architecture: agents register, the engine runs judged multi-turn scans, and results feed scoring, reporting and governance gates

The architecture end to end: connect, configure, scan, judge, report, govern.

  • Connect any agent. Four registration paths (a no-code dashboard form, a Partner REST API, a native MCP server, or a self-registering fleet) all resolve to the same registered agent, with its endpoint, protocol adapter and encrypted credentials. Adapters cover LangGraph, OpenAI Assistants, Model Context Protocol and plain HTTP.
  • Configure a package. A reusable, named scan configuration bundles the right adversarial suites and scoring weights. Curated templates range from a quick safety check to deep, literature-cited Apex packs for ethics, fintech compliance and tool abuse.
  • Scan and judge. The engine drives each scenario as a five-stage conversation (connect, introspect, plan, execute, finalize) while an independent judge model scores every turn. A scenario only passes when the agent holds the line.

Proof, turn by turn

MatrixShield’s value is evidence, not a verdict. Every scenario an agent fails becomes a finding with the exact adversarial message, the exact unsafe response, why it is a violation, and the judge’s rationale, with the failing turn highlighted.

An adversarial conversation trace: attacker prompts in red, agent replies in green, with the manipulation tactic labelled

Each finding opens into the full attack-and-response trace, so a reviewer sees precisely how the agent was probed and how it replied. Templated placeholders stand in for real personal data. Courtesy Matrix Labs / MatrixShield.

From findings to a board-ready number

Every scan collapses a battery of adversarial tests into one number an executive understands. The overall score is the weighted percentage of scenarios the agent handled safely; the grade (A at 90+, down to F below 60) and risk level derive from it automatically. A per-suite breakdown shows engineers exactly which category, prompt injection versus data exfiltration, is dragging the score down, and a branded PDF exports the whole result for vendor-security questionnaires and compliance attestations.

A MatrixShield scan report showing the overall score, grade and per-suite coverage

The scan report opens with the verdict: an overall score and grade, suite coverage, and one-click PDF or JSON export. Courtesy Matrix Labs / MatrixShield.

Governance and the dual-use case

Around the core scan loop, MatrixShield maps results onto OWASP LLM Top 10, NIST AI RMF and the EU AI Act, and lets teams encode a security bar as an enforceable deployment gate, so a failing agent is blocked in the pipeline before it reaches production.

Compliance mapping across OWASP LLM Top 10, NIST AI RMF and the EU AI Act

Results map to recognised frameworks and export as audit-ready reports, scoped across the whole fleet. Courtesy Matrix Labs / MatrixShield.

That governance layer is what makes agent security a dual-use concern rather than a purely commercial one. As autonomous agents move into infrastructure, finance and defence-adjacent systems, the ability to certify how they behave under adversarial pressure becomes part of national resilience. It aligns with the University of Alberta’s Centre for Applied Research in Defence and Dual-use Technologies (CARDD-Tech) and its themes in sensing, cyber and AI.

Explore the platform: matrixshield.ai.