Skip to content
AI SECURITY / LLM · RAG · AGENTS

AI red teaming for LLMs and agents

We test the entire system — data, tools, permissions, memory, orchestration and actions — rather than judging model output in isolation.

Not every undesirable response is a vulnerability. Priority follows achieved capability: accessing another user’s data, performing an action, bypassing policy or persistently changing system behaviour.

LLM · RAG · AGENTS
System layers
IMPACT FIRST
Impact rating
MULTI-TURN
Scenarios
RETEST
Control verification
01 / DECISION CONTEXT

When to run an AI red team

  • 01Before public launch of an assistant, agent or AI feature.
  • 02When the model can access internal data, tools or operations.
  • 03After changing models, prompts, RAG, memory or permissions.
  • 04When customers or regulators need evidence of AI risk control.
02 / TEST SURFACE

Assessment coverage

01

Instructions and data

We test boundaries between trusted control and untrusted content.

  • Direct and indirect prompt injection
  • RAG poisoning and cross-tenant retrieval
  • Prompt, data and secret extraction
02

Agents and tools

We assess what the system can do, not just what it can say.

  • Tool abuse, excessive agency and confused deputy
  • Operation authorisation and confirmations
  • SSRF, injection and unsafe model output handling
03

Product resilience

Behaviour is tested across time and dependency failures.

  • Memory, multi-turn and delayed triggers
  • Model fallbacks, limits and cost abuse
  • Monitoring, audit trail and response
03 / DELIVERY

AI testing method

  1. BR / 01

    Threat model

    We identify assets, roles, untrusted-content sources and unacceptable outcomes.

  2. BR / 02

    Evaluation set

    Normal, edge, adversarial and dependency-failure scenarios become a repeatable test set.

  3. BR / 03

    Adversarial testing

    Manual testing, automation and review of the surrounding application are combined.

  4. BR / 04

    Evidence, remediation and retest

    We report achieved impact, required conditions and the control expected to stop the scenario.

04 / EVIDENCE STANDARD

Indirect prompt injection triggers a tool outside user intent

FINDING MODEL / NO CLIENT DATA

A changed response is lower impact than a confirmed operation performed with agent permissions.

This model explains our reporting format and is not a client story.

E-1

The agent retrieves a document containing an untrusted instruction

E-2

Orchestration does not separate data from control instructions

E-3

The model selects a write-capable tool

E-4

The fix enforces policy outside the model and requires confirmation

05 / OUTPUT

What the AI team receives

01
Threat model and trust-boundary map
Included deliverable
02
Versioned evaluation scenarios
Included deliverable
03
Confirmed findings with impact evidence
Included deliverable
04
Application, IAM, RAG and guardrail recommendations
Included deliverable
05
Team workshop and retest
Included deliverable
06 / QUESTIONS

Common questions

← All services
01Do you test the model itself?+

We can evaluate model behaviour, but application, data and tools often create the greatest risk. The full flow is covered when access permits.

02Is a jailbreak automatically a vulnerability?+

No. We rate impact, repeatability and context. A style change is different from accessing another tenant’s data or executing an operation.

03Do you provide regression tests?+

Yes when the architecture supports safe automation. Scenarios are versioned with the model, prompt, tools and acceptance criteria.

BR / NEXT STEP

Does your AI system access data or tools?

We will map trust boundaries and design tests around material business outcomes.

NDA · clear scope · direct communication