LLM red teaming for Canadian AI products
LLM red teaming is an attacker's view of your AI feature: a tester spends days trying to make the model leak data, ignore its rules or misuse the tools it can call, then shows you exactly how they did it.
LLM red teaming costs $4,000 to $20,000 CAD depending on how much the model can do, and it is the right test when your AI feature can take actions: call tools, send messages, read documents, query a database or spend money. A tester runs direct and indirect prompt injection, multi-turn jailbreaks, attempts to extract the system prompt, and attempts to make an agent act on instructions it found in content rather than from the user. You get a report of what worked, with the exact prompts, and a fix for each.
Red teaming or an assessment
An AI security assessment is structured. It walks the OWASP Top 10 for LLM applications and checks each risk against your design. Red teaming is open-ended. It asks what a determined attacker can get the system to do, and keeps going. A chatbot that only answers questions from public documents needs the assessment. An agent with tool access needs red teaming, and usually both. The comparison is set out on assessment versus pentest.
What gets tested
| Attack | What the tester tries | OWASP LLM Top 10 (2025) |
|---|---|---|
| Direct prompt injection | Instructions in the chat that override your system prompt | LLM01 |
| Indirect prompt injection | Instructions hidden in a web page, email, PDF or ticket the model reads | LLM01 |
| Jailbreaks | Multi-turn and role-play attacks that walk the model past its policies | LLM01 |
| Data disclosure | Getting the model to reveal another user's data, internal documents or keys | LLM02 |
| Output handling | Model output that becomes a script, a query or a link in your app | LLM05 |
| Excessive agency | Getting an agent to call a tool, send a message or change a record it should not | LLM06 |
| System prompt leakage | Extracting instructions, internal rules or credentials placed in the prompt | LLM07 |
| Retrieval weaknesses | Poisoning or reading across the vector store that feeds retrieval | LLM08 |
| Unbounded consumption | Prompts that run up tokens, loop an agent or exhaust your quota | LLM10 |
Each item is explained in plain language on the OWASP Top 10 for LLM applications.
Agents are the reason this exists
A chatbot that says something wrong is embarrassing. An agent that does something wrong has side effects. Once a model can call a function, every piece of text it reads is a possible instruction, including the ones an attacker wrote into a support ticket or a shared document. The fix is rarely a better prompt. It is narrower tool permissions, confirmation before actions with consequences, and treating model output as untrusted input. See excessive agency.
What it costs
| Scope | Typical range |
|---|---|
| Chatbot over public content, no tools | $4,000 to $7,000 |
| Assistant with retrieval over customer or internal data | $6,000 to $12,000 |
| Agent with tool access, multi-tenant | $10,000 to $20,000 |
TrazTech lists LLM red teaming from $4,000 CAD, delivered with testing partners, with a findings report, reproduction steps and fixes.
Before you book one
- Write down every tool the model can call and what each one can change.
- List every source of text the model reads that a user or outsider could influence.
- Decide what "bad" means for your business: leaking a document, sending an email, changing a price, answering off-topic.
- Have a staging copy with realistic but fake data, so testers can push hard.
Find out what your model will do for an attacker
Describe the feature, the model and the tools it can call. You get a scope back.
Get matchedCommon questions
Can a guardrail product replace red teaming?
No. Guardrail filters reduce some attacks and are worth having. Red teaming is how you find out which attacks still get through, including the ones written specifically to pass the filter.
Does red teaming test the model or my app?
Your app. You are not going to retrain a hosted model. What you control is the prompt, the retrieval, the tools and what your code does with the output, and that is where findings get fixed.
How often should we red team?
Before launch, and again when you add a tool, a data source or a new model. Those are the changes that open new paths.