Skip to main content
AI Security Hub
04AI Security Pillar

Model Behaviour Guardrails

Set the policies and boundaries governing what an AI system may say, recommend, or refuse.

Business focus

Define acceptable behaviour

Why it matters

Security starts with a defined business boundary

Model behaviour guardrails translate business rules into system instructions, refusal conditions, output boundaries, and evaluation criteria.

Business risk

A well-written prompt is not an access control or a guarantee. Businesses need explicit rules for acceptable use, uncertain answers, restricted topics, and situations that require human judgment.

What this pillar covers
  • System policy rules
  • Refusal and safe-fallback logic
  • Output boundaries
  • Reasoning and evidence constraints
Operating model

Turn define acceptable behaviour into repeatable controls

A policy is only the starting point. For each AI use case, name the business owner, define the allowed boundary, configure the relevant technical controls, and decide what evidence proves those controls are working. Repeat the review when the model, data, connected tools, or business purpose changes.

Start with a single high-value workflow instead of trying to govern every experimental use at once. That makes it possible to test the controls with real users, find exceptions, and create a pattern the rest of the business can reuse.

StepDecisionEvidence to retain
1Scope the workflowOwner, purpose, approved data, users, and connected systems.
2Apply the controlsConfiguration, access rules, approval points, and test cases.
3Operate and reviewLogs, review results, exceptions, incidents, and change records.

Questions for leadership

  • Which outcomes are acceptable, restricted, or prohibited?
  • When should the AI refuse, defer, or ask for approval?
  • Who can change system policies and prompts?
  • How are guardrails tested after model or workflow changes?
Practical control checklist

Put the pillar into practice

  1. 1Document approved and prohibited uses in business language.
  2. 2Define safe fallback behaviour for uncertainty and missing evidence.
  3. 3Keep secrets and permissions out of system prompts.
  4. 4Version policy, prompt, model, and evaluation changes.
  5. 5Test intended behaviour with normal, unsafe, and ambiguous requests.
Authoritative guidance

Use recognised guidance to validate the control design

These resources help teams translate AI-specific risks into documented, testable business and technical controls. Apply them to the actual data, permissions, and actions in the workflow rather than treating them as a one-time compliance exercise.

Review a real workflow

Turn this pillar into operating controls

Map the data, access, approvals, monitoring, and evidence around one important AI use case before expanding it.

Explore AI services