
Capabilities
- Severity-Based Filtering: 4-level severity classification (Safe, Low, Medium, High)
- Multi-Category Detection: Hate, sexual, violence, self-harm content
- Prompt Shield: Advanced jailbreak and injection detection
- Indirect Attack Detection: Identify hidden malicious instructions
- Protected Material: Detect copyrighted content (output only)
- Custom Blocklists: Define organization-specific blocked terms
Streaming output: When this profile is used in an
output or both rule, Bifrost accumulates the stream until the model response is complete, then checks the full response. It does not check individual stream chunks. See Streaming Output Guardrails for details.Configuration Fields
Collecting your API key and URL
Navigate to Azure foundry dashboard
- Copy API key to use it in the Azure content moderation config form
- Copy project endpoint and use base URL as endpoint in the form. e.g. (
https://xxx-resource.services.ai.azure.com)
Severity Threshold Levels
Detection Categories
- Hate and fairness
- Sexual content
- Violence
- Self-harm
Input-only features: Jailbreak Shield and Indirect Attack Shield only apply to input validation. Output-only
features: Copyright detection only applies to output validation.

