| WonderFence | A real-time AI guardrails platform for monitoring, controlling, and enforcing safe behavior in AI systems. |
| Policy | A rule definition that describes a category of content to detect, backed by a detection model and/or keyword rules. |
| Policy Catalog | The collection of all pre-built and custom policies available for enforcement. |
| Policy Group | A high-level classification for policies: Safety, Privacy, or Security. |
| System Policy | A pre-built policy provided by the platform with built-in detection capabilities. |
| Custom Policy | A user-created policy with custom guidelines and examples, backed by a generated SLM. |
| SLM (Small Language Model) | A purpose-built detection model generated from your guidelines and examples to detect custom violation types. |
| Enforcement Action | What happens when a violation is detected: Block, Redact (Mask), Warn (Detect), or Log. |
| Block Message | The custom message returned to users when content is blocked by a policy. |
| Audio Input Block Message | For blocked audio input, whether the caller gets the written block message or a spoken recording. Text input always gets the written message. |
| Confidence Level | The minimum detection threshold for a policy: Undetected, Low, Medium, or High. |
| Message Type | Whether a policy targets user input (Only Prompt), AI output (Only Response), or both (All). |
| Keyword Detection | Supplementary rule-based detection using keyword matching, in addition to model-based detection. |
| Violation Type | A specific category of harmful or prohibited behavior (e.g., hate speech, prompt injection, PII exposure). |
| Data Explorer | The investigation interface for browsing, filtering, and inspecting analyzed content and detected violations. |
| Workflows | Automated rules that trigger actions (webhooks, integrations) when specific policy conditions are met. |
| Overview | The analytics interface providing aggregated insights across policies, violations, and system performance. |
| Playground | An interactive testing environment for validating guardrail behavior before deploying to production. |
| OWASP LLM Top 10 | A risk framework identifying the most critical vulnerabilities in LLM applications. |
| MITRE ATLAS | A framework cataloging adversarial threats and techniques targeting AI systems. |
| False Positive | A detection event where benign content is incorrectly flagged as a violation. |
| False Negative | A case where violating content passes through undetected. |
| Guardrail | A policy-driven control that inspects AI traffic and takes action to enforce safety, security, or compliance. |
| Input Guardrail | A guardrail applied to user messages before they reach the AI model. |
| Output Guardrail | A guardrail applied to AI-generated responses before they reach the user. |