Skip to main content

WonderFence Glossary

TermDefinition
WonderFenceA real-time AI guardrails platform for monitoring, controlling, and enforcing safe behavior in AI systems.
PolicyA rule definition that describes a category of content to detect, backed by a detection model and/or keyword rules.
Policy CatalogThe collection of all pre-built and custom policies available for enforcement.
Policy GroupA high-level classification for policies: Safety, Privacy, or Security.
System PolicyA pre-built policy provided by the platform with built-in detection capabilities.
Custom PolicyA user-created policy with custom guidelines and examples, backed by a generated SLM.
SLM (Small Language Model)A purpose-built detection model generated from your guidelines and examples to detect custom violation types.
Enforcement ActionWhat happens when a violation is detected: Block, Redact (Mask), Warn (Detect), or Log.
Block MessageThe custom message returned to users when content is blocked by a policy.
Audio Input Block MessageFor blocked audio input, whether the caller gets the written block message or a spoken recording. Text input always gets the written message.
Confidence LevelThe minimum detection threshold for a policy: Undetected, Low, Medium, or High.
Message TypeWhether a policy targets user input (Only Prompt), AI output (Only Response), or both (All).
Keyword DetectionSupplementary rule-based detection using keyword matching, in addition to model-based detection.
Violation TypeA specific category of harmful or prohibited behavior (e.g., hate speech, prompt injection, PII exposure).
Data ExplorerThe investigation interface for browsing, filtering, and inspecting analyzed content and detected violations.
WorkflowsAutomated rules that trigger actions (webhooks, integrations) when specific policy conditions are met.
OverviewThe analytics interface providing aggregated insights across policies, violations, and system performance.
PlaygroundAn interactive testing environment for validating guardrail behavior before deploying to production.
OWASP LLM Top 10A risk framework identifying the most critical vulnerabilities in LLM applications.
MITRE ATLASA framework cataloging adversarial threats and techniques targeting AI systems.
False PositiveA detection event where benign content is incorrectly flagged as a violation.
False NegativeA case where violating content passes through undetected.
GuardrailA policy-driven control that inspects AI traffic and takes action to enforce safety, security, or compliance.
Input GuardrailA guardrail applied to user messages before they reach the AI model.
Output GuardrailA guardrail applied to AI-generated responses before they reach the user.