Skip to main content

WonderFence Core Workflows

Defining Policies and Rules​

  1. Start with the catalog — Review the pre-built policies and activate those that align with your organization's content policies.
  2. Configure each policy — Set the message type, confidence level (when enabled for your account), and action appropriate to each violation category and your risk tolerance.
  3. Create custom policies — For organization-specific requirements not covered by the catalog, create custom SLM-based policies with guidelines and examples.
  4. Add keyword rules — Augment model-based detection with keyword lists for terms, phrases, or patterns unique to your domain.
  5. Define workflows — Set up automated rules to trigger webhooks or integrations when specific policy conditions are met.

→ Go to Enforcement Policies page

Applying Guardrails to an AI Application​

  1. Choose your integration method — Select API, SDK, or proxy mode based on your architecture.
  2. Connect your application — Configure the integration to route traffic through WonderFence.
  3. Activate policies — Enable the enforcement policies relevant to your application.
  4. Test with the Playground — Validate that guardrails behave correctly for your specific use cases before sending production traffic.
  5. Go live — Begin routing production traffic through WonderFence.

Custom Fields​

By default, WonderFence is provided with a wide variety of fields to describe your content. Your organization can add additional customized fields that describe the content to be analyzed in order to enhance detection accuracy and enable more sophisticated policy enforcement.

Editable and Collected Custom Fields​

Two types of Custom Fields can be created in WonderFence using an API — Editable (by a moderator) and Collected (read-only). Both types appear throughout the WonderFence user interface after they are created, can be defined to appear in content data, and can be used to define filters or build Automated Workflows based on their values.

  • Editable Custom Fields — Values can be entered and changed by a user in the WonderFence interface. Editable field values can also be sent through an API request.
  • Collected Custom Fields — Read-only fields. This is useful when you want to send data via API but don't want users to edit it.

Adding Custom Fields​

  1. Click the Account Settings button in the top-right corner.
  2. Select Moderation Capabilities → Custom Fields. A list of the custom fields that were defined is displayed.
  3. To add a new field, click the Add Field button.

Defining a Collected Custom Field:

  1. In the Choose Custom Field type window, click COLLECTED then Next.
  2. Enter a free-text Title (spaces allowed).
  3. Enter a Key for the field's API identifier (no spaces or periods; underscores and dashes allowed).
  4. Select the Type: Text, Number, Date, Boolean, Single choice Dropdown, Multi choice Dropdown, or Formula. (Formula is only available for Collected fields and computes from other fields.)
  5. Click Finish.

Note — The data type of fields sent to WonderFence in a Custom Field will be verified, and an error is generated if incorrect.

Defining an Editable Custom Field:

  1. In the Choose Custom Field type window, click EDITABLE then Next.
  2. Enter the Title (spaces allowed).
  3. Enter the Key (no spaces or periods).
  4. Select the Type: Text, Number, Date, Boolean, Single choice Dropdown, or Multi choice Dropdown.
  5. Optionally set a default value (the input depends on the selected type).
  6. Enter a Tooltip to display when a user hovers over this field.
  7. For dropdown types, type the dropdown values into the Add values to dropdown area, pressing Enter after each.
  8. Click Finish.

These fields are now available to be used in policies, workflows, and API requests.

Building Automated Workflows​

  1. Select WonderFence → Workflows in the left pane.
  2. Click the + Add New button.
  3. The left pane provides a selection of building blocks: Initiators (green), Conditions (orange), and Actions (blue). Each Automated Workflow can have a single initiator, multiple conditions (with AND or OR relationships), and multiple actions.
  4. Define an initiator — Drag an initiator card onto the Insert initiator card. A form appears on the right where you define the parameters.
  5. Define one or more conditions (optional) — Drag Condition cards into the flow diagram. Drag multiple conditions for AND relationships, or place them on parallel branches for OR relationships.
  6. Define one or more actions — Drag Action cards into the flow diagram in the order they should run.
  7. Save the workflow with the Save button in the top-right corner.

Activating Automated Workflows: After saving, the workflow appears in the Workflows page (WonderFence → Workflows). Click the toggle to activate it so it starts examining content items and automatically performing its actions when its initiator condition is matched.

Managing the Workflows list: The Workflows page shows every workflow — active and inactive — in a single list. Use the search box to filter by name and the All / Active / Inactive control to narrow the list. Because workflows are evaluated top-to-bottom, drag a workflow to change its priority. Toggle the switch (or use the kebab menu) to activate / deactivate a workflow. All of these changes — activation, priority, and the Always Run flag — take effect immediately; there is no separate Save step. Each workflow's kebab menu offers Edit, Activate / Deactivate, Always Run / Run conditionally, Duplicate, and Delete.

When are Automated Workflows triggered? Every Automated Workflow examines the properties of each content item as it arrives in WonderFence (via API request). It compares each item with the Initiators that have been defined and triggers the actions defined in the relevant Workflow upon a match. In addition, Automated Workflows are triggered by the change of content properties caused by another Automated Workflow, and by manual changes made by users in the WonderFence interface. Workflows are evaluated in the order they appear in the Automated Workflows list — from top to bottom — and multiple workflows may fire for a single content item, one after another.

Monitoring Live Traffic​

Once guardrails are active on production traffic:

  1. Check the Overview — Monitor real-time violation counts, action distributions, and traffic volume.
  2. Review trends — Look for spikes in specific violation types that may indicate coordinated attacks or emerging misuse patterns.
  3. Monitor latency — Verify that guardrail evaluation is not adding unacceptable latency to your application's response times.
  4. Set up workflow alerts — Configure webhook-based notifications for high-severity events so your team is notified immediately.

Reviewing Flagged Events​

  1. Open Data Explorer — Navigate to the item browser to see all analyzed content and detected violations.
  2. Filter by priority — Start with high-confidence Block actions, then review Warn/Detect events.
  3. Inspect event details — Click individual events to see the full conversation context, detection metadata, and the action that was taken.
  4. Identify false positives — Flag events where the policy fired incorrectly. Use these insights to adjust confidence levels or refine keyword rules.
  5. Identify false negatives — If violations are getting through undetected, consider lowering the confidence threshold, adding keywords, or creating a custom policy.

→ Go to Data Explorer page

Data Explorer​

The Data Explorer page lets you browse the raw analyzed content and sessions that have flowed through WonderFence, so you can inspect individual items, debug policy decisions, or audit specific traffic.

The table lists one row per analyzed content item — a single message. To browse traffic one row per conversation instead, see Session Explorer.

You can:

  • Search by free text.
  • Apply structured filters across the columns (e.g., violation type, custom-field values).
  • Filter by Media Type (Text, Image, or Audio) to narrow the table to one kind of content.
  • See how serious each item was judged to be in the Severity column — High, Medium or Low — and filter by it. The column appears only when the feature is enabled for your account, and it sits beside Action so you can see what a severity actually enforced. Items that arrived without a severity show a dash, which includes every item on an application whose escalation adjudication is turned off; severity is how serious the violation is, while Confidence is how certain the detection is that there was a violation at all.
  • Filter by Group (Safety, Privacy, or Security) to narrow the table to the policy groups the Overview's Violation Group Distribution chart buckets violations into. Selecting a group keeps every item that tripped at least one policy in it, and combines with the Violations filter: an item must match both, so picking a policy outside the selected group returns no rows.
  • Filter by Risk Score to keep only items above or below a threshold. Pick a comparison — greater than (>), at least (≥), less than (<), or at most (≤) — type a score between 0 and 100, and select Apply; the filter button then shows the comparison in use (for example Risk Score > 60). Clear removes it, as does Reset All.
  • Sort and paginate through results (up to 1,000 results per page).
  • Click a row to open a side drawer with the full content item or session, including risk scores and any actions that were taken. The conversation is shown in the same message format the Playground uses, with each flagged message carrying the policies it tripped under it, so a session reads the same way wherever you open it.
  • Inspect non-text content: image items show a thumbnail in the Content column, and audio and video items play back in the side drawer with a built-in player — play/pause, scrub, volume and fullscreen, starting paused on the item's own thumbnail.
  • Download the current page of results (after any search, filters, and sorting) as a CSV file using the download button in the filter bar.
  • Refresh the table with the refresh button in the filter bar to pull in the latest items. The page also refreshes automatically each time you open it.

→ Go to Data Explorer page

Session Explorer​

Session Explorer is a proof of concept, off for every account by default. Ask your Alice contact to switch it on for your project; until then the sidebar item is hidden.

Where Data Explorer lists one row per message, Session Explorer lists one row per conversation — time, session ID, application, user, the violations raised anywhere in the session, and the highest risk score any of its turns reached. Those last two are read from the session's own turns, so they answer "what happened anywhere in this conversation" rather than describing a single message. Search by session ID or user, filter by application, and sort or page through the results. Times are shown on a 24-hour clock, to the second, because an agentic session's turns often land inside the same minute. A risk score is shown with its band — Low, Medium or High — beside the number, so the severity reads the same whatever your display or colour vision.

If the table is empty it says which kind of empty: No sessions recorded yet when the project has had no traffic, and No Results Found when a search or filter matched nothing. If the list cannot be loaded, the table says so and offers Try Again in place of the rows.

Selecting a row opens the session as a trace, in two panes:

  • The turn list, on the left — every stored turn in order, each showing whether it was a prompt, a response or a tool call, the time, its risk score, and how many violations it raised. If the turns cannot be loaded, the pane says so and offers Try Again. A turn a guardrail acted on is tagged BLOCK, MASK or DETECT. Order the list by Time to read the conversation, or by Risk to bring the worst turns to the top. Collapse the pane with the button in the header to give the detail more room.
  • The session and the selected turn, on the right — a stack of sections, described below. Step through turns with the up and down buttons in the header, or with the K and J keys.

The right pane's sections, in order:

  • Session summary — collapsed, with its headline figures on the header. Opened, it shows turns, prompts, responses, tool calls, turns a guardrail acted on, highest risk score, tokens counted, and the span from the session's first stored turn to its last, then Violations — every one the session raised, listed in full, where the table's row collapses them behind a +n. A figure with nothing behind it reads — rather than 0, so "not recorded" never looks like "none".
  • Agents — who spoke. An agentic client runs several agents down one conversation, and each is listed with whether it was the main agent or a subagent, how many turns it took, and which one produced the turn you are looking at.
  • History — collapsed. Opened, it shows every turn that came before the selected one, each as its own card that starts expanded, so you can read the conversation as it stood at that moment and fold away a long tool result without closing the rest.
  • The selected turn — its text, its guardrail verdict, and the policies it tripped with each policy's own risk score. Pretty shows the turn as it read; JSON shows the stored record behind it.
  • System prompt — the prompt each agent of the session actually ran with, one block per agent.
  • Tools — one card per tool, with how many times the session called it. Open a card for the tool's description and its full parameter schema.

Prompts, responses, system prompts, tool output and tool descriptions are rendered as Markdown, since that is how the model was given them and how it wrote them back.

Two limits are worth knowing:

  • Only turns a policy flagged are indexed. A session's row and its violation and token figures are built from those turns, so clean traffic is not counted.
  • The transcript is a window, not the whole conversation. It is capped around the flagged turn, so a very long session shows the turns nearest the one that fired.

The system prompts and tool definitions above are read from the session's own turns, and are recorded only for traffic arriving through the LiteLLM integration. A session that recorded none — every session predating the integration — falls back to the application's current configuration in Application Inventory and says so on the panel, which means a prompt edited since the session ran reads as it is now rather than as it was.

→ Go to Session Explorer page

Tuning and Improving Guardrails Over Time​

Guardrails are not set-and-forget. Continuous tuning is essential for optimal performance:

  1. Review Overview analytics weekly — Track detection rates, false positive rates, and action distributions over time.
  2. Adjust confidence thresholds — Lower thresholds for categories with observed false negatives; raise thresholds for categories with excessive false positives.
  3. Refine keyword lists — Add new terms as you discover emerging patterns; remove terms that generate noise.
  4. Update custom policies — Revise guidelines and examples for custom SLMs as your understanding of edge cases improves.
  5. Test changes in the Playground — Always validate policy changes before applying them to production.
  6. Coordinate with red teaming — Use WonderBuild to systematically test your guardrails against adversarial inputs and verify they hold up.