WonderFence Key Features
Policy Catalog
WonderFence ships with a comprehensive catalog of pre-built detection models covering the most common AI safety and security risks. Each policy is backed by purpose-built models that are ready to enforce immediately — no training required to get started.
Policies are organized into the following groups:
| Group | Coverage |
|---|---|
| Security | Prompt injection, jailbreaking, system exploitation, data exfiltration attempts. |
| Privacy | PII detection (names, emails, phone numbers, addresses, financial data), sensitive data exposure. |
| Safety | Hate speech, violence, self-harm, sexual content, illegal activities, harassment. |
The catalog can be viewed as a card grid or a list using the view toggle above it; your choice is remembered for next time. In both views, each policy shows at a glance its detection group, name, description, source (System Policy or Custom Policy), its enforcement action, the number of applications the policy is applied to (hover the count to see the application names), and its associated risk-framework indicators (OWASP LLM Top 10, MITRE ATLAS). A policy that is not applied to any application is grayed out in the catalog, so unused policies are easy to spot.
Policy Configuration
Every policy in the catalog is fully configurable, allowing you to tailor enforcement behavior to your organization's risk tolerance and operational requirements.
Enforcement actions — Each policy supports one of the following:
| Action | Behavior |
|---|---|
| Block | Prevents the content from being delivered. Returns a configurable block message to the user. |
| Redact (Mask) | Replaces sensitive information with masked placeholders. The message is delivered with sensitive data removed. Available for Privacy policies. |
| Warn (Detect) | Flags the content for review without interrupting the conversation. The content is delivered unchanged — nothing is blocked or masked — but the violation is logged and surfaced for analysis. Available for every policy type, Privacy included, so you can monitor a policy before enforcing against it. |
| Log | Records the event for analytics and audit purposes with no user-visible action. |
Action by violation severity — When the feature is enabled for your account, a checkbox appears beside this heading, below the Action field. It is off by default, and the policy behaves exactly as described above while it is off.
Tick it to override the action above with a layered configuration based on the severity of the violation — one row per severity level:
| Severity | Applies when |
|---|---|
| High | The violation was judged to be serious. |
| Medium | The violation was judged to be moderate. |
| Low | The violation was judged to be minor. |
All three rows must be set, and they start on whatever the Action above is — so ticking the box changes nothing until you actually change a level.
The Action above stays in effect for everything the rows do not cover — any violation that arrives without a severity, which includes every violation on an application whose escalation adjudication is turned off. Unticking the checkbox removes the rows and returns the policy to that single Action.
The block message is set once for the policy, and appears whenever anything is set to Block — the Action above, or any severity row. It is pre-filled with the default block message either way.
Audio input block message — When the action is Block, and the feature is enabled for your account, choose what the caller gets back when audio input is blocked:
- Text — Returns the written block message you typed. This is the default, and what every policy does when the feature is off.
- Audio — Returns a spoken recording instead of the written message, for voice applications that cannot render text. The Policy Editor plays the recording so you can hear it; the recording is provided by the platform and cannot be replaced.
This choice applies only to audio input. A blocked text message always returns the written block message, whichever option is selected.
Message type targeting — Scope each policy to specific parts of the conversation:
- All — Applies to both user inputs and AI responses.
- Only Prompt — Applies only to user messages (input guardrails).
- Only Response — Applies only to AI-generated responses (output guardrails).
Confidence thresholds — Set the sensitivity of each policy's detection. This control is shown in the Policy Editor when it is enabled for your account:
| Level | Behavior |
|---|---|
| High | Only clear-cut violations are flagged. Maximizes precision, minimizes false positives. |
| Medium | Balances precision and recall for general-purpose enforcement. |
| Low | Catches more potential violations, accepting a higher rate of false positives. |
| Undetected | Flags even minimal or indirect signals related to a violation. |
Keyword augmentation — Supplement model-based detection with keyword rules for additional precision.
Custom Model Creation
When the pre-built policy catalog does not cover a violation type specific to your organization, WonderFence allows you to create custom detection policies.
How it works:
- Define the policy — Provide a descriptive name, detailed guidelines explaining what should be detected, and example texts of both acceptable and violating messages.
- Generate the model — WonderFence builds a custom Small Language Model (SLM) tailored to your guidelines and examples. Model generation completes within approximately 24 hours.
- Configure and enable — Configure message type targeting, confidence thresholds (when enabled for your account), and enforcement action, and enable the policy for the applications it should protect using the per-application toggles.
Custom models integrate seamlessly with the rest of the policy catalog. They appear alongside system policies and support the same configuration options, analytics, and workflow integrations.
Create a policy:
Clicking Create Policy on the Enforcement Policies page opens a full-page, two-step wizard: Create Policy followed by Train Policy. Step 1 collects the policy definition and your examples:
| Field | Description |
|---|---|
| Policy Name | A concise name representing the policy. |
| Type | The detection group the policy belongs to. Required — choose Safety, Privacy, or Security. The available enforcement actions follow your choice: Privacy policies offer Redact (Mask) in addition to Warn (Detect) and Block, and start on Redact. |
| Policy Guidelines | Specific guidelines or criteria that define what constitutes a violation of this policy. Must be at least 100 characters — a live character counter under the field shows your progress toward the minimum. Use Upload from File to populate the guidelines from a .txt or .pdf file. |
| Benign Examples | Exactly 5 example messages that represent acceptable, non-violating behavior. See Entering your examples below. |
| Violative Examples | Exactly 5 example messages that violate this policy. See Entering your examples below. |
Below the definition, the same Policy Setup controls described under Policy Configuration are available on step 1 — message type targeting, confidence threshold (when enabled for your account), enforcement action and its block message, keyword augmentation, and the per-application toggles — so a new policy can be fully configured before it is ever created.
Entering your examples:
Each set of examples is a deck of 5 cards that you fill in one at a time. The cards still to come are stacked behind the one you are writing in, and the counter under the deck (for example, 2/5) shows where you are. Next becomes available once the current card has text, and Previous takes you back to review or edit an earlier card — nothing you have typed is lost as you move between them.
Instead of typing all 5, you can use Upload from File to import them from a .csv, .xlsx, or .tsv file. The first column of the file is used, a header row such as "example" or "prompt" is ignored, and the imported values replace the whole deck and return you to the first card. Files with fewer than 5 examples leave the remaining cards blank for you to complete; anything beyond the first 5 is ignored.
Once the definition, examples and setup are valid, Next advances to Train Policy. Arriving there starts the work of drafting your policy and generating candidate responses for it, which takes about a minute; the step shows what it is doing and then counts the responses off as they are labelled. If that work cannot be completed, the step says so and offers Try again — nothing has been created at that point, and you can also go back to adjust the policy definition and start over.
Once the responses are ready, the system asks "Do you agree with each verdict about the following responses?" and shows you a round of them, each with the verdict it proposed — Verdict: Benign or Verdict: Violative — below the text. Confirm (👍) or reject (👎) each one to teach the policy. Longer responses are shortened to two lines with a Show More link that expands them in place. The number of responses is decided per policy, so the rounds are not always the same size.
Training runs as a short sequence of rounds, and the round you are on is shown above the list as Round 1 of 2. Each round is built from the one before it: when you have judged every response in a round, Next folds your verdicts into the policy and generates a fresh round against the improved version, which takes about a minute. The step tells you what it is doing while it works and counts the responses off as they are labelled, and Next stays unavailable until that finishes. On the final round the button becomes Create instead.
Rejecting opens a small floating box where you can say what the verdict got wrong. The note is optional — the rejection counts either way, and you can close the box with Done or by clicking anywhere outside it. What you type is kept as you type it, so nothing is lost when the box closes. Once written, the note is shown under the verdict as Your note:, and clicking 👎 again reopens the box so you can change it. Switching that response back to 👍 clears the note. After you review every round, Create is enabled; it asks you to confirm, because submitting locks the policy and starts model generation, and then returns you to the catalog.
Your verdicts are not just training signal — they are used to rewrite the policy itself, so the guidelines saved on the created policy are a fuller version of what you wrote on step 1, expanded with the criteria your confirmations and rejections established. You can read and edit them afterwards in the Policy Editor. If that rewrite cannot be completed, the policy is still created, keeping the guidelines exactly as you typed them.
Viewing a created policy: opening a policy from the catalog shows its definition the way you entered it — the guidelines, then Benign Examples and Violative Examples as two separate read-only decks, in the same order the create form collected them. A policy that carries no benign examples — one created before the two-step wizard — shows its violations as a single Examples list instead.
If model generation fails: the policy card shows an alert indicator, and a Rerun Training action becomes available in the card's menu (the ⋮ icon). Selecting it re-initiates model generation using the policy's existing guidelines and examples — there is no need to recreate the policy. If training keeps failing, contact support.
Event Investigation
The Data Explorer is where you investigate the content and sessions that flowed through your guardrails.
Key capabilities:
- Item browser — Browse every analyzed content item and conversation session in a unified table, across the Messages and Sessions tabs.
- Search and filtering — Search by free text and apply structured filters across the columns (violation group, violation type, custom-field values, media type, and more), then sort and paginate the results.
- Item detail drawer — Click any row to open the full content item or session, including its risk scores and the actions that were taken. Image items show a thumbnail; audio items play back in the drawer.
- Export — Download the current page of results, with your search, filters, and sorting applied, as a CSV file.
Workflows
Workflows allow you to define automated rules that trigger webhook actions when specific conditions are met, extending WonderFence's enforcement capabilities beyond the built-in policy actions.
Key capabilities:
- Rule definition — Create rules based on policy violations, confidence levels, event volume thresholds, or combinations of conditions.
- Webhook triggers — Send notifications to external systems (e.g., Slack, PagerDuty, custom endpoints) when rules are triggered.
- Client-side integrations — Trigger actions in your own infrastructure, such as escalating to a human reviewer, suspending a user session, or logging to your SIEM.
- Conditional logic — Define rules with AND/OR conditions across multiple policies for sophisticated enforcement scenarios (e.g., "block and notify if both jailbreak AND PII detected in the same session").
- Three-strikes escalation — Follow repeat offenders and sanction users who are deliberately violating policy.
Overview
The WonderFence Overview provides aggregated analytics and insights across all your policies, violations, and system performance in a single view. Use the Last 7 / 30 / 90 Days selector in the top-right corner to adjust the time window for all metrics; it opens on Last 90 Days.
Filters (top-right): the Application picker scopes every card on the page — or choose All Applications (the default) for a project-wide view — and the date-range selector does the same.
Message Type sits on the Enforcement Distribution panel itself, because it scopes that panel alone: choose Prompt or Response to see the traffic in one direction, or All Message Types (the default) for both. Every other card stays on all message types. Prompt and response figures start on September 7, 2026, so a longer date range counts only from then, and Total Tokens Saved counts only tokens never spent on a blocked prompt, so it reads zero under Response.
Key metrics (top row):
| Metric | What it shows |
|---|---|
| Violation Group Distribution | Total violations detected, broken down by policy group. Select a segment or a legend row to open Data Explorer filtered to that group, over the same period and application. |
| Protected Sessions | Number of unique user sessions scanned by WonderFence guardrails in the selected period. |
| Messages Scanned | Total individual messages evaluated by guardrails. |
| Violation Rate | Percentage of scanned messages that triggered at least one guardrail (violations ÷ messages × 100). |
| Security Readiness | Inverse of the violation rate — reflects how much of your traffic is passing all guardrails cleanly. |
Charts:
- Risk Trend Over Time — Line chart showing daily violation counts over the selected period, useful for spotting spikes or improving trends.
- Actions Distribution — Donut chart breaking down the enforcement actions taken (Block, Redact, Warn, Log) across real-time and event-based responses.
- Alice Recommends — Upcoming AI-powered recommendations panel for tuning your guardrail configuration.
Playground
The Playground is an interactive testing environment where you can evaluate model responses and guardrail behavior in real time.
Key capabilities:
- Live policy testing — Enter sample prompts and responses to see how your active policies would evaluate them. View detection results, confidence scores, and actions that would be triggered.
- Policy experimentation — Test different confidence levels, action settings, and message type configurations to understand their impact before saving changes.
- Custom content testing — Paste real-world examples, edge cases, or adversarial inputs to validate policy behavior against known scenarios.
- Batch prompt files — Attach a CSV of up to 200 prompts and evaluate them all at once instead of sending them one at a time. See Testing a batch of prompts below.
- Example prompts to start from — Before the first message, a row of example types sits under the message box. Pressing one loads a sample prompt of that kind into the box for you to read, edit and send. See Starting from an example below.
- Sample prompt bank — Full Examples CSV, the last item in that same row, downloads a ready-made CSV of test prompts matched to your project's policies, to review, edit, and send back through the batch flow. See Starting from the sample prompt bank below.
- Image prompts — Models that accept images take one alongside the text prompt, so a policy can be tested against what an image carries and not only against wording. See Attaching a file below.
- Audio prompts — Attach a WAV or MP3 and the model answers it while your policies evaluate what it says. See Attaching a file below.
- One model list, grouped by provider — The Model dropdown is the only model control: it lists every provider's models together, under a heading naming the provider that serves each run of them, so picking a model is one step rather than picking a provider first and then a model. Provider is recorded with the message from whichever model you chose.
- Live audio — Where it is enabled for your account, hold a spoken back-and-forth with a speech-to-speech model instead of sending one clip at a time. See Live audio below.
- Every model listed can be run — The Model dropdown offers only the models the Playground can send to, rather than the provider's whole catalog with the rest shown greyed out. The list is short, and anything in it is a valid choice.
- What each model accepts — Every row in the list carries icons for what that model takes: a picture icon for a model that accepts images, a play icon for one that accepts audio, and a document icon on every row, since they all take text. Attaching is never blocked by them — pick by the icons, or attach first and let the send button tell you the model cannot take what you have staged.
- Rapid iteration — Adjust policy settings and re-test immediately, enabling fast feedback loops during policy development.
- Per-message verdict — Every evaluated message is labelled at the top of its own card with what the guardrails did to it: Blocked where the message was refused, Masked where it was rewritten and still delivered, Detected where a policy matched and the message went through anyway, and Not-violative where nothing fired. A message that was acted on is outlined in red. Under the card, each policy that detected something is listed with its risk score out of 100, named as it is in your Policy Catalog — so a verdict here reads the same as the same policy does in Data Explorer and on the policy pages.
- Which prompts were evaluated — A prompt sent with Guardrails on carries a shield beside User, and the Not-violative label when nothing fired, so a message that came back clean is distinguishable from one no policy ever looked at. A prompt sent with the toggle off carries neither. The mark is recorded per prompt, so a chat where you turned the toggle on or off part-way through still shows it correctly for each prompt when you reopen it.
- When a message gets no answer — A prompt whose request timed out or failed says so in place of the answer, rather than sitting there looking like one still being written. Where the fault was the model provider's — it timed out, rate limited us, or returned an error — the message says so, so a slow or unhappy model is not mistaken for a problem with the platform. Send it again to retry, or pick a different model. How long a model is given depends on the model: the fast text models are cut off after about 20 seconds, while the reasoning, image and audio models are given about 90, since they legitimately take longer to answer.
- An application is required — A message is always evaluated against a specific application, so with none to choose from the message box stays blocked with No applications available to you. Add one under Application Inventory, or, on an account that uses group application access, ask an administrator to give one of your groups an application under Groups.
- Chat list — Past chats are listed on the left, under a Chats tab, where you can search them and reopen one. Delete Chats, below the list, clears them all. Use the toggle at the right of the tab strip to collapse the pane into a narrow rail — each chat becomes an icon whose tooltip shows its name. The choice is remembered the next time you open the Playground.
- Active Policies — The second tab beside Chats lists the policies attached to the application selected in the top bar, so what is guarding this conversation is readable without leaving the page. Each one is a card carrying its name, whether it is a System or Custom policy, the message types it governs, and its action — Block, Detect or Mask. Search narrows the list by name, and changing the application in the top bar re-lists it. The cards are read-only; a policy is changed on Enforcement Policies. The tab is also offered on a demo link, listing the policies of the one application that link is pinned to. It is not offered where there is no application to scope it to — a project that assigns policies account-wide rather than per application.
- Starting a new chat — New Chat sits on the right of the top bar, so it is in reach whether the chat list is expanded or collapsed. The chat on screen is saved to the list before the new one opens.
- One history per account and project — Chats belong to the account and project they were sent in, and are listed only there. Switching workspace gives you that workspace's own list, and switching back brings yours here. Chats are stored in your browser, so they do not follow you to another browser or machine.
- Deleting one chat — Hovering a chat in the list reveals an actions menu at its right edge; choose Delete there and confirm to remove that chat on its own, leaving the rest of the list alone. Deleting the chat you are reading clears the thread and starts you on a new one. The actions menu is not offered in the collapsed rail, and deletion cannot be undone.
- Chats reopen in context — A chat is saved with the settings it was held under: model, application, and whether Guardrails was on. Reopening it puts those back in the top bar, so what you are reading is labelled with how it was produced — a verdict only means something next to the policy set that produced it. A chat saved before this was recorded leaves the top bar as you have it.
- The top bar is remembered — What you last selected with no chat open is what the next new chat — and your next visit to this account and project — starts from, so a reload does not send you back to picking a model and an application again. Reopening a saved chat changes the top bar for that chat only; it does not become your starting point. A remembered model or application that no longer exists is dropped rather than left in place, since it could not be sent with.
Attaching a file
The plus on the leading edge of the message box is where every attachment starts, whatever kind of file it is.
- What the plus offers — Prompt file (CSV), Image and Audio. Each option states what it takes (max 200 prompts; PNG, JPEG, GIF or WebP up to 5 MB; WAV or MP3 up to 7 MB).
- Every kind can always be attached — Image and Audio are offered whatever the selected model accepts, so you never have to guess a model before you can pick the file. What the model accepts is checked when you send, not when you attach.
- The send button is what refuses a file the model cannot take — Stage an image on a model that takes no images, or audio on one that takes no audio, and Send greys out. Hover it and it names the mismatch and tells you to switch to a model that accepts that kind of input. Switching the model in the top bar releases the send; the staged file stays put.
- The model list shows what each model accepts — every model in the dropdown carries small icons for text, image and audio. Hover an icon to read it. Picking a model marked for the kind you are about to send saves the round trip.
- One thing at a time — A typed message, one image, one audio file, or a CSV — never a combination. While one of them is in the box the others are greyed out, and the plus itself goes inactive once you start typing. An attached image or audio file can still be swapped for another; remove it with the ✕ on its card.
- What an attachment looks like — Whatever the kind, it stages as a card inside the message box carrying the file name and what it holds — the number of prompts for a CSV, or simply Image or Audio. A CSV and an image offer Preview; audio offers a play button that plays it where it stands.
- What comes back for audio — The model answers it, and your policies evaluate it as well when Guardrails is on. Only a model that accepts audio can be sent it, and only in the formats it accepts, so an audio turn is not left without a reply.
- Playing it back — A sent image or audio file stays in the conversation, and audio plays from the message itself, so you can hear what was evaluated.
- Sending a long prompt — The message box grows with what you type, up to about six lines, then scrolls. Enter sends; Shift + Enter starts a new line, so a multi-line prompt can be composed without sending it half-written.
Starting from an example
Before the first message of a chat, a row under the message box offers one example per policy area your project moderates — Test prompt injection, Test PII leakage, Test adult content, and so on.
- What pressing one does — It loads a sample prompt of that kind into the message box and puts the cursor at the end of it. Nothing is sent: read it, edit it, then send it yourself.
- Which examples you see — The same narrowing the sample prompt bank uses, so the row covers what your project actually moderates. Every example is meant to trip something: the harmless control prompts are in the CSV download, not here, since an empty message box already invites you to type anything you like.
- One prompt per example per visit — Pressing the same example twice reloads the same prompt. Reload the page for a different draw.
- When it is not offered — The row belongs to the empty chat only; once a chat has a message, the message box stands alone. The examples also grey out while an attachment is in the box, since they write into a box that is holding something else.
Live audio
Where it is enabled for your account, a waveform button sits beside Send. It is for talking with the model rather than sending it a clip: the conversation runs in both directions at once.
- Starting and stopping — The button replaces the chat with a live panel. Start talking asks for your microphone the first time, then streams it to the model and plays the reply back; what each of you says is transcribed into the panel as it is spoken. Stop ends the session, Start again opens a fresh one, and Back to chat returns you to the composer, ending anything still running on the way out.
- Its own model list — Live audio runs on speech-to-speech models, so the Model dropdown offers those while the panel is open, grouped by provider like the chat list. The button is greyed out, with the reason on hover, while that list is loading or when there is none to offer.
- A status you can read at a glance — The panel carries a label for where the session is: Not started, Connecting, Listening, Speaking, Ended, or Error. An error is spelled out in the panel rather than left to the status alone.
- Guardrails do not apply yet — The Guardrails toggle is shown off and locked while the panel is open, and the panel says the same. Nothing spoken in a live session is evaluated against your policies, and the conversation is not kept in your chat history.
- An application is still required — The same rule as the message box: with no application selected, Start talking stays disabled and the panel says why.
- Limits — A session ends itself after seven and a half minutes; start another to carry on. Only a couple can run at once under the same user, so a session opened in a third tab is refused.
- When it is not offered — The panel says so on a browser that cannot capture live audio, and it is not offered at all on a demo link.
Starting from the sample prompt bank
Full Examples CSV, the last item in the examples row, gives you a starting set of test prompts, so you do not have to write one from scratch.
- What you get — A
sample-prompts.csvwith a singlepromptcolumn, ready to attach back as-is. - Matched to your policies — The file only contains prompts for the policy areas your project has enabled, so you are not testing for things you do not moderate. Everyday, harmless prompts are always included as a control group: they show whether your policies over-block ordinary traffic.
- When it cannot be matched — Custom policies are named by you, so they cannot be mapped to a policy area. If your project moderates on one — or has no policies enabled at all — you get the full bank instead of a narrowed one, since a file of nothing but harmless prompts would come back entirely clean and read as guardrails that never fire.
- Yours to edit — Review the file, delete what is not relevant, and add your own prompts before uploading it.
Testing a batch of prompts
Attach a .csv of prompts (up to 2 MB) through the plus button — see Attaching a file above.
- File format — The first row must be a header row. If the file has a single column, that column is used. Otherwise the column named
prompt,text, orinputis used; if none of those is present, the message box asks which column holds the prompts, under the attachment card. Other columns are ignored. - Before sending — The attachment card shows the file name and how many prompts were detected, so a wrong file is caught before spending model calls. Preview reads the full list, headed by the column the prompts came from. Repeated prompts are kept, not removed — a deliberate duplicate is still evaluated.
- Results — Each prompt is evaluated on its own, as an independent single-turn conversation, so no prompt is influenced by the one before it. The chat shows one result card per file, headed the same way the results view is — the file, its size and model, and the run's split as All, Violative, Not-violative and Failed. Under the header it previews the first couple of prompts; click a group to preview that group's first couple instead, so a run whose violative prompts are all at the foot of the file can be seen from the chat without opening it. Open the card for a row per prompt with its result, risk score, and any detections.
- The file header — The results view opens on a header naming the file, with how many prompts it holds, its size and the model they ran against, and the run's split on its trailing edge: All, then Violative and Not-violative where guardrails were on, and Failed where a prompt never got a verdict. Those counts are also the filter — see below.
- Filtering — Click a group in the file header — All, Violative, Not-violative, Failed — to narrow the list to it, or search across prompts, responses, and detection names. Each option carries its own count, taken over the whole run, and the three groups add up to All. Failed is not a verdict — it collects the prompts the run never managed to evaluate, so they can be pulled out of the list and re-run rather than read as prompts that passed. A group with nothing in it is not offered: a run's results are fixed once it finishes, so an empty option could never become useful. A run where no group can be picked still states its total, since that is the run's size either way — it just isn't offered as a choice, and search is the only narrowing left.
- Sorting — The list opens in file order, so the row number matches your CSV. Click a column header — #, Prompt, Result or Risk score — to sort by it, click again to reverse it, and a third time to drop the sort and go back to file order. Sorting applies to the rows the filter and search have left, and the whole run stays on one scrolling list — there are no pages to step through.
- Result column — A prompt where any action fired — BLOCK, MASK or DETECT — is tagged Violative; a prompt that was evaluated and came back clean is tagged Not-violative. Both are labelled on purpose: an untagged row would read equally as "clean" and as "never evaluated". The only untagged rows are the ones with no verdict at all, which the Failed filter accounts for. Which action was taken, against which policies, is on the row's detail: open a row to see the same ACTION: line and detections a single chat message shows. Click a row — or arrow up and down the list — to move through the prompts.
- Risk score — Each row shows its risk as a dial scored out of 100, coloured by band (high, medium, low), drawn the same way here and one click deeper. A prompt that tripped several policies shows the highest of their scores, not an average, so one severe hit is never diluted by mild ones — hover the dial to see that spelled out. A not-violative prompt still shows its score, since a prompt that scored high and went through anyway is worth reviewing.
- Sent with guardrails off — A batch sent with the Guardrails toggle off was never evaluated, so there is no verdict to report: the card and the results view both drop the two verdict groups and the per-prompt tags, leaving the run's total. Failed still appears if any prompt failed, since a prompt can fail to send whether or not guardrails were on. Nothing is hidden — every prompt and response is still there. A run that WAS evaluated is marked with a shield on its card and on its results header, the same mark a single evaluated prompt carries; a run sent with the toggle off carries none.
- Downloading the results — A download button on the result card in the chat, and on the file header inside the results view, writes the whole run to a CSV — so verdicts can be shared, compared between runs, or analysed outside the platform, without having to open the run first. The file has one row per prompt you uploaded, in file order and keyed by the row number in your own CSV — including the prompts the run never evaluated — and carries the prompt, the response, the verdict, the policies that fired, the risk score, and the error where there is one. It is the whole run, not the list you are looking at: the filter, the search and the sort do not narrow it. The file is named after the file you uploaded and the run it came from, so two runs of the same prompts do not overwrite each other. A run sent with Guardrails off leaves the verdict, policy and risk columns empty, since nothing was evaluated to report.
- Kept with the chat — A file's full breakdown is saved alongside the chat, so reopening that chat from the sidebar brings the result card and every per-prompt row back. Results live in your browser only, so clearing the chat history discards them.
- Long-running files — Evaluation is capped at about a minute. If the file does not finish in time, the results still open with a row for every prompt in your CSV. A prompt that was not evaluated carries no verdict — it is counted as neither violative nor not-violative, and the card does not count it as either. Inside the results view it is grouped under Failed, and opening it says whether it timed out or the request failed. Re-attach the file to try those prompts again.