What runs on every reply
Hallucination grader
Checks every reply for factual claims that nothing supports: prices, hours, availability, contact
details. Deterministic, on the live path, with no model call, so it costs no latency. A second
asynchronous pass catches what the rules miss.
Forbidden phrases
Matches every reply against your forbidden list, which merges your vertical pack’s defaults with
anything you add.
Say-guard
Strips reasoning, tool narration and deliberation out of the reply. The customer gets the answer,
never the working out.
Booking circuit breaker
Catches a reply that claims an action succeeded when no tool committed it, for example announcing
a confirmed appointment that was never booked.
Hallucination grading
The grader extracts factual claims from the reply and verifies each one against the evidence available in that turn: tool results, your active catalog, your configured working hours, and what the customer already said. Flag kinds:Threshold
Action
Choosing a setting
- Clinics, legal, financial services:
mediumandhandoff. A medium-confidence fabrication should never reach a customer. You pay for it in handoff volume. - Retail and hospitality:
highandwarn. The default. - A new workspace still calibrating:
lowandwarn. Maximally noisy, so you can see the real flag distribution before committing to a policy.
Forbidden phrases
Your list merges two sources: the phrases your vertical pack ships, and anything you add. Matching is case-insensitive substring matching, sodiagnose catches “I diagnose”, “let me
diagnose” and “I cannot diagnose”. This is deliberate. Near-misses are usually the model trying to
talk around a term you banned.
block gives the agent a graceful exit and keeps it in the conversation. handoff means a human
must take it from here.
Pack-level forbidden phrases are a floor. You can add to the list, you cannot remove what the pack
ships. This is what stops a clinic workspace accidentally switching off the diagnosis block.
Confidence calibration
Beyond catching fabrications, the platform tracks whether the agent’s confidence matches its accuracy. An agent that hedges when it is right and asserts when it is wrong is badly calibrated even if its hallucination rate looks fine. Calibration is reported alongside quality scores.Tool containment
Guardrails cover what the agent says. The tool permissions cover what it can do. Turning a tool off removes it from the agent’s surface entirely rather than instructing the agent not to use it. See The agent.What is not a guardrail
- Profanity filtering. Add terms to your forbidden list if you need them. There is no generic profanity list.
- PII redaction in replies. The agent has no access to another customer’s data, so there is nothing to redact at the reply layer. Redaction applies to exports and audit records instead.
- Topic restriction. “Only talk about property, not the weather” is handled by your vertical pack’s redirect copy rather than as a separate guardrail.
