Help Center/Call Handling

Watchtower — AI Call Quality Monitoring

Watchtower checks every customer call against your own knowledge and rules using a fixed list of pass/fail checks plus any expectations you add, verifies every fail with a second pass, and rolls the results into a 30-day health score — so you know how your AI agent is actually performing without listening to every recording.

Watchtower is the quality-monitoring layer that runs on every customer call. After each call ends, a scorer reads the transcript together with the knowledge, rules and lookup results Allison actually had on that call, runs a fixed list of pass/fail checks, and a second pass verifies every fail before it counts. The results roll into your organization's 30-day health score. You see them on the Watchtower dashboard, in a weekly digest email, on individual calls, and through Allison herself.

The point: you don't have to listen to recordings to know whether your agent is doing a good job, and when something goes wrong you can see the exact line it went wrong on and whose fix it is.

How It Works

Every customer call (a real conversation on your AI agent's number) ends, and four things happen:

  1. Engagement check. Was this a real call worth measuring? Robocalls, wrong-number hangups, silent dialer attempts, and very short non-engagements are filtered out. (See "What doesn't get scored" below.)
  2. Checks. For real calls, the scorer is given the transcript, the knowledge and escalation rules Allison had on that call, the lookups she ran and what they returned, and how the call ended (who hung up). It judges every check as pass, fail, or not applicable. Pass is the default: a check can only fail if the scorer quotes the transcript line that shows it.
  3. Verification. Every fail goes to a second, independent pass that reads the same evidence and either confirms or overturns it. A fail whose quoted line is not actually in the transcript is thrown out automatically. A confirmed fail counts. A fail the second pass could not reach stands as not yet verified and counts until it is re-checked.
  4. Aggregation. Several times a day, your health score is recomputed from the last 30 days of scored calls. It is the share of applicable checks passed, with recent calls weighted more (the last 7 days count fully, week 2 counts 75%, weeks 3 and 4 count 50%). A health score needs at least 5 scored calls in the window; until then the dashboard shows an empty state.

The Checks

Twelve core checks in four categories. Every business is scored against the same twelve, so scores mean the same thing everywhere.

CategoryCheckWhat passes
ResolutionUnderstood what the caller wantedAllison worked out what the caller was after
ResolutionNeed met, or clearly scoped with a next stepThe caller got what they needed from your knowledge, or was told clearly what you don't offer and given a next step. Saying "we don't do that" when you don't is a pass.
ResolutionCaptured what the flow requiredA message, callback number, order, or booking detail was captured and the caller confirmed it
AccuracyNo unsupported business factsEvery hours, price, service, policy, address, or staff claim came from your knowledge or a lookup
AccuracyNo self-contradictionAllison did not contradict herself, your knowledge, or a lookup result
AccuracyHonest when knowledge ran outWhen your knowledge did not cover a question, she said so rather than guessing
ProfessionalismNothing a caller would complain aboutNo rudeness, dismissiveness, talking over the caller, or sustained confusion
ProfessionalismClear and on topicNo loops, no proceeding before the caller answered, no reading internal rules aloud. The short hold messages Allison plays automatically while she works on an answer never count against her
ProfessionalismEnded the call cleanlyWhen Allison ended the call, the next step was confirmed and she said goodbye
EscalationEscalated when a rule matchedWhen one of your escalation rules applied, she did what it says
EscalationNo unnecessary escalationShe did not hand off something she could have handled from your knowledge
EscalationExplained the handoffWhen transferring or arranging a callback, the caller was told what happens next

Not applicable is a real verdict, not a dodge. If no escalation happened, the escalation checks are not applicable and neither help nor hurt. If the caller hung up mid-conversation, "Ended the call cleanly" is not applicable: a caller leaving is never counted against Allison.

Your expectations. You can add your own checks in plain language: "Every caller who asks about service is offered an appointment before the call ends," or "Never quote a price; refer to an estimate." They are judged with the same quote-and-verify process and count in your score exactly like the core checks. Add them on the Watchtower page under "What Allison is scored on," or just tell Allison in the support chat. Core checks can be switched off for your business if one does not fit how you work; they cannot be reworded.

What the Numbers Mean

Each call's score is the share of its applicable checks that passed, 0 to 100. A call handled correctly scores 100, including a call where the caller wanted something you don't offer and Allison said so.

Category bars on the dashboard show the share of checks passed in that category over 30 days, 0 to 10. A category with nothing applicable in the window shows no bar.

Score rangeLabelColor
90+ExcellentGreen
75-89GoodBlue
60-74Needs attentionYellow
Below 60CriticalRed

The health score is pooled across calls rather than averaged: a short call with two applicable checks does not weigh the same as a long call with fourteen. So the hero number is not the arithmetic mean of the call list, and that is deliberate.

Caller mood is still measured (1 to 10) and shown next to the bars, but it is not part of the score. A caller disappointed by a truthful "no" does not make Allison wrong.

Whose Fix Is It

Every confirmed fail carries a gap type:

  • Knowledge gap. Something is missing or wrong in your knowledge. The fix is a fact, an FAQ, or an intent.
  • Rule gap. An escalation rule is missing or not matching. The fix is the rule.
  • Allison's behaviour. She had what she needed and still got it wrong: a stall, proceeding before you answered, a wrong fact she had the right information for. That is on us. Allison files it as a ticket when you ask her about it.

This is the distinction that keeps a low score honest. A gap in your setup and a fault in the agent look different, and the dashboard says which is which.

Flagged Calls

A call is flagged when it fails a check a business owner would want to look at: a need left unmet, a missed capture, an unsupported or contradictory fact, a guess presented as fact, complaint-worthy behaviour, a broken escalation rule, or an unnecessary transfer. A coherence hiccup or a rough close costs score but does not flag; those are ours to fix, not yours to review.

Flagged calls show up:

  • On the Watchtower dashboard, marked with an orange icon and the checks that failed. The "N flagged for review" count in the health header is a filter.
  • On the call list, alongside the call's score badge.
  • In the weekly digest, grouped by the check that failed, with one quoted example.

What You See on the Dashboard

Visit /dashboard/watchtower to see:

  • Health score hero: the 0 to 100 score, its label and trend, and the category bars.
  • Weakest category callout: when a category is below 7 or calls have been flagged, a card names it and offers Ask Allison about this. Allison reads which checks failed, with the quoted lines, tells you whether each is a knowledge gap, a rule gap, or her own behaviour, and can make the configuration change for you once you confirm it.
  • Score distribution: four tiles, each a filter for the call list.
  • Call list: every scored call from the last 30 days. Flagged rows show the failed checks inline. Clicking a row opens the call: transcript, recording, and every check with its verdict, the quoted line under each fail, the gap type, and whether the second pass confirmed it.
  • What Allison is scored on: the core checks with on/off switches, and your expectations with add, edit, enable, and delete.

If you have multiple locations, the sidebar location selector filters the counts and the list to that location; expectations you add while a location is selected apply to that location only. The health score and category bars are always organization-wide.

Manual Overrides

Sometimes you'll disagree with the engagement check, in either direction. Open the call and look for Health score inclusion. You'll see the AI's verdict and an option to override it: exclude a call that counts, or re-include one that was filtered out. Overrides are reversible.

Beneath the title, the dashboard shows how many calls were scored in the last 30 days and how many were excluded.

Weekly Digest

Once a week, at the start of the week, Watchtower sends:

  • Current health score with label and trend
  • Category pass rates (rolling 30 days), including your expectations once you have any, with caller mood noted for context
  • Total calls scored
  • Failed checks this week, one row per check with how many calls failed it, the gap mix, and one quoted example. The 30-day flagged total is shown alongside.
  • Your weakest category, with a recent example

The digest goes to recipients in your Quality and performance notification category (Settings → Notifications).

Real-Time Alerts

Below threshold: your health score drops below 70 for the first time (or the first time after a recovery). Score drop: an already-below-threshold score drops by 3 or more points. Recovery: the score crosses back above 70. Trend decline: three consecutive declining snapshots. Each fires at most once per 24 hours.

What Doesn't Get Scored

  • Calls under 10 seconds
  • Calls where the caller never spoke
  • Calls under 5 caller-words with no lookups
  • Calls the engagement check reads as robocalls, wrong-number hangups, IVR-keypress probes, or pure non-engagement

The filter is conservative: when in doubt, the call is included. The manual override is your fix for the rare miss.

What's Not Yet Supported

  • Discovery-call scoring: Watchtower scores customer calls on your AI agent's number, not the discovery calls prospects make to Allison Voice itself.
  • Rewording core checks: they can be switched off per business, not edited. Add an expectation for anything the core checks don't say.
  • Configurable thresholds: the 90 / 75 / 60 labels are fixed.
  • Historical exports: snapshots are kept indefinitely but there is no download. Reach out if you need the series.
  • Daily or custom-cadence digests: weekly only.
  • Audio review: checks are judged from the transcript and the knowledge Allison had, not the audio.

Still have questions? Log in to chat with Allison.

Log In to Chat