checklist
AI chatbot weekly review scorecard for small teams
A practical weekly review scorecard for customer-facing AI chatbots, covering answer quality, handoff, tool actions, source freshness, privacy, incidents, customer impact, owner decisions, and next-week actions.
Use this weekly scorecard after a chatbot launch, after a high-risk change, during a pilot, or any time a customer-facing AI chatbot is handling real support, sales, account, billing, product, or data questions.
A weekly review keeps the team from treating “the bot is still running” as success. The review should decide whether the bot can expand, stay limited, pause one risky path, fix sources, tighten handoff, disable a tool action, or send corrections. If the team cannot explain the bot’s current risk, run the AI Tool Risk Checker and attach the result to the weekly record.
Bottom line
A small team should review these eight chatbot areas every week:
- Answer quality and customer corrections.
- Human handoff and escalation.
- Tool actions and connector behavior.
- Knowledge source freshness.
- Privacy, data requests, and sensitive inputs.
- Prompt injection, abuse, and unsafe behavior.
- Incidents, pauses, customer messages, and vendor tickets.
- Next-week scope decision.
Use the Small Team AI Security Checklist to confirm owners, access review, incident routing, and evidence storage. This page gives the weekly scorecard.
When to use this scorecard
| Scenario | Use weekly? | Why |
|---|---|---|
| First month after launch | Yes | Early failures often show up in real customer traffic. |
| Active pilot | Yes | Expansion should depend on evidence, not excitement. |
| New source, connector, prompt, or tool action | Yes | Changes can shift answer quality and risk. |
| Bot handles support or customer data | Yes | Customer impact and data handling need recurring review. |
| Bot is paused or in fallback mode | Yes | The team needs a restart or continued fallback decision. |
| Bot has no customer traffic | Maybe | Review monthly unless settings, sources, or access changed. |
| Static public FAQ bot with no data and no actions | Maybe | A lighter monthly review may be enough after stabilization. |
Run the review on a calendar. Do not wait for the next incident.
Weekly scorecard
Score each area from 0 to 3.
| Score | Meaning |
|---|---|
| 3 | Healthy. Evidence supports current scope. |
| 2 | Acceptable with minor fixes. Keep scope but track action. |
| 1 | Risky. Limit scope or fix before expansion. |
| 0 | Unacceptable. Pause affected capability or route to human fallback. |
| Area | Score | Evidence to review |
|---|---|---|
| Answer quality | Sampled conversations, corrections, customer disputes, repeated wrong answers. | |
| Human handoff | Handoff success, sensitive-topic routing, customer wait time, missed escalations. | |
| Tool actions | Attempted/completed/denied/failed actions, duplicate actions, rollback events. | |
| Sources | Source freshness, conflicts, stale documents, missing owners, retrieval failures. | |
| Privacy and data | Sensitive inputs, deletion/export/correction requests, retention exceptions. | |
| Prompt injection and abuse | Abuse attempts, suspicious prompts, guardrail failures, source manipulation. | |
| Incident communication | Customer messages, internal updates, vendor tickets, status page decisions. | |
| Monitoring and logs | Log coverage, missing event details, alert quality, dashboard gaps. | |
| Customer impact | Affected tickets, support load, wait time, complaints, corrections sent. | |
| Next-week readiness | Open risks, owner capacity, test coverage, approval for expansion. |
If any area is 0, do not expand the bot. If two or more areas are 1, keep the bot limited until fixes close.
Review inputs
Collect these before the meeting.
| Input | Source owner |
|---|---|
| Conversation sample | Support/product owner. |
| Customer disputes and corrections | Support owner. |
| Escalation and human handoff report | Support lead. |
| Tool action log | Product/admin owner. |
| Connector and source sync status | Engineering/source owner. |
| Privacy/data request log | Privacy/support owner. |
| Prompt injection and abuse notes | Security owner. |
| Incidents, pauses, and fallback events | Incident lead. |
| Vendor tickets and status | Admin/vendor owner. |
| Monitoring dashboard gaps | Bot/admin owner. |
The review should use redacted excerpts and controlled links, not broad copies of transcripts or customer data.
30-minute agenda
| Time | Topic |
|---|---|
| 0-5 min | Confirm scope, traffic level, owner attendance, and previous actions. |
| 5-10 min | Review answer quality, corrections, disputes, and source gaps. |
| 10-15 min | Review handoff, escalation, sensitive topics, and customer impact. |
| 15-20 min | Review tool actions, connectors, privacy/data requests, and abuse attempts. |
| 20-25 min | Score each area and pick keep/limit/pause/expand decision. |
| 25-30 min | Assign next-week actions, owners, due dates, and monitoring focus. |
If the meeting needs more than 45 minutes, the bot probably needs a narrower scope or better review inputs.
Answer quality review
| Signal | Green | Red flag |
|---|---|---|
| Sampled answers | Mostly correct, sourced, and within approved scope. | Repeated wrong answers on the same customer-impacting topic. |
| Corrections | Few, tracked, and closed. | Corrections are sent without source or prompt fixes. |
| Disputes | Routed to support and resolved. | Customers argue with the bot without human handoff. |
| No-answer behavior | Bot declines or routes when unsure. | Bot invents answers or gives unsupported advice. |
| Source citation | Source-backed answers map to current sources. | Source conflicts, stale documents, or missing source owner. |
Use the AI chatbot answer correction workflow template when corrections appear repeatedly.
Handoff and escalation review
| Signal | Green | Red flag |
|---|---|---|
| Human handoff | Customers can reach a person when needed. | Customer asks for human and bot keeps responding. |
| Sensitive topics | Legal, privacy, billing, access, HR, health, finance, and security topics route correctly. | Bot answers sensitive topics without approved wording or owner routing. |
| Escalation severity | S1/S2 items have owner and timeline. | High-impact tickets stay in normal queue. |
| Customer wait time | Fallback queue is manageable. | Fallback queue grows after bot pause or failure. |
| Evidence | Ticket has enough context to review. | Full transcript copied broadly or key context missing. |
Use the AI chatbot customer escalation workflow template for disputed or sensitive customer conversations.
Tool action and connector review
| Signal | Green | Red flag |
|---|---|---|
| Action approvals | High-impact actions require human confirmation. | Bot can change account, billing, access, deletion, or outbound messages without review. |
| Action logs | Attempts, approvals, failures, retries, and rollbacks are visible. | Logs cannot reconstruct what the bot tried to do. |
| Connector scope | Minimum required read/write permissions. | Broad connector scopes remain after pilot. |
| Failures | Failed actions are routed to humans. | Bot retries unsafe or duplicate actions. |
| Changes | New actions go through approval. | Tool action changed without owner review. |
Use the AI chatbot tool action approval checklist before expanding action scope.
Source and knowledge review
| Signal | Green | Red flag |
|---|---|---|
| Source owner | Each source has a named owner. | Bot answers from ownerless docs. |
| Freshness | Key sources are current and synced. | Stale docs drive live answers. |
| Conflicts | Conflicting answers are resolved before restart or expansion. | Bot chooses between conflicting sources without rules. |
| Sensitive sources | Customer, internal, or restricted sources are scoped. | Sensitive content appears in customer-facing answers. |
| Retrieval tests | Known questions still return correct sources. | Retrieval fails after source or prompt change. |
Use the AI chatbot knowledge base review checklist when source-backed answers score below 2.
Privacy and data review
| Signal | Green | Red flag |
|---|---|---|
| Sensitive inputs | Passwords, payment data, private keys, and regulated data are routed safely. | Sensitive inputs remain in uncontrolled logs or tickets. |
| Data requests | Deletion, export, correction, opt-out, and data-use questions route to owner. | Bot invents privacy commitments or misroutes requests. |
| Retention | Conversation retention follows approved policy. | Retention settings are unknown or changed without review. |
| Training/product improvement | Admin settings and vendor terms are reviewed. | Team cannot say whether chats are used for training or product improvement. |
| Access | Admin and reviewer access is minimum necessary. | Broad staff access to chatbot logs continues after pilot. |
Use the AI chatbot deletion and export request workflow for customer data requests.
Incident and pause review
| Signal | Green | Red flag |
|---|---|---|
| Pauses | Pause scope, owner, and restart gate are recorded. | Bot was paused but no one knows why or when to restart. |
| Fallback | Customers had a safer route during pause. | Customers were stranded or sent in loops. |
| Communication | Customer messages were factual and approved when needed. | Team said “no risk” before evidence review. |
| Vendor ticket | Vendor got redacted, useful technical context. | Vendor received unnecessary customer data or no useful evidence. |
| Learning | New test, source fix, monitor, or owner rule was added. | Same incident pattern repeats. |
Use the AI chatbot incident communication template when customer or vendor messages were sent.
Decision rules
| Score pattern | Decision |
|---|---|
| All areas 2-3 and no unresolved S1/S2 issues | Keep current scope. Consider limited expansion only if owners agree. |
| One area at 1 | Keep current scope but assign fix and review next week. |
| Two or more areas at 1 | Limit scope. No expansion until fixes close. |
| Any area at 0 | Pause affected capability or route to human fallback. |
| Missing logs for high-impact action or incident | Pause affected high-risk capability until evidence improves. |
| Repeated customer-impacting wrong answer | Pause affected answer category and fix source/test set. |
| Repeated handoff failure | Force human route for sensitive topics until fixed. |
| Broad connector or tool-action concern | Disable or restrict the connector/action until reviewed. |
The weekly decision should be boring and explicit: keep, limit, pause, restart, expand, or retire.
Weekly review record
Copy this into the bot operations record.
| Field | Entry |
|---|---|
| Review week | |
| Bot name and channels | |
| Review owner | |
| Traffic level | |
| Conversation sample size | |
| Answer quality score | |
| Handoff score | |
| Tool action score | |
| Source score | |
| Privacy/data score | |
| Prompt injection/abuse score | |
| Incident communication score | |
| Monitoring/log score | |
| Customer impact score | |
| Final decision | Keep, limit, pause, restart, expand, or retire. |
| Top risk this week | |
| Top fix next week | |
| Owner and due date | |
| Evidence location |
Keep this record short enough that the team will actually complete it weekly.
Next-week action tracker
| Action | Owner | Due | Evidence of completion |
|---|---|---|---|
| Add repeated wrong answer to regression test set | |||
| Update or remove stale source | |||
| Tighten human handoff rule | |||
| Disable or restrict risky tool action | |||
| Update customer notice or fallback message | |||
| Review connector permissions | |||
| Add alert or dashboard check | |||
| Close vendor ticket or preserve vendor response |
Every action should have one owner. Shared ownership often means no ownership.
Metrics to track
| Metric | Why it matters |
|---|---|
| Weekly conversation count | Shows whether the sample is meaningful. |
| Sampled answer pass rate | Tracks answer quality over time. |
| Customer corrections | Shows repair workload and source quality. |
| Human handoff success rate | Shows whether customers can escape automation. |
| Sensitive-topic handoffs | Shows how often the bot meets high-risk topics. |
| Tool action attempts and failures | Shows automation exposure. |
| Source freshness failures | Shows knowledge quality. |
| Privacy/data requests routed correctly | Shows data request reliability. |
| Prompt injection or abuse attempts | Shows attack pressure. |
| Pauses and fallback events | Shows operational stability. |
| Time to close next-week actions | Shows whether review leads to improvement. |
Do not optimize only for deflection. A bot that deflects support but creates wrong answers, hidden risk, or unhappy customers is not healthy.
Evidence checked
This scorecard is aligned with:
- NIST AI RMF Core, which emphasizes documented roles, monitoring, measurement, incident identification, response, recovery, communication, third-party review, and periodic review of risks and controls.
- NIST AI 800-4 monitoring report summary, which identifies post-deployment monitoring as crucial and describes functionality, operational, security, compliance, impact, and human-factors monitoring categories.
- NIST AI 800-4 publication page, which describes monitoring deployed AI systems to validate real-world reliability, track unforeseen outputs, and gain visibility into unexpected consequences.
- CISA JCDC AI Cybersecurity Collaboration Playbook alert, which emphasizes voluntary information-sharing processes for AI cybersecurity incidents and vulnerabilities.
- CISA joint guidance on deploying AI systems securely, which emphasizes protecting AI systems and related data/services, detecting malicious activity, and responding to incidents.
- OWASP Top 10 for LLM Applications, which covers prompt injection, sensitive information disclosure, insecure plugin design, excessive agency, misinformation, overreliance, and related LLM application risks.
- FTC artificial intelligence guidance, which tracks FTC guidance and enforcement activity related to AI claims, accuracy, privacy, confidentiality, chatbot monitoring, and consumer protection.
- Cybergiz templates for chatbot launch, handoff, source review, tool action approval, monitoring, customer escalation, pause and fallback, incident communication, and change approval.
This page is practical operating guidance, not legal, privacy, compliance, audit, certification, customer-support, incident-response, product-management, or security assurance advice.
FAQ
How many conversations should we sample each week?
Start with 20 conversations or 5 percent of weekly chatbot conversations, whichever is smaller. Sample all S1/S2 escalations, all tool-action attempts, and all customer disputes.
Who should attend the weekly review?
At minimum: support owner, bot/product owner, and one security or privacy owner. Add source, engineering, legal, or vendor owners when their area scored below 2.
Should we review only bad conversations?
No. Include random normal conversations, customer disputes, handoffs, tool actions, and sensitive-topic routes. Reviewing only failures can hide broad quality drift, and reviewing only happy paths can hide risk.
What score means the bot can expand?
Expansion should require all areas at 2 or 3, no unresolved S1/S2 issues, working handoff, current sources, visible logs, and owner approval. Expansion is a decision, not a default reward.
What if the team cannot find enough evidence to score an area?
Score it 1 or 0 depending on risk. Missing evidence is itself a risk, especially for tool actions, data requests, incidents, and customer-impacting answers.
Should this replace incident review?
No. Incident review happens when something goes wrong. Weekly review catches patterns, recurring weak signals, and unfinished corrective actions.
How long should we keep weekly records?
Keep them according to the team’s evidence retention schedule. Avoid storing unnecessary raw transcripts, credentials, payment data, regulated data, or full customer exports in the weekly record.
What is the most important weekly decision?
Whether to keep, limit, pause, restart, expand, or retire a chatbot capability. A weekly review without a scope decision becomes reporting, not governance.