checklist
AI chatbot quarterly governance review scorecard for small teams
A quarterly governance scorecard for customer-facing AI chatbots, covering ownership, customer impact, access, data, model and prompt changes, connectors, tool actions, incidents, exceptions, vendors, decisions, and follow-up.
Use this scorecard once each quarter to decide whether a customer-facing AI chatbot should continue, be limited, be changed, or be paused.
The quarterly review is not a larger weekly status meeting. It is a governance decision based on evidence from the last period: ownership, customer impact, access, data handling, model and prompt changes, connectors, tool actions, incidents, exceptions, vendor changes, and unresolved control gaps. Run the AI Tool Risk Checker before the meeting when the route, data, or action scope has changed.
Bottom line
Keep a customer-facing chatbot in service only when the team can explain who owns it, what it is allowed to do, which data it can reach, how users can reach a human, what changed, what went wrong, and what residual risk remains. A good quarterly result may be continue, limit, redesign, replace, or retire. It does not have to be an unconditional approval.
Use the Small Team AI Security Checklist for baseline owners, access, incident routing, and evidence storage. Pair this scorecard with the AI chatbot weekly review scorecard for recurring operating signals, the AI chatbot access recertification checklist for permissions, and the AI chatbot exception register template for time-bound deviations.
When to use this scorecard
| Situation | Use the quarterly scorecard? | Additional action |
|---|---|---|
| Customer-facing chatbot has operated for a full quarter | Yes | Review the full evidence set and make a continuation decision. |
| Chatbot launched or materially changed during the quarter | Yes | Compare the approved baseline with the current route. |
| Route has high-impact account, billing, order, or outbound actions | Yes | Require action-owner and security or privacy review. |
| A major incident or near miss occurred | Yes | Link the post-incident review and reassess residual risk. |
| The team wants to expand sources, users, channels, or actions | Yes | Make expansion a separate go/no-go decision. |
| The route is low-risk and internal-only | Optional | Use a lighter scorecard with documented rationale. |
| The route is being retired | Yes | Use the scorecard to verify closure, access removal, and customer continuity. |
Do not use this page as legal, regulatory, or product safety advice. Adapt the evidence and approval path to the business, customers, data, and requirements in scope.
Quarterly scorecard
Score each area using evidence from the quarter. Do not use a green status when the team simply did not collect the evidence.
| Area | Evidence reviewed | Status | Owner | Decision or gap |
|---|---|---|---|---|
| Purpose and scope | Current use cases, route, audience, and prohibited uses. | |||
| Ownership and accountability | Business, technical, support, security, and privacy owners. | |||
| Customer impact | Complaints, escalations, corrections, handoffs, and affected decisions. | |||
| Access and identity | User, admin, service, vendor, guest, and break-glass review. | |||
| Data and retention | Data classes, sources, exports, retention, deletion, and requests. | |||
| Model and prompt | Version history, fixed tests, refusals, and change approvals. | |||
| Knowledge and sources | Source ownership, freshness, access boundary, and retrieval findings. | |||
| Connectors and tools | Scopes, approvals, actions, downstream authorization, and retries. | |||
| Monitoring and logs | Coverage, samples, alerts, access to evidence, and review cadence. | |||
| Incidents and exceptions | Incident records, near misses, open exceptions, and overdue actions. | |||
| Vendor and subprocessor | Security changes, support access, terms, and deletion obligations. | |||
| Continuity and recovery | Pause, fallback, handoff, restore, and customer communication. | |||
| Final risk decision | Residual risk, conditions, approver, and next review. |
Use green, yellow, red, or unknown only when the team has defined what each status means. Treat unknown as a control gap, not as green.
Review inputs
Collect a bounded evidence packet before the meeting.
| Input | What to bring | Do not bring |
|---|---|---|
| Usage and outcome summary | Counts, trends, refusals, handoffs, corrections, and error rates. | Unredacted customer transcripts. |
| Access review | Recertification result, removals, exceptions, and negative tests. | Full employee directory export. |
| Audit logs | Event coverage, action results, alerts, and review findings. | Tokens, secrets, or raw private payloads. |
| Change history | Model, prompt, policy, source, connector, tool, and channel changes. | Unsupported claims about vendor internals. |
| Incident and near-miss review | Severity, scope, containment, correction, and follow-up. | Private incident details in the public repository. |
| Exception register | Active, expired, renewed, closed, and overdue exceptions. | Permanent exceptions without a decision. |
| Vendor review | Security notices, support access, contract terms, and changes. | Confidential contract contents unless protected. |
| Customer feedback | Themes, escalations, notice and handoff issues. | Identifying customer information. |
| Recovery evidence | Pause, fallback, restore, and communication test results. | Credentials or operational secrets. |
- Set the evidence window and review owner.
- Record missing evidence before discussing the final status.
- Use counts, redacted references, hashes, or protected record IDs where possible.
- Separate observed evidence from management interpretation.
- Link high-impact findings to the protected incident, change, or exception record.
Governance and ownership review
| Question | Evidence | Pass condition |
|---|---|---|
| Who owns the business outcome? | Named owner and current responsibility. | The owner can approve continue, limit, or stop decisions. |
| Who owns technical operation? | On-call, deployment, and vendor contacts. | Someone can change, pause, and recover the route. |
| Who reviews customer impact? | Support, product, privacy, or safety review. | Customer harm and handoff issues have a clear route. |
| Who can approve high-impact actions? | Action owner and independent approver. | The model is not the only authorization layer. |
| Who can stop the route? | Pause owner, runbook, and escalation path. | The team can stop safely without waiting for a vendor. |
| Who maintains evidence? | Evidence owner, location, retention rule. | Records are current, protected, and findable. |
| Who approves residual risk? | Named role and decision record. | Risk acceptance is accountable and time-bound. |
- Review whether owners still work on the relevant workflow.
- Check that decision rights are documented rather than assumed.
- Confirm the support and privacy paths are reachable outside engineering.
- Check for shared accounts, stale groups, orphaned service identities, and unowned vendor access.
- Record a deputy for critical operations and recovery.
NIST AI RMF Core calls for periodic review, clear roles and responsibilities, inventory, documentation, and accountability structures. Use those outcomes to test whether the chatbot is governed in practice, not only whether a policy exists.
Use case and customer impact review
| Review area | Questions | Evidence |
|---|---|---|
| Intended use | Did actual requests match the approved use case? | Use-case summary and sample categories. |
| Prohibited use | Were sensitive or out-of-scope requests received? | Refusal, escalation, and abuse trends. |
| Accuracy and correction | Which answers required correction or customer follow-up? | Redacted correction records. |
| Human handoff | Did users reach the right human path within the target time? | Handoff samples and support metrics. |
| Sensitive topics | Were billing, privacy, security, legal, health, or safety topics routed correctly? | Escalation evidence and review results. |
| Notice and choice | Was the AI notice understandable and available in the relevant channel? | Notice version and customer feedback. |
| Accessibility and language | Did fallback and notice remain usable for supported audiences? | Sample results and support observations. |
| Customer harm | Did an answer or action create actual or possible harm? | Incident or customer-impact record. |
- Sample normal answers, refusals, sensitive topics, corrections, handoffs, and abuse attempts.
- Compare customer-impact trends with the previous quarter.
- Separate model error, source error, workflow error, access error, and human-review error.
- Escalate repeated high-impact corrections even when the aggregate error rate looks small.
- Record whether the customer can recover when the chatbot is wrong.
Access and data boundary review
Use the latest access recertification evidence and verify that the current route still matches the approved data boundary.
| Control | Review question | Red flag |
|---|---|---|
| Human users | Does each user still have a legitimate business need? | Stale, shared, or unexplained access. |
| Admins | Are admin roles named, MFA-protected, and separated from routine use? | Shared admin or unmonitored exports. |
| Service accounts | Does each workload have an owner and minimum downstream role? | Generic high-privilege identity. |
| Sources | Does retrieval stay within the approved tenant and fields? | Cross-tenant or stale private data. |
| Connectors | Are grants, scopes, consent, and expiry current? | Read-only workflow with write or delete scope. |
| Tool actions | Are high-impact actions approved downstream? | The prompt is the only control. |
| Retention | Are transcripts, traces, exports, and backups covered? | Untracked copies or expired retention. |
| Offboarding | Are removed identities and grants verified by negative test? | Access removal is assumed, not tested. |
| Exceptions | Are all deviations owned, approved, and time-bound? | Exception renewal without new evidence. |
- Check direct users, groups, guests, vendors, service accounts, connectors, tools, and break-glass access.
- Review effective permissions, not only role labels.
- Confirm MFA, SSO, session, export, and audit controls where applicable.
- Verify deletion, retention, and customer request workflows.
- Link open gaps to the exception register or remediation tracker.
Source, model, and prompt review
| Change or signal | Review question | Decision evidence |
|---|---|---|
| Model or provider | Did output, refusal, latency, data handling, or cost change? | Version comparison and fixed test set. |
| System prompt or policy | Did scope, notice, handoff, or action behavior change? | Approved diff and regression result. |
| Retrieval source | Is content current, approved, and permission-aware? | Source inventory and freshness review. |
| Source conflict | How did the route resolve conflicting or stale material? | Sample and correction result. |
| Customer channel | Did abuse exposure, notice, or fallback change? | Channel review and test result. |
| Guardrail | Did a control prevent or merely hide the behavior? | Pass/fail sample and failure path. |
| Evaluation set | Does the set represent real high-impact routes? | Test coverage and owner. |
- Link every material change to an approver and deployment record.
- Re-run fixed normal, refusal, sensitive, handoff, and abuse tests.
- Include multilingual and accessibility cases when the route supports them.
- Record unresolved source, model, or prompt uncertainty as a finding.
- Do not infer vendor data handling or training behavior without current official evidence.
Connector and tool action review
| Action class | Quarterly question | Stop condition |
|---|---|---|
| Read | Is the source, identity, field scope, and tenant boundary necessary? | The action can read unrelated customer data. |
| Create or update | Is the requested change authorized and reconciled? | No downstream policy or result confirmation. |
| Delete | Is deletion required, confirmed, and reversible? | Delete is bundled into a broad extension. |
| Export or send | Who approves destination and content? | Raw data can leave without review. |
| Billing or account | What customer state can change? | No rollback or independent approval. |
| Webhook or queue | How are retries and duplicates handled? | Unknown results can repeat a side effect. |
| Admin setting | Who can change and review the capability? | One person can change and self-approve. |
OWASP LLM06:2025 identifies excessive functionality, permissions, and autonomy as causes of excessive agency. It recommends minimum extensions and permissions, downstream authorization, user-context checks, approval for high-impact actions, logging, and rate limiting.
- Review successful, denied, retried, timed-out, and unknown action results.
- Compare action arguments with the approved data classification.
- Confirm downstream authorization is independent from the model output.
- Run a safe negative test for removed scope, denied action, and expired grant.
- Disable or limit a capability when the result cannot be reconstructed.
Incident, exception, and change review
| Evidence type | Questions | Required decision |
|---|---|---|
| Incident | What happened, who was affected, and what changed afterward? | Confirm closure, residual risk, and retest. |
| Near miss | Which control prevented impact and is it reliable? | Keep, strengthen, or redesign the control. |
| Exception | Why was the deviation needed and is it still active? | Close, reduce, renew, or convert to baseline change. |
| Change | Did the release change data, access, output, action, or notice? | Confirm approval, test result, and rollback. |
| Vendor event | Did terms, subprocessors, support access, or security change? | Reassess the route or pause vendor access. |
| Customer correction | Is the same failure recurring? | Add a source, workflow, evaluation, or training fix. |
- Review all high-severity incidents and open findings.
- Review repeated low-severity findings for a systemic pattern.
- Review overdue exceptions and remediation actions before approving continuation.
- Confirm every material change had a test set, owner, and rollback or pause path.
- Record lessons learned in the next quarter’s control plan.
Vendor and subprocessor review
| Review item | Question | Evidence |
|---|---|---|
| Service scope | What data and actions does the vendor support? | Contract summary and route inventory. |
| Access | Which named vendor identities can access the workflow? | Access review and support windows. |
| Security change | Did the vendor report a relevant change or event? | Notice, response, and owner assessment. |
| Subprocessors | Did the data path or service provider list change? | Current vendor documentation and review. |
| Retention | How long are prompts, transcripts, traces, or exports kept? | Current terms or approved vendor answer. |
| Deletion | How can the team delete or export required records? | Request result or protected evidence. |
| Incident response | How does the vendor notify and coordinate? | Support path and tested contact. |
| Exit | Can the team revoke access and recover operations? | Exit, fallback, or replacement plan. |
- Confirm vendor access is need-to-know and time-bound.
- Verify vendor security controls instead of relying only on brand familiarity.
- Record current terms, security notices, and unresolved questions.
- Link vendor changes to the chatbot risk decision.
- Keep customer-safe continuity and exit options available.
The FTC’s small-business guidance recommends written vendor security expectations, verification, need-to-know access, MFA, and changes when threats or vendor practices change. Use the vendor review to test the actual chatbot data path, not just the procurement record.
Control maturity decision table
| Result | Meaning | Action |
|---|---|---|
| Continue | Intended use, owners, evidence, access, monitoring, and recovery are working. | Keep cadence and track improvements. |
| Continue with conditions | Known gaps are bounded and owned. | Set conditions, expiry, and next review. |
| Limit | Data, users, sources, channels, or actions exceed evidence or tolerance. | Reduce scope and retest before expansion. |
| Pause | A high-impact control is missing, failed, or unverified. | Stop the affected route and use a safe fallback. |
| Redesign | Repeated failures show the workflow or control model is wrong. | Change architecture, authorization, data path, or human review. |
| Replace | Vendor, model, or route cannot meet the required boundary. | Run a controlled replacement and migration plan. |
| Retire | The value no longer justifies risk, cost, or operating burden. | Remove access, data, connectors, tools, and customer dependencies. |
Do not use a green status to avoid a difficult business decision. A clear limit, pause, redesign, replacement, or retirement decision is better evidence of governance than an unconditional approval with hidden gaps.
Quarterly review agenda
Use a 45-minute agenda for a small team.
| Minutes | Activity | Output |
|---|---|---|
| 0-5 | Confirm route, owners, quarter, decision rights, and missing evidence. | Scope and open questions. |
| 5-12 | Review customer impact, handoffs, corrections, and incidents. | Impact findings. |
| 12-20 | Review access, data, sources, connectors, and tool actions. | Boundary findings. |
| 20-28 | Review model, prompt, policy, vendor, and change history. | Change findings. |
| 28-35 | Review exceptions, remediation, monitoring, and recovery tests. | Control status. |
| 35-42 | Select continue, condition, limit, pause, redesign, replace, or retire. | Decision and conditions. |
| 42-45 | Assign owners, dates, evidence, and next review. | Action tracker. |
If the evidence shows an active high-impact issue, stop the meeting and follow the incident or pause path instead of finishing the agenda for appearance.
Decision record
| Field | Entry |
|---|---|
| Review quarter | |
| Chatbot route and environment | |
| Business and technical owners | |
| Evidence window | |
| Material changes | |
| Customer impact summary | |
| Access and data result | |
| Connector and tool result | |
| Incident and exception result | |
| Vendor result | |
| Final decision | Continue, condition, limit, pause, redesign, replace, or retire. |
| Conditions and expiry | |
| Residual risk summary | |
| Approver | |
| Next review date |
NIST AI RMF recommends documenting risk responses and residual risk while assigning clear roles and periodic review. The decision record should state what is known, what is uncertain, what will change, and who is accountable for the remaining risk.
Action tracker
| Action | Source finding | Severity | Owner | Due date | Closure evidence | Status |
|---|---|---|---|---|---|---|
- Separate urgent containment from long-term improvement.
- Give every action a closure test, not only a due date.
- Re-run the AI Tool Risk Checker when the data, access, connector, or action scope changes.
- Escalate overdue high-risk actions before the next review.
- Reopen actions when the evidence does not prove the control works.
Final review checklist
- The review period, route, environment, owners, and decision rights are recorded.
- Purpose, intended use, prohibited use, customer impact, notice, and human handoff were reviewed.
- Users, admins, vendors, guests, service accounts, connectors, tools, and break-glass access were reviewed.
- Data classes, sources, retention, deletion, exports, and customer requests were checked.
- Model, prompt, policy, source, channel, connector, tool, and vendor changes were correlated.
- Normal, refusal, sensitive, correction, handoff, multilingual, and abuse samples were reviewed where relevant.
- High-impact actions have downstream authorization, approval, logging, and safe negative tests.
- Incidents, near misses, exceptions, vendor events, and overdue remediation were reviewed.
- Pause, fallback, recovery, and customer communication paths were tested or evidence-backed.
- The final decision includes conditions, residual risk, approver, and next review date.
- Every follow-up has an owner, due date, closure test, and evidence location.
- The next quarter’s review inputs and success criteria are already defined.
Metrics to track
- Percentage of quarterly review areas with current evidence rather than unknown status.
- Percentage of customer-impacting requests that reached the correct human or policy path.
- Percentage of high-impact actions with independent authorization and complete logs.
- Number of access, data, source, connector, tool, vendor, incident, and exception findings.
- Number of material changes linked to tests, approvals, and rollback or pause paths.
- Time from a red signal to containment and from finding to verified closure.
- Number of active, expired, renewed, and repeated exceptions.
- Negative-test pass rate for revoked access, denied actions, reduced scopes, and recovery paths.
- Percentage of overdue actions at the next quarterly review.
- Final decisions by route: continue, condition, limit, pause, redesign, replace, or retire.
Evidence checked
- NIST AI RMF Core includes periodic review, clear roles, inventory, documentation, accountability, incident response, change management, and residual-risk outcomes.
- NIST AI RMF Playbook provides adaptable Govern, Map, Measure, and Manage actions and is not intended as a one-size-fits-all checklist.
- FTC Cybersecurity for Small Business covers governance, access control, monitoring, incident response, recovery, training, and vendor security.
- FTC Safeguards Rule guidance discusses risk reassessment, access review, data inventory, MFA, logging, service providers, and change management for covered financial institutions.
- OWASP LLM06:2025 Excessive Agency covers excessive functionality, permissions, autonomy, downstream authorization, human approval, logging, and rate limiting.
- Cybergiz internal templates for chatbot weekly review, access recertification, exception management, audit logs, incidents, and tool actions were checked for overlap and linked workflow boundaries.
FAQ
Is a quarterly review enough for a customer-facing chatbot?
No. Use weekly or monthly operating reviews for errors, handoffs, actions, access, and alerts. The quarterly review adds governance: ownership, customer impact, residual risk, material change, vendor, exception, and continue or stop decisions.
What if the chatbot has very little traffic?
Low traffic does not prove low risk. Review the route’s permissions, data, actions, fallback, and evidence. A rarely used admin or deletion capability can still create high impact. Use traffic as one signal, not as the final decision.
Who should attend the quarterly review?
At minimum, include the business owner and technical owner. Add support, security, privacy, or legal owners when the route handles sensitive data, customer decisions, high-impact actions, vendors, or incidents. The people who can pause the route should know the decision and escalation path.
Should a quarterly scorecard have only green or red statuses?
No. Use statuses that expose conditions: green, yellow, red, and unknown. The decision should separately say continue, continue with conditions, limit, pause, redesign, replace, or retire.
How should we handle an open exception at the quarterly review?
Confirm the current need, scope, compensating controls, use, alerts, owner, and expiry. Then close, reduce, renew with a new decision, or convert it into a baseline change. Do not approve a renewal only because the exception already exists.
What is the most important evidence for a chatbot with tool actions?
The team should be able to link the user or service identity, request, approved scope, tool call, downstream authorization, result, retries, human approval where required, and final customer or system state. If that chain cannot be reconstructed, limit or pause the action.
Does this scorecard replace a vendor security review?
No. It helps the chatbot owner identify vendor-related changes, access, retention, support, and incident questions. Use the vendor’s current security documentation and your contractual or regulatory review process for the actual assessment.
Where should we store the completed scorecard?
Store it in a protected governance or evidence system with need-to-know access, retention, and deletion rules. Keep raw customer data and private incident material out of the public repository and out of the scorecard unless the protected system explicitly requires it.