checklist

AI chatbot quarterly governance review scorecard for small teams

A quarterly governance scorecard for customer-facing AI chatbots, covering ownership, customer impact, access, data, model and prompt changes, connectors, tool actions, incidents, exceptions, vendors, decisions, and follow-up.

Audience: Founders, support leads, product owners, engineering owners, security owners, privacy owners, and managers governing customer-facing AI chatbots Risk: High Evidence: NIST AI RMF Core and Manage Playbook, FTC Safeguards Rule guidance, FTC small-business cybersecurity guidance, OWASP LLM06:2025 Excessive Agency, and Cybergiz chatbot operations templates

Use this scorecard once each quarter to decide whether a customer-facing AI chatbot should continue, be limited, be changed, or be paused.

The quarterly review is not a larger weekly status meeting. It is a governance decision based on evidence from the last period: ownership, customer impact, access, data handling, model and prompt changes, connectors, tool actions, incidents, exceptions, vendor changes, and unresolved control gaps. Run the AI Tool Risk Checker before the meeting when the route, data, or action scope has changed.

Bottom line

Keep a customer-facing chatbot in service only when the team can explain who owns it, what it is allowed to do, which data it can reach, how users can reach a human, what changed, what went wrong, and what residual risk remains. A good quarterly result may be continue, limit, redesign, replace, or retire. It does not have to be an unconditional approval.

Use the Small Team AI Security Checklist for baseline owners, access, incident routing, and evidence storage. Pair this scorecard with the AI chatbot weekly review scorecard for recurring operating signals, the AI chatbot access recertification checklist for permissions, and the AI chatbot exception register template for time-bound deviations.

When to use this scorecard

SituationUse the quarterly scorecard?Additional action
Customer-facing chatbot has operated for a full quarterYesReview the full evidence set and make a continuation decision.
Chatbot launched or materially changed during the quarterYesCompare the approved baseline with the current route.
Route has high-impact account, billing, order, or outbound actionsYesRequire action-owner and security or privacy review.
A major incident or near miss occurredYesLink the post-incident review and reassess residual risk.
The team wants to expand sources, users, channels, or actionsYesMake expansion a separate go/no-go decision.
The route is low-risk and internal-onlyOptionalUse a lighter scorecard with documented rationale.
The route is being retiredYesUse the scorecard to verify closure, access removal, and customer continuity.

Do not use this page as legal, regulatory, or product safety advice. Adapt the evidence and approval path to the business, customers, data, and requirements in scope.

Quarterly scorecard

Score each area using evidence from the quarter. Do not use a green status when the team simply did not collect the evidence.

AreaEvidence reviewedStatusOwnerDecision or gap
Purpose and scopeCurrent use cases, route, audience, and prohibited uses.
Ownership and accountabilityBusiness, technical, support, security, and privacy owners.
Customer impactComplaints, escalations, corrections, handoffs, and affected decisions.
Access and identityUser, admin, service, vendor, guest, and break-glass review.
Data and retentionData classes, sources, exports, retention, deletion, and requests.
Model and promptVersion history, fixed tests, refusals, and change approvals.
Knowledge and sourcesSource ownership, freshness, access boundary, and retrieval findings.
Connectors and toolsScopes, approvals, actions, downstream authorization, and retries.
Monitoring and logsCoverage, samples, alerts, access to evidence, and review cadence.
Incidents and exceptionsIncident records, near misses, open exceptions, and overdue actions.
Vendor and subprocessorSecurity changes, support access, terms, and deletion obligations.
Continuity and recoveryPause, fallback, handoff, restore, and customer communication.
Final risk decisionResidual risk, conditions, approver, and next review.

Use green, yellow, red, or unknown only when the team has defined what each status means. Treat unknown as a control gap, not as green.

Review inputs

Collect a bounded evidence packet before the meeting.

InputWhat to bringDo not bring
Usage and outcome summaryCounts, trends, refusals, handoffs, corrections, and error rates.Unredacted customer transcripts.
Access reviewRecertification result, removals, exceptions, and negative tests.Full employee directory export.
Audit logsEvent coverage, action results, alerts, and review findings.Tokens, secrets, or raw private payloads.
Change historyModel, prompt, policy, source, connector, tool, and channel changes.Unsupported claims about vendor internals.
Incident and near-miss reviewSeverity, scope, containment, correction, and follow-up.Private incident details in the public repository.
Exception registerActive, expired, renewed, closed, and overdue exceptions.Permanent exceptions without a decision.
Vendor reviewSecurity notices, support access, contract terms, and changes.Confidential contract contents unless protected.
Customer feedbackThemes, escalations, notice and handoff issues.Identifying customer information.
Recovery evidencePause, fallback, restore, and communication test results.Credentials or operational secrets.
  • Set the evidence window and review owner.
  • Record missing evidence before discussing the final status.
  • Use counts, redacted references, hashes, or protected record IDs where possible.
  • Separate observed evidence from management interpretation.
  • Link high-impact findings to the protected incident, change, or exception record.

Governance and ownership review

QuestionEvidencePass condition
Who owns the business outcome?Named owner and current responsibility.The owner can approve continue, limit, or stop decisions.
Who owns technical operation?On-call, deployment, and vendor contacts.Someone can change, pause, and recover the route.
Who reviews customer impact?Support, product, privacy, or safety review.Customer harm and handoff issues have a clear route.
Who can approve high-impact actions?Action owner and independent approver.The model is not the only authorization layer.
Who can stop the route?Pause owner, runbook, and escalation path.The team can stop safely without waiting for a vendor.
Who maintains evidence?Evidence owner, location, retention rule.Records are current, protected, and findable.
Who approves residual risk?Named role and decision record.Risk acceptance is accountable and time-bound.
  • Review whether owners still work on the relevant workflow.
  • Check that decision rights are documented rather than assumed.
  • Confirm the support and privacy paths are reachable outside engineering.
  • Check for shared accounts, stale groups, orphaned service identities, and unowned vendor access.
  • Record a deputy for critical operations and recovery.

NIST AI RMF Core calls for periodic review, clear roles and responsibilities, inventory, documentation, and accountability structures. Use those outcomes to test whether the chatbot is governed in practice, not only whether a policy exists.

Use case and customer impact review

Review areaQuestionsEvidence
Intended useDid actual requests match the approved use case?Use-case summary and sample categories.
Prohibited useWere sensitive or out-of-scope requests received?Refusal, escalation, and abuse trends.
Accuracy and correctionWhich answers required correction or customer follow-up?Redacted correction records.
Human handoffDid users reach the right human path within the target time?Handoff samples and support metrics.
Sensitive topicsWere billing, privacy, security, legal, health, or safety topics routed correctly?Escalation evidence and review results.
Notice and choiceWas the AI notice understandable and available in the relevant channel?Notice version and customer feedback.
Accessibility and languageDid fallback and notice remain usable for supported audiences?Sample results and support observations.
Customer harmDid an answer or action create actual or possible harm?Incident or customer-impact record.
  • Sample normal answers, refusals, sensitive topics, corrections, handoffs, and abuse attempts.
  • Compare customer-impact trends with the previous quarter.
  • Separate model error, source error, workflow error, access error, and human-review error.
  • Escalate repeated high-impact corrections even when the aggregate error rate looks small.
  • Record whether the customer can recover when the chatbot is wrong.

Access and data boundary review

Use the latest access recertification evidence and verify that the current route still matches the approved data boundary.

ControlReview questionRed flag
Human usersDoes each user still have a legitimate business need?Stale, shared, or unexplained access.
AdminsAre admin roles named, MFA-protected, and separated from routine use?Shared admin or unmonitored exports.
Service accountsDoes each workload have an owner and minimum downstream role?Generic high-privilege identity.
SourcesDoes retrieval stay within the approved tenant and fields?Cross-tenant or stale private data.
ConnectorsAre grants, scopes, consent, and expiry current?Read-only workflow with write or delete scope.
Tool actionsAre high-impact actions approved downstream?The prompt is the only control.
RetentionAre transcripts, traces, exports, and backups covered?Untracked copies or expired retention.
OffboardingAre removed identities and grants verified by negative test?Access removal is assumed, not tested.
ExceptionsAre all deviations owned, approved, and time-bound?Exception renewal without new evidence.
  • Check direct users, groups, guests, vendors, service accounts, connectors, tools, and break-glass access.
  • Review effective permissions, not only role labels.
  • Confirm MFA, SSO, session, export, and audit controls where applicable.
  • Verify deletion, retention, and customer request workflows.
  • Link open gaps to the exception register or remediation tracker.

Source, model, and prompt review

Change or signalReview questionDecision evidence
Model or providerDid output, refusal, latency, data handling, or cost change?Version comparison and fixed test set.
System prompt or policyDid scope, notice, handoff, or action behavior change?Approved diff and regression result.
Retrieval sourceIs content current, approved, and permission-aware?Source inventory and freshness review.
Source conflictHow did the route resolve conflicting or stale material?Sample and correction result.
Customer channelDid abuse exposure, notice, or fallback change?Channel review and test result.
GuardrailDid a control prevent or merely hide the behavior?Pass/fail sample and failure path.
Evaluation setDoes the set represent real high-impact routes?Test coverage and owner.
  • Link every material change to an approver and deployment record.
  • Re-run fixed normal, refusal, sensitive, handoff, and abuse tests.
  • Include multilingual and accessibility cases when the route supports them.
  • Record unresolved source, model, or prompt uncertainty as a finding.
  • Do not infer vendor data handling or training behavior without current official evidence.

Connector and tool action review

Action classQuarterly questionStop condition
ReadIs the source, identity, field scope, and tenant boundary necessary?The action can read unrelated customer data.
Create or updateIs the requested change authorized and reconciled?No downstream policy or result confirmation.
DeleteIs deletion required, confirmed, and reversible?Delete is bundled into a broad extension.
Export or sendWho approves destination and content?Raw data can leave without review.
Billing or accountWhat customer state can change?No rollback or independent approval.
Webhook or queueHow are retries and duplicates handled?Unknown results can repeat a side effect.
Admin settingWho can change and review the capability?One person can change and self-approve.

OWASP LLM06:2025 identifies excessive functionality, permissions, and autonomy as causes of excessive agency. It recommends minimum extensions and permissions, downstream authorization, user-context checks, approval for high-impact actions, logging, and rate limiting.

  • Review successful, denied, retried, timed-out, and unknown action results.
  • Compare action arguments with the approved data classification.
  • Confirm downstream authorization is independent from the model output.
  • Run a safe negative test for removed scope, denied action, and expired grant.
  • Disable or limit a capability when the result cannot be reconstructed.

Incident, exception, and change review

Evidence typeQuestionsRequired decision
IncidentWhat happened, who was affected, and what changed afterward?Confirm closure, residual risk, and retest.
Near missWhich control prevented impact and is it reliable?Keep, strengthen, or redesign the control.
ExceptionWhy was the deviation needed and is it still active?Close, reduce, renew, or convert to baseline change.
ChangeDid the release change data, access, output, action, or notice?Confirm approval, test result, and rollback.
Vendor eventDid terms, subprocessors, support access, or security change?Reassess the route or pause vendor access.
Customer correctionIs the same failure recurring?Add a source, workflow, evaluation, or training fix.
  • Review all high-severity incidents and open findings.
  • Review repeated low-severity findings for a systemic pattern.
  • Review overdue exceptions and remediation actions before approving continuation.
  • Confirm every material change had a test set, owner, and rollback or pause path.
  • Record lessons learned in the next quarter’s control plan.

Vendor and subprocessor review

Review itemQuestionEvidence
Service scopeWhat data and actions does the vendor support?Contract summary and route inventory.
AccessWhich named vendor identities can access the workflow?Access review and support windows.
Security changeDid the vendor report a relevant change or event?Notice, response, and owner assessment.
SubprocessorsDid the data path or service provider list change?Current vendor documentation and review.
RetentionHow long are prompts, transcripts, traces, or exports kept?Current terms or approved vendor answer.
DeletionHow can the team delete or export required records?Request result or protected evidence.
Incident responseHow does the vendor notify and coordinate?Support path and tested contact.
ExitCan the team revoke access and recover operations?Exit, fallback, or replacement plan.
  • Confirm vendor access is need-to-know and time-bound.
  • Verify vendor security controls instead of relying only on brand familiarity.
  • Record current terms, security notices, and unresolved questions.
  • Link vendor changes to the chatbot risk decision.
  • Keep customer-safe continuity and exit options available.

The FTC’s small-business guidance recommends written vendor security expectations, verification, need-to-know access, MFA, and changes when threats or vendor practices change. Use the vendor review to test the actual chatbot data path, not just the procurement record.

Control maturity decision table

ResultMeaningAction
ContinueIntended use, owners, evidence, access, monitoring, and recovery are working.Keep cadence and track improvements.
Continue with conditionsKnown gaps are bounded and owned.Set conditions, expiry, and next review.
LimitData, users, sources, channels, or actions exceed evidence or tolerance.Reduce scope and retest before expansion.
PauseA high-impact control is missing, failed, or unverified.Stop the affected route and use a safe fallback.
RedesignRepeated failures show the workflow or control model is wrong.Change architecture, authorization, data path, or human review.
ReplaceVendor, model, or route cannot meet the required boundary.Run a controlled replacement and migration plan.
RetireThe value no longer justifies risk, cost, or operating burden.Remove access, data, connectors, tools, and customer dependencies.

Do not use a green status to avoid a difficult business decision. A clear limit, pause, redesign, replacement, or retirement decision is better evidence of governance than an unconditional approval with hidden gaps.

Quarterly review agenda

Use a 45-minute agenda for a small team.

MinutesActivityOutput
0-5Confirm route, owners, quarter, decision rights, and missing evidence.Scope and open questions.
5-12Review customer impact, handoffs, corrections, and incidents.Impact findings.
12-20Review access, data, sources, connectors, and tool actions.Boundary findings.
20-28Review model, prompt, policy, vendor, and change history.Change findings.
28-35Review exceptions, remediation, monitoring, and recovery tests.Control status.
35-42Select continue, condition, limit, pause, redesign, replace, or retire.Decision and conditions.
42-45Assign owners, dates, evidence, and next review.Action tracker.

If the evidence shows an active high-impact issue, stop the meeting and follow the incident or pause path instead of finishing the agenda for appearance.

Decision record

FieldEntry
Review quarter
Chatbot route and environment
Business and technical owners
Evidence window
Material changes
Customer impact summary
Access and data result
Connector and tool result
Incident and exception result
Vendor result
Final decisionContinue, condition, limit, pause, redesign, replace, or retire.
Conditions and expiry
Residual risk summary
Approver
Next review date

NIST AI RMF recommends documenting risk responses and residual risk while assigning clear roles and periodic review. The decision record should state what is known, what is uncertain, what will change, and who is accountable for the remaining risk.

Action tracker

ActionSource findingSeverityOwnerDue dateClosure evidenceStatus
  • Separate urgent containment from long-term improvement.
  • Give every action a closure test, not only a due date.
  • Re-run the AI Tool Risk Checker when the data, access, connector, or action scope changes.
  • Escalate overdue high-risk actions before the next review.
  • Reopen actions when the evidence does not prove the control works.

Final review checklist

  • The review period, route, environment, owners, and decision rights are recorded.
  • Purpose, intended use, prohibited use, customer impact, notice, and human handoff were reviewed.
  • Users, admins, vendors, guests, service accounts, connectors, tools, and break-glass access were reviewed.
  • Data classes, sources, retention, deletion, exports, and customer requests were checked.
  • Model, prompt, policy, source, channel, connector, tool, and vendor changes were correlated.
  • Normal, refusal, sensitive, correction, handoff, multilingual, and abuse samples were reviewed where relevant.
  • High-impact actions have downstream authorization, approval, logging, and safe negative tests.
  • Incidents, near misses, exceptions, vendor events, and overdue remediation were reviewed.
  • Pause, fallback, recovery, and customer communication paths were tested or evidence-backed.
  • The final decision includes conditions, residual risk, approver, and next review date.
  • Every follow-up has an owner, due date, closure test, and evidence location.
  • The next quarter’s review inputs and success criteria are already defined.

Metrics to track

  • Percentage of quarterly review areas with current evidence rather than unknown status.
  • Percentage of customer-impacting requests that reached the correct human or policy path.
  • Percentage of high-impact actions with independent authorization and complete logs.
  • Number of access, data, source, connector, tool, vendor, incident, and exception findings.
  • Number of material changes linked to tests, approvals, and rollback or pause paths.
  • Time from a red signal to containment and from finding to verified closure.
  • Number of active, expired, renewed, and repeated exceptions.
  • Negative-test pass rate for revoked access, denied actions, reduced scopes, and recovery paths.
  • Percentage of overdue actions at the next quarterly review.
  • Final decisions by route: continue, condition, limit, pause, redesign, replace, or retire.

Evidence checked

  • NIST AI RMF Core includes periodic review, clear roles, inventory, documentation, accountability, incident response, change management, and residual-risk outcomes.
  • NIST AI RMF Playbook provides adaptable Govern, Map, Measure, and Manage actions and is not intended as a one-size-fits-all checklist.
  • FTC Cybersecurity for Small Business covers governance, access control, monitoring, incident response, recovery, training, and vendor security.
  • FTC Safeguards Rule guidance discusses risk reassessment, access review, data inventory, MFA, logging, service providers, and change management for covered financial institutions.
  • OWASP LLM06:2025 Excessive Agency covers excessive functionality, permissions, autonomy, downstream authorization, human approval, logging, and rate limiting.
  • Cybergiz internal templates for chatbot weekly review, access recertification, exception management, audit logs, incidents, and tool actions were checked for overlap and linked workflow boundaries.

FAQ

Is a quarterly review enough for a customer-facing chatbot?

No. Use weekly or monthly operating reviews for errors, handoffs, actions, access, and alerts. The quarterly review adds governance: ownership, customer impact, residual risk, material change, vendor, exception, and continue or stop decisions.

What if the chatbot has very little traffic?

Low traffic does not prove low risk. Review the route’s permissions, data, actions, fallback, and evidence. A rarely used admin or deletion capability can still create high impact. Use traffic as one signal, not as the final decision.

Who should attend the quarterly review?

At minimum, include the business owner and technical owner. Add support, security, privacy, or legal owners when the route handles sensitive data, customer decisions, high-impact actions, vendors, or incidents. The people who can pause the route should know the decision and escalation path.

Should a quarterly scorecard have only green or red statuses?

No. Use statuses that expose conditions: green, yellow, red, and unknown. The decision should separately say continue, continue with conditions, limit, pause, redesign, replace, or retire.

How should we handle an open exception at the quarterly review?

Confirm the current need, scope, compensating controls, use, alerts, owner, and expiry. Then close, reduce, renew with a new decision, or convert it into a baseline change. Do not approve a renewal only because the exception already exists.

What is the most important evidence for a chatbot with tool actions?

The team should be able to link the user or service identity, request, approved scope, tool call, downstream authorization, result, retries, human approval where required, and final customer or system state. If that chain cannot be reconstructed, limit or pause the action.

Does this scorecard replace a vendor security review?

No. It helps the chatbot owner identify vendor-related changes, access, retention, support, and incident questions. Use the vendor’s current security documentation and your contractual or regulatory review process for the actual assessment.

Where should we store the completed scorecard?

Store it in a protected governance or evidence system with need-to-know access, retention, and deletion rules. Keep raw customer data and private incident material out of the public repository and out of the scorecard unless the protected system explicitly requires it.