checklist
AI chatbot post-cutover review checklist for small teams
A practical first-7-days review checklist for a customer-facing AI chatbot after a vendor or platform cutover, covering answer quality, handoff, data, sources, connectors, actions, abuse signals, rollback, and operating decisions.
Use this checklist during the first seven days after replacing a customer-facing AI chatbot vendor, model, platform, or orchestration layer.
The cutover is not proven by a successful deployment. Review real operational signals, customer handoffs, data boundaries, source retrieval, connector and action logs, abuse attempts, and rollback readiness before expanding the new route. Run the AI Tool Risk Checker and attach the result to the post-cutover record.
Bottom line
Keep the new chatbot at its approved scope until the first-week review shows that answer quality, sensitive-topic routing, customer data handling, source access, tool actions, logging, support load, and fallback behavior are acceptable. If one area fails, narrow that topic, connector, channel, or action first. Pause or roll back when the failure could expose data, create customer-impacting actions, hide a high-impact event, or leave customers without a safe route.
Use the Small Team AI Security Checklist for baseline ownership, access review, and incident routing. Pair this page with the AI chatbot replacement cutover checklist for pre-launch gates and the AI chatbot weekly review scorecard after the first week.
When to use this checklist
| Situation | Use this checklist? | Review window |
|---|---|---|
| New chatbot vendor went live for customers | Yes | Daily for the first seven days. |
| Model or orchestration layer changed | Yes | Daily for high-risk routes; at least several representative review sessions. |
| A new connector or tool action was enabled | Yes | Review every attempt and failure during the first week. |
| A new channel or customer segment was added | Yes | Review by channel and segment, not only in aggregate. |
| The target ran in shadow mode only | Not yet | Use the cutover checklist before customer traffic. |
| A cosmetic disclosure copy change was made | Usually no | Use normal change review unless behavior changed. |
Do not expand scope just because the first day was quiet. Low volume can hide missing evidence.
First-7-days review matrix
| Review area | Daily question | Escalate when |
|---|---|---|
| Answer quality | Did the bot answer approved questions correctly and with current sources? | Wrong or unsupported answers repeat or affect customers. |
| No-answer behavior | Did it refuse or route questions outside scope? | It guesses, loops, or hides the human route. |
| Sensitive topics | Were privacy, security, billing, access, legal, HR, health, and safety topics handled safely? | A sensitive topic receives unsafe advice or misses handoff. |
| Customer data | Did the bot use only approved fields and contexts? | Unapproved data appears in input, retrieval, output, logs, or exports. |
| Sources | Are retrieval, freshness, conflicts, and access behaving as expected? | Private, stale, poisoned, or conflicting sources affect answers. |
| Connectors | Are read and write scopes limited and owned? | Unknown access, scope drift, or unapproved identity appears. |
| Tool actions | Were actions approved, logged, idempotent, and reversible? | An action is unauthorized, duplicated, irreversible, or unexplained. |
| Abuse and security | Did prompt injection or abuse change behavior? | The bot reveals data, changes policy, or triggers unsafe action. |
| Operations | Can owners see errors, latency, queue load, and vendor events? | High-impact events cannot be reconstructed. |
| Customer impact | Can customers reach a person and get timely help? | Wait time, disputes, or fallback volume exceed the limit. |
The matrix is a review aid, not a substitute for the team’s risk tolerance or incident process.
Review intake form
Copy this into the post-cutover record.
| Field | Entry |
|---|---|
| Review date | |
| Cutover date | |
| Review owner | |
| Old and new system | |
| Approved scope | |
| Channels and customer groups | |
| Data classes approved | Public, internal, customer, account, billing, regulated, or credential-like. |
| Connectors and tool actions enabled | |
| First-week review period | |
| Rollback trigger | |
| Support fallback owner | |
| Security/privacy owner | |
| Customer impact owner | |
| Evidence location | |
| Decision deadline |
If the new scope, owners, or rollback trigger are unclear, keep the bot in limited mode while the record is completed.
Baseline and comparison table
Compare the target with the accepted pre-cutover baseline. A new tool may change user behavior, routing, latency, or action volume even when the headline answer score looks similar.
| Signal | Old baseline | New result | Acceptable range or decision |
|---|---|---|---|
| Approved-topic answer pass rate | |||
| No-answer/refusal pass rate | |||
| Sensitive-topic handoff rate | |||
| Customer dispute rate | |||
| Source retrieval success | |||
| Source conflict rate | |||
| Unapproved data-field observations | |||
| Tool action attempts | |||
| Tool action failures or duplicates | |||
| Prompt injection and abuse signals | |||
| Median and tail latency | |||
| Human fallback wait time | |||
| Rollback readiness test result |
Record the sample size and exclusions. Do not compare percentages without enough context to understand volume and severity.
Answer quality sample
Sample conversations by risk and not only by popularity.
- Include representative normal questions from each approved topic.
- Include questions where the expected answer is “I do not know” or a human route.
- Include recent customer corrections, disputes, and escalations.
- Include questions that depend on current policy, pricing, account state, or source freshness.
- Include multilingual, accessibility, and channel-specific cases where relevant.
- Record source trace, reviewer decision, issue severity, and owner.
- Separate model error, source error, routing error, and policy/configuration error.
- Retest a fixed set every day so the trend is comparable.
| Finding | Severity | Immediate boundary | Owner | Retest date |
|---|---|---|---|---|
Do not put real customer secrets, credentials, or unredacted transcripts into a public issue or repository. Store protected evidence elsewhere and link only to an approved location.
Handoff and customer impact review
| Check | Pass condition | Result |
|---|---|---|
| Human request | Customer can ask for a person without a loop. | |
| Sensitive topic | Privacy, security, billing, access, and other high-impact topics route correctly. | |
| Context transfer | Human receives enough context without unnecessary sensitive data. | |
| Queue ownership | A named team owns the request and response target. | |
| Customer notice | Disclosure and fallback wording match actual behavior. | |
| Dispute handling | Customer can challenge an answer and get correction. | |
| Open conversations | Requests started before cutover are not dropped. | |
| Support load | Fallback volume is within staffing and response limits. |
- Review every high-severity handoff from the previous day.
- Review customer complaints and correction requests separately from normal traffic.
- Check whether the new bot creates support work that the old bot did not.
- Confirm customer-facing routes do not expose internal error details or vendor identifiers unnecessarily.
Data and privacy review
Inspect actual fields and contexts sent to the new system, not only the intended schema.
| Review item | Evidence | Decision |
|---|---|---|
| Input fields | Request schema sample and field inventory. | |
| Retrieved context | Source trace and access scope. | |
| Output data | Sample outputs and sensitive-data review. | |
| Conversation logs | Retention, access, export, and deletion settings. | |
| Analytics events | Event names, identifiers, and retention. | |
| Temporary migration files | Location, owner, access, and deletion date. | |
| Vendor support access | Role, approval, logging, and current status. | |
| Customer requests | Deletion, export, correction, and opt-out routing. |
- Confirm only approved data classes enter the new route.
- Check for accidental raw exports, credentials, or unrelated customer records.
- Verify retention and deletion settings against the approved policy.
- Reconcile the new system’s data with the migration boundary.
- Escalate any unapproved data observation immediately.
The FTC’s small-business guidance recommends specifying vendor data use, sharing, retention, deletion, verification, and need-to-know access. Treat an unexpected field as a control finding even if no incident is confirmed.
Source and retrieval review
- Review source freshness failures and retrieval misses.
- Review answers that cite a source outside the approved scope.
- Test conflicting documents and stale policy pages.
- Check whether the target indexes private sources with broader access than the old route.
- Verify source owners and update dates.
- Search for prompt or retrieval changes made during the cutover.
- Retest any source that produced a customer-impacting answer.
| Source finding | Impacted topic | Containment | Owner | Retest |
|---|---|---|---|---|
Do not fix a source silently during the first-week review. Record the change, test it, and preserve enough before/after evidence to explain the result.
Connector and action review
Review every attempt, not only successful actions.
| Event | Questions to answer |
|---|---|
| Connector read | Was the data source and scope approved for this route? |
| Connector write | Who or what approved the change before it happened? |
| Action denial | Did the denial fail safely and reach a human when needed? |
| Retry | Could the retry duplicate a customer-impacting action? |
| Timeout | Did the user receive a safe status without an unknown side effect? |
| Rollback | Can the action be corrected or escalated with evidence? |
| Identity change | Did the new service identity use only the expected permissions? |
| Queue or webhook | Did any old route continue to create events? |
- Review connector permission changes since cutover.
- Review all high-impact action arguments and outcomes.
- Test one safe denial from each relevant action route.
- Verify idempotency and duplicate handling.
- Confirm old credentials, queues, webhooks, and service identities are no longer active where intended.
- Disable an action or connector if logs cannot explain its behavior.
OWASP’s current LLM risk materials cover prompt injection, sensitive information disclosure, supply-chain risk, and excessive agency. Treat a connector or action boundary as a security control, not just an integration setting.
Security and abuse review
| Test or signal | Daily review |
|---|---|
| Direct prompt injection | Attempts to override policy or reveal protected instructions. |
| Indirect prompt injection | Untrusted source content that changes behavior. |
| Sensitive information disclosure | Secrets, customer data, internal instructions, or cross-customer context. |
| Excessive agency | Unexpected tool action, broad permissions, or missing approval. |
| Source poisoning | Untrusted or changed content influencing answers. |
| Abuse and spam | High-volume, oversized, manipulative, or automated requests. |
| Model or vendor event | Model change, outage, safety notice, or support-access event. |
- Preserve representative attack inputs and outputs in protected evidence storage.
- Confirm the incident owner knows how to pause the affected route.
- Confirm a prompt or source fix cannot silently widen permissions.
- Review vendor notices and changes during the observation window.
- Escalate any evidence of data exposure or unsafe action as an incident, not just a quality ticket.
Operational monitoring
NIST’s 2026 report on monitoring deployed AI systems describes post-deployment monitoring as important because deployed AI can behave variably and produce unforeseen outcomes. Use a small set of owner-assigned signals rather than a dashboard nobody reviews.
| Signal | Owner | Threshold or trigger | Response |
|---|---|---|---|
| Answer quality failures | |||
| Sensitive-topic handoff failures | |||
| Customer disputes | |||
| Unapproved data observation | |||
| Tool action failure or duplicate | |||
| Prompt injection or abuse | |||
| Source freshness or conflict | |||
| Vendor outage or change | |||
| Fallback wait time |
- Assign a person for each signal.
- Define what is sampled manually and what is alert-driven.
- Set a review time and record decisions daily.
- Keep monitoring for the full approved observation window even if the first days are quiet.
Decision rules
| Finding | Decision |
|---|---|
| All gates pass and no material trend is worsening | Keep the approved scope and move to weekly review. |
| One low-impact topic or source fails | Limit that boundary, fix it, and retest. |
| Sensitive-topic handoff fails | Force human handoff and pause the affected topic. |
| Unapproved customer data appears | Pause the affected route and escalate to security/privacy. |
| High-impact action is unauthorized or not reconstructable | Disable the action and reconcile downstream state. |
| Support fallback is overloaded | Reduce automation or customer traffic until staffing is ready. |
| Multiple high-severity findings recur | Roll back, pause, or return to human-only operation. |
| Evidence is missing but no failure is observed | Keep limited and close the evidence gap before expansion. |
Do not close a finding because the bot’s overall satisfaction score is good. Severity and customer impact matter more than aggregate averages.
Post-cutover review record
| Field | Entry |
|---|---|
| Review period | |
| Traffic and sample size | |
| Approved scope reviewed | |
| Answer quality decision | |
| Handoff decision | |
| Data/privacy decision | |
| Source/retrieval decision | |
| Connector/action decision | |
| Security/abuse decision | |
| Customer impact decision | |
| Rollback readiness result | |
| Open findings | |
| Final outcome | Keep, limit, pause, rollback, or human-only. |
| Owners and due dates | |
| Approvers | |
| Next review date |
Store the record with the cutover tests, vendor exit packet, access review, and any incident evidence. Do not store customer exports, API keys, tokens, or private credentials in the public repository.
Next-week action tracker
| Action | Finding | Owner | Due date | Evidence of closure |
|---|---|---|---|---|
- Every high-severity finding has a containment decision.
- Every accepted exception has an expiry date and compensating control.
- Every source or prompt fix has a retest.
- Every connector or action change has an owner and verification.
- Every customer-impacting correction has a support follow-up.
Final first-week checklist
- The reviewed traffic represents each approved channel and customer group.
- Normal, no-answer, sensitive, abuse, and prompt-injection cases were sampled.
- Customer data fields and retrieved context stayed within the approved boundary.
- Sources were current, access-scoped, and tested for conflict.
- Connector permissions and tool actions were reviewed from logs.
- High-impact events can be reconstructed from available evidence.
- Human handoff works and support load is acceptable.
- Rollback or pause was tested without causing an unapproved write.
- Old routes, credentials, queues, and webhooks are disabled where intended.
- Customer notices and support scripts match actual behavior.
- Open findings have owners, due dates, and containment.
- The decision to keep, limit, pause, rollback, or go human-only is recorded.
Metrics to track
| Metric | Why it matters |
|---|---|
| Answer pass rate by topic | Finds weak areas hidden by aggregate averages. |
| No-answer and handoff success | Shows whether the bot knows its boundaries. |
| Customer disputes and corrections | Shows customer-impacting quality failures. |
| Unapproved data observations | Tests the input and retrieval boundary. |
| Connector/action attempts and failures | Shows automation exposure and control health. |
| Prompt injection and abuse signals | Shows adversarial pressure after launch. |
| Source freshness and conflict findings | Shows knowledge-base stability. |
| Fallback volume and wait time | Shows customer continuity and support load. |
| Rollback or pause events | Shows operational stability. |
| Findings closed by due date | Shows whether review produces control improvements. |
Do not optimize for deflection alone. A lower human handoff rate can be a warning if customers are getting wrong answers instead.
Evidence checked
This checklist is aligned with:
- NIST AI RMF Core, which includes monitoring, accountability, third-party risk, decommissioning, incident response, recovery, and change management outcomes.
- NIST AI RMF Manage Playbook, which describes monitoring third-party AI systems, applying documented controls, contingency processes, and decommissioning systems that exceed risk tolerances.
- NIST report on challenges to monitoring deployed AI systems, which describes post-deployment monitoring as important for real-world reliability and unforeseen outcomes.
- NIST AI 800-4 monitoring report, which organizes deployed AI monitoring questions and categories for operational review.
- FTC cybersecurity guidance for small businesses, which recommends vendor data-use, retention, deletion, verification, and least-necessary-access controls.
- CISA and UK NCSC secure AI system development guidance, which covers secure practices across AI design, development, deployment, and operation.
- OWASP Top 10 for LLM Applications, which covers current risks including prompt injection, sensitive information disclosure, supply chain, data poisoning, and excessive agency.
- Cybergiz templates for chatbot cutover, vendor exit, retirement, red-team testing, tool action approval, human handoff, and weekly review.
This page is practical operating guidance, not legal, privacy, compliance, audit, certification, customer-support, incident-response, procurement, or security assurance advice.
FAQ
Is a successful deployment enough to end the review?
No. Deployment proves that the software was released. The first-week review tests how the system behaves with real traffic, customer data, source changes, handoffs, actions, abuse, and support load.
How many conversations should we sample?
Use enough traffic to cover every approved topic, channel, and risk category. For low-volume bots, extend the window rather than treating a small sample as proof of safety. Always include known no-answer and sensitive-topic cases.
Should we roll back after one wrong answer?
Not automatically. Triage the severity, customer impact, data involved, frequency, and whether the issue is isolated. Roll back or pause immediately when the answer exposes data, triggers an unsafe action, or leaves a customer without a safe route.
What if the target has better answers but worse handoff?
Keep the scope limited or return to the old route/human-only fallback until handoff is fixed. Answer quality does not compensate for a failed escape path on sensitive or disputed topics.
How do we review customer data without exposing it further?
Use redacted or synthetic examples where possible, restrict access to protected evidence, record data classes and field names instead of values, and never put raw customer exports or secrets in a public repository.
Should the old vendor remain available during the review?
Only if it can be isolated and the team can prevent duplicate writes, stale routes, and unclear ownership. Keep the rollback path tested and follow the old vendor’s exit plan once the new route is accepted.
What is a good outcome after seven days?
A documented decision to keep the approved scope, limit a defined boundary, pause or roll back a failed area, or move to human-only operation. The record should show evidence, owners, open findings, and the next review date.
When should we move from daily to weekly review?
After the agreed first-week observation window passes, high-risk findings are contained, customer handoff and actions are stable, and the owners approve the move. Keep event-driven reviews for incidents, vendor changes, source changes, connector changes, and material customer impact.