checklist
AI chatbot answer provenance and citation traceability checklist for small teams
A practical checklist for tracing chatbot answers to approved sources, recording model and retrieval versions, checking citations, and keeping review evidence private.
Use this checklist when an AI chatbot presents an answer, source link, quote, recommendation, or generated summary that someone may rely on. It is especially useful for retrieval-augmented chatbots, customer support routes, internal knowledge assistants, and workflows that pass an answer into a ticket, CRM, email, document, or tool.
The short answer: make every material answer traceable to the route, model, prompt, retrieval source, source version, citation decision, reviewer decision, and downstream action that produced it. Verify that the cited source actually supports the claim, make uncertainty visible, keep authorization outside the model, and retain only the minimum redacted evidence needed to reproduce or investigate the decision. A citation link by itself is not proof that an answer is correct, current, authorized, or safe to act on.
Start with the AI Tool Risk Checker and attach its result to the provenance review. Use the Small Team AI Security Checklist for baseline identity, data, access, logging, and incident controls. Pair this page with the AI chatbot output validation checklist for destination handling, the AI chatbot human review sampling and answer quality scorecard for representative sampling, and the AI chatbot knowledge base review checklist for source approval and freshness.
Bottom line
Do not approve a chatbot answer because it contains a URL or a footnote. A defensible provenance process can show:
- Which route, tenant, role, model, provider, prompt, policy, and retrieval configuration produced the answer.
- Which source passages were selected, when they were current, what version was used, and whether the user was authorized to see them.
- Which answer claims were supported, unsupported, stale, ambiguous, or outside the source scope.
- Whether the citation was shown to the user, kept internal, or removed because it disclosed private or misleading context.
- What validation, human review, abstention, or escalation happened before any durable or high-impact action.
- Which redacted evidence, owner, decision, and follow-up make the review reproducible without collecting a private conversation archive.
OWASP LLM09:2025 describes misinformation as false or misleading output that can appear credible and points to validation, human oversight, risk communication, and clear limitations. OWASP LLM05:2025 treats improper handling of model output as a security risk when output reaches a browser, database, tool, or other system. NIST’s AI RMF Core asks teams to document test sets and metrics, monitor deployed behavior, track risks over time, and document management decisions. This checklist turns those principles into an operational evidence chain; it is not a universal accuracy score or a guarantee that a cited answer is safe.
This is operational guidance, not legal advice, a compliance certification, or a security assurance opinion.
When to use this checklist
| Situation or change | Use this checklist? | Minimum review |
|---|---|---|
| New customer-facing chatbot route | Yes | Trace a representative answer set, source claims, access boundary, and release decision. |
| New model, provider, prompt, or system policy | Yes | Record versions and compare a regression set against the prior route. |
| New retrieval source, index, file store, or connector | Yes | Verify source authority, freshness, permissions, citation behavior, and deletion path. |
| Source update or content owner change | Yes | Recheck affected claims and the route’s stale-source behavior. |
| New language, market, tenant, or user role | Yes | Test locale, scope, source coverage, and unauthorized cross-boundary requests. |
| Answer sent to a ticket, CRM, email, document, or tool | Yes | Validate destination handling and separately authorize every downstream action. |
| Complaint, correction, or disputed citation | Yes | Preserve a redacted record, reproduce the route, correct the source or answer, and re-test. |
| Internal prototype with synthetic data and no durable action | Maybe | Record owner, environment, scope, and the path to a production review. |
If the team cannot identify the source version or explain what happens when evidence is missing, keep the route read-only, draft-only, limited to approved users, or paused.
Provenance review intake form
Copy this form into the release, change, incident, or recurring review record. Keep customer content and restricted source material in an approved system; this record should contain identifiers and redacted evidence, not a raw transcript.
| Field | Entry |
|---|---|
| Review ID | |
| Review date and timezone | |
| Chatbot and environment | |
| Route and output destination | Customer UI, ticket, CRM, email, document, database, tool, or internal draft. |
| Tenant, user, role, and locale scope | |
| Model, provider, region, and model version | |
| Prompt, policy, and safety configuration version | |
| Retrieval, index, connector, and source-set version | |
| Review trigger | Launch, change, incident, complaint, periodic review, or expansion. |
| Approved source owner | |
| Provenance reviewer | |
| Privacy and security reviewer | |
| Release or pause owner | |
| Restricted evidence location | Use an approved access-controlled location; do not paste credentials or raw customer data here. |
| Next review trigger |
Assign a stable review ID and a stable answer or sample ID. These identifiers should let a reviewer join the minimum records needed for investigation without making the identifier itself contain customer names, ticket text, or sensitive values.
Source and claim map
Create a row for each material claim, recommendation, number, quote, or policy instruction. A single answer may need multiple rows. Do not mark an answer as supported merely because one source is generally related to the topic.
| Claim ID | Answer claim or redacted summary | Source ID and passage | Source version/date | Support result | Reviewer note |
|---|---|---|---|---|---|
| Supported / partial / unsupported / conflicting / not applicable | |||||
Use a source ID instead of copying a restricted document into the answer record. The source register should separately identify the owner, approval status, classification, canonical location, version or hash, last review date, expected update cadence, and retirement state.
Citation quality matrix
Score the citation decision by dimension, but do not collapse a failed dimension into a reassuring average. A citation that points to an authoritative source but does not support the exact claim should fail the support test.
| Dimension | Pass question | Fail or escalation signal | Default action |
|---|---|---|---|
| Authority | Is the source approved for this use and owned by a responsible team? | Unknown owner, unofficial copy, or conflicting authority. | Remove or route to a reviewer. |
| Identity | Can the reviewer identify the exact document, page, passage, and version? | Link resolves to a moving page with no captured version or passage. | Record a stable locator or retain a redacted excerpt. |
| Entailment | Does the cited passage support the claim as written? | The answer adds a conclusion, scope, number, or guarantee the source does not make. | Rewrite, qualify, or abstain. |
| Completeness | Are all material claims cited or clearly labeled as analysis? | One citation is used to imply support for several unrelated claims. | Split the claims and review each one. |
| Freshness | Was the source current for the decision date? | Expired policy, stale index, withdrawn guidance, or unknown update time. | Re-retrieve, label stale, or pause. |
| Scope | Does the source apply to this product, role, tenant, region, and workflow? | A source for one plan, market, or environment is generalized to another. | Narrow the answer or escalate. |
| Authorization | May this user and route access or reveal the cited material? | Citation leaks restricted title, snippet, URL parameter, or tenant content. | Redact, deny, and investigate. |
| User clarity | Can the user tell what is source-backed, uncertain, or generated? | Citation placement creates false confidence or hides a limitation. | Add a limitation or human handoff. |
The pass question is a review aid, not a mathematical safety threshold. High-impact, customer-visible, or irreversible actions should require explicit gates even when the citation dimensions pass.
Answer traceability record
Record metadata and decisions at the answer level. Use a redacted excerpt, structured fields, or a synthetic fixture. Do not retain a raw customer conversation simply because it is convenient for debugging.
| Field | Entry |
|---|---|
| Answer ID | |
| Request class or synthetic fixture ID | |
| Request timestamp and timezone | |
| Route and tenant boundary | |
| User role and authorization result | |
| Model and provider version | |
| Prompt and policy version | |
| Retrieval query or intent class | Store a minimized representation where possible. |
| Source IDs and passage locators | |
| Answer claim IDs | |
| Citation shown to user? | Yes / no / internal only / redacted. |
| Validation result | |
| Human review result | |
| Downstream action | None / draft / ticket / CRM / email / tool / other. |
| Abstain or handoff result | |
| Finding IDs | |
| Evidence retention class |
Separate the answer record from the access decision. A source may be relevant but not visible to the requesting user. The application or policy gateway, not the language model, must enforce that boundary.
Retrieval and version controls
| Control | Minimum implementation | Evidence to retain |
|---|---|---|
| Source inventory | Assign a source ID, owner, classification, canonical location, and approval status. | Source register row and owner confirmation. |
| Version identity | Record document version, effective date, retrieval time, or content hash. | Version field or immutable snapshot reference. |
| Index identity | Record index, embedding, parser, chunking, and ingestion job versions where they can change results. | Build or ingestion ID and change record. |
| Freshness | Define a review or refresh rule by source class; do not use one universal interval. | Freshness result, exception, and next refresh date. |
| Deletion | Remove retired or unauthorized source material from retrieval and caches. | Deletion job result and post-deletion test. |
| Conflict handling | Detect conflicting sources and route ambiguous answers to review. | Conflict ID and resolution owner. |
| Tenant scope | Filter and test retrieval by tenant, role, document, and record permissions. | Boundary test result and authorization decision. |
| Failure behavior | Abstain or use a safe fallback when retrieval is unavailable, stale, or unauthorized. | Failure fixture and observed response. |
Do not represent a source hash or ingestion timestamp as proof of truth. It proves which artifact was used, not whether the artifact was authoritative or whether the model interpreted it correctly.
Prompt, model, and route record
The same source can produce different answers after a model, prompt, tool, temperature, parser, or routing change. Record the factors needed to reproduce the result.
- Record the deployed route or application version.
- Record the model and provider version, region, and any fallback route.
- Record the system prompt, policy, guardrail, and output contract version.
- Record retrieval parameters that can change evidence selection, such as filters, top-k, score threshold, or reranking version.
- Record tool definitions, authorization checks, and downstream destination rules.
- Record whether the answer was generated in a normal, degraded, fallback, or incident state.
- Record the evaluation fixture and expected decision without storing unnecessary private input.
- Record an explanation of any missing or unavailable metadata.
If the route cannot expose enough metadata for a safe review, add a gateway-generated correlation record. Do not ask the model to self-report its own provenance and treat that response as authoritative.
Data minimization and privacy boundary
Provenance logs can become a second copy of the sensitive data they are meant to protect. Use the smallest evidence set that supports the operational decision.
| Data element | Default handling | Review question |
|---|---|---|
| Customer name, email, phone, or account ID | Mask, tokenize, or replace with a stable local sample ID. | Can the reviewer complete the decision without the real identifier? |
| Raw customer prompt | Prefer intent class, fixture ID, or a redacted excerpt. | Is the full prompt necessary to reproduce a material risk? |
| Raw model answer | Store only the relevant claim or output fragment when possible. | Can a structured finding replace the full answer? |
| Source document body | Keep in the approved source system; store source ID and passage locator. | Does the review record need a restricted excerpt? |
| Access token, secret, password, or signed URL | Never store in provenance records. | Has the value been removed from logs, traces, exports, and screenshots? |
| Sensitive category or regulated data | Use a synthetic or approved masked fixture. | Is the route allowed to process this class at all? |
| Reviewer identity | Record a team or controlled user ID as required by policy. | Who may access the review and why? |
Keep access controls, retention, deletion, and export rules on the provenance record itself. Do not assume that an internal log is safe because the chatbot is internal.
Tool and downstream action traceability
An answer may be well cited and still unsafe to send, write, purchase, delete, or execute. Keep action authorization separate from citation quality.
| Downstream result | Required record | Default gate |
|---|---|---|
| Customer-visible answer only | Answer ID, claim map, citation decision, uncertainty label, and route boundary. | Publish only supported and appropriately scoped claims. |
| Draft ticket or CRM note | Destination, fields written, actor, source scope, and edit history. | Review before external or durable use. |
| Email or customer message | Recipient scope, approved content, citation visibility, and sender authorization. | Human approval for material claims or sensitive context. |
| Tool read | Tool, arguments, authorization, returned data scope, and error result. | Allow only the minimum read scope. |
| Tool write or external action | Exact action, arguments, actor, approval, idempotency, and result. | Human approval or a separately approved low-risk policy. |
| Delete, refund, access change, or other high-impact action | Business authorization, evidence, reviewer, and rollback or appeal path. | Fail closed until the action gate passes. |
The model should not be the only component deciding whether a cited answer authorizes a downstream action. Enforce permission, validation, confirmation, and audit requirements in application code or a policy gateway.
Uncertainty and abstention rules
Write the answer behavior before reviewing examples. This prevents reviewers from rewarding confident language when the evidence is incomplete.
- State when the route must say that evidence is missing, stale, conflicting, or outside scope.
- Require clarification when the request lacks the tenant, product, time period, or other fact needed to select the right source.
- Route high-impact, regulated, safety-sensitive, or disputed decisions to a named human owner.
- Do not let a citation create certainty that the source itself does not contain.
- Distinguish source-backed facts, model-generated synthesis, assumptions, and recommended next steps.
- Keep unsupported claims out of customer-visible or durable destinations.
- Test prompt injection, conflicting instructions, source poisoning, and citation fabrication as negative cases.
- Define pause, rollback, correction, and customer communication paths for material provenance failures.
Evidence retention and access
NIST AI RMF emphasizes documented measurement, monitoring, risk tracking, and management decisions. CISA and UK NCSC guidance also frames secure AI as a lifecycle concern spanning design, deployment, and operation. Use a retention design that supports those decisions without collecting everything forever.
| Evidence | Keep | Access rule | Deletion or review trigger |
|---|---|---|---|
| Review intake and decision | Review ID, scope, owner, decision, and dates. | Review team and accountable owner. | Policy review or retention expiry. |
| Source register | Source ID, owner, approval, version, and lifecycle state. | Source owners and reviewers. | Source retirement or owner change. |
| Claim map | Redacted claim and passage locator. | Need-to-know reviewers. | Correction, source retirement, or expiry. |
| Test result | Fixture ID, route version, expected and observed result. | Engineering and security reviewers. | Test set maintenance cycle. |
| Action audit | Authorization, action, result, and correlation ID. | Operators and incident responders. | Retention policy or incident closure. |
| Restricted excerpt | Only when needed to reproduce a material issue. | Explicitly approved restricted group. | Shorter retention and verified deletion. |
Protect provenance records from unauthorized modification. If a record is corrected, append the correction and preserve the original decision history under access control rather than silently overwriting the evidence.
Review workflow
| Step | Owner | Output |
|---|---|---|
| 1. Define the route and risk boundary | Product or service owner | Intake form, destination map, and risk-checker result. |
| 2. Register sources and versions | Source owner | Approved source register and freshness rule. |
| 3. Build a claim and citation sample | Reviewer | Claim map with supported, partial, and failed cases. |
| 4. Test access and failure behavior | Engineering and security owner | Tenant boundary, stale source, conflict, injection, and abstention results. |
| 5. Review downstream actions | Product and operations owner | Action authorization and human approval gates. |
| 6. Decide scope | Release owner | Approve, limit, draft-only, require review, rollback, or pause. |
| 7. Record remediation | Finding owner | Finding IDs, deadlines, evidence, and re-test plan. |
| 8. Recheck after change or signal | Assigned reviewer | Updated provenance record and next review trigger. |
Use a separate reviewer when the route is customer-facing or high impact. A developer who built the retrieval or prompt path should not be the only person deciding that its citations are sufficient.
Test set and evidence record
Build a small, stable test set that reflects the route rather than only the easiest questions. Keep fixtures synthetic or approved masked, and record limitations.
| Test class | Example coverage | Expected result | Evidence |
|---|---|---|---|
| Supported claim | Answer is directly stated in an approved current source. | Cite the exact source and passage with correct scope. | Fixture ID and claim map. |
| Partial support | Source covers the topic but not the answer’s extra conclusion. | Qualify or split the answer; do not overclaim. | Failed entailment record. |
| No support | Request is outside the approved source set. | Abstain, clarify, or route to a human. | Abstention result. |
| Conflicting source | Two approved sources disagree. | Surface uncertainty and escalate. | Conflict ID and owner. |
| Stale source | Source is expired or retired. | Avoid citation or use a safe current fallback. | Freshness result. |
| Unauthorized source | User role cannot access the selected record. | Deny, redact, or choose an authorized source. | Boundary test. |
| Prompt injection | Input asks the route to ignore source or access rules. | Keep policy and authorization boundaries. | Negative test result. |
| Citation fabrication | Locator or quote does not exist in the source. | Reject or regenerate; never publish the invented citation. | Citation failure. |
| Downstream action | Answer is passed to a write or external tool. | Separate action approval and validate arguments. | Action audit record. |
| Localization | Same intent in each supported language or market. | Preserve scope, citation meaning, and uncertainty. | Locale result. |
For every failed test, capture the observed behavior, risk, owner, containment, correction, and re-test date. Do not hide failures by removing the fixture from the set without documenting why.
Release gate
Approve only the defined scope when each applicable gate has evidence.
- The route, user, tenant, source, model, prompt, and retrieval versions are identifiable.
- The source register has an owner, approval state, version, freshness rule, and retirement path.
- Material claims were checked for authority, identity, entailment, completeness, freshness, scope, authorization, and user clarity.
- Unsupported, conflicting, stale, or unauthorized evidence produces a clear abstain, clarification, redaction, or human handoff.
- Customer-visible citations do not expose restricted source names, passages, identifiers, tokens, or signed URLs.
- Model output is validated for its destination and is not treated as authorization for a tool or write.
- The test set includes negative cases and the results are reproducible from redacted or synthetic evidence.
- High-impact actions require the documented human or policy-gateway approval.
- Findings have owners, severity, due dates, containment, and re-test evidence.
- The decision names the approved scope, limitations, monitoring owner, and pause or rollback trigger.
If a gate fails, do not make the average score look healthy by excluding the failed dimension. Limit the route, require review, remove the affected source or action, or pause until the owner accepts a documented residual risk.
Staged rollout plan
| Stage | Scope | Provenance checks | Exit condition |
|---|---|---|---|
| 0. Offline | Synthetic fixtures and approved masked data only. | Claim map, source versions, negative tests, and redaction review. | No unresolved critical boundary or citation fabrication issue. |
| 1. Internal read-only | Named reviewers and one route. | Live source freshness, access checks, abstention, and correction workflow. | Reviewers can reproduce decisions and route failures safely. |
| 2. Limited pilot | Small approved group or tenant set. | Monitor citation failures, corrections, handoffs, and unauthorized retrieval. | Signals stay within team-defined limits and owners respond. |
| 3. Controlled customer scope | Explicit routes and supported intents. | Human review for high-impact cases and sampled claim maps. | No material unresolved finding; rollback is tested. |
| 4. Ongoing operation | Approved production scope. | Scheduled source, version, route, and metric review. | Re-review on change, incident, complaint, or drift signal. |
Do not expand from internal read-only to customer-visible use merely because latency or answer preference improved. Provenance, authorization, privacy, and downstream action gates still need evidence.
Findings and remediation
| Finding ID | Failure or limitation | Impact and scope | Containment | Owner and due date | Re-test evidence |
|---|---|---|---|---|---|
| Unsupported or over-broad claim | Remove claim, narrow route, or require review. | ||||
| Stale or conflicting source | Refresh, label uncertainty, or pause source. | ||||
| Unauthorized citation or retrieval | Block access, redact, preserve evidence, investigate. | ||||
| Fabricated locator or quote | Disable citation path, add negative test, correct output. | ||||
| Downstream action not separately authorized | Disable write or require approval. | ||||
| Excessive provenance data retained | Minimize, restrict, delete, and verify logs. |
Classify severity by realistic impact, not by how easy the finding is to reproduce. A low-volume tenant-boundary failure may be more urgent than many low-impact citation formatting issues.
Decision table
| Review result | Default decision | Follow-up |
|---|---|---|
| All material claims are supported, scoped, authorized, and reproducible | Approve the defined route | Monitor and schedule the next review. |
| A narrow claim lacks support but no action or customer impact occurred | Limit the intent or remove the claim | Fix source or wording and re-test. |
| Source is stale, conflicting, or ownership is unknown | Abstain or require human review | Assign a source owner and establish a freshness rule. |
| Citation exposes restricted data or crosses a tenant boundary | Block and investigate | Restrict access, preserve evidence, and assess exposure. |
| Citation is fabricated or cannot be reproduced | Disable the citation path | Add a negative test, correct output, and review affected answers. |
| Downstream write or external action lacks independent authorization | Disable automatic action | Require human approval and re-test the gateway. |
| Provenance records contain unnecessary private data | Limit access and minimize | Delete excess data and verify retention controls. |
| Repeated material failures or unresolved incident | Pause or rollback | Use the incident, correction, and recovery owner path. |
Sign-off record
| Field | Entry |
|---|---|
| Review ID | |
| Approved chatbot and route | |
| Approved user, tenant, locale, and intent scope | |
| Approved source set and version | |
| Model, prompt, retrieval, and policy versions | |
| Test set and evidence record | |
| Open findings and accepted limitations | |
| Human review and action gates | |
| Monitoring and correction owner | |
| Pause, rollback, and incident owner | |
| Decision | Approve / limit / draft-only / require review / pause / rollback. |
| Decision date and timezone | |
| Release owner | |
| Reviewer | |
| Next review trigger |
The sign-off approves the documented scope only. A later model, prompt, source, permission, route, tenant, action, or policy change requires a new review or a documented change assessment.
Action tracker
| Action ID | Action | Reason | Owner | Priority | Due date | Status | Evidence link or record ID |
|---|---|---|---|---|---|---|---|
| Open / blocked / done | |||||||
Keep the action tracker connected to the review ID and finding ID. A task marked done should point to the re-test or deletion evidence that closed it.
Final provenance checklist
- I can identify the route, model, provider, prompt, policy, retrieval, source, tenant, and user scope.
- Each material claim has a source locator, version, support result, and reviewer note.
- Citation authority, entailment, completeness, freshness, scope, authorization, and user clarity were checked.
- The answer distinguishes source-backed facts, synthesis, assumptions, uncertainty, and recommendations.
- Missing, stale, conflicting, injected, or unauthorized evidence causes abstention, redaction, clarification, or handoff.
- Provenance data is minimized, access-controlled, retained by policy, and removable.
- No password, API key, token, signed URL, raw customer export, or unredacted private transcript is in the record.
- Downstream tool reads, writes, messages, and high-impact actions have separate authorization and validation.
- Negative tests cover citation fabrication, stale sources, conflicting sources, prompt injection, and tenant boundaries.
- The release decision names scope, limitations, owner, monitoring, correction, pause, and rollback conditions.
- Findings and actions have owners, dates, status, and re-test evidence.
- I ran the AI Tool Risk Checker and linked the result to the review.
- I completed the Small Team AI Security Checklist for the surrounding controls.
Metrics to track
Track trends by route, model, source class, tenant scope, and review period. Avoid publishing a single score as a safety guarantee.
| Metric | Why it matters | Review question |
|---|---|---|
| Supported-claim rate | Shows whether material claims were backed by the selected sources. | Is the rate changing by intent or source class? |
| Partial or unsupported claim rate | Surfaces overclaiming and missing evidence. | Are answers narrowing, abstaining, or hiding uncertainty? |
| Citation locator failure rate | Detects broken, fabricated, or non-reproducible citations. | Can the reviewer reach the exact passage and version? |
| Stale-source rate | Shows whether freshness controls are working. | Which source owners or ingestion jobs need attention? |
| Unauthorized retrieval or citation attempts | Indicates boundary pressure or possible exposure. | Are access filters and alert routes working? |
| Conflict and abstention rate | Makes uncertainty visible. | Are conflicts resolved or repeatedly resurfacing? |
| Human review and correction rate | Measures operational load and answer reliability signals. | Are reviewers seeing a new failure pattern? |
| Downstream action block rate | Separates answer quality from action safety. | Are automatic actions too broad or correctly bounded? |
| Provenance record minimization exceptions | Finds unnecessary data retention. | Can the exception be removed or shortened? |
| Time to reproduce a finding | Measures whether evidence is usable. | Can another reviewer investigate without raw private data? |
Metrics need a defined population, time window, route version, denominator, owner, and limitation. Retain enough context to interpret a change, but do not turn analytics into a hidden transcript archive.
Evidence checked
- OWASP LLM09:2025 Misinformation describes credible false or misleading output and the need for validation, human oversight, risk communication, and clear limitations.
- OWASP LLM05:2025 Improper Output Handling explains why model output needs validation, sanitization, and safe handling at its destination.
- NIST AI RMF Core describes documented measurement, monitoring, risk tracking, and management decisions across the AI lifecycle.
- NIST AI RMF Generative AI Profile provides a generative-AI-focused companion profile for identifying and managing lifecycle risks.
- CISA and UK NCSC Guidelines for Secure AI System Development covers secure design, development, deployment, and operation of AI systems.
- AI Tool Risk Checker for a route-level risk record.
- Small Team AI Security Checklist for identity, data, access, logging, and incident foundations.
FAQ
Is a citation enough to prove that a chatbot answer is correct?
No. The cited source must be authoritative for the use case, current for the decision, authorized for the user, and actually support the claim as written. The review also needs to consider model interpretation, missing evidence, downstream handling, and human escalation.
Should we store every prompt and answer for traceability?
Usually no. Start with a stable answer ID, intent or fixture ID, claim map, source locator, version metadata, decision, and redacted evidence. Retain a restricted excerpt only when it is needed to reproduce a material issue and permitted by the team’s data policy.
Can the model generate the provenance record itself?
The model can propose a structured record, but the application should populate authoritative route, model, access, source, version, and action fields. Treat model-reported citations as untrusted until the system verifies them against the source and authorization boundary.
What if the source is relevant but does not answer the exact question?
Mark the claim as partial or unsupported. Narrow the answer, ask a clarifying question, identify the limitation, or hand off to a human. Do not let a generally related citation imply a conclusion the source does not state.
How often should provenance be reviewed?
Review at launch and after model, provider, prompt, retrieval, source, permission, route, tenant, locale, tool, or policy changes. Add recurring review based on risk and monitor complaints, corrections, conflicts, stale sources, and unauthorized attempts. A universal interval is not a substitute for a risk-based trigger.
Does this checklist create a compliance certification?
No. It is a practical operating aid. Legal, privacy, records, sector, and contractual requirements vary by jurisdiction and workflow; obtain qualified advice for those decisions.
Recommended next step
Run the AI Tool Risk Checker for the chatbot route, create one provenance review intake, and trace a small synthetic or approved masked sample through source selection, claim support, citation display, access checks, and downstream action gates. Then record the result in the Small Team AI Security Checklist and keep the first release read-only or human-reviewed until the evidence chain is reproducible.