How to keep an AI knowledge base up to date
Keep company AI answers current with source owners, update rules, deletion checks and tests that catch obsolete facts, broken syncs and access changes.
To keep an AI knowledge base up to date, assign owners to its sources, define how changes and removals reach the assistant, and test the answers after each important update. A scheduled sync helps move information. It does not establish that the information is approved, complete or still appropriate for the person asking.
The file can be correct while the answer is wrong.
Suppose a company changes a return window from 14 to 30 days. The policy owner updates the document, and the connector reports success. An employee asks the assistant about returns and still gets 14 days. The old statement could survive in a leftover text segment, an exported PDF, a generated summary or the conversation itself.
That is a fictional example, not a reported LeanOrchestr customer incident. It illustrates the acceptance test this guide uses: the new answer must work, and the superseded answer must stop appearing as current.
Separate a recent file from a current answer
“Last updated” can refer to several different events. Record which one you mean.
The source changed when someone edited the document. It was verified when an accountable person checked its accuracy and approval. The search copy changed when the assistant’s retrieval system processed it. The answer was verified when a user-facing test showed the correct result.
Those events do not necessarily happen together. A spelling correction makes a file recent without making an old policy accurate. An approved policy may be written before its effective date. A successful indexing job can finish before the new material is available to queries.
AWS documents that last distinction for Bedrock ingestion: after a job completes, query availability can still take several minutes for the documented stores, with an Aurora exception. Treat that as a product-specific reason to verify retrieval, not as a universal waiting period.
For important sources, keep a small register containing:
- The stable source ID, canonical location and accountable owner.
- What the source is authoritative for and who may use it.
- Its approval state, effective date and last substantive review.
- The latest detected change and successfully processed version.
- The last answer check, its result and any unresolved exception.
Keep the source’s own dates separate from your system’s collection time. Reading a document today does not mean its contents were verified today.
This register need not become another large application. A maintained table can serve a small pilot. Review an exception list of overdue approvals, failed updates and unverified answers instead of asking someone to inspect every unchanged file each morning.
Decide how much delay each source can tolerate
Do not begin by scheduling everything nightly. First ask what happens if someone receives yesterday’s answer.
A company-history page may remain useful between periodic reviews. A changed escalation procedure may need to reach the assistant before the next shift. An inventory balance may change too quickly for a document-indexing workflow to be the right source.
Set a maximum acceptable source-to-answer delay for each important class of information. Allocate that time across change detection, approval where needed, ingestion and verification. Leave room for recovery. A scheduler that starts at the end of the allowed window cannot meet it.
For example, an internal team might choose a two-hour update target for an ordinary procedure change. That is a hypothetical operating choice, not a recommended standard. If the connector runs every two hours, indexing takes additional time and nobody notices a failed job until tomorrow, the design cannot satisfy the target.
For live operational facts, consider querying the permission-controlled system of record at answer time. The connector guide explains why connecting a system is different from retrieving the right records. Still show the query time and handle unavailability; a direct connection is not a guarantee of a current result.
Define an urgent correction route separately from routine review. If a source is known to contain harmful or materially wrong guidance, the immediate action may be to exclude that topic from answers while its replacement is approved. Continuing to serve a known bad instruction until the next convenient sync is not a safe fallback.
Give content review an owner and time to happen
The technical owner can verify that a file was processed. They may not be qualified to approve what it says.
Assign content responsibility to the person or team that controls the underlying process. Assign synchronization, monitoring and recovery to the person operating the assistant. One person can hold both roles in a small business, but both responsibilities must be explicit.
If a group owns a topic, identify who handles its review queue and escalation. Plan what happens when that person leaves or changes roles. “Owned by Operations” is not sufficient if every member expects someone else to approve the correction.
Practitioners in the ServiceNow knowledge-maintenance discussion describe different approaches to review intervals, author capacity and reassignment. These are experiences, not a controlled comparison. They do expose a practical constraint: adding reminders does not create time for the work.
Attach review to events that already change the business. A process release can require its related instructions to be checked before the release closes. A corrected customer answer can generate a review item for the source that produced the error.
Use age, feedback and usage as review signals, not automatic verdicts. A rarely opened emergency procedure may still be essential. Before retiring it, ask whether it is still required and accurate, not just whether it received enough views.
AI can flag likely contradictions or draft a proposed correction. The responsible person still needs to decide which instruction the business endorses.
Specify what happens for each kind of change
A useful update design distinguishes an addition from a replacement, a removal and a change of access. Test each one.
| Source event | Intended behavior | Failure to check for |
|---|---|---|
| Approved material added | Make the approved version available to its permitted audience | The file exists, but extraction missed a table or section |
| Existing material edited or shortened | Replace affected searchable text and remove superseded segments | Old paragraphs remain beside the replacement |
| File moved or renamed | Preserve its identity or reconcile the replacement explicitly | A second record is created and the first survives |
| Material retired | Remove it from current-answer retrieval and handle dependent copies | A PDF export or summary keeps supplying the old fact |
| Access removed | Stop retrieval and delivery for the affected user | A broad connector or cached answer bypasses the change |
| Future policy approved | Make it current only under the agreed effective-date rule | A newer file overrides the policy still in force |
An assistant often searches smaller text segments rather than whole documents. If a replacement is shorter, overwriting only its new segments can leave old ones behind. Keep a way to associate each segment with its parent source and version so you can check the replacement as a whole.
Existing guides already cover parts of this. MindStudio’s semantic-search guide discusses changed-document embedding and removal of old chunks. Your acceptance test should also inspect the copies and dependent outputs outside that immediate index.
Platform details matter. Microsoft’s storage-indexer documentation explains that deletion detection needs an appropriate strategy. Its described policies do not cover one-to-many indexing, which needs explicit index-document deletion. Do not assume a parent file disappearing proves every derived record disappeared.
Keep retirement distinct from authorized historical access. An old policy may remain useful for explaining what applied at an earlier date. If retained, separate it from current guidance and label its valid period and audience. A historical record should not silently win a question about today’s procedure.
Check interrupted replacements too. If only half a document is processed, the assistant should not quietly combine the new opening with the old ending. Where the platform supports it, validate a replacement version before making it current. Otherwise, define how the affected source is withheld during an incomplete update.
Preserve version order when jobs retry. A delayed job for yesterday’s document must not overwrite today’s approved replacement. The operator needs to be able to identify which source version a successful job actually applied.
Follow the copies, including generated summaries
Return to the fictional 14-day policy. Replacing its main document does not repair a training PDF that still repeats 14 days. Nor does it correct an AI-written onboarding note copied into a different folder.
Keep a source-to-dependent-record map for material you deliberately derive. At minimum, record which source versions produced each maintained summary. When a source changes, send the affected summaries for review or exclude them until they are checked.
A citation is useful, but it does not keep the cited statement current. A summary can still link to the right document while preserving its old wording. The review needs to compare the claim, not merely confirm that the link opens.
Avoid feeding generated answers back into the authoritative collection automatically. An assistant can otherwise encounter its own earlier mistake in several documents and treat that repetition as apparent support. Preserve generated work only when there is a reason to maintain it, and make its approval and lineage visible.
Also define how the application handles cached responses and existing conversations. A new chat might retrieve the corrected policy while an old chat continues from context containing the earlier answer. Test both paths. Where the application cannot reliably refresh sensitive context, restrict the affected workflow and tell users to begin a verified new session.
Revoking access cannot erase a copy someone already received. Prevent further retrieval or delivery, handle retained conversation data under the application’s access and retention rules, and do not describe a successful revocation test as recall of previously disclosed information.
The document-preparation guide covers source selection and extraction before launch. Maintenance needs to repeat the relevant checks when the documents change, not assume the first clean import solved them permanently.
Choose a maintenance method you can operate
For a small collection, a reviewed folder and a manual replacement checklist may be sufficient. Keep the source register, remove superseded copies deliberately and run the answer checks. Measure the actual maintenance effort before adding automation.
A managed connector is useful when it supports the source types and change behavior you need. Verify those details for the account and configuration you will use.
For example, DigitalOcean’s current indexing documentation distinguishes connected sources from uploaded files, whose changes require removing and adding the file again. It also documents skipped scheduled runs when indexing is already in progress. A failed run does not cancel future schedules. Those distinctions make job-result monitoring necessary even when a schedule exists.
A custom workflow makes sense when you need controls the managed route does not provide. It also makes you responsible for retries, reconciliation, access handling and operational support. Do not choose it solely because a tutorial makes the first successful upload look simple.
In Shafik Soleh’s self-updating knowledge-base demonstration, the workflow adds a record, but the video ends before showing the live agent answer from it. That is an observation about the demonstration, not a claim that the platform cannot do more. Validate the missing steps in your own environment.
Whichever route you choose, treat access changes as updates too. AWS’s S3 connector documentation warns that crawled material is available to identities with its documented retrieval permission despite source restrictions. Do not assume every connector automatically reproduces each end user’s source permissions.
Use access enforcement before retrieval and an explicit revocation test. A prompt asking the assistant to keep information confidential cannot repair an overly broad data connection.
Test the new answer and the answer that must disappear
Use a test collection without real customer or employee data. Choose a few representative questions, record the expected source and answer conditions, then change the material deliberately.
Here is a fictional test for a return-window policy changed from 14 to 30 days, effective 10 September. The document was edited on 9 September. That difference lets you test effective-date handling as well as file recency.
| Check | Expected result | Evidence to retain |
|---|---|---|
| Ask the ordinary question after the effective date | The answer uses 30 days and cites the approved policy | Test time, user scope, source ID/version and answer |
| Ask using wording from the old paragraph | Fourteen days is not presented as the current rule | Retrieved source identities and response |
| Shorten the document to remove a test exception | The removed exception no longer appears as current guidance | Old/new version IDs and removal result |
| Ask from a user who lost access | Restricted policy content is not retrieved or disclosed | Test identity, permissions and result, without unnecessary sensitive text |
| Repeat in an existing conversation | Old context does not override the effective policy | Conversation-state condition and result |
| Interrupt synchronization | The failed update is visible and the affected topic follows its fallback rule | Job result, alert and recovery check |
Also test a future-dated draft and a genuine conflict. An assistant should not pick a disputed rule just because it is expressed more clearly than the approved one. Define when it must withhold a definitive answer and identify the person who can resolve the issue.
A source link alone is not a pass. Open it and confirm that the exact passage supports the answer for the relevant date and audience. Check whether retrieval selected the right material before deciding that a wording change to the prompt will fix the problem.
Keep the test small enough to repeat after meaningful changes. Expand it when an actual failure exposes another path, and preserve that regression case. Never turn a one-time clean result into a claim that the assistant will always be current.
Make failed updates visible
Monitor the last successfully processed source version, not just the last time a scheduler started. Preserve the pending change until its required steps succeed.
Distinguish no change from failed collection. If a connector cannot reach the source, an empty result should not trigger mass deletion. Confirm retirement through a reliable deletion signal or a successful reconciliation of the authorized source scope.
Reconcile periodically even when event-driven updates appear healthy. Compare the intended source set with indexed parent records and investigate missing or unexpected versions. Keep removal signals available until their processing is confirmed. A dropped event should not leave an obsolete answer in place indefinitely.
Decide what an operator should do when an update is late: retry within agreed limits, repair extraction, request content approval or temporarily remove the affected topic from answers. Name an escalation owner and record the recovery check.
Keep operational logs proportionate. Source identifiers, versions, job outcomes and test results may be enough; copying complete private documents into alerts creates another collection to protect and maintain. Restrict logs and set a retention period.
Using a previous version may be acceptable when it remains valid and permitted. It is not an acceptable fallback after a known critical correction or access revocation. Do not roll those changes back merely to restore a green status.
A daily business briefing can surface overdue reviews and failed updates, provided it labels collection gaps. Keep urgent alerts separate so a serious stale-answer problem does not wait for tomorrow’s digest.
Start with one source you can prove is current
Select one useful, low-risk process document and its owner. Record who may use it, what makes it authoritative and how long an update may take.
Then edit a fact, shorten a section, retire a copy and change a test user’s access. Check the live answers after each event. Record the elapsed time and the manual work needed to reach the expected result.
If a step fails, repair that path before expanding the collection. This gives you a more useful basis for investment than counting connected documents.
A Business Brain needs knowledge the company can maintain. The first milestone is one source whose correct use, replacement and retirement you can demonstrate.
Questions people ask
How often should an AI knowledge base be updated?
Set the interval according to how long an old answer can safely remain in use. Stable reference material may need periodic review; operational changes can require event-triggered updates or direct queries. Include collection, indexing, verification and failure recovery in the allowed delay. A daily schedule does not guarantee daily freshness.
Does uploading a replacement file remove the old information?
Not necessarily. Check whether the tool replaces the original source or creates another record, and whether it removes all old text segments. Test a fact deleted from the replacement file. Also check exported copies, generated summaries and cached answers that could preserve it.
Should AI-generated answers be added back into the knowledge base?
Only after an accountable reviewer checks and approves them. Keep links to the original evidence, a review date and a clear label distinguishing the summary from its sources. When a source changes, identify affected summaries for review. Repeated generated statements are not independent corroboration.
What should happen when two sources disagree?
Use the agreed source authority, approval state and effective date to determine whether one supersedes the other. Do not choose solely by file modification time. If a material conflict remains unresolved, withhold a definitive answer for that topic and route the conflict to its owner.
How do I know an update actually worked?
Check the live answer and its supporting source as the intended user. Confirm the new fact is used, the obsolete claim is no longer presented as current, and access restrictions still hold. Test both a new conversation and an existing one. A successful ingestion job is only one part of this check.