Reliability continues after launch

A retrieval-augmented generation system does not become reliable once and remain that way.

Documents change. Policies are superseded. Permissions move with employees. Users ask questions the original evaluation set did not anticipate. Models, prompts, retrieval settings, and third-party services are updated.

A RAG assistant that worked at launch can therefore become less dependable without producing an obvious technical error.

The practical answer is to treat it as a maintained knowledge product. Give people explicit ownership, manage the source collection throughout its lifecycle, evaluate material changes before release, monitor failures in production, and maintain a safe way to fall back or disengage.

Assign ownership before discussing monitoring

A dashboard cannot compensate for unclear responsibility. A dependable RAG service needs three forms of ownership.

One person may cover more than one role in a smaller organisation. The important requirement is that none of these responsibilities remain implicit.

When a poor answer is reported, the team should know whether the problem belongs to the source material, retrieval configuration, answer generation, permissions, or the workflow surrounding the assistant. Otherwise every failure becomes an unstructured technical investigation.

  • Knowledge ownership: Who decides which documents are authoritative, resolves conflicts, and removes obsolete material?
  • Product or operational ownership: Who defines the intended users, supported questions, acceptable failure behaviour, and escalation path?
  • Technical ownership: Who maintains ingestion, retrieval, access controls, evaluation, observability, providers, and deployment?

Define the approved knowledge boundary

The first operational question is not, “How many documents are indexed?” It is, “Which information is the assistant currently allowed to treat as authoritative?” Maintain a source register with the details below.

This prevents the knowledge base from becoming a convenient dumping ground. A RAG system can retrieve from thousands of documents and still be unreliable if users cannot tell which version governs, two policies contradict each other, or deleted material remains searchable.

The UK Government Data Quality Framework describes quality as fitness for purpose and recommends managing quality throughout the data lifecycle, addressing issues at their source, maintaining current metadata, and regularly monitoring data that changes. Although written for government data, the operating principles apply well to an organisational knowledge base.

  • The source collection or business area
  • Its knowledge owner
  • Intended audience
  • Access restrictions
  • Authoritative location
  • Effective and review dates
  • Status such as draft, approved, superseded, or archived
  • Expected update frequency
  • What should replace it when it is retired

Make freshness measurable

“Updated regularly” is not an operating requirement.

Different knowledge collections need different freshness expectations. A frequently changing product catalogue may need updates within minutes. An internal procedure might allow one business day. A stable historical archive may be refreshed much less often.

Measure the complete delay from the authoritative source changing to the new retrieval behaviour being verified. A successful indexing job alone does not prove that the assistant now returns the right version.

Test additions, updates, deletions, and permission changes separately. They often follow different technical paths and can fail independently. For each important collection, define the following.

  • How changes are detected
  • How quickly additions become searchable
  • How quickly corrections replace previous versions
  • How quickly deleted or revoked material disappears
  • How permission changes propagate
  • What happens when ingestion fails
  • Who receives the failure notification

Evaluate every material change

Changes to a RAG system should be treated like product releases.

A new model can alter how retrieved evidence is interpreted. A prompt change can affect refusal behaviour. New chunking or ranking settings can improve common questions while breaking another department’s queries. A document update can introduce a contradiction. A connector change can affect permissions.

Run the relevant evaluation before releasing changes to sources, retrieval, prompts, models, permissions, or orchestration. Compare against the current production version rather than asking only whether the new version looks acceptable.

OpenAI’s current evaluation guidance recommends evaluating every change, monitoring production use for new nondeterministic cases, and expanding the evaluation set over time. The principle is useful beyond any one model provider. Maintain a versioned evaluation set containing the following.

  • Common user questions
  • High-consequence questions
  • Ambiguous questions requiring clarification
  • Questions requiring several sources
  • Questions whose answers changed between document versions
  • Unanswerable questions
  • Conflicting-source cases
  • Permission-restricted cases
  • Previously reported failures

Monitor the workflow, not only the infrastructure

Server uptime does not tell you whether employees are receiving supported answers. A useful operating view includes the signals below.

Do not collect sensitive prompts and retrieved passages by default merely because they may help debugging. Decide what telemetry is necessary, restrict access, establish retention periods, and redact or minimise sensitive content where possible.

The UK National Cyber Security Centre recommends monitoring system inputs and behaviour in line with privacy and data-protection requirements. It also advises treating changes to data, models, or prompts as behavioural changes that need suitable testing and evaluation.

  • Source-ingestion success and delay
  • Documents approaching or exceeding their review date
  • Deleted documents still appearing in retrieval
  • Permission-sync failures
  • Questions that retrieve no useful evidence
  • Questions that retrieve conflicting evidence
  • Refusal and escalation frequency
  • User corrections and reported poor answers
  • Retrieval and answer quality on the regression set
  • Latency and cost per completed task
  • Changes in question categories over time

Turn user feedback into engineering evidence

A thumbs-down count is not enough to improve a knowledge assistant. When practical, let users report why an answer failed.

Review reported failures with both the technical and knowledge owners. Fix the problem at the correct layer, then add a representative case to the regression set.

  • The wrong source was retrieved
  • The source was outdated
  • The answer contradicted the source
  • The answer omitted an important condition
  • The user lacked access to the cited material
  • The assistant should have refused
  • The question was misunderstood
  • The answer was correct but not useful

Make each correction repeatable

This creates a useful cycle. Without the added evaluation, the same failure can quietly return during a later change.

  • Observe a production failure.
  • Classify its cause and consequence.
  • Correct the source, retrieval, prompt, permission, or workflow.
  • Add an evaluation that reproduces it.
  • Verify the correction without breaking existing behaviour.
  • Release and continue monitoring.

Protect deletion and permission behaviour

For an internal RAG system, removing access can be more important than adding content. Test the situations below.

Access restrictions should be applied before restricted content is retrieved and passed to the model. Hiding a citation in the interface is not an access-control mechanism.

Keep enough evidence to verify that permission and deletion updates have propagated without retaining unnecessary sensitive content.

  • An employee changes role
  • A document becomes restricted
  • A source is deleted
  • A team or folder is renamed
  • A cached result refers to material the user can no longer access
  • Permissions change while an indexing process is incomplete

Hold a recurring reliability review

The appropriate cadence depends on how quickly the knowledge changes and how consequential the answers are. Avoid imposing one arbitrary schedule on every system.

NIST’s voluntary AI Risk Management Framework Playbook recommends post-deployment monitoring that includes user feedback, change management, incident response, recovery, override, and decommissioning. It also recommends documenting and regularly monitoring risks arising from third-party AI components. The review should answer the following.

  • Are the sources still authoritative and sufficiently current?
  • Which questions are failing most often?
  • Are failures concentrated in one collection or user group?
  • Have permissions and deletion workflows been tested recently?
  • What changed in models, prompts, retrieval, connectors, or source structure?
  • Did those changes pass regression evaluation?
  • Are users relying on the assistant for decisions outside its intended boundary?
  • Which incidents or near-misses need corrective action?
  • Does the system still justify its operating cost and risk?
  • Is any collection or capability ready to be retired?

Require an operable handoff

If an external partner builds the system, the handoff should include more than source code and deployment credentials.

If only the original developer can explain why retrieval works, investigate a failed answer, or replace a provider, the project has not yet reached an operational handoff. Ask for the following.

  • The approved source and permission model
  • Document ingestion, update, and deletion procedures
  • The versioned evaluation set and current results
  • Known limitations and unsupported uses
  • Monitoring and alert ownership
  • A failure-triage process
  • Provider and component inventory
  • Change and release procedures
  • Backup, fallback, and recovery instructions
  • A safe shutdown or decommissioning process

Reliability is an operating capability

A production RAG system is not simply a chat interface connected to a vector database. It is a relationship between changing organisational knowledge, user questions, access rules, retrieval behaviour, and accountable people.

Its long-term quality depends less on whether the launch demonstration looked intelligent and more on whether the organisation can detect change, investigate failure, verify corrections, and preserve the knowledge boundary over time.

If your RAG pilot works but no one yet owns its maintenance, I can help turn it into an operable knowledge product with clear source governance, evaluation, monitoring, and handoff.

Primary sources