DEVELOPER GUIDE / FIELD GUIDE 12

AI LLM Disk Browser

Design file-aware retrieval around provenance, authorization, and answers that lead back to inspectable evidence.

Build the evidence path first

An AI LLM disk browser architecture should explain how a source becomes a cited answer. Our reference workflow separates approved sources, extraction, passage creation, indexing, authorized retrieval, and answer presentation. Keep write tools outside the first implementation. Each stage should produce evidence that a developer can inspect when a result is wrong.

Ground the retrieval concept

The original Retrieval-Augmented Generation paper describes combining retrieved information with generation. Our file-oriented guide applies that idea as a proposed design, not as an implementation or benchmark attributed to the paper. Evaluate the design against your own authorized collection and question set.

Track source identity and coverage

Propose records for source identity, path, version, extraction status, access scope, and passage location. Keep failed extraction distinct from an empty or irrelevant document. Explain how moves, replacements, edits, and deletions affect the index. Do not silently present an older source passage as current after the underlying material changes.

Evaluate the stages independently

Test whether the expected source was admitted, successfully parsed, retrieved, and accurately represented in the answer. Include questions the collection cannot answer and permission-revocation cases. Report observed outcomes under defined conditions rather than inventing accuracy figures. Keep release approval scoped to the corpus, permissions, formats, and use cases actually evaluated.

Questions worth asking

Does a retrieval layer guarantee a correct answer?

No guarantee follows from the architecture alone. Evaluate source coverage, extraction, relevance, attribution, uncertainty, and authorization in the actual implementation.

Where should a prototype begin?

Use a small authorized text collection, a read-only task, and questions with independently known evidence. Add complexity only when the new input or capability has its own acceptance tests.