Retrieval-augmented generation (RAG) gives a language model evidence to work with: ingest documents, split their text into chunks, embed those chunks, retrieve relevant passages for a question, and generate an answer that cites them.
The interesting engineering question appears when the first search returns only part of what the user needs. Should the application answer from those passages, admit the gap, or search again with a different question?
I explored that choice through two local document-chat prototypes: Document Chat, built with Streamlit, and Folio, a Next.js/FastAPI application with a LangGraph research workflow. They share the same basic document-library design. Their main difference is how they decide what to retrieve and when to stop.
Case-study implementations
The working implementations discussed in this article are available on GitHub:
- Document Chat — conventional Streamlit RAG
- Folio — bounded agentic RAG with FastAPI, Next.js, and LangGraph
The examples and implementation details below refer directly to these repositories.
My default is to start with conventional RAG and earn the additional complexity of agentic orchestration through measured retrieval failures. A longer workflow can gather better evidence, but it also creates more opportunities to make a bad decision.
The shared foundation
Both prototypes accept PDF, DOCX, TXT, and Markdown files. They extract text, create chunks of up to 200 tokens with 40-token overlap within each extracted section, and embed them locally with MiniLM. Chroma stores the text, vectors, and source metadata in a temporary collection for each session. Both use Groq’s openai/gpt-oss-20b for language-model calls.
That shared foundation makes the orchestration choices easier to examine. Neither application gains access to a better corpus simply because its interface or workflow is more elaborate.
These are local prototypes: document collections live in memory, idle sessions expire, and there is no authenticated production tenancy model. Session isolation is useful, but it does not replace access control. Questions and retrieved excerpts are sent to the model provider; locally stored vectors do not make the entire answering flow local.
Throughout this post, conventional RAG means the Streamlit application’s single retrieval pass. Bounded agentic RAG means Folio’s model-directed search planning and evidence feedback inside a fixed graph. These are two points on a spectrum. Conventional systems can also use query rewriting, hybrid search, or reranking; agentic systems can range from constrained workflows to agents with many tools. LangChain’s retrieval architecture guide similarly distinguishes fixed retrieval, agent-directed retrieval, and hybrid workflows.
Conventional RAG: retrieve once, then answer
Document Chat follows a short request path. It takes the question, resolves conversational references when necessary, retrieves up to five nearest chunks from the session’s collection, and asks the model to answer from those excerpts.
For a first question such as “Who owns the release checklist?”, the search query is the question itself. For a follow-up such as “Who covers for them?”, a separate model call rewrites it into a standalone search query using recent conversation. There is still only one retrieval pass and one answer-generation pass. Single-pass RAG does not necessarily mean a single model call.
Conceptually, the flow is:
query = resolve_follow_up_if_needed(question, recent_history)
evidence = search(session_library, query, limit=5)
if evidence is empty:
return insufficient_information
answer = generate_from(question, evidence, recent_history)
return check_citation_references_or_allow_abstention(answer, evidence)
The generator is instructed to use the excerpts as evidence and to abstain when they cannot support an answer. The application rejects answers with missing or out-of-range citation numbers. It does not make a second search to fill a gap.
This makes a strong baseline for documentation lookup, FAQs, and questions whose answer is concentrated in one or two passages. There are fewer sequential model calls, a more predictable execution path, and fewer intermediate decisions to debug. Actual latency still depends on inference, network conditions, and retrieval performance.
Its limitation is the fixed evidence window. The five nearest passages might all describe the same part of a topic. If another part of the question requires a different vocabulary or document, the generator cannot retrieve that missing evidence. Increasing the chunk limit can help coverage, but also increases context size and can introduce distractions.
One question, two request paths
The diagram shows the orchestration difference. In Folio, “Accepted” means its routing policy permits generation; “Stop” means evidence is empty or the final round rejects it. The stopping policy is explained below. On small screens, scroll the diagram horizontally.
flowchart TB
accTitle: Conventional and bounded agentic RAG request paths
accDescr: Conventional RAG retrieves once before generation. Folio plans searches, retrieves evidence, and validates it, looping within a three-round budget or requesting missing documents.
subgraph conventional[Conventional RAG]
direction TB
C[Question + history] --> W[Optional query rewrite]
W --> R[Retrieve up to five]
R --> G[Answer or abstain]
G --> A[Check citations]
end
subgraph agentic[Bounded agentic RAG: Folio]
direction TB
Q[Question + history] --> P[Plan searches]
P --> S[Retrieve evidence]
S --> V{Coverage?}
V -->|Accepted| F[Answer + citations]
V -->|Retry| P
V -->|Stop| U[Request documents]
end
Consider an illustrative question: “What must be completed before Project Cedar can launch, who owns each item, and what happens if the security review slips?”
Suppose the corpus contains a launch checklist, an ownership document, and an exception policy. A single search could retrieve everything needed. It could also return five similar checklist passages and miss the exception policy entirely. The question’s difficulty comes from covering several evidence needs, not merely from its length.
Folio can decompose it into searches for launch prerequisites, owners, and security-review exceptions. If the first results cover only the first two, a validator can identify the missing exception policy and feed that gap into another planning round. This is a plausible benefit of the architecture, not a measured result from a benchmark of these prototypes.
How Folio’s agentic loop works
1. Plan searches around evidence needs
The planner receives the question, bounded recent conversation, and feedback from the previous validation round. It produces up to three focused search queries, each with a stated purpose.
The planner’s job is to express what information to look for. The application validates its structured output with Pydantic and allows one repair attempt if the model returns malformed JSON. Schema validation checks the shape of a plan; it cannot establish whether the searches are useful.
2. Retrieve through a constrained operation
Retrieval is ordinary application code. Each query searches only the current session’s Chroma collection and returns up to three chunks. Folio deduplicates chunks by identity and accumulates at most ten across the request.
The model cannot choose an arbitrary database, browse the web, or execute commands. These boundaries come from the available code path. The planner can vary the query, while the application controls where and how it runs.
The evidence cap bounds the downstream context, but introduces a trade-off: this implementation keeps the first admitted chunks rather than reranking and replacing them. Once ten chunks have accumulated, subsequent rounds cannot add a better passage. Additional planning and validation can therefore consume time without improving the evidence set.
3. Assess coverage and route the next step
The validator receives the original question, the current search plan, and the accumulated excerpts. It returns a sufficiency decision, a confidence estimate, and the types of evidence still missing.
The first two rounds accept a nonempty evidence set when the validator says it is sufficient. Otherwise, the graph can return to planning with the missing-evidence feedback. Empty evidence terminates with a request for more documents. The overall budget is three rounds, with up to three searches per round and three results per search, subject to the ten-chunk cap.
There is a significant implementation caveat. At the time of writing, the third round permits generation when evidence exists and the validator’s confidence is at least 0.1, even if its decision still says more evidence is needed. The README describes a threshold of 0.5; the source currently uses 0.1. This article follows the source behavior.
That threshold is a prototype policy, not a recommended confidence level. The score is the model’s own estimate, not a calibrated probability that an answer is correct. On this fallback path, the generator is instructed to answer only supported parts and identify unresolved gaps. If the threshold is not met, the application requests the missing documents or sections.
Conceptually:
evidence = []
feedback = []
for round in 1..3:
plan = plan_searches(question, recent_history, feedback, limit=3)
evidence = retrieve_and_accumulate(session_library, plan, evidence, cap=10)
assessment = assess_coverage(question, plan, evidence)
accepted = evidence is not empty and (
assessment.sufficient if round < 3
else assessment.confidence >= configured_final_threshold
)
if accepted:
answer = generate_supported_answer(question, evidence, assessment.gaps)
return check_citation_references_or_allow_abstention(answer, evidence)
feedback = assessment.gaps
if evidence is empty:
break
return request_missing_documents(feedback)
4. Generate from the accumulated evidence
Once routing permits generation, the generator receives the original question, numbered excerpts, and remaining evidence gaps. It is instructed to cite factual claims, and the application validates citation references before accepting an ordinary answer. It can also return an insufficient-information response.
Planner, validator, and generator are separate roles using the same model, rather than independently trained experts. Their errors may be correlated: a mistaken interpretation can survive more than one model call.
Folio exposes planned queries, retrieval counts, validation decisions, and the final outcome in a research panel. These are observable workflow events and explicit model outputs. They do not reveal private model reasoning or prove that a decision was sound. They do help an engineer see whether failure began with the plan, retrieval, evidence assessment, or generation.
What the extra control buys—and costs
Agentic orchestration provides a chance to respond to incomplete evidence before answering. It can target multiple aspects of a question and make unresolved gaps visible. It also adds decisions that can fail: the planner may drift from the user’s intent, the validator may accept weak evidence, or another round may retrieve the same irrelevant passages.
The call structure makes the latency and cost trade-off concrete. In the normal path, Document Chat uses one generation call, plus a rewrite call when history exists. Folio uses a planner and validator call per round, followed by generation when accepted: three model calls for a first-round answer or seven for a third-round answer, before structured-output repairs and any provider-client retries. These are source-derived call counts, not timing or price measurements. Cost also depends on tokens per call, especially as evidence grows.
Both approaches retain a fundamental limitation: valid citation numbers do not prove that each claim is supported by the cited passage. A pre-generation evidence check cannot validate claims that have not yet been written. Where stronger assurance is needed, evaluate claim support after generation and apply deterministic domain checks or human review where appropriate.
| Decision factor | Conventional RAG fits when… | Bounded agentic RAG fits when… |
|---|---|---|
| Query shape | Most questions are direct lookups with a clear search target. | Questions combine several evidence needs or require follow-up searches. |
| Corpus quality | Relevant passages are well structured and easy to retrieve. | Relevant evidence exists but is scattered across documents or terminology. Neither approach repairs missing or incorrect source material. |
| Latency and cost | A short, predictable path matters more than exploring alternatives. | Users can wait for extra model calls in exchange for potentially better coverage. |
| Correctness expectations | A retrieved answer or clear abstention meets the need. | Explicit coverage checks and partial-answer handling are valuable, alongside independent evaluation. |
| Operational complexity | A small team needs a workflow that is easy to test and support. | The team can own state, budgets, intermediate failures, and tracing. |
| Typical starting point | Documentation assistants, FAQs, and routine support lookup. | Research assistants, cross-document comparison, and investigation with ambiguous evidence needs. |
The richer interface has its own engineering cost. Folio keeps accepted operations running in Python when the browser disconnects, and the browser can recover status and results by polling. That requires operation state and concurrency controls. These choices support a longer-running research experience, but they are application-design decisions rather than requirements of RAG itself.
Improve the baseline before adding a loop
An agent cannot recover text that was never extracted from a scanned document. It cannot infer an exception policy that nobody uploaded. A poor embedding model or badly chosen chunk boundaries can defeat both implementations.
Before adding orchestration, examine the failed retrievals. Metadata filters can separate document versions; lexical and vector retrieval can complement each other; reranking can improve which passages reach the generator. These changes can remain part of a fixed workflow. They deserve evaluation before interpreting every failure as a need for an agent.
There is also a useful intermediate design: a conventional path for routine questions with an escalation path for demonstrated coverage gaps. This would be an extension beyond these prototypes. Its escalation policy needs evaluation too, because a model can confidently decide that incomplete evidence is sufficient.
A practical adoption sequence
Start with a representative question set and identify the passages required to answer each question. Include direct lookups, conversational follow-ups, questions requiring several documents, and unanswerable questions. Add conflicting or stale documents so that success means more than finding text with similar wording.
Run the same corpus and questions through both approaches. Measure evidence coverage, supported answer claims, appropriate abstention, and whether citations actually support their associated claims. Record median and tail latency, token usage, model-call counts, and failures. Segment results by question type: an overall average can hide regressions on the simple questions that make up most traffic.
For the agentic path, also record rounds taken, repeated searches, new chunks per round, validator acceptance, and how often the final-round fallback produces an answer. A loop that adds no evidence while increasing confidence deserves scrutiny. Test malformed plans, provider failures, unsupported questions, and evidence embedded with instructions intended to redirect the model. Prompt instructions help, but tool and data-access boundaries should be enforced in code.
Use those results to choose the stopping policy, including whether partial answers are acceptable at all. Round and chunk limits are a starting point; a production design may also need end-to-end deadlines and token budgets. Log enough to diagnose failures while controlling access to document excerpts and personal information.
Conventional RAG is a useful default because it gives you a small system whose failures can be inspected. Add planning and evidence feedback when the failures show that the system needs another search strategy, then verify that the added work improves supported answers enough to justify its cost. The goal is an assistant that knows when its evidence is adequate, when another search is useful, and when it should ask for information it does not have.