Conventional RAG vs. Agentic RAG: When Retrieval Becomes Research
Retrieval-augmented generation (RAG) gives a language model evidence to work with: ingest documents, split their text into chunks, embed those chunks, retrieve relevant passages for a question, and generate an answer that cites them. The interesting engineering question appears when the first search returns only part of what the user needs. Should the application answer from those passages, admit the gap, or search again with a different question? I explored that choice through two local document-chat prototypes: Document Chat, built with Streamlit, and Folio, a Next.js/FastAPI application with a LangGraph research workflow. They share the same basic document-library design. Their main difference is how they decide what to retrieve and when to stop. ...