Operating this agentic RAG pattern means treating the supplied template as the retrieval baseline, then placing an enforced citation gate and a rehearsed restart path around it. The template defines neither control: it tells the agent to search its knowledge and include sources, streams events, and renders citation URLs only when a streamed chunk provides them (application code). For us, done means the rendered answer has been checked against its evidence and, after an interruption, real state has been checked before any step is repeated.
What it solves
The original code for Agentic RAG with Reasoning provides a compact reference for joining retrieval, tool-enabled reasoning and streamed output. Its README describes a RAG system that exposes an agent’s step-by-step reasoning process through Agno, Gemini and OpenAI.
The configured roles are direct: Gemini 2.5 Flash handles language processing, an OpenAI embedding model supports vector search, and ReasoningTools provides analysis tools (template README). Agno’s search_knowledge=True adds a knowledge-search tool—the mechanism its documentation identifies as agentic RAG—and is enabled by default when an agent receives knowledge (Agno’s Agentic RAG with LanceDB guide).
The resulting architecture is useful for understanding the path from a stored document to a sourced answer. It does not solve citation enforcement, restart recovery or production governance by itself. The containing awesome-llm-apps repository uses the Apache License, Version 2.0.
How to run it
From the template directory, install the dependencies listed in requirements.txt:
pip install -r requirements.txt
The dependency file contains:
streamlit
agno>=2.2.10
lancedb
openai
python-dotenv
Those entries come directly from the template’s requirements file. Start the application with the documented command:
streamlit run rag_reasoning_agent.py
The README says to enter the Google API key in the first field and the OpenAI API key in the second; both keys are required. It defines these UI fields rather than environment variable names, so do not assume a .env contract that the README does not document.
The relevant knowledge configuration in the application code is:
kb = Knowledge(
vector_db=LanceDb(
uri="tmp/lancedb",
table_name="agno_docs",
search_type=SearchType.vector,
embedder=OpenAIEmbedder(
api_key=openai_key,
),
),
)
The agent settings are:
search_knowledge=True,
tools=[ReasoningTools(add_instructions=True)],
instructions=[
"Include sources in your response.",
"Always search your knowledge before answering the question.",
],
That creates a local LanceDB database at tmp/lancedb and selects vector search. Agno’s LanceDB documentation confirms that a filesystem path creates a local database. The code does not specify an embedding model, so the documented default for OpenAIEmbedder applies: text-embedding-3-small with 1536 dimensions (OpenAI Embedder documentation).
The README does not document a corpus-ingestion command. Check the repository’s current code before assuming that launching Streamlit populates agno_docs.
Check it worked
Start with one representative query whose answer should be present in a controlled corpus. Read the output where a user sees it, rather than relying only on the process status.
Check that the knowledge-search path ran, inspect any reasoning events that were emitted, and note whether the answer contains citation URLs. ReasoningTools lets the model decide when to produce planning and analysis notes, so reasoning output is diagnostic rather than a mandatory field on every query (Reasoning Tools documentation).
Also confirm that tmp/lancedb contains the expected corpus before judging answer quality. An empty or stale store can look like a reasoning failure even though the agent is working correctly. Finally, submit a question that your controlled corpus cannot answer; the citation gate described below should reject that case.
What we observed
This page is based on reading the template’s code and documentation, and we have not run it.
The code centres on an Agno Knowledge object backed by LanceDB, with an OpenAI embedder and vector search. The agent explicitly enables knowledge search, loads ReasoningTools, instructs the model to search before answering and to include sources, then calls agent.run with streaming and all events enabled (implementation).
Agno describes reasoning notes as ordinary tool data. Applications control whether tool events, stored messages, reasoning content and logs are displayed or retained; the toolkit is not a privacy boundary (Reasoning Tools documentation). That matters because streaming reasoning is not the same as validating it.
Citation handling is conditional. The implementation collects citations.urls when a streamed chunk exposes them and shows the citations only when at least one URL is available (application code). The code does not verify that a cited page exists, that its content supports the claim, or that an answer without citations should be rejected.
The template also does not define a recovery drill, a restart-state check, a hybrid-search configuration or a citation gate. Its code can inform those controls, but it does not establish their behaviour or provide measured reliability, cost or latency.
Where it breaks
An instruction to include sources is not enforcement. A displayed URL is also not proof: it may be stale, inaccessible, unrelated or merely adjacent to the claim. The citation path therefore needs an external pass/fail check.
The retrieval configuration is narrower than the underlying adapter. This template selects SearchType.vector. Agno’s LanceDB adapter also supports full-text and hybrid search, but that does not mean the template enables them (LanceDB overview). If you need hybrid search RAG, check the current Agno documentation for the supported configuration rather than guessing at a flag.
The local database path creates another operational boundary. Our sessions often share working directories, and we have had generic temporary files overwritten when separate sessions reused the same name. That incident does not establish a LanceDB defect; it does make an isolated working copy or an explicitly reviewed path change a control we require around concurrent runs. The README supplies no relocation flag.
Finally, agno>=2.2.10 is a minimum constraint while the other listed dependencies use plain package names (requirements file). A later installation can therefore differ from the dependency set reviewed here.
Agentic RAG operating checklist
Citation gates
We would place the citation gate outside the template’s prompt and rendering logic. For each answer intended for a reader, the gate should:
- Require a source for every material claim.
- Open the cited evidence and verify that the relevant passage supports the claim.
- Reject an answer when its sources are missing, inaccessible or unrelated.
- Return a clear failure state instead of silently accepting the model’s output.
- Test the gate with a known unsupported case that must fail.
The negative test is essential. We have had a copied verification check report success without testing the relevant behaviour. A gate that has never rejected anything is not evidence that it works.
Recovery drills
We would rehearse recovery with a controlled corpus and an answer that has no public side effect. A useful drill interrupts the run, restarts the environment and then asks what actually survived.
After restart, check the application, the LanceDB path, the expected table contents and any result that was expected to persist. Do not repeat a step merely because the previous process disappeared. In our own restart incidents, scratch files and background processes were gone while some external work had already landed; checking real state prevented unnecessary repetition.
Use an isolated scratch location for each concurrent run, record what finished and what remains, and rerun only the missing part. The drill passes only when the rendered answer and retained state have been read back and the next action is unambiguous.