Review/Trends
๐๐ ๐๐ง๐ญ๐ข๐ ๐๐๐ โจ new cookbook
Curiosity: What if RAG systems could think more like humansโquestioning their own retrievals, reformulating queries, and iterating until they find theโฆ
Agentic RAG Cookbook: Improving RAG with Agent Systems
Curiosity: What if RAG systems could think more like humansโquestioning their own retrievals, reformulating queries, and iterating until they find the right answer? What happens when we give RAG the ability to retrieve, critique, and retrieve again?
A new cookbook demonstrates how to easily improve RAG with an agent system using Transformers Agents. This approach addresses key limitations of vanilla RAG by making systems more intelligent and self-correcting.
Vanilla RAG Limitations
Retrieve: Vanilla RAG systems have fundamental limitations that impact performance.
Key Limitations:
| Limitation | Description | Impact |
|---|---|---|
| Single Retrieval | Retrieves documents only once | โ ๏ธ Poor quality if initial retrieval fails |
| Suboptimal Similarity | Uses user query as reference | โ ๏ธ Questions vs. statements mismatch |
| No Self-Correction | Cannot refine or re-retrieve | โ No improvement mechanism |
Problem Details:
- User queries are typically questions
- Relevant documents use affirmative statements
- Similarity scores are downgraded
- No opportunity for improvement
Vanilla RAG vs. Agentic RAG
| Aspect | Vanilla RAG | Agentic RAG |
|---|---|---|
| Retrieval Strategy | Single retrieval pass | Iterative retrieval with critique |
| Query Handling | Direct user query | Query reformulation & optimization |
| Self-Correction | โ No | โ Yes - can re-retrieve if needed |
| Performance | Baseline (70.0%) | Improved (+8.5% = 78.5%) |
| Latency | Lower (1 LLM call) | Higher (multiple LLM calls) |
| Quality | โ ๏ธ Limited | โฌ๏ธ Better |
Agentic RAG Solution
Innovate: Making a RAG agentโsimply, an agent armed with a retriever toolโalleviates both problems!
Key Capabilities:
- โ Query Reformulation: Agent formulates optimized queries
- โ Self-Query: Agent critiques content and re-retrieves if needed
Architecture:
graph TD
A[User Query] --> B[Agent: Query Reformulation]
B --> C[Retrieve Documents]
C --> D[Agent: Critique Retrieved Content]
D --> E{Content<br/>Relevant?}
E -->|No| B
E -->|Yes| F[Generate Answer]
F --> G[Final Response]
style B fill:#e1f5ff
style D fill:#fff3cd
style F fill:#d4edda
style E fill:#f8d7da
Performance Comparison
Retrieve: Evaluation with LLM-as-a-judge (Llama-3-70B) shows significant improvement.
| Metric | Vanilla RAG | Agentic RAG | Improvement |
|---|---|---|---|
| Accuracy Score | 70.0% | 78.5% | +8.5% ๐ช |
| LLM Calls | 1 | 3-5 | Higher latency |
| Self-Correction | โ | โ | Better quality |
| Query Optimization | โ | โ | Better retrieval |
Trade-offs:
- โฌ๏ธ Better quality (+8.5%)
- โ ๏ธ Higher latency (multiple LLM calls)
- โ๏ธ Balance quality vs. speed needed
Sample Agentic RAG Implementation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
from transformers import pipeline
from langchain.agents import create_react_agent
from langchain.tools import Tool
# Create retrieval tool
retrieval_tool = Tool(
name="retrieve_documents",
func=vector_store.similarity_search,
description="Retrieves relevant documents for a query"
)
# Create agent with retrieval tool
agent = create_react_agent(
llm=llm,
tools=[retrieval_tool],
prompt=agent_prompt
)
# Agent workflow
def agentic_rag(query):
# Step 1: Query reformulation
reformulated_query = agent.run(
f"Reformulate this query for better retrieval: {query}"
)
# Step 2: Retrieve documents
docs = retrieval_tool.run(reformulated_query)
# Step 3: Critique and potentially re-retrieve
critique = agent.run(
f"Critique these documents for relevance to: {query}\n{docs}"
)
if "not relevant" in critique.lower():
# Re-retrieve with different strategy
docs = retrieval_tool.run(query, k=10) # Get more docs
# Step 4: Generate answer
answer = llm.generate(
context=docs,
question=query
)
return answer
๐๐ถ๐๐ฐ๐ผ๐๐ฒ๐ฟ ๐๐ต๐ฒ ๐ฐ๐ผ๐ผ๐ธ๐ฏ๐ผ๐ผ๐ธ ๐
๐๐ด๐ฒ๐ป๐๐ถ๐ฐ ๐๐ฎ๐๐ฎ ๐ฎ๐ป๐ฎ๐น๐๐๐: ๐ฑ๐ฟ๐ผ๐ฝ ๐๐ผ๐๐ฟ ๐ฑ๐ฎ๐๐ฎ ๐ณ๐ถ๐น๐ฒ, ๐น๐ฒ๐ ๐๐ต๐ฒ ๐๐๐ ๐ฑ๐ผ ๐๐ต๐ฒ ๐ฎ๐ป๐ฎ๐น๐๐๐ถ๐ ๐โ๏ธ
Need to make quick exploratory data analysis? โก๏ธ Get help from an agent.
I was impressed by Llama-3.1โs capacity to derive insights from data. Given a csv file, it makes quick work of exploratory data analysis and can derive interesting insights.
On the data from the Kaggle titanic challenge, that records which passengers survived the Titanic wreckage, it was able by itself to derive interesting trends like โpassengers that paid higher fares were more likely to surviveโ or โsurvival rate was much higher for women than menโ.
The cookbook even lets the agent built its own submission to the challenge, and it ranks under 3,000 out of 17,000 submissions: ๐ not bad at all!
- Try it for yourself in this Space demo ๐ https://lnkd.in/gzaqQ3rT
- Read the cookbook to dive deeper ๐ https://lnkd.in/gXx3-AyH
Translate to Korean
๋ฐฉ๊ธ Transformers Agents๋ฅผ ์ฌ์ฉํ์ฌ ์์ด์ ํธ ์์คํ ์ผ๋ก RAG(Retrieval Augmented Generation)๋ฅผ ์ฝ๊ฒ ๊ฐ์ ํ๋ ๋ฐฉ๋ฒ์ ๋ณด์ฌ์ฃผ๋ ์๋ก์ด ์ฟก๋ถ์ ์ถํํ์ต๋๋ค.
Vanilla RAG์๋ ๋ค์๊ณผ ๊ฐ์ ์ ํ ์ฌํญ์ด ์์ต๋๋ค.
- โค ์์ค ๋ฌธ์๋ฅผ ํ ๋ฒ๋ง ๊ฒ์ํฉ๋๋ค: ๊ฒ์๋ ๋ฌธ์๊ฐ ์ถฉ๋ถํ ๊ด๋ จ์ฑ์ด ์์ผ๋ฉด ์์ฑ์ด ๋๋น ์ง ๊ฒ์ ๋๋ค.
- โค ์๋ฏธ๋ก ์ ์ ์ฌ์ฑ์ ์ฌ์ฉ์ ์ฟผ๋ฆฌ๋ฅผ ์ฐธ์กฐ๋ก ์ฌ์ฉํ์ฌ ๊ณ์ฐ๋๋ฉฐ, ์ด๋ ์ข ์ข ์ฐจ์ ์ฑ ์ ๋๋ค: ์๋ฅผ ๋ค์ด, ์ฌ์ฉ์ ์ฟผ๋ฆฌ๋ ๋๋ถ๋ถ ์ง๋ฌธ์ด๊ณ ์ค์ ๋ต๋ณ์ ํฌํจํ๋ ๋ฌธ์๋ ๊ธ์ ์์ฑ์ด๋ฏ๋ก ์ ์ฌ์ฑ ์ ์๋ ์๋ฌธ ํ์์ ๊ด๋ จ์ฑ์ด ๋ฎ์ ์์ค ๋ฌธ์์ ๋นํด ๋ค์ด๊ทธ๋ ์ด๋๋์ด ๊ด๋ จ ๋ฌธ์๋ฅผ ์ ํํ์ง ์์ ์ํ์ด ์์ต๋๋ค.
RAG ์์ด์ ํธ๋ฅผ ๋ง๋ค๋ฉด(์์ฃผ ๊ฐ๋จํ๊ฒ, ๋ฆฌํธ๋ฆฌ๋ฒ ๋๊ตฌ๋ก ๋ฌด์ฅํ ์์ด์ ํธ) ์ด ๋ ๊ฐ์ง ๋ฌธ์ ๋ฅผ ๋ชจ๋ ์ํํ ์ ์์ต๋๋ค!
- โ ์ฟผ๋ฆฌ ์์ฒด๋ฅผ ๊ณต์ํํฉ๋๋ค(์ฟผ๋ฆฌ ์ฌ๊ตฌ์ฑ).
- โ ํ์ํ ๊ฒฝ์ฐ ๋ค์ ๊ฒ์ํ ์ฝํ ์ธ ๋นํ(์์ฒด ์ฟผ๋ฆฌ)Critique the content to re-retrieve if needed (self-query)
์ด ์์ด์ ํธ ์ค์ ์ด ๊ฒฐ๊ณผ๋ฅผ ์ผ๋ง๋ ๊ฐ์ ํฉ๋๊น? ์๋ฆฌ์ฑ ์ Llama-3-70B๋ฅผ ์ฌ์ฉํ๋ LLM-as-a-judge์ ํ๊ฐ ๋ถ๋ถ์ ์ถ๊ฐํ์ต๋๋ค. ๋ฐ๋๋ผ์์ ์์ด์ ํธ RAG๋ก ์ ํํ๋ฉด ์ ์๊ฐ 8.5% ์ฆ๊ฐํฉ๋๋ค! ๐ช (70.0%์์ 78.5%๋ก)
ํ์ง๋ง ํ ๊ฐ์ง ์ค์ํ ๋จ์ ์, ์์คํ ์ด 1์ด ์๋ ์ฌ๋ฌ LLM ํธ์ถ์ ํ๊ธฐ ๋๋ฌธ์ RAG ์์คํ ์ ๋ฐํ์๋ ์ฆ๊ฐํ๋ค๋ ๊ฒ์ ๋๋ค. ์ ์ ํ ์ ์ถฉ์์ ์ฐพ์์ผ ํฉ๋๋ค!
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me