RAG/Search
๐ก CRAG (Comprehensive RAG) is a new RAG benchmark dataset
Curiosity: How can we create a realistic benchmark for RAG systems? What makes CRAG more challenging than existing datasets?
CRAG: Comprehensive RAG Benchmark Dataset
Curiosity: How can we create a realistic benchmark for RAG systems? What makes CRAG more challenging than existing datasets?
CRAG (Comprehensive RAG) is a new benchmark dataset that provides robust and challenging test cases for evaluating RAG and QA systems. Even GPT-4 struggles, achieving less than 34% accuracy, highlighting the challenge.
The Problem
Retrieve: Existing RAG datasets have limitations.
| Issue | Description | Impact |
|---|---|---|
| Lack of Diversity | Limited question types | โ ๏ธ Incomplete evaluation |
| Complexity Gap | Donโt represent real-world QA | โ ๏ธ Suboptimal assessment |
| Evaluation Issues | Poor performance metrics | โ ๏ธ Unreliable results |
Result: Suboptimal performance evaluation of RAG systems.
CRAG Dataset Overview
Innovate: Comprehensive benchmark for RAG evaluation.
graph TB
A[CRAG Dataset] --> B[4,409 QA Pairs]
A --> C[5 Domains]
A --> D[8 Question Categories]
A --> E[Mock APIs]
A --> F[Score System]
E --> E1[Web Search]
E --> E2[KG Search]
F --> F1[Penalize Hallucinations]
F --> F2[Reliable Evaluation]
style A fill:#e1f5ff
style B fill:#fff3cd
style F fill:#d4edda
Dataset Features
Retrieve: CRAGโs comprehensive features.
| Feature | Details | Benefit |
|---|---|---|
| QA Pairs | 4,409 pairs | โฌ๏ธ Large scale |
| Domains | 5 domains | โฌ๏ธ Diversity |
| Categories | 8 question types | โฌ๏ธ Coverage |
| Complexity | Simple facts to complex queries | โฌ๏ธ Real-world |
| Mock APIs | Web and KG search | โฌ๏ธ Realistic |
| Score System | Penalizes hallucinations | โฌ๏ธ Reliable |
Evaluation Tasks
Innovate: Comprehensive task coverage.
Task Types:
- Web Retrieval: Realistic web search scenarios
- Structured Querying: Knowledge Graph queries
- Summarization: Multi-document summarization
Coverage: From simple facts to complex multi-hop queries.
Performance Results
Retrieve: CRAG reveals significant challenges.
| System | Accuracy | Notes |
|---|---|---|
| Advanced LLMs (GPT-4) | <34% | Highlights challenge |
| Direct RAG | 44% | Needs improvement |
| SOTA Industry RAG | 63% | Without hallucination |
Key Findings:
- Even best LLMs struggle (<34%)
- Direct RAG only reaches 44%
- Industry solutions achieve 63% (best case)
Score System Innovation
Innovate: Better evaluation through hallucination penalties.
Key Feature: Penalizes hallucinated answers more than missing answers
Benefits:
- โ Encourages accuracy over completeness
- โ Reduces false information
- โ More reliable evaluation
- โ Better reflects real-world needs
Key Takeaways
Retrieve: CRAG provides a comprehensive benchmark with 4,409 QA pairs across 5 domains and 8 categories, including realistic retrieval scenarios and a score system that penalizes hallucinations.
Innovate: By creating a challenging benchmark that even GPT-4 struggles with (<34% accuracy), CRAG encourages development of more advanced RAG solutions, with industry SOTA reaching 63% accuracy.
Curiosity โ Retrieve โ Innovation: Start with curiosity about RAG evaluation, retrieve insights from CRAGโs comprehensive approach, and innovate by building RAG systems that can handle the complexity and diversity of real-world QA tasks.
Next Steps:
- Read the full paper
- Test your RAG on CRAG
- Analyze performance gaps
- Improve your systems
Translate to Korean
๐ ๋ค์์ RAG ํ์ดํ๋ผ์ธ์ ํ ์คํธํ ์ ์๋ ์ด๋ ค์ด ์ค์ ๋ฒค์น๋งํฌ์ ๋๋ค! GPT-4์ ๊ฐ์ LLM์กฐ์ฐจ๋ 34% ๋ฏธ๋ง์ ์ ํ๋๋ฅผ ๋ฌ์ฑํ๋ ๋ฐ ์ด๋ ค์์ ๊ฒช๊ณ ์์ต๋๋ค.
๊ธฐ์กด RAG ๋ฐ์ดํฐ ์ธํธ๋ ๋ค์์ฑ์ด ๋ถ์กฑํ๊ณ ์ค์ QA ์์ ์ ๋ณต์ก์ฑ์ ๋ํ๋ด์ง ๋ชปํ์ฌ ์ฑ๋ฅ ํ๊ฐ๊ฐ ์ต์ ํ๋์ง ์์ต๋๋ค.
๐ก CRAG(Comprehensive RAG)๋ RAG ๋ฐ QA ์์คํ ์ ํ๊ฐํ๊ธฐ ์ํ ๊ฐ๋ ฅํ๊ณ ๋์ ์ ์ธ ํ ์คํธ ์ผ์ด์ค๋ฅผ ์ ๊ณตํ๋ ์๋ก์ด RAG ๋ฒค์น๋งํฌ ๋ฐ์ดํฐ ์ธํธ๋ก, ์ ๋ขฐํ ์ ์๋ LLM ๊ธฐ๋ฐ ์ง๋ฌธ ๋ต๋ณ์ ๋ฐ์ ์ ์ฅ๋ คํฉ๋๋ค.
- โณ CRAG์๋ 5๊ฐ ๋๋ฉ์ธ๊ณผ 8๊ฐ ์ง๋ฌธ ๋ฒ์ฃผ์ ๊ฑธ์ณ 4,409๊ฐ์ QA ์์ด ํฌํจ๋์ด ์์ผ๋ฉฐ, ๊ฐ๋จํ ์ฌ์ค๋ถํฐ ๋ณต์กํ ์ฟผ๋ฆฌ๊น์ง ๋ค๋ฃน๋๋ค.
- โณ ์น ๋ฐ KG(Knowledge Graph) ๊ฒ์์ ์ํ ๋ชจ์ API๋ฅผ ์ ๊ณตํ์ฌ ํ์ค์ ์ธ ๊ฒ์ ์๋๋ฆฌ์ค๋ฅผ ์ ๊ณตํฉ๋๋ค.
- โณ ๋ฏธ๊ฒฐ ๋ต๋ณ๋ณด๋ค ํ๊ฐ์ ๊ฑธ๋ฆฐ ๋ต๋ณ์ ๋ ๋ง์ ํ๋ํฐ๋ฅผ ์ฃผ๋ ์ ์ ์์คํ ์ ๋์ ํ์ฌ ์ ๋ขฐํ ์ ์๋ ํ๊ฐ๋ฅผ ๋ณด์ฅํฉ๋๋ค.
- โณ ์น ๊ฒ์, ๊ตฌ์กฐ์ ์ฟผ๋ฆฌ ๋ฐ ์์ฝ์ ์ํ ์์ ์ ์ ๊ณตํ์ฌ RAG ์๋ฃจ์ ์ ์ข ํฉ์ ์ผ๋ก ํ๊ฐํ ์ ์์ต๋๋ค.
๊ธฐ์ฌ
- ๐ ๊ฐ์ฅ ์ง๋ณด๋ LLM์ ๋ค์๊ณผ ๊ฐ์ ์ฑ๊ณผ๋ฅผ ๊ฑฐ๋ก๋๋ค. <34% accuracy on CRAG, highlighting the challenge.
- ๐ Direct application of RAG improves accuracy to only 44%, indicating the need for more advanced solutions.
- ๐ State-of-the-art industry RAG solutions reach 63% accuracy without hallucination.
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me