RAG/Search
RAG or Fine Tuning? Fine-tune Embedding models for Retrieval Augmented Generation (RAG)
For customizing LLMs, in addition to RAG, another optimization technique is fine-tuning.
RAG or Fine Tuning? A simple feature comparision to decide which technique you should use!
For customizing LLMs, in addition to RAG, another optimization technique is fine-tuning.
๐ฅ๐๐ is akin to providing a textbook to the model, allowing it to retrieve information based on specific queries. This approach is suitable for scenarios where the model needs to address particular information retrieval tasks. However, RAG is not suitable for teaching the model to understand broad domains or learn new languages, formats, or styles.
๐๐ถ๐ป๐ฒ-๐๐๐ป๐ถ๐ป๐ด is similar to enabling students to internalize knowledge through extensive learning. Fine-tuning can enhance the performance of non-fine-tuned models and make interactions more efficient. It is particularly suitable for emphasizing existing knowledge in the base model, modifying or customizing the modelโs output, and providing complex directives to the model.
Sometimes it may not seem straightforward to choose one approach or the other, thatโs why this guide will help you to differentiate which technique fits better your use case!
RAG in Production: The importance of a Solid Data Strategy ๐ฅ
Retrieval-Augmented Generation (RAG) has become one of the hottest topics in Generative AI, providing powerful ways to enhance model responses with real-world data. But letโs be honest, without a solid data strategy, youโre setting yourself up for a meme-worthy fail. ๐
๐ ๐ช๐ต๐ ๐ฅ๐๐ ๐ก๐ฒ๐ฒ๐ฑ๐ ๐ฎ ๐๐ฎ๐๐ฎ ๐ฆ๐๐ฟ๐ฎ๐๐ฒ๐ด๐:
- ๐๐ฎ๐๐ฎ ๐ค๐๐ฎ๐น๐ถ๐๐: Garbage in, garbage out. Your model is only as good as the data it retrieves.
- ๐ฅ๐ฒ๐น๐ฒ๐๐ฎ๐ป๐ฐ๐ฒ: Ensure your data is relevant to your use case.
- ๐ฆ๐ฐ๐ฎ๐น๐ฎ๐ฏ๐ถ๐น๐ถ๐๐: Manage and scale your data efficiently to keep up with growing demands.
Remember, a well-thought-out data strategy is the backbone of any successful RAG implementation.
๐ ๐๐ผ๐ป๐ฐ๐น๐๐๐ถ๐ผ๐ป: Donโt let your RAG use case fall flat. Invest in your data strategy and watch your AI soar! ๐
Fine-Tuning Embedding Models for RAG: Significant Performance Gains
Retrieve: How can we improve RAG performance through embedding model fine-tuning? What techniques enable domain-specific optimization?
Embedding models are crucial for RAG applications, but general models often fall short of domain-specific tasks. Fine-tuning embedding models can significantly boost retrieval performance, as demonstrated in a comprehensive study using financial RAG applications.
Performance Improvements
| Metric | Improvement | Impact |
|---|---|---|
| Overall Performance | 7.4% to 22.55% boost | โฌ๏ธ Significant |
| Training Samples | Only 6.3k samples needed | โฌ๏ธ Efficient |
| Training Time | ~5 minutes on consumer GPUs | โก Fast |
| Model Size | 6x smaller with Matryoshka | โฌ๏ธ Efficient |
| Dimension Efficiency | 128-dim > 768-dim baseline | โฌ๏ธ Better |
Fine-Tuning Workflow
graph TB
A[Base Embedding Model] --> B[Domain Data]
B --> C[Fine-Tuning Process]
C --> D[Matryoshka Learning]
D --> E[Optimized Embeddings]
F[Synthetic Data] --> C
G[Evaluation] --> C
H[Sentence Transformers v3] --> C
E --> I[RAG System]
I --> J[Improved Retrieval]
style A fill:#e1f5ff
style C fill:#fff3cd
style E fill:#d4edda
style J fill:#f8d7da
Key Techniques
1. Matryoshka Representation Learning (MRL)
Innovate: MRL enables variable-dimension embeddings that maintain performance at smaller sizes.
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
from sentence_transformers import SentenceTransformer, losses
from sentence_transformers.training_args import BatchSamplers
# Initialize model
model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
# Matryoshka loss enables variable dimensions
train_loss = losses.MultipleNegativesRankingLoss(model)
# Training with Matryoshka
model.fit(
train_objectives=[(train_dataloader, train_loss)],
epochs=3,
output_path='./fine-tuned-embedding-model'
)
# Use different dimensions
embeddings_128 = model.encode(texts, output_value='sentence_embedding',
convert_to_numpy=True)[:, :128]
embeddings_256 = model.encode(texts, output_value='sentence_embedding',
convert_to_numpy=True)[:, :256]
Benefits:
- ๐ช 99% performance at 6x smaller size
- ๐ 128-dim model outperforms 768-dim baseline by 6.51%
- ๐พ Reduced storage and compute requirements
2. Synthetic Data Generation
Retrieve: Generate training data automatically for fine-tuning:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
from sentence_transformers import InputExample
import random
def generate_synthetic_pairs(base_texts, num_pairs=1000):
"""Generate synthetic query-document pairs"""
pairs = []
for _ in range(num_pairs):
# Create variations
query = create_query_variation(random.choice(base_texts))
document = find_relevant_document(query, base_texts)
pairs.append(InputExample(texts=[query, document], label=1.0))
return pairs
# Use synthetic data for fine-tuning
synthetic_pairs = generate_synthetic_pairs(financial_documents, num_pairs=6300)
3. Baseline Creation & Evaluation
During Training:
- โ Continuous evaluation
- โ Performance tracking
- โ Early stopping
- โ Best model selection
Implementation Example
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
from sentence_transformers import SentenceTransformer, losses, evaluation
from sentence_transformers.datasets import NoDuplicatesDataLoader
# Load base model
model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
# Prepare training data
train_examples = [
InputExample(texts=["Query about revenue", "Document about financial revenue"]),
# ... more examples
]
train_dataloader = NoDuplicatesDataLoader(train_examples, batch_size=16)
train_loss = losses.MultipleNegativesRankingLoss(model)
# Evaluation during training
evaluator = evaluation.InformationRetrievalEvaluator(
queries=test_queries,
corpus=test_corpus,
relevant_docs=test_relevant_docs
)
# Fine-tune
model.fit(
train_objectives=[(train_dataloader, train_loss)],
epochs=3,
warmup_steps=100,
evaluator=evaluator,
evaluation_steps=500,
output_path='./financial-embedding-model'
)
Results Summary
| Dimension | Performance | vs. Baseline | Storage |
|---|---|---|---|
| 128-dim | 6.51% better | Outperforms 768-dim | 6x smaller |
| 256-dim | Near-optimal | 99% of full performance | 3x smaller |
| 512-dim | Optimal | Full performance | 1.5x smaller |
| 768-dim | Baseline | Reference | Full size |
Use Case: Financial RAG
Dataset: NVIDIAโs 2023 SEC Filing dataset
Results:
- ๐ 7.4% to 22.55% performance improvement
- โฑ๏ธ Fast training (5 minutes on consumer GPUs)
- ๐งฌ Synthetic data generation
- ๐ช Matryoshka efficiency
Resources
๐ Original Article: https://www.philschmid.de/fine-tune-embedding-model-for-rag
๐ Code Repository: https://github.com/philschmid/deep-learning-pytorch-huggingface/blob/main/training/fine-tune-embedding-model-for-rag.ipynb
๐ Hugging Face RAG Documentation: https://huggingface.co/docs/transformers/model_doc/rag
Key Takeaways
Retrieve: Fine-tuning embedding models with domain-specific data can boost RAG performance by 7-22%, even with small datasets (6.3k samples).
Innovate: Techniques like Matryoshka Representation Learning enable efficient embeddings that maintain performance at smaller sizes, reducing computational requirements while improving results.
Curiosity โ Retrieve โ Innovation: Start with curiosity about improving RAG performance, retrieve knowledge about embedding fine-tuning techniques, and innovate by applying these methods to your specific domain.
Next Steps:
- Explore the implementation code
- Fine-tune embeddings for your domain
- Experiment with Matryoshka learning
- Measure performance improvements
How to Select an Embedding Model for Your RAG Application?
Embeddings form the foundation for achieving precise and contextually relevant LLM outputs across different tasks.
Which encoder you select to generate embeddings is a critical decision, hugely impacting the overall success of the RAG system. Low quality embeddings lead to poor retrieval.
When selecting an embedding model, consider the vector dimension, average retrieval performance, and model size.
Companies such as OpenAI, Cohere, and Voyage consistently release enhanced embedding models.
Different types of embeddings are designed to address unique challenges and requirements in different domains.
โฎ Dense embeddings are continuous, real-valued vectors that represent information in a high-dimensional space.
Curiosity: In the context of RAG applications, dense embeddings, such as those generated by models like OpenAIโs Ada or sentence transformers, contain non-zero values for every element.
โฎ Sparse embeddings, on the other hand, are representations where most values are zero, emphasizing only relevant information.
In RAG applications, sparse vectors are essential for scenarios with many rare keywords or specialized terms.
โฎ Multi-vector embedding models like ColBERT feature late interaction, where the interaction between query and document representations occurs late in the process, after both have been independently encoded.
โฎ Long documents have always posed a particular challenge for embedding models.
The limitation on maximum sequence lengths, often rooted in architectures like BERT, leads to practitioners segmenting documents into smaller chunks. Unfortunately, this segmentation can result in fragmented semantic meanings and misrepresentation of entire paragraphs.
โฎ Variable dimension embeddings are a unique concept built on Matryoshka Representation Learning (MRL).
MRL learns lower-dimensional embeddings that are nested into the original embedding, akin to a series of Matryoshka Dolls.
โฎ Code embeddings are a recent development used to integrate AI-powered capabilities into Integrated Development Environments (IDEs), fundamentally transforming how developers interact with codebases.
Curiosity: What insights can we retrieve from this? How does this connect to innovation in the field?
There are several factors that need to be considered while selecting an embedding model.
Know more about embeddings and models in this article: https://www.rungalileo.io/blog/mastering-rag-how-to-select-an-embedding-model
Translate to Korean
RAG ๋๋ ๋ฏธ์ธ ์กฐ์ ? ์ด๋ค ๊ธฐ์ ์ ์ฌ์ฉํด์ผ ํ๋์ง ๊ฒฐ์ ํ๊ธฐ ์ํ ๊ฐ๋จํ ๊ธฐ๋ฅ ๋น๊ต!
LLM์ ์ปค์คํฐ๋ง์ด์งํ๊ธฐ ์ํด RAG ์ธ์๋ ๋ ๋ค๋ฅธ ์ต์ ํ ๊ธฐ์ ์ด ๋ฏธ์ธ ์กฐ์ ์ ๋๋ค.
RAG๋ ๋ชจ๋ธ์ ๊ต๊ณผ์๋ฅผ ์ ๊ณตํ๋ ๊ฒ๊ณผ ์ ์ฌํ์ฌ ํน์ ์ฟผ๋ฆฌ๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ์ ๋ณด๋ฅผ ๊ฒ์ํ ์ ์์ต๋๋ค. ์ด ๋ฐฉ๋ฒ์ ๋ชจ๋ธ์ด ํน์ ์ ๋ณด ๊ฒ์ ์์ ์ ์ฒ๋ฆฌํด์ผ ํ๋ ์๋๋ฆฌ์ค์ ์ ํฉํฉ๋๋ค. ๊ทธ๋ฌ๋ RAG๋ ๋ชจ๋ธ์ด ๊ด๋ฒ์ํ ๋๋ฉ์ธ์ ์ดํดํ๊ฑฐ๋ ์๋ก์ด ์ธ์ด, ํ์ ๋๋ ์คํ์ผ์ ํ์ตํ๋๋ก ํ์ต์ํค๋ ๋ฐ๋ ์ ํฉํ์ง ์์ต๋๋ค.
๋ฏธ์ธ ์กฐ์ ์ ํ์๋ค์ด ๊ด๋ฒ์ํ ํ์ต์ ํตํด ์ง์์ ๋ด๋ฉดํํ ์ ์๋๋ก ํ๋ ๊ฒ๊ณผ ์ ์ฌํฉ๋๋ค. ๋ฏธ์ธ ์กฐ์ ์ ๋ฏธ์ธ ์กฐ์ ๋์ง ์์ ๋ชจ๋ธ์ ์ฑ๋ฅ์ ํฅ์์ํค๊ณ ์ํธ ์์ฉ์ ๋ณด๋ค ํจ์จ์ ์ผ๋ก ๋ง๋ค ์ ์์ต๋๋ค. ๊ธฐ๋ณธ ๋ชจ๋ธ์ ๊ธฐ์กด ์ง์์ ๊ฐ์กฐํ๊ณ , ๋ชจ๋ธ์ ์ถ๋ ฅ์ ์์ ํ๊ฑฐ๋ ์ฌ์ฉ์ ์ง์ ํ๊ณ , ๋ชจ๋ธ์ ๋ณต์กํ ์ง์๋ฌธ์ ์ ๊ณตํ๋ ๋ฐ ํนํ ์ ํฉํฉ๋๋ค.
๋๋ก๋ ํ ๊ฐ์ง ์ ๊ทผ ๋ฐฉ์ ๋๋ ๋ค๋ฅธ ์ ๊ทผ ๋ฐฉ์์ ์ ํํ๋ ๊ฒ์ด ๊ฐ๋จํ์ง ์์ ๊ฒ์ฒ๋ผ ๋ณด์ผ ์ ์์ผ๋ฏ๋ก ์ด ๊ฐ์ด๋๋ ์ฌ์ฉ ์ฌ๋ก์ ๋ ์ ํฉํ ๊ธฐ์ ์ ๊ตฌ๋ณํ๋ ๋ฐ ๋์์ด ๋ ๊ฒ์ ๋๋ค!
์์ฐ ํ์ฅ์์์ RAG: ๊ฒฌ๊ณ ํ ๋ฐ์ดํฐ ์ ๋ต๐ฅ์ ์ค์์ฑ
RAG(Retrieval-Augmented Generation)๋ ์ ๋๋ ์ดํฐ๋ธ AI์์ ๊ฐ์ฅ ์ธ๊ธฐ ์๋ ์ฃผ์ ์ค ํ๋๊ฐ ๋์์ผ๋ฉฐ, ์ค์ ๋ฐ์ดํฐ๋ก ๋ชจ๋ธ ์๋ต์ ํฅ์์ํฌ ์ ์๋ ๊ฐ๋ ฅํ ๋ฐฉ๋ฒ์ ์ ๊ณตํฉ๋๋ค. ๊ทธ๋ฌ๋ ์์งํ ๋งํด์ ๊ฒฌ๊ณ ํ ๋ฐ์ดํฐ ์ ๋ต์ด ์์ผ๋ฉด ๋ฐ์ ์ด์ธ๋ฆฌ๋ ์คํจ๋ฅผ ๋ง์ดํ๊ฒ ๋ฉ๋๋ค. ๐
๐ RAG์ ๋ฐ์ดํฐ ์ ๋ต์ด ํ์ํ ์ด์ :
- ๋ฐ์ดํฐ ํ์ง: ์ฐ๋ ๊ธฐ ์ ์ , ์ฐ๋ ๊ธฐ ๋ฐฐ์ถ. ๋ชจ๋ธ์ ๊ฒ์ํ๋ ๋ฐ์ดํฐ๋งํผ๋ง ์ฐ์ํฉ๋๋ค.
- ๊ด๋ จ์ฑ: ๋ฐ์ดํฐ๊ฐ ์ฌ์ฉ ์ฌ๋ก์ ๊ด๋ จ์ด ์๋์ง ํ์ธํฉ๋๋ค.
- ํ์ฅ์ฑ: ์ฆ๊ฐํ๋ ์์๋ฅผ ๋ฐ๋ผ์ก๊ธฐ ์ํด ๋ฐ์ดํฐ๋ฅผ ํจ์จ์ ์ผ๋ก ๊ด๋ฆฌํ๊ณ ํ์ฅํฉ๋๋ค.
์ ์คํ ๋ฐ์ดํฐ ์ ๋ต์ ์ฑ๊ณต์ ์ธ RAG ๊ตฌํ์ ์ค์ถ๋ผ๋ ์ ์ ๊ธฐ์ตํ์ญ์์ค.
๐ ๊ฒฐ๋ก : RAG ์ฌ์ฉ ์ฌ๋ก๊ฐ ์คํจํ์ง ์๋๋ก ํ์ญ์์ค. ๋ฐ์ดํฐ ์ ๋ต์ ํฌ์ํ๊ณ AI๊ฐ ๊ธ์ฆํ๋ ๊ฒ์ ์ง์ผ๋ณด์ญ์์ค! ๐
๋ฏธ์ธ ์กฐ์ ์ ๊ฒ์ ์๋๋ฅผ ํฌ๊ฒ ๋์ผ ์ ์์ต๋๋ค. ๐
์๋ฒ ๋ฉ ๋ชจ๋ธ์ RAG(Retrieval-Augmented Generation) ์ ํ๋ฆฌ์ผ์ด์ ์ ๋งค์ฐ ์ค์ํ์ง๋ง ์ผ๋ฐ ๋ชจ๋ธ์ ๋๋ฉ์ธ๋ณ ์์ ์ ๋ฏธ์น์ง ๋ชปํ๋ ๊ฒฝ์ฐ๊ฐ ๋ง์ต๋๋ค.
Matryoshka Representation Learning๊ณผ ๊ฐ์ ์ต์ ์ฐ๊ตฌ๋ฅผ ์ฌ์ฉํ์ฌ NVIDIA์ 2023 SEC Filing ๋ฐ์ดํฐ ์ธํธ๋ฅผ ์ฌ์ฉํ์ฌ ๊ธ์ต RAG ์ ํ๋ฆฌ์ผ์ด์ ์ฉ ์๋ฒ ๋ฉ ๋ชจ๋ธ์ ๋ฏธ์ธ ์กฐ์ ํ๋ ๋ฐฉ๋ฒ์ ๋ํ ์๋ก์ด ๋ธ๋ก๊ทธ๋ฅผ ๊ณต์ ํ๊ฒ ๋์ด ๊ธฐ์ฉ๋๋ค.
- ๐ ๋ฏธ์ธ ์กฐ์ ์ผ๋ก ๋จ 6.3k ์ํ๋ก 7.4%์์ 22.55%๊น์ง ์ฑ๋ฅ ํฅ์
- โ ๊ธฐ์ค ์์ฑ + ํ์ต ์ค ํ๊ฐ
- ๐งฌ ๋ฏธ์ธ ์กฐ์ ์ ์ฌ์ฉ๋๋ ์์ฑ๋ ํฉ์ฑ ๋ฐ์ดํฐ
- โฑ๏ธ ~10,000์ ๋ํ ๊ต์ก, ์๋น์์ฉ GPU์์ ๋จ 5๋ถ
- ๐ช Matryoshka๋ 6๋ฐฐ ๋ ์์ ํฌ๊ธฐ๋ก 99%์ ์ฑ๋ฅ์ ์ ์งํฉ๋๋ค.
- ๐ ๋ฏธ์ธ ์กฐ์ ๋ 128-dim ๋ชจ๋ธ์ ๊ธฐ์ค 768-dim๋ณด๋ค 6.51% ๋ ์ฐ์ํฉ๋๋ค.
- ๐ ์๋ก์ด ๋ฌธ์ฅ ๋ณํ๊ธฐ v3 ์ฌ์ฉ
๋น๋ํ๋ฌ ๊ฐ์ธ์! ๐ค
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me
