LLM/Model & Papers
Release Chameleon Model
Curiosity: Will Chameleon be Meta Llama 4? ๐ฆ ๐ฆ Meta proposes โChameleon: Mixed-Modal Early-Fusion Foundation Modelsโ with a unified approach for fullyโฆ
![["Chameleon Architecture"]](/assets/img/llm/LLM_chameleon.jpeg)
Chameleon: Mixed-Modal Early-Fusion Foundation Models
Curiosity: Will Chameleon be Meta Llama 4? ๐ฆ ๐ฆ Meta proposes โChameleon: Mixed-Modal Early-Fusion Foundation Modelsโ with a unified approach for fully token-based representations of both image and text. No Encoders or connectors. ๐
Architecture Overview
Retrieve: Chameleon uses a unified token-based approach for multimodal understanding and generation.
graph TB
A[Input] --> B[Image Tokenizer]
A --> C[Text Tokenizer]
B --> D[1024 Image Tokens]
C --> E[Text Tokens]
D --> F[Unified Token Sequence]
E --> F
F --> G[Llama 2 Decoder]
G --> H[Output: Text/Image]
style A fill:#e1f5ff
style F fill:#fff3cd
style H fill:#d4edda
Implementation Details
| Step | Component | Details |
|---|---|---|
| 1. Tokenizers | Image + Text | Image: 512ร512 โ 1024 tokens (codebook 8192) Text: BPE vocab 65,536 (includes image tokens) |
| 2. Architecture | Llama 2 Decoder | Query-key normalization Layer norm reordering Stabilized mixed-modal training |
| 3. Pretraining Stage 1 | 80% of training | Text-only: 2.9T tokens Text-image: 1.4B pairs/1.5T tokens Interleaved: 400B tokens |
| 4. Pretraining Stage 2 | 20% of training | Higher quality data Instruction data Half dataset size |
| 5. Fine-tuning | Final stage | ~1.8M samples ~100k vision samples |
Training Data Breakdown
pie title Training Data Distribution
"Text-only (2.9T)" : 60
"Text-Image (1.5T)" : 31
"Interleaved (400B)" : 9
Key Insights
Retrieve: Chameleonโs unified token-based approach enables native multimodal understanding and generation.
| Insight | Description | Impact |
|---|---|---|
| Unified Tokens | No encoders/connectors | โฌ๏ธ Native multimodal generation |
| Training Scale | 9.2T tokens, 2.1 epochs | โฌ๏ธ Strong performance |
| Code Data | Improved reasoning | โฌ๏ธ Text-only tasks |
| Scaling Challenge | Difficult above 8B/1T | โ ๏ธ Training stability |
| High-Quality Data | Last 20% crucial | โฌ๏ธ Significant boost |
| Performance | Outperforms competitors | โฌ๏ธ Strong results |
Performance Comparison
Innovate: Chameleon-34B achieves competitive performance across benchmarks.
Text Tasks:
- Outperforms Llama2-70B
- Approaches Mixtral 8x7B/Gemini-Pro
- Strong on GSM8K, MATH, MMLU
Vision Tasks:
- Outperforms Flamingo-80B and IDEFICS-80B on MS-COCO
- Matches performance on Flickr30k
Multimodal Evaluation:
- 60.4% win rate vs. Gemini-Pro
- 51.6% win rate vs. GPT-4V
Comparison with Previous MLLMs
| Model | Architecture | Multimodal Generation |
|---|---|---|
| Idefics, GPT-4v, Flamingo | Encoders + Connectors | โ Limited |
| Chameleon | Unified Tokens | โ Native support |
Key Advantage: Chameleon can generate both text and images using discrete tokens, enabling true multimodal document generation.
Key Takeaways
Retrieve: Chameleon demonstrates that unified token-based representations can achieve strong multimodal performance without separate encoders or connectors.
Innovate: By using discrete tokens for both images and text, Chameleon enables native multimodal understanding and generation, approaching GPT-4oโs capabilities with a simpler architecture.
Curiosity โ Retrieve โ Innovation: Start with curiosity about unified multimodal models, retrieve insights from Chameleonโs token-based approach, and innovate by building applications that leverage native multimodal generation.
Next Steps:
- Read the full paper
- Explore Chameleon architecture
- Compare with GPT-4o
- Build multimodal applications
Note: Chameleon looks to be closer to OpenAI GPT-4o than Uni-MoE (shared yesterday) with its native multi-modal tokens. ๐ก
Translate to Korean
Chameleon: Mixed-Modal Early-Fusion Foundation Models
์นด๋ฉ๋ ์จ์ ๋ผ๋ง 4Meta ๋ ๊น์? ๐ฆ ๐ฆ Meta๋ ์ด๋ฏธ์ง์ ํ ์คํธ ๋ชจ๋๋ฅผ ์์ ํ ํ ํฐ ๊ธฐ๋ฐ์ผ๋ก ํํํ๊ธฐ ์ํ ํตํฉ ์ ๊ทผ ๋ฐฉ์์ ํตํด โChameleon: Mixed-Modal Early-Fusion Foundation Modelsโ๋ฅผ ์ ์ํฉ๋๋ค. ์ธ์ฝ๋ ๋๋ ์ปค๋ฅํฐ๊ฐ ์์ต๋๋ค. ๐
Implementation:
- 1๏ธโฃ ํ๋ จ๋ 2๊ฐ์ ํ ํฌ๋์ด์ , 512 ร 512 ์ด๋ฏธ์ง๋ฅผ ์ฝ๋๋ถ(8192)์์ 1024๊ฐ์ ํ ํฐ์ผ๋ก ์ธ์ฝ๋ฉํ๋ ์ด๋ฏธ์ง ํ ํฌ๋์ด์ ์ 8192 ์ด๋ฏธ์ง ์ฝ๋๋ถ ํ ํฐ์ ํฌํจํ๋ 65,536์ ์ดํ๋ฅผ ๊ฐ์ง BPE.
- 2๏ธโฃ๋ Llama 2๋ฅผ ๊ธฐ๋ฐ์ผ๋ก ํ๋ ๋์ฝ๋ ์ํคํ ์ฒ๋ฅผ ์ฌ์ฉํ์ง๋ง ์ฟผ๋ฆฌ ํค ์ ๊ทํ ๋ฐ ๋ ์ด์ด ๊ท๋ฒ์ ์ฌ์ ๋ ฌ์ ํตํฉํ์ฌ ํผํฉ ๋ชจ๋ฌ ์ค์ ์์ ํ๋ จ์ ์์ ํํฉ๋๋ค.
- 3๏ธโฃ ํ ์คํธ ์ ์ฉ(Llama 2, CodeLlama โ 2.9T ํ ํฐ), ํ ์คํธ ์ด๋ฏธ์ง(1.4B ์/1.5T ํ ํฐ), ํ ์คํธ/์ด๋ฏธ์ง ์ธํฐ๋ฆฌ๋ธ(400B ํ ํฐ)์ ๋ํ ์ฌ์ ํ์ต 1๋จ๊ณ(80%);
- 4๏ธโฃ ์ฌ์ ํ์ต 2๋จ๊ณ (20%) ์ฒซ ๋ฒ์งธ ๋จ๊ณ์ ๋ฐ์ดํฐ ์ธํธ๋ฅผ ์ ๋ฐ์ผ๋ก ์ค์ด๊ณ ๋ ๋์ ํ์ง์ ๋ฐ์ดํฐ์ ์ง์นจ ๋ฐ์ดํฐ๋ฅผ ํฌํจํฉ๋๋ค.
- 5๏ธโฃ ~100k ๋น์ ์ํ๋ก ~180๋ง ๊ฐ์ ์ํ์ ๋ฏธ์ธ ์กฐ์ .
Insights:
- ๐ ์ด์ MLLM(Idefics, GPT-4v, Flamingo)์ ๋ฉํฐ๋ชจ๋ฌ๋ฆฌํฐ๋ฅผ ์ํด ์ธ์ฝ๋์ ์ปค๋ฅํฐ๋ฅผ ์ฌ์ฉํ๊ธฐ ๋๋ฌธ์ ๋ฉํฐ๋ชจ๋ฌ ๋ฌธ์(์ด๋ฏธ์ง + ํ ์คํธ ์ถ๋ ฅ)๋ฅผ ์์ฑํ๋ ๊ธฐ๋ฅ์ด ์ ํ๋์์ต๋๋ค.
- ๐ฆ ์นด๋ฉ๋ ์จ์ ๊ฐ๋ณ ํ ํฐ์ ์ฌ์ฉํ์ฌ ํ ์คํธ์ ์ด๋ฏธ์ง๋ฅผ ๋ชจ๋ ์ดํดํ๊ณ ์์ฑํ ์ ์์ต๋๋ค
- ๐ Chameleon-34B๋ ์ด 9.2T ํ ํฐ์ ๋ํด ์ ์ฒด ํ๋ จ ๋ฐ์ดํฐ ์ธํธ์์ 2.1 epoch ๋์ ํ๋ จํ์ต๋๋ค.
- ๐ง ์ฝ๋ ๋ฐ์ดํฐ๋ ํ ์คํธ ์ ์ฉ ์ถ๋ก ์์ ์ฑ๋ฅ์ ๊ฐ์ ํ์ต๋๋ค.
- โ๏ธ ์นด๋ฉ๋ ์จ ๋ชจ๋ธ์ 8B ๋งค๊ฐ๋ณ์ ๋ฐ 1T ํ ํฐ ์ด์์ผ๋ก ํ์ฅํ ๋ ์์ ์ ์ธ ํ๋ จ์ ์ ์งํ๋ ๋ฐ ์ด๋ ค์์ด ์์ต๋๋ค.
- ๐ ๊ณ ํ์ง ๋ฐ์ดํฐ๋ฅผ ์ฌ์ฉํ ์ฌ์ ํ์ต์ ๋ง์ง๋ง 20%๋ ์ฑ๋ฅ์ ํฌ๊ฒ ํฅ์์์ผฐ์ต๋๋ค.
- ๐ Chameleon-34B๋ Llama2-70B๋ฅผ ๋ฅ๊ฐํ๋ฉฐ Mixtral 8x7B/Gemini-Pro, GSM8K, MATH ๋ฐ MMLU์ ๊ทผ์ ํฉ๋๋ค.
- ๐ Chameleon-34B๋ MS-COCO์์ Flamingo-80B ๋ฐ IDEFICS-80B๋ฅผ ๋ฅ๊ฐํ๋ฉฐ Flickr30k์์๋ ์ผ์นํฉ๋๋ค.
- ๐ฏ Chameleon-34B๋ Gemini-Pro๋ฅผ ์๋๋ก 60.4%, GPT-4V๋ฅผ ์๋๋ก 51.6%์ ์น๋ฅ ์ ๋ฌ์ฑํ์ต๋๋ค.
- โ๏ธ ๊ท ํ ์กํ ๋ชจ๋ฌ๋ฆฌํฐ ๋ฐ์ดํฐ ์ธํธ๋ ๋ฏธ์ธ ์กฐ์ ๋ฐ ์ ๋ ฌ์ ์ค์ํฉ๋๋ค.
์ฐธ๊ณ : ์นด๋ฉ๋ ์จ์ ๊ธฐ๋ณธ ๋ฉํฐ๋ชจ๋ฌ ํ ํฐ์ด ์๋ Uni-MoE(์ด์ ๊ณต์ )๋ณด๋ค OpenAI GPT-4o์ ๋ ๊ฐ๊น์ ๋ณด์ ๋๋ค. ๐ก
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me