Data Science/Algorithms
๐ ๐๐๐ ๐๐ซ๐จ๐ฏ๐ข๐๐๐ซ'๐ฌ ๐๐๐ฅ๐๐๐ฌ๐ (2024 1H) & ๐ 2024 ๐๐๐ ๐๐ฎ๐ซ๐ฏ๐๐ฒ (on Training / Data / RAG / Serving / Agent)
Curiosity: What patterns can we retrieve from the rapid pace of LLM releases in 2024?
![["LLM 2024 Procider"]](/assets/img/llm/LLM_Provider_2024.jpeg)
๐ LLM Providerโs Release (2024 1H): A Comprehensive Overview
Curiosity: What patterns can we retrieve from the rapid pace of LLM releases in 2024? How do these innovations connect to the broader evolution of the field?
2024โs first half witnessed an unprecedented surge in LLM releases, with 21 major models from leading providers. This comprehensive overview retrieves insights from release patterns, technical innovations, and market dynamics to understand where the field is heading.
Release Timeline Overview
gantt
title LLM Releases 2024 1H Timeline
dateFormat YYYY-MM-DD
section Major Releases
GPT-4o (OpenAI) :2024-05-13, 1d
Llama-3 (Meta) :2024-04-18, 1d
Claude-3 (Anthropic) :2024-03-04, 1d
Gemini-1.5 (Google) :2024-03-08, 1d
section Open Source
Qwen-2 (Alibaba) :2024-06-07, 1d
DeepSeek-V2 :2024-05-07, 1d
Phi-3 (Microsoft) :2024-04-22, 1d
section Specialized
Solar-Mini-ja (Upstage) :2024-05-22, 1d
Mistral-Large :2024-02-26, 1d
21 LLM Releases: Complete Catalog
| # | Model | Provider | Release Date | Key Features | News | Paper |
|---|---|---|---|---|---|---|
| 1 | Qwen-2 | Alibaba Group | 2024.06.07 | Multilingual, large-scale | Link | - |
| 2 | Solar-Mini-ja | Upstage | 2024.05.22 | Japanese-optimized | Link | - |
| 3 | Yi-Large | 01.AI | 2024.05.13 | Large-scale model | Link | - |
| 4 | Yi-1.5 | 01.AI | 2024.05.13 | Enhanced version | Link | arXiv |
| 5 | GPT-4o | OpenAI | 2024.05.13 | Omni-modal, faster | Link | - |
| 6 | Qwen-Max | Alibaba Group | 2024.05.11 | Maximum performance | Link | - |
| 7 | DeepSeek-V2 | DeepSeek | 2024.05.07 | Efficient architecture | Link | arXiv |
| 8 | Snowflake-Arctic | Snowflake | 2024.04.24 | Enterprise-focused | Link | - |
| 9 | Phi-3 | Microsoft | 2024.04.22 | Small language model | Link | arXiv |
| 10 | Llama-3 | Meta | 2024.04.18 | Open-source leader | Link | - |
| 11 | Mixtral-8x22B | Mistral AI | 2024.04.17 | Mixture of experts | Link | - |
| 12 | Reka-Core | Reka AI | 2024.04.15 | Multimodal | Link | arXiv |
| 13 | Command-R-Plus | Cohere | 2024.04.04 | Enterprise RAG | Link | - |
| 14 | DBRX | Databricks | 2024.03.27 | Open-source SOTA | Link | - |
| 15 | Gemini-1.5 | 2024.03.08 | Long context | Link | arXiv | |
| 16 | Claude-3 | Anthropic | 2024.03.04 | Safety-focused | Link | - |
| 17 | Mistral-Large | Mistral AI | 2024.02.26 | European leader | Link | - |
| 18 | Gemma | 2024.02.21 | Open models | Link | arXiv | |
| 19 | Qwen-1.5 | Alibaba Group | 2024.02.04 | Multilingual | Link | - |
| 20 | Solar-Mini | Upstage | 2024.01.25 | Efficient Korean | Link | - |
| 21 | Solar-10.7B | Upstage | 2023.12.23 | Top pre-trained | Link | arXiv |
Provider Distribution
pie title LLM Releases by Provider (2024 1H)
"Alibaba Group" : 3
"Upstage" : 3
"Google" : 2
"Mistral AI" : 2
"01.AI" : 2
"Others" : 9
Key Trends & Insights
Retrieve: Analysis of release patterns reveals several key trends:
- Open Source Acceleration: Major releases from Meta (Llama-3), Alibaba (Qwen series), and Databricks (DBRX)
- Multimodal Expansion: GPT-4o, Gemini-1.5, Reka-Core emphasize vision capabilities
- Efficiency Focus: Phi-3, Solar-Mini demonstrate small model excellence
- Regional Specialization: Solar-Mini-ja (Japanese), Qwen series (Chinese)
Innovate: These releases show the field moving toward:
- More efficient architectures (DeepSeek-V2, Phi-3)
- Better multilingual support (Qwen, Solar)
- Enterprise-ready solutions (Snowflake Arctic, Command-R-Plus)
- ๐๐ฐ๐๐ง-2 (โAlibaba Groupโ, 2024.06.07)
- โข ๐ฃNews: https://qwenlm.github.io/blog/qwen2/
- ๐๐จ๐ฅ๐๐ซ-๐๐ข๐ง๐ข-๐ฃ๐ (โUpstageโ, 2024.05.22)
- โข ๐ฃNews: https://www.upstage.ai/feed/tech/solar-mini-chat-ja
- ๐๐ข-๐๐๐ซ๐ ๐ (โ01.AIโ, 2024.05.13)
- โข ๐ฃNews: https://x.com/01AI_Yi/status/1789929378467426794
- ๐๐ข-1.5 (โ01.AIโ, 2024.05.13)
- โข ๐ฃNews: https://x.com/01AI_Yi/status/1789869537317540016
- โข ๐arXiv: https://arxiv.org/abs/2403.04652
- ๐๐๐-4๐จ (โOpenAIโ, 2024.05.13)
- โข ๐ฃNews: https://openai.com/index/hello-gpt-4o/
- ๐๐ฐ๐๐ง-๐๐๐ฑ (โAlibaba Groupโ, 2024.05.11)
- โข ๐ฃNews: https://qwenlm.github.io/blog/qwen-max-0428/
- ๐๐๐๐ฉ๐๐๐๐ค-๐2 (DeepSeek, 2024.05.07)
- โข ๐ฃNews: https://x.com/deepseek_ai/status/1787478986731429933
- โข ๐arXiv: https://arxiv.org/abs/2405.04434
- ๐๐ง๐จ๐ฐ๐๐ฅ๐๐ค๐-๐๐ซ๐๐ญ๐ข๐ (โSnowflakeโ, 2024.04.24)
- โข ๐ฃNews: https://www.snowflake.com/blog/arctic-open-efficient-foundation-language-models-snowflake/
- ๐๐ก๐ข-3 (โMicrosoftโ, 2024.04.22)
- โข ๐ฃNews: https://azure.microsoft.com/en-us/blog/introducing-phi-3-redefining-whats-possible-with-slms/
- โข ๐arXiv: https://arxiv.org/abs/2404.14219
- ๐๐ฅ๐๐ฆ๐-3 (โMeta Facebookโ, 2024.04.18)
- โข ๐ฃNews: https://ai.meta.com/blog/meta-llama-3/
- ๐๐ข๐ฑ๐ญ๐ซ๐๐ฅ-8๐ฑ22๐ (โMistral AIโ, 2024.04.17)
- โข ๐ฃNews: https://mistral.ai/news/mixtral-8x22b/
- ๐๐๐ค๐-๐๐จ๐ซ๐ (โReka AIโโ, 2024.04.15)
- โข ๐ฃNews: https://www.reka.ai/news/reka-core-our-frontier-class-multimodal-language-model
- โข ๐arXiv: https://arxiv.org/abs/2404.12387
- ๐๐จ๐ฆ๐ฆ๐๐ง๐-๐-๐๐ฅ๐ฎ๐ฌ (โCohereโ, 2024.04.04)
- โข ๐ฃNews: https://cohere.com/blog/command-r-plus-microsoft-azure
- ๐๐๐๐ (โDatabricksโ, 2024.03.27)
- ๐๐๐ฆ๐ข๐ง๐ข-1.5 (โGoogleโ, 2024.03.08)
- โข ๐ฃNews: https://blog.google/technology/ai/google-gemini-next-generation-model-february-2024/
- โข ๐arXiv: https://arxiv.org/abs/2403.05530
- ๐๐ฅ๐๐ฎ๐๐-3 (โAnthropicโ, 2024.03.04)
- โข ๐ฃNews: https://www.anthropic.com/news/claude-3-family
- ๐๐ข๐ฌ๐ญ๐ซ๐๐ฅ-๐๐๐ซ๐ ๐ (โMistral AIโ, 2024.02.26)
- โข ๐ฃNews: https://mistral.ai/news/mistral-large/
- ๐๐๐ฆ๐ฆ๐ (โGoogleโ, 2024.02.21)
- โข ๐ฃNews: https://blog.google/technology/developers/gemma-open-models/
- โข ๐arXiv: https://arxiv.org/abs/2403.08295
- ๐๐ฐ๐๐ง-1.5 (โAlibaba Groupโ, 2024.02.04)
- โข ๐ฃNews: https://qwenlm.github.io/blog/qwen1.5/
- ๐๐จ๐ฅ๐๐ซ-๐๐ข๐ง๐ข (โUpstageโ, 2024.01.25)
- ๐๐จ๐ฅ๐๐ซ-10.7๐ (โUpstageโ, 2023.12.23)
- โข ๐ฃNews: https://www.upstage.ai/feed/press/solar-10-7b-emerges-as-worlds-top-pre-trained-llm
- โข ๐arXiv: https://arxiv.org/abs/2312.15166
๐ 2024 LLM Survey: Comprehensive Research Overview
Retrieve: What are the latest research trends across training, data, RAG, serving, and agents? This section compiles essential survey papers that capture the state of the art.
Essential Reading: These surveys provide comprehensive overviews of rapidly evolving LLM research areas.
Survey Categories Overview
graph TB
A[2024 LLM Surveys] --> B[Training]
A --> C[Data]
A --> D[RAG]
A --> E[Serving]
A --> F[Agent]
B --> B1[Self-Evolution]
B --> B2[Continual Learning]
B --> B3[Pre-trained Models]
C --> C1[Datasets]
C --> C2[Data Selection]
C --> C3[Instruction Tuning]
D --> D1[RALM Survey]
D --> D2[AIGC RAG]
D --> D3[LLM RAG]
E --> E1[Inference]
E --> E2[Invocation Methods]
E --> E3[Resource Efficiency]
F --> F1[Multimodal Agents]
F --> F2[Multi-Agents]
F --> F3[Personal Agents]
style A fill:#e1f5ff
style B fill:#fff3cd
style C fill:#d4edda
style D fill:#f8d7da
style E fill:#e7d4f8
style F fill:#ffe5e5
๐ Training Surveys
Retrieve: How do LLMs evolve and adapt? These surveys explore self-evolution, continual learning, and transfer learning.
| Survey | Date | Focus | arXiv | GitHub |
|---|---|---|---|---|
| Self-Evolution of LLMs | 2024.04.22 | Autonomous improvement mechanisms | Link | Repo |
| Continual Learning of LLMs | 2024.04.25 | Lifelong learning approaches | Link | Repo |
| Continual Learning with PTMs | 2024.01.29 | Pre-trained model adaptation | Link | Repo |
๐ Data Surveys
Innovate: Data quality and selection are critical for LLM performance. These surveys explore dataset curation and optimization.
| Survey | Date | Focus | arXiv | GitHub |
|---|---|---|---|---|
| Datasets for LLMs | 2024.02.28 | Comprehensive dataset catalog | Link | Repo |
| Data Selection for LMs | 2024.02.26 | Selection strategies | Link | Repo |
| Data Selection for Instruction Tuning | 2024.02.04 | Instruction data curation | Link | Repo |
๐ RAG Surveys
Retrieve: Retrieval-Augmented Generation is transforming how LLMs access knowledge. These surveys cover the latest RAG research.
| Survey | Date | Focus | arXiv | GitHub |
|---|---|---|---|---|
| RAG and RAU Survey | 2024.04.30 | RALM in NLP | Link | Repo |
| RAG for AIGC | 2024.02.29 | AI-generated content | Link | Repo |
| RAG for LLMs | 2023.12.18 | Comprehensive RAG overview | Link | Repo |
โก Serving Surveys
Innovate: Efficient inference and serving are crucial for production deployment. These surveys explore optimization strategies.
| Survey | Date | Focus | arXiv | GitHub |
|---|---|---|---|---|
| LLM Inference Unveiled | 2024.02.26 | Roofline model insights | Link | Repo |
| Effective LLM Service Invocation | 2024.02.05 | LLMaaS strategies | Link | Repo |
| Resource-Efficient LLMs | 2024.01.01 | Efficiency optimization | Link | Repo |
๐ค Agent Surveys
Retrieve: AI agents represent the next frontier. These surveys explore multimodal, multi-agent, and personal agent systems.
| Survey | Date | Focus | arXiv | GitHub |
|---|---|---|---|---|
| Large Multimodal Agents | 2024.02.23 | Vision-language agents | Link | Repo |
| LLM-based Multi-Agents | 2024.01.21 | Multi-agent systems | Link | Repo |
| Personal LLM Agents | 2024.01.10 | Personalization & security | Link | Repo |
Research Trends Summary
graph LR
A[2024 LLM Research] --> B[Training<br/>3 surveys]
A --> C[Data<br/>3 surveys]
A --> D[RAG<br/>3 surveys]
A --> E[Serving<br/>3 surveys]
A --> F[Agent<br/>3 surveys]
style A fill:#e1f5ff
style B fill:#fff3cd
style C fill:#d4edda
style D fill:#f8d7da
style E fill:#e7d4f8
style F fill:#ffe5e5
Key Takeaways
Retrieve: These 15 comprehensive surveys cover the essential areas of LLM research: training methodologies, data strategies, RAG systems, serving optimization, and agent architectures.
Innovate: By studying these surveys, you can retrieve the latest research insights and innovate on your own LLM applications, staying at the forefront of this rapidly evolving field.
Curiosity โ Retrieve โ Innovation: Start with curiosity about LLM capabilities, retrieve knowledge from these surveys, and innovate by applying cutting-edge techniques to your projects.
Information about Tokens in LLsM
Why do we keep talking about โtokensโ in LLMs instead of words?
It happens to be much more efficient to break the words into sub-words (tokens) for model performance!
The typical strategy used in most modern LLMs since GPT-1 is the Byte Pair Encoding (BPE) strategy. The idea is to use, as tokens, sub-word units that appear often in the training data. The algorithm works as follows:
- We start with a character-level tokenization
- we count the pair frequencies
- We merge the most frequent pair
- We repeat the process until the dictionary is as big as we want it to be
The size of the dictionary becomes a hyperparameter that we can adjust based on our training data. For example, GPT-1 has a dictionary size of ~40K merges, GPT-2, GPT-3, and ChatGPT have a dictionary size of ~50K, and Llama 3 128K.
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me
