Review/Trends
๐ ๐ฎ๐ฟ๐ธ ๐ญ๐๐ฐ๐ธ๐ฒ๐ฟ๐ฏ๐ฒ๐ฟ๐ด ๐ท๐๐๐ ๐ฎ๐ป๐ป๐ผ๐๐ป๐ฐ๐ฒ๐ฑ ๐๐ต๐ฒ ๐๐ฃ๐ง-๐ฐ ๐ธ๐ถ๐น๐น๐ฒ๐ฟ, ๐๐น๐ฎ๐บ๐ฎ-๐ฏ.๐ญ ๐ฅ
LLM, Llama3.1
- ๐ Read our announcement blog post: https://huggingface.co/blog/llama31
- ๐ค Model card for the 405B on the Hub: https://huggingface.co/meta-llama/Meta-Llama-3.1-405B-FP8
Metaโs Llama-3.1
Curiosity: Metaโs Llama-3.1 patches the 8B and 70B Llama-3 models, already top performers in their weight class, to make them even better + gives us the strongest open-source model ever with the 405B.
Two main points:
๐ซ ๐ง๐ต๐ฒ ๐ป๐ฒ๐ ๐ธ๐ถ๐ป๐ด ๐ผ๐ณ ๐ข๐ฆ ๐บ๐ผ๐ฑ๐ฒ๐น๐: ๐๐น๐ฎ๐บ๐ฎ-๐ฏ.๐ญ-๐ฐ๐ฌ๐ฑ๐ on par or above GPT-4o on many benchmarks.
If confirmed on further testing, it is officially the first time that an OS model becomes the strongest model overall, on top of all models from anthropic and OpenAI!
Let me repeat this: ๐๐ต๐ฒ ๐๐๐ฟ๐ผ๐ป๐ด๐ฒ๐๐ ๐๐๐ ๐ฒ๐๐ฒ๐ฟ ๐ฐ๐ฎ๐ป ๐ฏ๐ฒ ๐ฑ๐ผ๐๐ป๐น๐ผ๐ฎ๐ฑ๐ฒ๐ฑ ๐ณ๐ฟ๐ผ๐บ ๐๐ต๐ฒ ๐๐๐ฏ.
๐ ๐ง๐ต๐ฒ ๐ด๐ ๐ฎ๐ป๐ฑ ๐ณ๐ฌ๐ ๐บ๐ผ๐ฑ๐ฒ๐น๐ ๐ฎ๐ฟ๐ฒ ๐ฒ๐ ๐๐ฒ๐ป๐ฑ๐ฒ๐ฑ ๐๐ผ ๐ญ๐ฎ๐ด๐ธ ๐๐ผ๐ธ๐ฒ๐ป๐ ๐ฐ๐ผ๐ป๐๐ฒ๐ ๐ ๐น๐ฒ๐ป๐ด๐๐ต.
The previous models were limited to 8k tokens, meaning they could process at max as much text as around 15 pages in a Word doc: this was a terrible blocker anytime you need a bit of memory, like for RAG or agent workflows.
Well, not anymore! โ Now we get a much more comfortable 128k context length for all sizes, which is great for most my agentic use-cases.
Both points above are huge and would be newsworthy, dropping these together in a โ3.1โ version is crazy! ๐คฏ
๐ง๐ฒ๐ฐ๐ต๐ป๐ถ๐ฐ๐ฎ๐น ๐ถ๐ป๐๐ถ๐ด๐ต๐๐:
๐ซ ๐ ๐ป๐ฒ๐ ๐ฐ๐ฌ๐ฑ๐, ๐ฝ๐ผ๐๐๐ถ๐ฏ๐น๐ ๐๐ต๐ฒ ๐๐๐ฟ๐ผ๐ป๐ด๐ฒ๐๐ ๐๐๐ ๐ฒ๐๐ฒ๐ฟ, with 128k context length, 88.6% on MMLU, a crazy 96.8% on GSM8K.
- โค The 405B has a FP8 quantized version. FP8 quantization was only applied to the major linear operators of the model, such as the gate and up and down projections for the FFNs (covering 75% of the inference FLOPs).
- โค You still need 8xH100 to run it with full context length.
๐ฆฃ Improved 8B & 70B models, with a ๐บ๐๐ฐ๐ต ๐น๐ฎ๐ฟ๐ด๐ฒ๐ฟ ๐ฐ๐ผ๐ป๐๐ฒ๐ ๐ ๐๐ถ๐๐ฒ ๐ผ๐ณ ๐ญ๐ฎ๐ด๐ธ ๐๐ ๐ด๐ธ โ ๐๐ต๐ถ๐ ๐ถ๐ ๐ฎ ๐ด๐ฎ๐บ๐ฒ-๐ฐ๐ต๐ฎ๐ป๐ด๐ฒ๐ฟ ๐ณ๐ผ๐ฟ ๐ฅ๐๐ ๐ฎ๐ป๐ฑ ๐๐ด๐ฒ๐ป๐๐.
- ๐ Pretrained on 15T tokens, a more diverse training dataset than Llama-3 to reinforce multilinguality: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
- ๐ License: same as Llama-3, and on top of that it allows using the output data from Llama-3.1 for training other models (distillation)
- โจ One new role for the instruct version: on top of System, User, and Assistant, Ipython lets you write the output of a code tool call! This should work really well with Transformers agents ๐
Thanks a lot to Meta for this release which will make our lives better! ๐ค
