Multimodal/Computer Vision
The giant leaps of open-source models for Vision Models
Curiosity: What insights can we retrieve from this? How does this connect to innovation in the field?
๐๐ข๐ฌ๐ข๐จ๐ง ๐ฅ๐๐ง๐ ๐ฎ๐๐ ๐ ๐ฆ๐จ๐๐๐ฅ๐ฌ
Curiosity: What insights can we retrieve from this? How does this connect to innovation in the field?
Curiosity: Andrew Reed built a cool space that shows that OS LLMs are catching up with closed source LLMs in ELO ranking in the Arena (link below).
For vision, the same dynamic is happening: the field is still evolving fast, but soon OS models will be able to match GPT-4oโs vision skills.
I witnessed the Idefics teamโs work and their many late nights before their publishing of Idefics-2-8b. Now they just published a paper that summarizes their insights!
๐๐๐ง๐โ๐จ ๐ ๐จ๐ช๐ข๐ข๐๐ง๐ฎ ๐ค๐ ๐ฌ๐๐๐ฉ ๐ฉ๐๐๐ฎ ๐๐ค๐ช๐ฃ๐:
โค ๐ฃ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ ๐ผ๐ณ ๐ฉ๐๐ ๐ ๐ถ๐ ๐น๐ฎ๐ฟ๐ด๐ฒ๐น๐ ๐ฑ๐ฟ๐ถ๐๐ฒ๐ป ๐ฏ๐ ๐ฝ๐ฒ๐ฟ๐ณ ๐ผ๐ณ ๐๐ต๐ฒ๐ถ๐ฟ ๐๐ฒ๐ ๐-๐ผ๐ป๐น๐ ๐ฏ๐ฎ๐ฐ๐ธ๐ฏ๐ผ๐ป๐ฒ๐. In ablation studies, replacing the llama-1-7b with Mistral-7b directly brings +7% performance ๐คฏ
โค ๐ง๐ต๐ฒ๐ ๐ฐ๐ผ๐บ๐ฝ๐ฎ๐ฟ๐ฒ๐ฑ ๐๐๐ผ ๐ฐ๐ผ๐บ๐ฝ๐ฒ๐๐ถ๐ป๐ด ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ๐:
- ๐ ๐๐ฟ๐ผ๐๐ ๐ฎ๐๐๐ฒ๐ป๐๐ถ๐ผ๐ป ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ: images are encoded through the vision backbone, and their information is inserted within the text processing at various places
- ๐ข ๐๐๐น๐น๐ ๐ฎ๐๐๐ผ๐ฟ๐ฒ๐ด๐ฟ๐ฒ๐๐๐ถ๐๐ฒ ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ: the output is directly concatenated to the sequence of text embeddings, and entire sequence passed as input to the LM (cf image) The comparisonโs outcome is the following โ ๐๐๐น๐น๐ ๐ฎ๐๐๐ผ๐ฟ๐ฒ๐ด๐ฟ๐ฒ๐๐๐ถ๐๐ฒ ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ ๐ผ๐๐๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ ๐ฐ๐ฟ๐ผ๐๐-๐ฎ๐๐๐ฒ๐ป๐๐ถ๐ผ๐ป ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ when you fine-tune the whole system using LoRA
โก๏ธ ๐ง๐ต๐ฒ๐๐ฒ ๐ณ๐ถ๐ป๐ฑ๐ถ๐ป๐ด๐ ๐น๐ฒ๐ฑ ๐๐ผ ๐๐ฒ๐๐ฒ๐ฟ๐ฎ๐น ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฎ๐น ๐ถ๐บ๐ฝ๐ฟ๐ผ๐๐ฒ๐บ๐ฒ๐ป๐ ๐ถ๐ป ๐๐ฑ๐ฒ๐ณ๐ถ๐ฐ๐-๐ฎ: โค Replaced cross-attention architecture with fully autoregressive architecture
โค Enable treating images with varying aspect ratio
โค Allow to split an image in 4, to be encoded on 320 vision tokens instead of 64, if you want to increase perf at the cost of more compute
โจ As a result, Idefics-2 reaches state-of-the-art performance for this model size! Now just a few more steps to catch up to GPT-4o!
Congrats for this great release Lรฉo Tronchon Hugo Laurenรงon Victor Sanh! ๐
๐ ๐ฅ๐ฒ๐ฎ๐ฑ ๐๐ต๐ฒ ๐๐ฑ๐ฒ๐ณ๐ถ๐ฐ๐-๐ฎ ๐ฝ๐ฎ๐ฝ๐ฒ๐ฟ: https://huggingface.co/papers/2405.02246
๐ ๐๐ป๐ฑ๐ฟ๐ฒ๐โ๐ ๐๐ฝ๐ฎ๐ฐ๐ฒ ๐๐ต๐ฎ๐ ๐๐ต๐ผ๐๐ ๐ข๐ฆ ๐บ๐ผ๐ฑ๐ฒ๐น๐ ๐ฐ๐ฎ๐๐ฐ๐ต๐ถ๐ป๐ด ๐๐ฝ (๐ณ๐ผ๐ฟ ๐๐ฒ๐ ๐ ๐บ๐ผ๐ฑ๐ฒ๐น๐): https://huggingface.co/spaces/andrewrreed/closed-vs-open-arena-elo
โ๏ธ ๐๐ผ๐บ๐ฝ๐ฎ๐ฟ๐ฒ ๐๐ถ๐๐ถ๐ผ๐ป ๐บ๐ผ๐ฑ๐ฒ๐น๐ ๐ถ๐ป ๐๐ต๐ฒ ๐ฉ๐ถ๐๐ถ๐ผ๐ป ๐ฎ๐ฟ๐ฒ๐ป๐ฎ: https://huggingface.co/spaces/WildVision/vision-arena
Translate to Korean
์คํ ์์ค ๋ชจ๋ธ์ ๊ฑฐ๋ํ ๋์ฝ
Andrew Reed ์๋ ๋์ ELO ์์์์ OS LLM์ด ํด๋ก์ฆ๋ ์์ค LLM์ ๋ฐ๋ผ์ก๊ณ ์์์ ๋ณด์ฌ์ฃผ๋ ๋ฉ์ง ๊ณต๊ฐ์ ๊ตฌ์ถํ์ต๋๋ค(์๋ ๋งํฌ). ๋น์ ์ ๊ฒฝ์ฐ์๋ ๋์ผํ ์ญํ์ด ์ผ์ด๋๊ณ ์์ต๋๋ค: ์ด ๋ถ์ผ๋ ์ฌ์ ํ ๋น ๋ฅด๊ฒ ์งํํ๊ณ ์์ง๋ง ๊ณง OS ๋ชจ๋ธ์ด GPT-4o์ ๋น์ ๊ธฐ์ ๊ณผ ์ผ์นํ ์ ์๊ฒ ๋ ๊ฒ์ ๋๋ค.
๋๋ Idefics ํ์ ์์ ๊ณผ Idefics-2-8b๋ฅผ ์ถํํ๊ธฐ ์ ์ ๋ง์ ๋ฆ์ ๋ฐค์ ๋ชฉ๊ฒฉํ์ต๋๋ค. ์ด์ ๊ทธ๋ค์ ๊ทธ๋ค์ ํต์ฐฐ๋ ฅ์ ์์ฝํ ๋ ผ๋ฌธ์ ๋ฐํํ์ต๋๋ค!
๊ทธ๋ค์ด ๋ฐ๊ฒฌํ ๋ด์ฉ์ ์์ฝํ๋ฉด ๋ค์๊ณผ ๊ฐ์ต๋๋ค.
โค VLM์ ์ฑ๋ฅ์ ์ฃผ๋ก ํ ์คํธ ์ ์ฉ ๋ฐฑ๋ณธ์ ์ฑ๋ฅ์ ์ํด ์ข์ฐ๋ฉ๋๋ค. ์ ์ ์ฐ๊ตฌ์์ llama-1-7b๋ฅผ Mistral-7b๋ก ์ง์ ๋์ฒดํ๋ฉด +7%์ ์ฑ๋ฅ์ ๐คฏ ์ป์ ์ ์์ต๋๋ค.
โค ๋ ๊ฐ์ง ๊ฒฝ์ ์ํคํ ์ฒ๋ฅผ ๋น๊ตํ์ต๋๋ค.
- ๐ ํฌ๋ก์ค ์ดํ ์ ์ํคํ ์ฒ: ์ด๋ฏธ์ง๋ ๋น์ ๋ฐฑ๋ณธ์ ํตํด ์ธ์ฝ๋ฉ๋๊ณ ํด๋น ์ ๋ณด๋ ๋ค์ํ ์์น์์ ํ ์คํธ ์ฒ๋ฆฌ ๋ด์ ์ฝ์ ๋ฉ๋๋ค.
- ๐ข ์์ ์๋ ํ๊ท ์ํคํ ์ฒ: ์ถ๋ ฅ์ ํ ์คํธ ์๋ฒ ๋ฉ ์ํ์ค์ ์ง์ ์ฐ๊ฒฐ๋๊ณ ์ ์ฒด ์ํ์ค๋ LM์ ์ ๋ ฅ์ผ๋ก ์ ๋ฌ๋ฉ๋๋ค(cf ์ด๋ฏธ์ง). ๋น๊ต ๊ฒฐ๊ณผ๋ ๋ค์๊ณผ ๊ฐ์ต๋๋คโ ์์ ์๋ ํ๊ท ์ํคํ ์ฒ๋ LoRA๋ฅผ ์ฌ์ฉํ์ฌ ์ ์ฒด ์์คํ ์ ๋ฏธ์ธ ์กฐ์ ํ ๋ ๊ต์ฐจ ์ฃผ์ ์ํคํ ์ฒ๋ณด๋ค ์ฑ๋ฅ์ด ๋ฐ์ด๋ฉ๋๋ค.
โก๏ธ ์ด๋ฌํ ๋ฐ๊ฒฌ์ Idefics-2์ ๋ช ๊ฐ์ง ์ํคํ ์ฒ ๊ฐ์ ์ผ๋ก ์ด์ด์ก์ต๋๋ค. โค cross-attention ์ํคํ ์ฒ๋ฅผ ์์ ์๋ ํ๊ท ์ํคํ ์ฒ๋ก ๋์ฒดํ์ต๋๋ค.
โค ๋ค์ํ ์ข ํก๋น๋ก ์ด๋ฏธ์ง ์ฒ๋ฆฌ ๊ฐ๋ฅ
โค ๋ ๋ง์ ์ปดํจํ ๋น์ฉ์ผ๋ก ์ฑ๋ฅ์ ๋์ด๋ ค๋ฉด ์ด๋ฏธ์ง๋ฅผ 4๊ฐ๋ก ๋ถํ ํ์ฌ 64๊ฐ ๋์ 320๊ฐ์ ๋น์ ํ ํฐ์ผ๋ก ์ธ์ฝ๋ฉํ ์ ์์ต๋๋ค.
โจ ๊ฒฐ๊ณผ์ ์ผ๋ก Idefics-2๋ ์ด ๋ชจ๋ธ ํฌ๊ธฐ์ ๋ํด ์ต์ฒจ๋จ ์ฑ๋ฅ์ ๋๋ฌํ์ต๋๋ค! ์ด์ GPT-4o๋ฅผ ๋ฐ๋ผ์ก๊ธฐ ์ํ ๋ช ๋จ๊ณ๋ง ๋ ๊ฑฐ์น๋ฉด ๋ฉ๋๋ค!
์ด ๋ฉ์ง ๋ฆด๋ฆฌ์ค Lรฉo Tronchon Hugo Laurenรงon Victor Sanh ์ถํํฉ๋๋ค! ๐
๐ Idefics-2 ๋ ผ๋ฌธ ์ฝ๊ธฐ: https://huggingface.co/papers/2405.02246
๐ OS ๋ชจ๋ธ์ด ๋ฐ๋ผ์ก๋ ๊ฒ์ ๋ณด์ฌ์ฃผ๋ Andrew์ ๊ณต๊ฐ (ํ ์คํธ ๋ชจ๋ธ์ ๊ฒฝ์ฐ) : https://huggingface.co/spaces/andrewrreed/closed-vs-open-arena-elo
โ๏ธ ๋น์ ๋ถ์ผ์ ๋น์ ๋ชจ๋ธ ๋น๊ต: https://huggingface.co/spaces/WildVision/vision-arena
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me