Multimodal/Computer Vision
Top Papers in Computer Vision, NLP, Speech, Multimodal AI, Core ML, RecSys, and Graph ML
๐๐ผ Iโve put together a summary of key papers in #AI and categorized them into (i) need-to-know and (ii) good-to-know.
๐ Top Papers in Computer Vision, NLP, Speech, Multimodal AI, Core ML, RecSys, and Graph ML โข
Distilled AI : https://aman.ai/papers/ aman AI : https://aman.ai/
๐๐ผ Iโve put together a summary of key papers in ํด์ํ๊ทธ#AI and segregated them into (i) need-to-know and (ii) good-to-know.
๐น Vision
- Image Classification (CNN architectures such as AlexNet, VGGNet, InceptionNet, ResNet to Transformer architectures such as ViT, DeiT, BEiT, MAE)
- Object Detection (YOLO v1-v8, Fast/er R-CNN, Mask R-CNN, CenterNet, Pix2Seq, DETR, Detic, Focal Loss)
- Semantic/Instance Segmentation (U-Net, Mask R-CNN, Segment Anything)
- NeRF (InstantNeRF, BlockNeRF)
- SSL Contrastive Learning (SimCLR, MoCo, DINO v1 & v2)
๐น NLP
- Transformers (original paper)
- Semantic Representation Encoders (BERT and its variants: RoBERTa, DistillBERT, ELECTRA, XLNet, MPNet, ALBERT)
- Autoregressive Decoders (GPT-n, Llama 1/2/3, Alpaca, Vicuna)
- Augmented LMs (RAG, Toolformer, HuggingGPT, Gorilla)
- Supervised Fine-tuning (Instruction tuning/FLAN, LIMA, LESS)
- LLM Alignment (RLHF/InstructGPT, PPO, DPO, KTO, GPO, IPO)
- Encoder + Decoder Architectures (T0, T5, BART)
- Machine Translation (M2M-100, NLLB-200)
- Contrastive Learning (SNCSE, InfoNCE, Sentence-BERT)
- Prompting (CoT, Auto-CoT, Self-Consistency, ToT, GoT, ReAct, APE, ART)
- PEFT (Prefix-tuning, Adapters, LoRA, LLaMA-Adapter v1 and v2, QLoRA, QA-LoRA, DoRA, NOLA)
๐น Speech
- SSL Pre-Training (WavLM, AudioMAE, HuBERT)
- Automatic Speech Recognition/Keyword Spotting (GMM-HMM, DNN-HMM, all-neural architectures such as LAS/Whisper, streaming architectures such as RNN-T/Transformer-T)
- Speaker Identification (i/d/x-vectors, GE2E loss, AAM loss)
- Text-to-Speech (HiFi-GAN, Tacotron v1 and v2, Voicebox)
- Text-to-Audio/Music (MusicGen, AudioGen)
๐น Multimodal
- SSL Pre-Training (ViLT, MLIM, UNiTER, LXMERT, VisualBERT, Data2Vec v1 and v2, I-Code, VL-BEIT, ImageBind)
- V+L Prompting (Flamingo, Frozen, InstructBLIP)
- Text-to-Image (DALL-E 1/2/3, Imagen, Latent Diffusion, Make-A-Scene, Make-a-Video)
- Translation (SeamlessM4T)
- Contrastive Learning (InfoNCE, CLIP, CLAP, AudioCLIP)
๐น Core ML
- Training Regularizer (Dropout)
- Training/Inference Efficiency (ZeRO, ZeRO-Infinity, FlashAttention, FlashAttention-2)
- Training Stability (Batch/Layer/Group/Instance Norm, Residual/Skip Connections)
- Explainable AI (Guided Backprop, Grad-CAM, CAV, Influence functions, Representer points, TracIn)
๐น RecSys
- ML-based Collaborative Filtering (Factorization Machines)
- DL-based Algorithms (Collaborative Deep Learning, Wide & Deep, DNNs for YouTube Recommendations, Product-based DNNs, NCF, Deep & Cross v1 and v2, DeepFM, Deep Interest Network, Behavior Sequence Transformer)
๐น Graph ML
- Factorization-based Algorithms LLE (LLE, LAP, HOPE)
- Random Walk-based Algorithms (Node2vec)
- Deep Learning-based Algorithms (SDNE, GraphSAGE, EGNN, GCN, GAT)
Translate to Korean
๐ ์ปดํจํฐ ๋น์ , NLP, ์์ฑ, ๋ฉํฐ๋ชจ๋ฌ AI, Core ML, RecSys ๋ฐ Graph ML ๋ถ์ผ์ ์ฃผ์ ๋ ผ๋ฌธ โข
Distilled AI : https://aman.ai/papers/ aman AI : https://aman.ai/
๐๐ผ ํด์ํ๊ทธ#AI ์ ์ฃผ์ ๋ ผ๋ฌธ์ ์์ฝํ์ฌ (i) ์์์ผ ํ ์ฌํญ๊ณผ (ii) ์์๋๋ฉด ์ข์ ๋ด์ฉ์ผ๋ก ๊ตฌ๋ถํ์ต๋๋ค.
๐น ์๋ ฅ
- ์ด๋ฏธ์ง ๋ถ๋ฅ(AlexNet, VGGNet, InceptionNet, ResNet๊ณผ ๊ฐ์ CNN ์ํคํ ์ฒ์์ ViT, DeiT, BEiT, MAE์ ๊ฐ์ Transformer ์ํคํ ์ฒ๊น์ง)
- ๋ฌผ์ฒด ๊ฐ์ง(YOLO v1-v8, Fast/er R-CNN, Mask R-CNN, CenterNet, Pix2Seq, DETR, Detic, Focal Loss)
- ์๋ฏธ๋ก ์ /์ธ์คํด์ค ๋ถํ (U-Net, Mask R-CNN, Segment Anything)
- NeRF (InstantNeRF, BlockNeRF)
- SSL ๋์กฐ ํ์ต(SimCLR, MoCo, DINO v1 ๋ฐ v2)
๐น NLP (์์ด)
- ๋ณ์๊ธฐ (์๋ณธ ์ฉ์ง)
- ์๋ฏธ๋ก ์ ํํ ์ธ์ฝ๋(BERT ๋ฐ ๊ทธ ๋ณํ: RoBERTa, DistillBERT, ELECTRA, XLNet, MPNet, ALBERT)
- ์๋ ํ๊ท ๋์ฝ๋(GPT-n, Llama 1/2/3, Alpaca, Vicuna)
- ์ฆ๊ฐ LM(RAG, Toolformer, HuggingGPT, Gorilla)
- ๊ฐ๋ ๋ฏธ์ธ ์กฐ์ (๋ช ๋ น ํ๋/FLAN, LIMA, LESS)
- LLM ์ผ๋ผ์ธ๋จผํธ (RLHF/InstructGPT, PPO, DPO, KTO, GPO, IPO)
- ์ธ์ฝ๋ + ๋์ฝ๋ ์ํคํ ์ฒ(T0, T5, BART)
- ๊ธฐ๊ณ ๋ฒ์ญ (M2M-100, NLLB-200)
- ๋์กฐ ํ์ต(SNCSE, InfoNCE, Sentence-BERT)
- ํ๋กฌํํธ(CoT, Auto-CoT, Self-Consistency, ToT, GoT, ReAct, APE, ART)
- PEFT(์ ๋์ฌ ํ๋, ์ด๋ํฐ, LoRA, LLaMA-์ด๋ํฐ v1 ๋ฐ v2, QLoRA, QA-LoRA, DoRA, NOLA)
๐น ์ฐ์ค
- SSL ์ฌ์ ๊ต์ก(WavLM, AudioMAE, HuBERT)
- ์๋ ์์ฑ ์ธ์/ํค์๋ ์คํฌํ (GMM-HMM, DNN-HMM, LAS/Whisper์ ๊ฐ์ ์ ์ฒด ์ ๊ฒฝ ์ํคํ ์ฒ, RNN-T/Transformer-T์ ๊ฐ์ ์คํธ๋ฆฌ๋ฐ ์ํคํ ์ฒ)
- ํ์ ์๋ณ(i/d/x-๋ฒกํฐ, GE2E ์์ค, AAM ์์ค)
- ํ ์คํธ ์์ฑ ๋ณํ(HiFi-GAN, Tacotron v1 ๋ฐ v2, Voicebox)
- ํ ์คํธ-์ค๋์ค/์์ (MusicGen, AudioGen)
๐น ๋ณตํฉ
- SSL ์ฌ์ ํ์ต(ViLT, MLIM, UNiTER, LXMERT, VisualBERT, Data2Vec v1 ๋ฐ v2, I-Code, VL-BEIT, ImageBind)
- V+L ํ๋กฌํํธ (Flamingo, Frozen, InstructBLIP)
- ํ ์คํธ-์ด๋ฏธ์ง(DALL-E 1/2/3, ์์, ์ ์ฌ ํ์ฐ, Make-A-Scene, Make-A-Video)
- ๋ฒ์ญ(SeamlessM4T)
- ๋์กฐ ํ์ต(InfoNCE, CLIP, CLAP, AudioCLIP)
๐น ์ฝ์ด ML
- ๊ต์ก ์ ๊ทํ๊ธฐ(๋๋กญ์์)
- ํ๋ จ/์ถ๋ก ํจ์จ์ฑ(ZeRO, ZeRO-Infinity, FlashAttention, FlashAttention-2)
- ํ์ต ์์ ์ฑ(๋ฐฐ์น/๋ ์ด์ด/๊ทธ๋ฃน/์ธ์คํด์ค ํ์ค, ์์ฐจ/์คํต ์ฐ๊ฒฐ)
- ์ค๋ช ๊ฐ๋ฅํ AI(์ ๋ ๋ฐฑํ๋กญ, Grad-CAM, CAV, ์ํฅ๋ ฅ ๊ธฐ๋ฅ, ๋ฐํ์ ํฌ์ธํธ, TracIn)
๐น ๋ ํฌ์์ค
- ML ๊ธฐ๋ฐ ํ์ ํํฐ๋ง(Factorization Machine)
- DL ๊ธฐ๋ฐ ์๊ณ ๋ฆฌ์ฆ (Collaborative Deep Learning, Wide & Deep, YouTube Recommendations์ฉ DNN, ์ ํ ๊ธฐ๋ฐ DNN, NCF, Deep & Cross v1 ๋ฐ v2, DeepFM, Deep Interest Network, Behavior Sequence Transformer)
๐น ๊ทธ๋ํ ML
- ์ธ์๋ถํด ๊ธฐ๋ฐ ์๊ณ ๋ฆฌ์ฆ LLE(LLE, LAP, HOPE)
- ๋๋ค ์ํฌ ๊ธฐ๋ฐ ์๊ณ ๋ฆฌ์ฆ(Node2vec)
- ๋ฅ๋ฌ๋ ๊ธฐ๋ฐ ์๊ณ ๋ฆฌ์ฆ (SDNE, GraphSAGE, EGNN, GCN, GAT)
Working on something like this?
I take a small number of paid, scoped reviews: AI agent/RAG architecture diagnosis, Unity CI & build-automation audits, and multimodal QA design review. Each one ends in a written findings document.
Work with me