iburenko/multimodal-reading-group

8

36 commits

updated Jun 13, 2025

See the code

README

multimodal-reading-group


DatePaperAuthorsCodeDemoComments
01.02.2024Visual Instruction TuningH. Liu, C. Li, Q. Wu, Y. J. LeeGitHub Project PageDemo
08.02.2024When and why vision-language models behave like bags-of-words, and what to do about it?M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, J. Zouhttps://github.com/mertyg/vision-language-models-are-bowsColabWhy did they expect that CLIP will take a word order into account given that CLIP is trained to match a bag-of-words with a corresponding image?
22.02.2024Learning Transferable Visual Models From Natural Language SupervisionA. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G.Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. SutskeverGitHub Project PageColabSee also open source implementation of CLIP; Scaling laws for contrastive language-image learning
29.02.2024ContinueFig. 2 is unclear. How do they obtain a vector for a bag-of-words?
07.03.2024Still (sic!) continueIt seems that they train using BoW, even though their inference pipeline does not reflect this.
14.03.2024Sigmoid Loss for Language Image Pre-TrainingX. Zhai, B. Mustafa, A. Kolesnikov, L. BeyerHuggingFace
21.03.2024Continue
28.03.2024Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningW. Liang, Y. Zhang, Y. Kwon, S. Yeung, J. ZouGitHub Project Page
04.04.2024What Makes Training Multi-modal Classification Networks Hard?Wang, Tran, Feiszli
11.04.2024MultiBench: Multiscale Benchmarks for Multimodal Representation LearningLiang, Lyu, Fan, Wu, Cheng, Wu, Chen, Wu, Lee, Zhu, Salakhutdinaov, MorencyGitHub Project PageDemos
16.04.2024Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsTong, Liu, Zhai, Ma, LeCun, XieGitHub Project PageHuggingFace
23.04.2024Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training StrategiesLi, Xie, Cubuk
30.04.2024Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsLu, Peng, Cheng, Galley, Chang, Wu, Zhu, GaoGitHub, Project Page
07.05.2024Many-Shot In-Context LearningAgarwal, Singh, Zhang, Bohnet, Chan, Anand, Abbas, Nova, Co-Reyes, Chu, Behbahani, Faust, LarochelleNot Provided
28.05.2024BABILong: a long-context needle-in-a-haystack benchmark for LLMsKuratob, Bulatov, Anokhin, Sorokin, Sorokin, BurtsevGitHub
04.06.2024Continue
11.06.20244M: Massively Multimodal Masked ModelingMizrahi, Bachmann, Kar, Yeo, Gao, Dehghan, ZamirGitHub Project Page
18.06.2024Continue
25.06.2024GLaMM: Pixel Grounding Large Multimodal ModelRasheed, Maaz, Shaji, Shaker, Khan, Cholakkal, Anwer, Xing, Yang, KhanGitHub Project PageDemo
02.07.2024Code Reading Group
09.07.2024Knowledge DistillationGemma 2 (pdf), MobileLLM, Knowledge distillation, On-Policy distillation of Language Models
16.07.2024Whiteboard-of-Thought: Thinking Step-by-Step Across ModalitiesMenon, Zemel, VondrickProject Page
23.07.2024Multimodal Neurons in Artificial Neural NetworksGoh, Cammarata, Voss, Carter, Petrov, Schubert, Radford, Olah
30.07.2024Continue + (very briefly) CLIPPO
06.08.2024Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!Hessel, Lee
13.08.2024Graph of Thoughts and Monte Carlo Tree SearchMonte Carlo Tree Search from Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B; Graph of Thoughts; Large Language Monkeys; STaR: Self-Taught Reasoner; Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents; Bonus! DeepSeek-Prover-V1.5Tinygrad example of MCTS
15.10.2024
22.10.2024Calibration Multimodal LearningMa, Zhang, Wu, Fu, Hu
08.11.2024Towards Mamba: the S4 model and topic around: HiPPo, S4 paper, Annotated S4 blog postGu, Goel, RéGitHub
15.11.2024Continue: HiPPO and S4
22.11.2024Continue: Mamba & differences between Transformers and SSMs
07.02.2025RLHF Block. Part 1. Motivation and RLHF fine-tuningStiennon and friends Learning to summarize from human feedback (RLHF for summarisation); Ouyang and friends Training language models to follow instructions with human feedback (RLHF for wide range of NLP-related tasks)
14.02.2025RLHF Block. Part 2. Motivation and RLHF fine-tuningStiennon et al., Ouyand et al. + Ziegler et al. Fine-Tuning Language Models from Human Preferences
21.02.2025RLHF Block. Part 3. Motivation and RLHF fine-tuningFinish discussing Ouyand et al.
24.02.2025RLHF Block. Part 4. Reinforcement Learning. Q-Learning and PPOSlides
21.03.2025RLHF Block. Part 5. DPO and oher alternatives to RLHF.Rafailov et al. The DPO paper, Transition from formula 3 to formula 4 from the DPO paper, the Rejection sampling paper, the Llama 2 paper, RLAIF
11.04.2025Reminder on Chain-of-ThoughtArXiv
02.05.2025DeepSeek-MathShao, Wang, Zhu, Xu, Song, Bi, Zhang, Zhang, Li, Wu, GuoGitHub
16.05.2025DeepSeek-ProverXin, Guo, Shao, Ren, Zhu, Liu, Ruan, Li, LiangProver-V2, Prover-V2-GitHub
30.05.2025DeepSeek-R1DeepSeek-AI

Datasets and benchmarks
Surveys
Representation Learning
Latent Space Structure
Fusion
Modality Competition. Quantitative Methods of Detection of Suboptimality.

Contributors

iburenko

36 commits

iburenko/multimodal-reading-group

8

36 commits

updated Jun 13, 2025

See the code

README

multimodal-reading-group


DatePaperAuthorsCodeDemoComments
01.02.2024Visual Instruction TuningH. Liu, C. Li, Q. Wu, Y. J. LeeGitHub Project PageDemo
08.02.2024When and why vision-language models behave like bags-of-words, and what to do about it?M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, J. Zouhttps://github.com/mertyg/vision-language-models-are-bowsColabWhy did they expect that CLIP will take a word order into account given that CLIP is trained to match a bag-of-words with a corresponding image?
22.02.2024Learning Transferable Visual Models From Natural Language SupervisionA. Radford, J.W. Kim, C. Hallacy, A. Ramesh, G.Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, I. SutskeverGitHub Project PageColabSee also open source implementation of CLIP; Scaling laws for contrastive language-image learning
29.02.2024ContinueFig. 2 is unclear. How do they obtain a vector for a bag-of-words?
07.03.2024Still (sic!) continueIt seems that they train using BoW, even though their inference pipeline does not reflect this.
14.03.2024Sigmoid Loss for Language Image Pre-TrainingX. Zhai, B. Mustafa, A. Kolesnikov, L. BeyerHuggingFace
21.03.2024Continue
28.03.2024Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningW. Liang, Y. Zhang, Y. Kwon, S. Yeung, J. ZouGitHub Project Page
04.04.2024What Makes Training Multi-modal Classification Networks Hard?Wang, Tran, Feiszli
11.04.2024MultiBench: Multiscale Benchmarks for Multimodal Representation LearningLiang, Lyu, Fan, Wu, Cheng, Wu, Chen, Wu, Lee, Zhu, Salakhutdinaov, MorencyGitHub Project PageDemos
16.04.2024Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsTong, Liu, Zhai, Ma, LeCun, XieGitHub Project PageHuggingFace
23.04.2024Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training StrategiesLi, Xie, Cubuk
30.04.2024Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsLu, Peng, Cheng, Galley, Chang, Wu, Zhu, GaoGitHub, Project Page
07.05.2024Many-Shot In-Context LearningAgarwal, Singh, Zhang, Bohnet, Chan, Anand, Abbas, Nova, Co-Reyes, Chu, Behbahani, Faust, LarochelleNot Provided
28.05.2024BABILong: a long-context needle-in-a-haystack benchmark for LLMsKuratob, Bulatov, Anokhin, Sorokin, Sorokin, BurtsevGitHub
04.06.2024Continue
11.06.20244M: Massively Multimodal Masked ModelingMizrahi, Bachmann, Kar, Yeo, Gao, Dehghan, ZamirGitHub Project Page
18.06.2024Continue
25.06.2024GLaMM: Pixel Grounding Large Multimodal ModelRasheed, Maaz, Shaji, Shaker, Khan, Cholakkal, Anwer, Xing, Yang, KhanGitHub Project PageDemo
02.07.2024Code Reading Group
09.07.2024Knowledge DistillationGemma 2 (pdf), MobileLLM, Knowledge distillation, On-Policy distillation of Language Models
16.07.2024Whiteboard-of-Thought: Thinking Step-by-Step Across ModalitiesMenon, Zemel, VondrickProject Page
23.07.2024Multimodal Neurons in Artificial Neural NetworksGoh, Cammarata, Voss, Carter, Petrov, Schubert, Radford, Olah
30.07.2024Continue + (very briefly) CLIPPO
06.08.2024Does my multimodal model learn cross-modal interactions? It's harder to tell than you might think!Hessel, Lee
13.08.2024Graph of Thoughts and Monte Carlo Tree SearchMonte Carlo Tree Search from Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B; Graph of Thoughts; Large Language Monkeys; STaR: Self-Taught Reasoner; Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents; Bonus! DeepSeek-Prover-V1.5Tinygrad example of MCTS
15.10.2024
22.10.2024Calibration Multimodal LearningMa, Zhang, Wu, Fu, Hu
08.11.2024Towards Mamba: the S4 model and topic around: HiPPo, S4 paper, Annotated S4 blog postGu, Goel, RéGitHub
15.11.2024Continue: HiPPO and S4
22.11.2024Continue: Mamba & differences between Transformers and SSMs
07.02.2025RLHF Block. Part 1. Motivation and RLHF fine-tuningStiennon and friends Learning to summarize from human feedback (RLHF for summarisation); Ouyang and friends Training language models to follow instructions with human feedback (RLHF for wide range of NLP-related tasks)
14.02.2025RLHF Block. Part 2. Motivation and RLHF fine-tuningStiennon et al., Ouyand et al. + Ziegler et al. Fine-Tuning Language Models from Human Preferences
21.02.2025RLHF Block. Part 3. Motivation and RLHF fine-tuningFinish discussing Ouyand et al.
24.02.2025RLHF Block. Part 4. Reinforcement Learning. Q-Learning and PPOSlides
21.03.2025RLHF Block. Part 5. DPO and oher alternatives to RLHF.Rafailov et al. The DPO paper, Transition from formula 3 to formula 4 from the DPO paper, the Rejection sampling paper, the Llama 2 paper, RLAIF
11.04.2025Reminder on Chain-of-ThoughtArXiv
02.05.2025DeepSeek-MathShao, Wang, Zhu, Xu, Song, Bi, Zhang, Zhang, Li, Wu, GuoGitHub
16.05.2025DeepSeek-ProverXin, Guo, Shao, Ren, Zhu, Liu, Ruan, Li, LiangProver-V2, Prover-V2-GitHub
30.05.2025DeepSeek-R1DeepSeek-AI

Datasets and benchmarks
Surveys
Representation Learning
Latent Space Structure
Fusion
Modality Competition. Quantitative Methods of Detection of Suboptimality.

Contributors

iburenko

36 commits