Curated list of research papers published in 2024 related to Large Language Models (LLM)
50
7 commits
updated Apr 14, 2026
Hello, curious reader! ๐ Are you someone who loves exploring the cutting-edge world of AI and large language models (LLMs)? Or maybe you're a researcher, developer, or simply an enthusiast wondering about how these incredible technologies are evolving? Either way, you're in the right place. This repository is your gateway to 600+ groundbreaking research papers that capture the exciting developments in AI (in recent months)
These papers cover a wide spectrum of topics, ranging from improving model efficiency to aligning AI with human values, and even diving into the magic of multimodal models. The field of AI is advancing faster than ever, and these works showcase the brilliant ideas and innovations shaping the future.
Papers that introduce innovative architectures, explore scaling laws, or improve the efficiency of LLMs. These works aim to make AI faster, cheaper, and better.
Research exploring how to align AI with human preferences, ensuring outputs are not just accurate but also ethical, safe, and aligned with societal values.
Ever wondered how AI can see, read, and understand all at once? These papers focus on models that combine multiple types of data, like images, text, and audio, to build richer, more versatile systems.
Tackling the challenges of long-context understanding, these papers discuss how to extend the memory of LLMs, making them capable of reasoning across longer documents or conversations.
From enabling AI to solve complex puzzles to improving its reasoning capabilities, these papers explore how LLMs "think" and manage the vast knowledge they've been trained on.
Smaller, faster, and more efficientโthese papers delve into techniques like quantization and compression to make AI models more practical for real-world applications.
A great model needs great evaluation. These papers propose new benchmarks and methodologies to assess the capabilities of AI systems more rigorously.
Want your AI to follow your instructions perfectly? These papers refine how we align LLMs to understand and execute instructions across various tasks.
Meta-level research analyzing trends, methods, and applications in AI. Perfect for getting a birdโs-eye view of where the field is heading.
Papers showcasing how LLMs are applied in diverse domains, from healthcare to coding and everything in between. These highlight the transformative power of AI in the real world.
These papers represent the state of the art in AI research. Whether you're here to stay updated, seek inspiration, or find practical tools, this collection is a treasure trove for AI enthusiasts. Dive in, explore, and let your curiosity guide you! ๐โจ
Ready to start? ๐ Browse the list and explore the exciting breakthroughs shaping the future of AI! ๐
| ๐ข S.No. | ๐ Paper Title | ๐ Link |
|---|---|---|
| 1 | LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning | Link |
| 2 | Knowledge Fusion of Large Language Models | Link |
| 3 | A Comprehensive Study of Knowledge Editing for Large Language Models | Link |
| 4 | DiffusionGPT: LLM-Driven Text-to-Image Generation System | Link |
| 5 | Tuning Language Models by Proxy | Link |
| 6 | An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models | Link |
| 7 | Soaring from 4K to 400K: Extending LLMโs Context with Activation Beacon | Link |
| 8 | VMamba: Visual State Space Model | Link |
| 9 | LLaMA Beyond English: An Empirical Study on Language Capability Transfer | Link |
| 10 | DeepSeek LLM: Scaling Open-Source Language Models with Longtermism | Link |
| 11 | Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM | Link |
| 12 | LLaMA Pro: Progressive LLaMA with Block Expansion | Link |
| 13 | RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation | Link |
| 14 | Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text | Link |
| 15 | Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling | Link |
| 16 | WARM: On the Benefits of Weight Averaged Reward Models | Link |
| 17 | SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection | Link |
| 18 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity | Link |
| 19 | Mixtral of Experts | Link |
| 20 | MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts | Link |
| 21 | Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering | Link |
| 22 | EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty | Link |
| 23 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization | Link |
| 24 | A Closer Look at AUROC and AUPRC under Class Imbalance | Link |
| 25 | Transformers are Multi-State RNNs | Link |
| 26 | LLM Augmented LLMs: Expanding Capabilities through Composition | Link |
| 27 | Rethinking Patch Dependence for Masked Autoencoders | Link |
| 28 | Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models | Link |
| 29 | Pix2gestalt: Amodal Segmentation by Synthesizing Wholes | Link |
| 30 | RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture | Link |
| 31 | An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models | Link |
| 32 | Knowledge Fusion of Large Language Models | Link |
| 33 | Scalable Pre-training of Large Autoregressive Image Models | Link |
| 34 | SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities | Link |
| 35 | Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities | Link |
| 36 | MambaByte: Token-free Selective State Space Model | Link |
| 37 | ReFT: Reasoning with Reinforced Fine-Tuning | Link |
| 38 | Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models | Link |
| 39 | Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training | Link |
| 40 | Denoising Vision Transformers | Link |
| 41 | Self-Rewarding Language Models | Link |
| 42 | LoRA+: Efficient Low Rank Adaptation of Large Models | Link |
| 43 | MobileVLM V2: Faster and Stronger Baseline for Vision Language Model | Link |
| 44 | Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization? | Link |
| 45 | ODIN: Disentangled Reward Mitigates Hacking in RLHF | Link |
| 46 | Genie: Generative Interactive Environments | Link |
| 47 | A Phase Transition Between Positional and Semantic Learning in a Solvable Model of Dot-Product Attention | Link |
| 48 | Neural Network Diffusion | Link |
| 49 | More Agents Is All You Need | Link |
| 50 | Scaling Laws for Downstream Task Performance of Large Language Models | Link |
| -------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------ |
| 51 | Repeat After Me: Transformers are Better than State Space Models at Copying | Link |
| 52 | Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models | Link |
| 53 | AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling | Link |
| 54 | BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of Data | Link |
| 55 | LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration | Link |
| 56 | LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens | Link |
| 57 | Policy Improvement using Language Feedback Models | Link |
| 58 | DoRA: Weight-Decomposed Low-Rank Adaptation | Link |
| 59 | FindingEmo: An Image Dataset for Emotion Recognition in the Wild | Link |
| 60 | TinyLLaVA: A Framework of Small-scale Large Multimodal Models | Link |
| 61 | Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models | Link |
| 62 | When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method | Link |
| 63 | Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping | Link |
| 64 | Suppressing Pink Elephants with Direct Principle Feedback | Link |
| 65 | The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | Link |
| 66 | LiPO: Listwise Preference Optimization through Learning-to-Rank | Link |
| 67 | Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model | Link |
| 68 | The Boundary of Neural Network Trainability is Fractal | Link |
| 69 | Recovering the Pre-Fine-Tuning Weights of Generative Models | Link |
| 70 | Scaling Laws for Fine-Grained Mixture of Experts | Link |
| 71 | Direct Language Model Alignment from Online AI Feedback | Link |
| 72 | CARTE: Pretraining and Transfer for Tabular Learning | Link |
| 73 | Grandmaster-Level Chess Without Search | Link |
| 74 | Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs | Link |
| 75 | Reformatted Alignment | Link |
| 76 | OLMo: Accelerating the Science of Language Models | Link |
| 77 | FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models | Link |
| 78 | Mixtures of Experts Unlock Parameter Scaling for Deep RL | Link |
| 79 | Generative Representational Instruction Tuning | Link |
| 80 | World Model on Million-Length Video And Language With RingAttention | Link |
| 81 | Efficient Exploration for LLMs | Link |
| 82 | YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information | Link |
| 83 | Towards Cross-Tokenizer Distillation: The Universal Logit Distillation Loss for LLMs | Link |
| 84 | Mixtures of Experts Unlock Parameter Scaling for Deep RL | Link |
| 85 | Sora Generates Videos with Stunning Geometrical Consistency | Link |
| 86 | BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of Data | Link |
| 87 | Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs | Link |
| 88 | Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models | Link |
| 89 | DoRA: Weight-Decomposed Low-Rank Adaptation | Link |
| 90 | The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | Link |
| 91 | Efficient Exploration for LLMs | Link |
| 92 | Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training | Link |
| 93 | You Only Cache Once: Decoder-Decoder Architectures for Language Models | Link |
| 94 | gzip Predicts Data-dependent Scaling Laws | Link |
| 95 | Self-Play Preference Optimization for Language Model Alignment | Link |
| 96 | PHUDGE: Phi-3 as Scalable Judge | Link |
| 97 | What Matters When Building Vision-Language Models? | Link |
| 98 | Towards Modular LLMs by Building and Reusing a Library of LoRAs | Link |
| 99 | Contextual Position Encoding: Learning to Count Whatโs Important | Link |
| 100 | RLHF Workflow: From Reward Modeling to Online RLHF | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 101 | Attention as an RNN | Link |
| 102 | AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability | Link |
| 103 | Instruction Tuning With Loss Over Instructions | Link |
| 104 | LoRA Learns Less and Forgets Less | Link |
| 105 | Trans-LoRA: Towards Data-free Transferable Parameter Efficient Finetuning | Link |
| 106 | VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections | Link |
| 107 | MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning | Link |
| 108 | LLaMA-NAS: Efficient Neural Architecture Search for Large Language Models | Link |
| 109 | SimPO: Simple Preference Optimization with a Reference-Free Reward | Link |
| 110 | The Road Less Scheduled | Link |
| 111 | Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models | Link |
| 112 | Is Flash Attention Stable? | Link |
| 113 | Value Augmented Sampling for Language Model Alignment and Personalization | Link |
| 114 | Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? | Link |
| 115 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention | Link |
| 116 | A Careful Examination of Large Language Model Performance on Grade School Arithmetic | Link |
| 117 | Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models | Link |
| 118 | DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model | Link |
| 119 | Dense Connector for MLLMs | Link |
| 120 | xLSTM: Extended Long Short-Term Memory | Link |
| 121 | SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization | Link |
| 122 | Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model | Link |
| 123 | Chameleon: Mixed-Modal Early-Fusion Foundation Models | Link |
| 124 | Is Bigger Edit Batch Size Always Better? An Empirical Study on Model Editing with Llama-3 | Link |
| 125 | AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability | Link |
| 126 | The Prompt Report: A Systematic Survey of Prompting Techniques | Link |
| 127 | Creativity Has Left the Chat: The Price of Debiasing Language Models | Link |
| 128 | Show, Donโt Tell: Aligning Language Models with Demonstrated Feedback | Link |
| 129 | WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild | Link |
| 130 | Scalable MatMul-free Language Modeling | Link |
| 131 | Never Miss A Beat: An Efficient Recipe for Context Window Extension of Large Language Models | Link |
| 132 | Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach | Link |
| 133 | Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models | Link |
| 134 | MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding | Link |
| 135 | Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling | Link |
| 136 | An Image is Worth 32 Tokens for Reconstruction and Generation | Link |
| 137 | Block Transformer: Global-to-Local Language Modeling for Fast Inference | Link |
| 138 | 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination | Link |
| 139 | Transformers Need Glasses! Information Over-Squashing in Language Tasks | Link |
| 140 | The Geometry of Categorical and Hierarchical Concepts in Large Language Models | Link |
| 141 | BERTs are Generative In-Context Learners | Link |
| 142 | An Empirical Study of Mamba-based Language Models | Link |
| 143 | Discovering Preference Optimization Algorithms with and for Large Language Models | Link |
| 144 | Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing | Link |
| 145 | Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation | Link |
| 146 | Are We Done with MMLU? | Link |
| 147 | OLoRA: Orthonormal Low-Rank Adaptation of Large Language Models | Link |
| 148 | Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each Step | Link |
| 149 | Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning | Link |
| 150 | Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 151 | Simple and Effective Masked Diffusion Language Models | Link |
| 152 | The Prompt Report: A Systematic Survey of Prompting Techniques | Link |
| 153 | Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs | Link |
| 154 | Be like a Goldfish, Donโt Memorize! Mitigating Memorization in Generative LLMs | Link |
| 155 | TextGrad: Automatic โDifferentiationโ via Text | Link |
| 156 | Large Language Models Must Be Taught to Know What They Donโt Know | Link |
| 157 | What If We Recaption Billions of Web Images with LLaMA-3? | Link |
| 158 | Discovering Preference Optimization Algorithms with and for Large Language Models | Link |
| 159 | An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels | Link |
| 160 | Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models | Link |
| 161 | CRAG โ Comprehensive RAG Benchmark | Link |
| 162 | Margin-aware Preference Optimization for Aligning Diffusion Models Without Reference | Link |
| 163 | Mixture-of-Agents Enhances Large Language Model Capabilities | Link |
| 164 | Large Language Model Unlearning via Embedding-Corrupted Prompts | Link |
| 165 | Bootstrapping Language Models with DPO Implicit Rewards | Link |
| 166 | THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response Generation | Link |
| 167 | Task Me Anything | Link |
| 168 | Nemotron-4 340B Technical Report | Link |
| 169 | Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges | Link |
| 170 | How Do Large Language Models Acquire Factual Knowledge During Pretraining? | Link |
| 171 | mDPO: Conditional Preference Optimization for Multimodal Large Language Models | Link |
| 172 | Unveiling Encoder-Free Vision-Language Models | Link |
| 173 | HARE: HumAn pRiors, a key to small language model Efficiency | Link |
| 174 | Measuring Memorization in RLHF for Code Completion | Link |
| 175 | DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence | Link |
| 176 | Iterative Length-Regularized Direct Preference Optimization: Improving 7B Language Models to GPT-4 Level | Link |
| 177 | From RAGs to Rich Parameters: Probing How Language Models Utilize External Knowledge Over Parametric Information | Link |
| 178 | DataComp-LM: In Search of the Next Generation of Training Sets for Language Models | Link |
| 179 | Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More? | Link |
| 180 | Instruction Pre-Training: Language Models are Supervised Multitask Learners | Link |
| 181 | Can LLMs Learn by Teaching? A Preliminary Study | Link |
| 182 | A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems | Link |
| 183 | LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs | Link |
| 184 | MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression | Link |
| 185 | Efficient Continual Pre-training by Mitigating the Stability Gap | Link |
| 186 | Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers | Link |
| 187 | WARP: On the Benefits of Weight Averaged Rewarded Policies | Link |
| 188 | Adam-mini: Use Fewer Learning Rates To Gain More | Link |
| 189 | The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale | Link |
| 190 | LongIns: A Challenging Long-context Instruction-based Exam for LLMs | Link |
| 191 | Following Length Constraints in Instructions | Link |
| 192 | A Closer Look into Mixture-of-Experts in Large Language Models | Link |
| 193 | RouteLLM: Learning to Route LLMs with Preference Data | Link |
| 194 | Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs | Link |
| 195 | Dataset Size Recovery from LoRA Weights | Link |
| 196 | From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data | Link |
| 197 | Changing Answer Order Can Decrease MMLU Accuracy | Link |
| 198 | Direct Preference Knowledge Distillation for Large Language Models | Link |
| 199 | LLM Critics Help Catch LLM Bugs | Link |
| 200 | Scaling Synthetic Data Creation with 1,000,000,000 Personas | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 201 | Tokenization Falling Short: The Curse of Tokenization | Link |
| 202 | Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs | Link |
| 203 | Bootstrapping Language Models with DPO Implicit Rewards | Link |
| 204 | Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges | Link |
| 205 | Learning and Leveraging World Models in Visual Representation Learning | Link |
| 206 | Improving LLM Code Generation with Grammar Augmentation | Link |
| 207 | The Hidden Attention of Mamba Models | Link |
| 208 | Training-Free Pretrained Model Merging | Link |
| 209 | Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures | Link |
| 210 | The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning | Link |
| 211 | Evolution Transformer: In-Context Evolutionary Optimization | Link |
| 212 | Enhancing Vision-Language Pre-training with Rich Supervisions | Link |
| 213 | Scaling Rectified Flow Transformers for High-Resolution Image Synthesis | Link |
| 214 | Design2Code: How Far Are We From Automating Front-End Engineering? | Link |
| 215 | ShortGPT: Layers in Large Language Models are More Redundant Than You Expect | Link |
| 216 | Backtracing: Retrieving the Cause of the Query | Link |
| 217 | Learning to Decode Collaboratively with Multiple Language Models | Link |
| 218 | SaulLM-7B: A pioneering Large Language Model for Law | Link |
| 219 | Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning | Link |
| 220 | 3D Diffusion Policy | Link |
| 221 | MedMamba: Vision Mamba for Medical Image Classification | Link |
| 222 | GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection | Link |
| 223 | Stop Regressing: Training Value Functions via Classification for Scalable Deep RL | Link |
| 224 | How Far Are We from Intelligent Visual Deductive Reasoning? | Link |
| 225 | Common 7B Language Models Already Possess Strong Math Capabilities | Link |
| 226 | Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context | Link |
| 227 | Is Cosine-Similarity of Embeddings Really About Similarity? | Link |
| 228 | LLM4Decompile: Decompiling Binary Code with Large Language Models | Link |
| 229 | Algorithmic Progress in Language Models | Link |
| 230 | Stealing Part of a Production Language Model | Link |
| 231 | Chronos: Learning the Language of Time Series | Link |
| 232 | Simple and Scalable Strategies to Continually Pre-train Large Language Models | Link |
| 233 | Language Models Scale Reliably With Over-Training and on Downstream Tasks | Link |
| 234 | BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences | Link |
| 235 | LocalMamba: Visual State Space Model with Windowed Selective Scan | Link |
| 236 | GiT: Towards Generalist Vision Transformer through Universal Language Interface | Link |
| 237 | MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training | Link |
| 238 | RAFT: Adapting Language Model to Domain Specific RAG | Link |
| 239 | TnT-LLM: Text Mining at Scale with Large Language Models | Link |
| 240 | Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression | Link |
| 241 | PERL: Parameter Efficient Reinforcement Learning from Human Feedback | Link |
| 242 | RewardBench: Evaluating Reward Models for Language Modeling | Link |
| 243 | LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models | Link |
| 244 | RakutenAI-7B: Extending Large Language Models for Japanese | Link |
| 245 | SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time Series | Link |
| 246 | Can Large Language Models Explore In-Context? | Link |
| 247 | LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement | Link |
| 248 | LLM Agent Operating System | Link |
| 249 | The Unreasonable Ineffectiveness of the Deeper Layers | Link |
| 250 | BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 251 | ViTAR: Vision Transformer with Any Resolution | Link |
| 252 | Long-form Factuality in Large Language Models | Link |
| 253 | Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models | Link |
| 254 | LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning | Link |
| 255 | Mechanistic Design and Scaling of Hybrid Architectures | Link |
| 256 | MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions | Link |
| 257 | Model Stock: All We Need Is Just a Few Fine-Tuned Models | Link |
| 258 | Do Language Models Plan Ahead for Future Tokens? | Link |
| 259 | Bigger is not Always Better: Scaling Properties of Latent Diffusion Models | Link |
| 260 | The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis | Link |
| 261 | Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models | Link |
| 262 | Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models | Link |
| 263 | Long-context LLMs Struggle with Long In-context Learning | Link |
| 264 | Emergent Abilities in Reduced-Scale Generative Language Models | Link |
| 265 | Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks | Link |
| 266 | On the Scalability of Diffusion-based Text-to-Image Generation | Link |
| 267 | BAdam: A Memory Efficient Full Parameter Training Method for Large Language Models | Link |
| 268 | Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion Models | Link |
| 269 | Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences | Link |
| 270 | Training LLMs over Neurally Compressed Text | Link |
| 271 | CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues | Link |
| 272 | ReFT: Representation Finetuning for Language Models | Link |
| 273 | Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data | Link |
| 274 | Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation | Link |
| 275 | AutoCodeRover: Autonomous Program Improvement | Link |
| 276 | Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence | Link |
| 277 | CodecLM: Aligning Language Models with Tailored Synthetic Data | Link |
| 278 | MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies | Link |
| 279 | Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models | Link |
| 280 | LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders | Link |
| 281 | Adapting LLaMA Decoder to Vision Transformer | Link |
| 282 | Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention | Link |
| 283 | LLoCO: Learning Long Contexts Offline | Link |
| 284 | JetMoE: Reaching Llama2 Performance with 0.1M Dollars | Link |
| 285 | Best Practices and Lessons Learned on Synthetic Data for Language Models | Link |
| 286 | Rho-1: Not All Tokens Are What You Need | Link |
| 287 | Pre-training Small Base LMs with Fewer Tokens | Link |
| 288 | Dataset Reset Policy Optimization for RLHF | Link |
| 289 | LLM In-Context Recall is Prompt Dependent | Link |
| 290 | State Space Model for New-Generation Network Alternative to Transformers: A Survey | Link |
| 291 | Chinchilla Scaling: A Replication Attempt | Link |
| 292 | Learn Your Reference Model for Real Good Alignment | Link |
| 293 | Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study | Link |
| 294 | Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies | Link |
| 295 | How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMsโ Internal Prior | Link |
| 296 | A Survey on Retrieval-Augmented Text Generation for Large Language Models | Link |
| 297 | When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes | Link |
| 298 | Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing | Link |
| 299 | OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data | Link |
| 300 | The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 301 | How Good Are Low-bit Quantized LLaMA3 Models? An Empirical Study | Link |
| 302 | Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone | Link |
| 303 | OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework | Link |
| 304 | A Survey on Self-Evolution of Large Language Models | Link |
| 305 | Multi-Head Mixture-of-Experts | Link |
| 306 | NExT: Teaching Large Language Models to Reason about Code Execution | Link |
| 307 | Graph Machine Learning in the Era of Large Language Models (LLMs) | Link |
| 308 | Retrieval Head Mechanistically Explains Long-Context Factuality | Link |
| 309 | Layer Skip: Enabling Early Exit Inference and Self-Speculative Decoding | Link |
| 310 | Make Your LLM Fully Utilize the Context | Link |
| 311 | LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report | Link |
| 312 | Better & Faster Large Language Models via Multi-token Prediction | Link |
| 313 | RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language Processing | Link |
| 314 | A Primer on the Inner Workings of Transformer-based Language Models | Link |
| 315 | When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively | Link |
| 316 | KAN: KolmogorovโArnold Networks | Link |
| 317 | LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives | Link |
| 318 | Searching for Best Practices in Retrieval-Augmented Generation | Link |
| 319 | Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models | Link |
| 320 | Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion | Link |
| 321 | Eliminating Position Bias of Language Models: A Mechanistic Approach | Link |
| 322 | JMInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention | Link |
| 323 | TokenPacker: Efficient Visual Projector for Multimodal LLM | Link |
| 324 | Reasoning in Large Language Models: A Geometric Perspective | Link |
| 325 | RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs | Link |
| 326 | AgentInstruct: Toward Generative Teaching with Agentic Flows | Link |
| 327 | HEMM: Holistic Evaluation of Multimodal Foundation Models | Link |
| 328 | Mixture of A Million Experts | Link |
| 329 | Learning to (Learn at Test Time): RNNs with Expressive Hidden States | Link |
| 330 | Vision Language Models Are Blind | Link |
| 331 | Self-Recognition in Language Models | Link |
| 332 | Inference Performance Optimization for Large Language Models on CPUs | Link |
| 333 | Gradient Boosting Reinforcement Learning | Link |
| 334 | FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision | Link |
| 335 | SpreadsheetLLM: Encoding Spreadsheets for Large Language Models | Link |
| 336 | New Desiderata for Direct Preference Optimization | Link |
| 337 | Context Embeddings for Efficient Answer Generation in RAG | Link |
| 338 | Qwen2 Technical Report | Link |
| 339 | The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism | Link |
| 340 | From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank Gradients | Link |
| 341 | GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression | Link |
| 342 | Scaling Diffusion Transformers to 16 Billion Parameters | Link |
| 343 | NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window? | Link |
| 344 | Patch-Level Training for Large Language Models | Link |
| 345 | LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models | Link |
| 346 | A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks | Link |
| 347 | Spectra: A Comprehensive Study of Ternary, Quantized, and FP16 Language Models | Link |
| 348 | Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation | Link |
| 349 | Weak-to-Strong Reasoning | Link |
| 350 | Understanding Reference Policies in Direct Preference Optimization | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 351 | Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies | Link |
| 352 | BOND: Aligning LLMs with Best-of-N Distillation | Link |
| 353 | Compact Language Models via Pruning and Knowledge Distillation | Link |
| 354 | LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference | Link |
| 355 | Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training | Link |
| 356 | DDK: Distilling Domain Knowledge for Efficient Large Language Models | Link |
| 357 | Generation Constraint Scaling Can Mitigate Hallucination | Link |
| 358 | Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach | Link |
| 359 | Course-Correction: Safety Alignment Using Synthetic Preferences | Link |
| 360 | Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data? | Link |
| 361 | Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge | Link |
| 362 | Improving Retrieval Augmented Language Model with Self-Reasoning | Link |
| 363 | Apple Intelligence Foundation Language Models | Link |
| 364 | ThinK: Thinner Key Cache by Query-Driven Pruning | Link |
| 365 | The Llama 3 Herd of Models | Link |
| 366 | Gemma 2: Improving Open Language Models at a Practical Size | Link |
| 367 | SAM 2: Segment Anything in Images and Videos | Link |
| 368 | POA: Pre-training Once for Models of All Sizes | Link |
| 369 | RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework | Link |
| 370 | A Survey of Mamba | Link |
| 371 | MiniCPM-V: A GPT-4V Level MLLM on Your Phone | Link |
| 372 | RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation | Link |
| 373 | Self-Taught Evaluators | Link |
| 374 | BioMamba: A Pre-trained Biomedical Language Representation Model Leveraging Mamba | Link |
| 375 | EXAONE 3.0 7.8B Instruction Tuned Language Model | Link |
| 376 | 1.5-Pints Technical Report: Pretraining in Days, Not Months โ Your Language Model Thrives on Quality Data | Link |
| 377 | Conversational Prompt Engineering | Link |
| 378 | Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP | Link |
| 379 | The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Link |
| 380 | Hermes 3 Technical Report | Link |
| 381 | Customizing Language Models with Instance-wise LoRA for Sequential Recommendation | Link |
| 382 | Enhancing Robustness in Large Language Models: Prompting for Mitigating the Impact of Irrelevant Information | Link |
| 383 | To Code, or Not To Code? Exploring Impact of Code in Pre-training | Link |
| 384 | LLM Pruning and Distillation in Practice: The Minitron Approach | Link |
| 385 | Jamba-1.5: Hybrid Transformer-Mamba Models at Scale | Link |
| 386 | Controllable Text Generation for Large Language Models: A Survey | Link |
| 387 | Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time | Link |
| 388 | A Practitionerโs Guide to Continual Multimodal Pretraining | Link |
| 389 | Building and Better Understanding Vision-Language Models: Insights and Future Directions | Link |
| 390 | CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation | Link |
| 391 | The Mamba in the Llama: Distilling and Accelerating Hybrid Models | Link |
| 392 | ReMamba: Equip Mamba with Effective Long-Sequence Modeling | Link |
| 393 | Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling | Link |
| 394 | LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models | Link |
| 395 | Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis | Link |
| 396 | X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models | Link |
| 397 | Free Process Rewards without Process Labels | Link |
| 398 | Scaling Image Tokenizers with Grouped Spherical Quantization | Link |
| 399 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models | Link |
| 400 | Perception Tokens Enhance Visual Reasoning in Multimodal Language Models | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 401 | Evaluating Language Models as Synthetic Data Generators | Link |
| 402 | Best-of-N Jailbreaking | Link |
| 403 | PaliGemma 2: A Family of Versatile VLMs for Transfer | Link |
| 404 | VisionZip: Longer is Better but Not Necessary in Vision Language Models | Link |
| 405 | Evaluating and Aligning CodeLLMs on Human Preference | Link |
| 406 | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale | Link |
| 407 | Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling | Link |
| 408 | LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods | Link |
| 409 | Does RLHF Scale? Exploring the Impacts From Data, Model, and Method | Link |
| 410 | Unraveling the Complexity of Memory in RL Agents: An Approach for Classification and Evaluation | Link |
| 411 | Training Large Language Models to Reason in a Continuous Latent Space | Link |
| 412 | AutoReason: Automatic Few-Shot Reasoning Decomposition | Link |
| 413 | Large Concept Models: Language Modeling in a Sentence Representation Space | Link |
| 414 | Phi-4 Technical Report | Link |
| 415 | Byte Latent Transformer: Patches Scale Better Than Tokens | Link |
| 416 | SCBench: A KV Cache-Centric Analysis of Long-Context Methods | Link |
| 417 | Cultural Evolution of Cooperation among LLM Agents | Link |
| 418 | DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding | Link |
| 419 | No More Adam: Learning Rate Scaling at Initialization is All You Need | Link |
| 420 | Precise Length Control in Large Language Models | Link |
| 421 | The Open Source Advantage in Large Language Models (LLMs) | Link |
| 422 | A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges | Link |
| 423 | Are Your LLMs Capable of Stable Reasoning? | Link |
| 424 | LLM Post-Training Recipes: Improving Reasoning in LLMs | Link |
| 425 | Hansel: Output Length Controlling Framework for Large Language Models | Link |
| 426 | Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning | Link |
| 427 | Alignment Faking in Large Language Models | Link |
| 428 | SCOPE: Optimizing Key-Value Cache Compression in Long-Context Generation | Link |
| 429 | LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-Context Multitasks | Link |
| 430 | Offline Reinforcement Learning for LLM Multi-Step Reasoning | Link |
| 431 | Mulberry: Empowering MLLM with O1-like Reasoning and Reflection via Collective Monte Carlo Tree Search | Link |
| 432 | Titans: Learning to Memorize at Test Time | Link |
| 433 | Addition is All You Need for Energy-efficient Language Models | Link |
| 434 | Quantifying Generalization Complexity for Large Language Models | Link |
| 435 | When a Language Model is Optimized for Reasoning, Does It Still Show Embers of Autoregression? | Link |
| 436 | Were RNNs All We Needed? | Link |
| 437 | Selective Attention Improves Transformer | Link |
| 438 | LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations | Link |
| 439 | LLaVA-Critic: Learning to Evaluate Multimodal Models | Link |
| 440 | Differential Transformer | Link |
| 441 | GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models | Link |
| 442 | ARIA: An Open Multimodal Native Mixture-of-Experts Model | Link |
| 443 | O1 Replication Journey: A Strategic Progress Report โ Part 1 | Link |
| 444 | Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG | Link |
| 445 | From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning | Link |
| 446 | KV Prediction for Improved Time to First Token | Link |
| 447 | Baichuan-Omni Technical Report | Link |
| 448 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models | Link |
| 449 | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models | Link |
| 450 | AFlow: Automating Agentic Workflow Generation | Link |
| -------------- | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 451 | Toward General Instruction-Following Alignment for Retrieval-Augmented Generation | Link |
| 452 | Pre-training Distillation for Large Language Models: A Design Space Exploration | Link |
| 453 | MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models | Link |
| 454 | Scalable Ranked Preference Optimization for Text-to-Image Generation | Link |
| 455 | Scaling Diffusion Language Models via Adaptation from Autoregressive Models | Link |
| 456 | Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback | Link |
| 457 | Counting Ability of Large Language Models and Impact of Tokenization | Link |
| 458 | A Survey of Small Language Models | Link |
| 459 | Accelerating Direct Preference Optimization with Prefix Sharing | Link |
| 460 | Mind Your Step (by Step): Chain-of-Thought Can Reduce Performance on Tasks Where Thinking Makes Humans Worse | Link |
| 461 | LongReward: Improving Long-context Large Language Models with AI Feedback | Link |
| 462 | ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference | Link |
| 463 | Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications | Link |
| 464 | CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation | Link |
| 465 | What Happened in LLMs Layers When Trained for Fast vs. Slow Thinking: A Gradient Perspective | Link |
| 466 | GPT or BERT: Why Not Both? | Link |
| 467 | Language Models Can Self-Lengthen to Generate Long Texts | Link |
| 468 | OLMoE: Open Mixture-of-Experts Language Models | Link |
| 469 | In Defense of RAG in the Era of Long-Context Language Models | Link |
| 470 | Attention Heads of Large Language Models: A Survey | Link |
| 471 | LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA | Link |
| 472 | How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data | Link |
| 473 | Theory, Analysis, and Best Practices for Sigmoid Self-Attention | Link |
| 474 | LLaMA-Omni: Seamless Speech Interaction with Large Language Models | Link |
| 475 | What is the Role of Small Models in the LLM Era: A Survey | Link |
| 476 | Policy Filtration in RLHF to Fine-Tune LLM for Code Generation | Link |
| 477 | RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval | Link |
| 478 | Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement | Link |
| 479 | Qwen2.5-Coder Technical Report | Link |
| 480 | Instruction Following without Instruction Tuning | Link |
| 481 | Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis | Link |
| 482 | The Perfect Blend: Redefining RLHF with Mixture of Judges | Link |
| 483 | Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations | Link |
| 484 | Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation | Link |
| 485 | Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models | Link |
| 486 | Sample-Efficient Alignment for LLMs | Link |
| 487 | A Comprehensive Survey of Small Language Models in the Era of Large Language Models | Link |
| 488 | โGive Me BF16 or Give Me Deathโ? Accuracy-Performance Trade-Offs in LLM Quantization | Link |
| 489 | Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation | Link |
| 490 | HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems | Link |
| 491 | Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination | Link |
| 492 | Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding | Link |
| 493 | Number Cookbook: Number Understanding of Language Models and How to Improve It | Link |
| 494 | Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models | Link |
| 495 | BitNet a4.8: 4-bit Activations for 1-bit LLMs | Link |
| 496 | Scaling Laws for Precision | Link |
| 497 | Energy Efficient Protein Language Models | Link |
| 498 | Balancing Pipeline Parallelism with Vocabulary Parallelism | Link |
| 499 | Toward Optimal Search and Retrieval for RAG | Link |
| 500 | Large Language Models Can Self-Improve in Long-context Reasoning | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 501 | Stronger Models are NOT Stronger Teachers for Instruction Tuning | Link |
| 502 | Direct Preference Optimization Using Sparse Feature-Level Constraints | Link |
| 503 | Cut Your Losses in Large-Vocabulary Language Models | Link |
| 504 | Does Prompt Formatting Have Any Impact on LLM Performance? | Link |
| 505 | SymDPO: Boosting In-Context Learning of Large Multimodal Models | Link |
| 506 | SageAttention2 Technical Report | Link |
| 507 | Bi-Mamba: Towards Accurate 1-Bit State Space Models | Link |
| 508 | RedPajama: An Open Dataset for Training Large Language Models | Link |
| 509 | Hymba: A Hybrid-head Architecture for Small Language Models | Link |
| 510 | Loss-to-Loss Prediction: Scaling Laws for All Datasets | Link |
| 511 | When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training | Link |
| 512 | Multimodal Autoregressive Pre-training of Large Vision Encoders | Link |
| 513 | Natural Language Reinforcement Learning | Link |
| 514 | Large Multi-modal Models Can Interpret Features in Large Multi-modal Models | Link |
| 515 | TรLU 3: Pushing Frontiers in Open Language Model Post-Training | Link |
| 516 | MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs | Link |
| 517 | LLMs Do Not Think Step-by-step In Implicit Reasoning | Link |
| 518 | O1 Replication Journey โ Part 2 | Link |
| 519 | Star Attention: Efficient LLM Inference over Long Sequences | Link |
| 520 | Low-Bit Quantization Favors Undertrained LLMs | Link |
| 521 | Rethinking Token Reduction in MLLMs | Link |
| 522 | Reverse Thinking Makes LLMs Stronger Reasoners | Link |
| 523 | Critical Tokens Matter | Link |
| 524 | Foundations of Large Language Models | Link |
| 525 | A Survey of Research in Large Language Models for Electronic Design Automation | Link |
| 526 | Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models | Link |
| 527 | Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions | Link |
| 528 | The Future of AI: Exploring the Potential of Large Concept Models | Link |
| 529 | Investigating Numerical Translation with Large Language Models | Link |
| 530 | CALM: Curiosity-Driven Auditing for Large Language Models | Link |
๐ Explore cutting-edge research โจ and stay updated! ๐ง ๐
| ๐ข S.No. | ๐ Paper Title | ๐ Link |
|---|---|---|
| 531 | DeepSeek-V3 Technical Report | Link |
| 532 | DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning | Link |
| 533 | Kimi k1.5: Scaling Reinforcement Learning with LLMs | Link |
| 534 | Reasoning Language Models: A Blueprint | Link |
| 535 | Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought | Link |
| 536 | The Lessons of Developing Process Reward Models in Mathematical Reasoning | Link |
| 537 | LIMO: Less is More for Reasoning | Link |
| 538 | Demystifying Long Chain-of-Thought Reasoning in LLMs | Link |
| 539 | Competitive Programming with Large Reasoning Models | Link |
| 540 | LLMs Can Easily Learn to Reason from Demonstrations: Structure, Not Content, Is What Matters | Link |
| 541 | Training Language Models to Reason Efficiently | Link |
| 542 | Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning | Link |
| 543 | SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution | Link |
| 544 | On the Emergence of Thinking in LLMs: Searching for the Right Intuition | Link |
| 545 | Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning | Link |
| 546 | Teaching Language Models to Critique via Reinforcement Learning | Link |
| 547 | A Review of DeepSeek Models' Key Innovative Techniques | Link |
| 548 | R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning | Link |
| 549 | Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning | Link |
| 550 | Understanding R1-Zero-Like Training: A Critical Perspective | Link |
| 551 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning | Link |
| 552 | Open-Reasoner-Zero: An Open Source Approach to Scaling Up RL on the Base Model | Link |
| 553 | Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't | Link |
| 554 | Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in LLMs | Link |
| 555 | JudgeLRM: Large Reasoning Models as a Judge | Link |
| 556 | Concise Reasoning via Reinforcement Learning | Link |
| 557 | Absolute Zero: Reinforced Self-play Reasoning with Zero Data | Link |
| 558 | Qwen3 Technical Report | Link |
| 559 | MiMo: Unlocking the Reasoning Potential of Language Models โ From Pretraining to Posttraining | Link |
| 560 | Llama-Nemotron: Efficient Reasoning Models | Link |
| 561 | INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning | Link |
| 562 | AdaptThink: Reasoning Models Can Learn When to Think | Link |
| 563 | Thinkless: LLM Learns When to Think | Link |
| 564 | General-Reasoner: Advancing LLM Reasoning Across All Domains | Link |
| 565 | Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents | Link |
| 566 | Reinforcement Pre-Training | Link |
| 567 | Magistral | Link |
| 568 | AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery | Link |
| 569 | SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training | Link |
| 570 | Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond The Base Model? | Link |
| 571 | From System 1 to System 2: A Survey of Reasoning Large Language Models | Link |
| 572 | Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning LLMs | Link |
| 573 | Multi-Agent Collaboration Mechanisms: A Survey of LLMs | Link |
| 574 | Search-o1: Agentic Search-Enhanced Large Reasoning Models | Link |
| 575 | Reasoning Models Can Be Effective Without Thinking | Link |
| 576 | RM-R1: Reward Modeling as Reasoning | Link |
| 577 | QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning | Link |
| 578 | Enigmata: Scaling Logical Reasoning in LLMs with Synthetic Verifiable Puzzles | Link |
| 579 | ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in LLMs | Link |
| 580 | Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective RL for LLM Reasoning | Link |
| 581 | Spurious Rewards: Rethinking Training Signals in RLVR | Link |
| 582 | Tina: Tiny Reasoning Models via LoRA | Link |
| 583 | Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math | Link |
| 584 | VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning | Link |
| 585 | Gemini Robotics: Bringing AI Into the Physical World | Link |
| 586 | Gemini 2.0: The Era of Multimodal Agentic AI | Link |
| 587 | RARE: Retrieval-Augmented Reasoning Modeling | Link |
| 588 | Learning from Failures in Multi-Attempt Reinforcement Learning | Link |
| 589 | R1-VL: Learning to Reason with Multimodal LLMs via Step-wise Group Relative Policy Optimization | Link |
| 590 | Diffusion-Based Language Models: A Survey | Link |
| 591 | FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching | Link |
| 592 | DAPO: An Open-Source LLM Reinforcement Learning System at Scale | Link |
| 593 | Scaling Laws for Inference-Time Compute | Link |
| 594 | LLM Post-Training: A Deep Dive into Reasoning Large Language Models | Link |
| 595 | Long-VITA: Scaling Large Vision-Language Models for Long Video Understanding | Link |
| 596 | Wan: Open and Advanced Large-Scale Video Generative Models | Link |
| 597 | Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models | Link |
| 598 | Seed1.5-VL Technical Report | Link |
| 599 | A Survey on LLM-based Agents: Recent Advances and New Frontiers | Link |
| 600 | RLVR Is Not RL: Revisiting Reinforcement Learning for LLMs | Link |
| 601 | RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning | Link |
| 602 | Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning | Link |
| 603 | LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL | Link |
| 604 | Reinforcement Learning for Reasoning in Large Language Models with One Training Example | Link |
| 605 | Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability | Link |
| 606 | The First Few Tokens Are All You Need: Unsupervised Prefix Fine-Tuning for Reasoning Models | Link |
| 607 | Learning to Reason without External Rewards | Link |
| 608 | Genius: A Generalizable and Purely Unsupervised Self-Training Framework for Advanced Reasoning | Link |
| 609 | Reinforcement Learning Teachers of Test Time Scaling | Link |
| 610 | Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening | Link |
7 commits
Curated list of research papers published in 2024 related to Large Language Models (LLM)
50
7 commits
updated Apr 14, 2026
Hello, curious reader! ๐ Are you someone who loves exploring the cutting-edge world of AI and large language models (LLMs)? Or maybe you're a researcher, developer, or simply an enthusiast wondering about how these incredible technologies are evolving? Either way, you're in the right place. This repository is your gateway to 600+ groundbreaking research papers that capture the exciting developments in AI (in recent months)
These papers cover a wide spectrum of topics, ranging from improving model efficiency to aligning AI with human values, and even diving into the magic of multimodal models. The field of AI is advancing faster than ever, and these works showcase the brilliant ideas and innovations shaping the future.
Papers that introduce innovative architectures, explore scaling laws, or improve the efficiency of LLMs. These works aim to make AI faster, cheaper, and better.
Research exploring how to align AI with human preferences, ensuring outputs are not just accurate but also ethical, safe, and aligned with societal values.
Ever wondered how AI can see, read, and understand all at once? These papers focus on models that combine multiple types of data, like images, text, and audio, to build richer, more versatile systems.
Tackling the challenges of long-context understanding, these papers discuss how to extend the memory of LLMs, making them capable of reasoning across longer documents or conversations.
From enabling AI to solve complex puzzles to improving its reasoning capabilities, these papers explore how LLMs "think" and manage the vast knowledge they've been trained on.
Smaller, faster, and more efficientโthese papers delve into techniques like quantization and compression to make AI models more practical for real-world applications.
A great model needs great evaluation. These papers propose new benchmarks and methodologies to assess the capabilities of AI systems more rigorously.
Want your AI to follow your instructions perfectly? These papers refine how we align LLMs to understand and execute instructions across various tasks.
Meta-level research analyzing trends, methods, and applications in AI. Perfect for getting a birdโs-eye view of where the field is heading.
Papers showcasing how LLMs are applied in diverse domains, from healthcare to coding and everything in between. These highlight the transformative power of AI in the real world.
These papers represent the state of the art in AI research. Whether you're here to stay updated, seek inspiration, or find practical tools, this collection is a treasure trove for AI enthusiasts. Dive in, explore, and let your curiosity guide you! ๐โจ
Ready to start? ๐ Browse the list and explore the exciting breakthroughs shaping the future of AI! ๐
| ๐ข S.No. | ๐ Paper Title | ๐ Link |
|---|---|---|
| 1 | LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning | Link |
| 2 | Knowledge Fusion of Large Language Models | Link |
| 3 | A Comprehensive Study of Knowledge Editing for Large Language Models | Link |
| 4 | DiffusionGPT: LLM-Driven Text-to-Image Generation System | Link |
| 5 | Tuning Language Models by Proxy | Link |
| 6 | An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models | Link |
| 7 | Soaring from 4K to 400K: Extending LLMโs Context with Activation Beacon | Link |
| 8 | VMamba: Visual State Space Model | Link |
| 9 | LLaMA Beyond English: An Empirical Study on Language Capability Transfer | Link |
| 10 | DeepSeek LLM: Scaling Open-Source Language Models with Longtermism | Link |
| 11 | Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLM | Link |
| 12 | LLaMA Pro: Progressive LLaMA with Block Expansion | Link |
| 13 | RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust Adaptation | Link |
| 14 | Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text | Link |
| 15 | Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling | Link |
| 16 | WARM: On the Benefits of Weight Averaged Reward Models | Link |
| 17 | SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token Detection | Link |
| 18 | A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity | Link |
| 19 | Mixtral of Experts | Link |
| 20 | MoE-Mamba: Efficient Selective State Space Models with Mixture of Experts | Link |
| 21 | Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering | Link |
| 22 | EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty | Link |
| 23 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization | Link |
| 24 | A Closer Look at AUROC and AUPRC under Class Imbalance | Link |
| 25 | Transformers are Multi-State RNNs | Link |
| 26 | LLM Augmented LLMs: Expanding Capabilities through Composition | Link |
| 27 | Rethinking Patch Dependence for Masked Autoencoders | Link |
| 28 | Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models | Link |
| 29 | Pix2gestalt: Amodal Segmentation by Synthesizing Wholes | Link |
| 30 | RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture | Link |
| 31 | An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models | Link |
| 32 | Knowledge Fusion of Large Language Models | Link |
| 33 | Scalable Pre-training of Large Autoregressive Image Models | Link |
| 34 | SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities | Link |
| 35 | Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities | Link |
| 36 | MambaByte: Token-free Selective State Space Model | Link |
| 37 | ReFT: Reasoning with Reinforced Fine-Tuning | Link |
| 38 | Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models | Link |
| 39 | Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training | Link |
| 40 | Denoising Vision Transformers | Link |
| 41 | Self-Rewarding Language Models | Link |
| 42 | LoRA+: Efficient Low Rank Adaptation of Large Models | Link |
| 43 | MobileVLM V2: Faster and Stronger Baseline for Vision Language Model | Link |
| 44 | Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization? | Link |
| 45 | ODIN: Disentangled Reward Mitigates Hacking in RLHF | Link |
| 46 | Genie: Generative Interactive Environments | Link |
| 47 | A Phase Transition Between Positional and Semantic Learning in a Solvable Model of Dot-Product Attention | Link |
| 48 | Neural Network Diffusion | Link |
| 49 | More Agents Is All You Need | Link |
| 50 | Scaling Laws for Downstream Task Performance of Large Language Models | Link |
| -------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------ |
| 51 | Repeat After Me: Transformers are Better than State Space Models at Copying | Link |
| 52 | Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models | Link |
| 53 | AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling | Link |
| 54 | BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of Data | Link |
| 55 | LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration | Link |
| 56 | LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens | Link |
| 57 | Policy Improvement using Language Feedback Models | Link |
| 58 | DoRA: Weight-Decomposed Low-Rank Adaptation | Link |
| 59 | FindingEmo: An Image Dataset for Emotion Recognition in the Wild | Link |
| 60 | TinyLLaVA: A Framework of Small-scale Large Multimodal Models | Link |
| 61 | Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models | Link |
| 62 | When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method | Link |
| 63 | Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via Bootstrapping | Link |
| 64 | Suppressing Pink Elephants with Direct Principle Feedback | Link |
| 65 | The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | Link |
| 66 | LiPO: Listwise Preference Optimization through Learning-to-Rank | Link |
| 67 | Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model | Link |
| 68 | The Boundary of Neural Network Trainability is Fractal | Link |
| 69 | Recovering the Pre-Fine-Tuning Weights of Generative Models | Link |
| 70 | Scaling Laws for Fine-Grained Mixture of Experts | Link |
| 71 | Direct Language Model Alignment from Online AI Feedback | Link |
| 72 | CARTE: Pretraining and Transfer for Tabular Learning | Link |
| 73 | Grandmaster-Level Chess Without Search | Link |
| 74 | Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs | Link |
| 75 | Reformatted Alignment | Link |
| 76 | OLMo: Accelerating the Science of Language Models | Link |
| 77 | FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models | Link |
| 78 | Mixtures of Experts Unlock Parameter Scaling for Deep RL | Link |
| 79 | Generative Representational Instruction Tuning | Link |
| 80 | World Model on Million-Length Video And Language With RingAttention | Link |
| 81 | Efficient Exploration for LLMs | Link |
| 82 | YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information | Link |
| 83 | Towards Cross-Tokenizer Distillation: The Universal Logit Distillation Loss for LLMs | Link |
| 84 | Mixtures of Experts Unlock Parameter Scaling for Deep RL | Link |
| 85 | Sora Generates Videos with Stunning Geometrical Consistency | Link |
| 86 | BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of Data | Link |
| 87 | Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs | Link |
| 88 | Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models | Link |
| 89 | DoRA: Weight-Decomposed Low-Rank Adaptation | Link |
| 90 | The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits | Link |
| 91 | Efficient Exploration for LLMs | Link |
| 92 | Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training | Link |
| 93 | You Only Cache Once: Decoder-Decoder Architectures for Language Models | Link |
| 94 | gzip Predicts Data-dependent Scaling Laws | Link |
| 95 | Self-Play Preference Optimization for Language Model Alignment | Link |
| 96 | PHUDGE: Phi-3 as Scalable Judge | Link |
| 97 | What Matters When Building Vision-Language Models? | Link |
| 98 | Towards Modular LLMs by Building and Reusing a Library of LoRAs | Link |
| 99 | Contextual Position Encoding: Learning to Count Whatโs Important | Link |
| 100 | RLHF Workflow: From Reward Modeling to Online RLHF | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 101 | Attention as an RNN | Link |
| 102 | AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability | Link |
| 103 | Instruction Tuning With Loss Over Instructions | Link |
| 104 | LoRA Learns Less and Forgets Less | Link |
| 105 | Trans-LoRA: Towards Data-free Transferable Parameter Efficient Finetuning | Link |
| 106 | VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections | Link |
| 107 | MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning | Link |
| 108 | LLaMA-NAS: Efficient Neural Architecture Search for Large Language Models | Link |
| 109 | SimPO: Simple Preference Optimization with a Reference-Free Reward | Link |
| 110 | The Road Less Scheduled | Link |
| 111 | Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models | Link |
| 112 | Is Flash Attention Stable? | Link |
| 113 | Value Augmented Sampling for Language Model Alignment and Personalization | Link |
| 114 | Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations? | Link |
| 115 | vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention | Link |
| 116 | A Careful Examination of Large Language Model Performance on Grade School Arithmetic | Link |
| 117 | Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models | Link |
| 118 | DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model | Link |
| 119 | Dense Connector for MLLMs | Link |
| 120 | xLSTM: Extended Long Short-Term Memory | Link |
| 121 | SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch Normalization | Link |
| 122 | Xmodel-VLM: A Simple Baseline for Multimodal Vision Language Model | Link |
| 123 | Chameleon: Mixed-Modal Early-Fusion Foundation Models | Link |
| 124 | Is Bigger Edit Batch Size Always Better? An Empirical Study on Model Editing with Llama-3 | Link |
| 125 | AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability | Link |
| 126 | The Prompt Report: A Systematic Survey of Prompting Techniques | Link |
| 127 | Creativity Has Left the Chat: The Price of Debiasing Language Models | Link |
| 128 | Show, Donโt Tell: Aligning Language Models with Demonstrated Feedback | Link |
| 129 | WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild | Link |
| 130 | Scalable MatMul-free Language Modeling | Link |
| 131 | Never Miss A Beat: An Efficient Recipe for Context Window Extension of Large Language Models | Link |
| 132 | Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven Approach | Link |
| 133 | Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models | Link |
| 134 | MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding | Link |
| 135 | Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling | Link |
| 136 | An Image is Worth 32 Tokens for Reconstruction and Generation | Link |
| 137 | Block Transformer: Global-to-Local Language Modeling for Fast Inference | Link |
| 138 | 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination | Link |
| 139 | Transformers Need Glasses! Information Over-Squashing in Language Tasks | Link |
| 140 | The Geometry of Categorical and Hierarchical Concepts in Large Language Models | Link |
| 141 | BERTs are Generative In-Context Learners | Link |
| 142 | An Empirical Study of Mamba-based Language Models | Link |
| 143 | Discovering Preference Optimization Algorithms with and for Large Language Models | Link |
| 144 | Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing | Link |
| 145 | Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation | Link |
| 146 | Are We Done with MMLU? | Link |
| 147 | OLoRA: Orthonormal Low-Rank Adaptation of Large Language Models | Link |
| 148 | Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each Step | Link |
| 149 | Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning | Link |
| 150 | Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 151 | Simple and Effective Masked Diffusion Language Models | Link |
| 152 | The Prompt Report: A Systematic Survey of Prompting Techniques | Link |
| 153 | Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs | Link |
| 154 | Be like a Goldfish, Donโt Memorize! Mitigating Memorization in Generative LLMs | Link |
| 155 | TextGrad: Automatic โDifferentiationโ via Text | Link |
| 156 | Large Language Models Must Be Taught to Know What They Donโt Know | Link |
| 157 | What If We Recaption Billions of Web Images with LLaMA-3? | Link |
| 158 | Discovering Preference Optimization Algorithms with and for Large Language Models | Link |
| 159 | An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels | Link |
| 160 | Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models | Link |
| 161 | CRAG โ Comprehensive RAG Benchmark | Link |
| 162 | Margin-aware Preference Optimization for Aligning Diffusion Models Without Reference | Link |
| 163 | Mixture-of-Agents Enhances Large Language Model Capabilities | Link |
| 164 | Large Language Model Unlearning via Embedding-Corrupted Prompts | Link |
| 165 | Bootstrapping Language Models with DPO Implicit Rewards | Link |
| 166 | THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response Generation | Link |
| 167 | Task Me Anything | Link |
| 168 | Nemotron-4 340B Technical Report | Link |
| 169 | Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges | Link |
| 170 | How Do Large Language Models Acquire Factual Knowledge During Pretraining? | Link |
| 171 | mDPO: Conditional Preference Optimization for Multimodal Large Language Models | Link |
| 172 | Unveiling Encoder-Free Vision-Language Models | Link |
| 173 | HARE: HumAn pRiors, a key to small language model Efficiency | Link |
| 174 | Measuring Memorization in RLHF for Code Completion | Link |
| 175 | DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence | Link |
| 176 | Iterative Length-Regularized Direct Preference Optimization: Improving 7B Language Models to GPT-4 Level | Link |
| 177 | From RAGs to Rich Parameters: Probing How Language Models Utilize External Knowledge Over Parametric Information | Link |
| 178 | DataComp-LM: In Search of the Next Generation of Training Sets for Language Models | Link |
| 179 | Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More? | Link |
| 180 | Instruction Pre-Training: Language Models are Supervised Multitask Learners | Link |
| 181 | Can LLMs Learn by Teaching? A Preliminary Study | Link |
| 182 | A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG Systems | Link |
| 183 | LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs | Link |
| 184 | MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression | Link |
| 185 | Efficient Continual Pre-training by Mitigating the Stability Gap | Link |
| 186 | Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range Transformers | Link |
| 187 | WARP: On the Benefits of Weight Averaged Rewarded Policies | Link |
| 188 | Adam-mini: Use Fewer Learning Rates To Gain More | Link |
| 189 | The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale | Link |
| 190 | LongIns: A Challenging Long-context Instruction-based Exam for LLMs | Link |
| 191 | Following Length Constraints in Instructions | Link |
| 192 | A Closer Look into Mixture-of-Experts in Large Language Models | Link |
| 193 | RouteLLM: Learning to Route LLMs with Preference Data | Link |
| 194 | Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs | Link |
| 195 | Dataset Size Recovery from LoRA Weights | Link |
| 196 | From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data | Link |
| 197 | Changing Answer Order Can Decrease MMLU Accuracy | Link |
| 198 | Direct Preference Knowledge Distillation for Large Language Models | Link |
| 199 | LLM Critics Help Catch LLM Bugs | Link |
| 200 | Scaling Synthetic Data Creation with 1,000,000,000 Personas | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 201 | Tokenization Falling Short: The Curse of Tokenization | Link |
| 202 | Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs | Link |
| 203 | Bootstrapping Language Models with DPO Implicit Rewards | Link |
| 204 | Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges | Link |
| 205 | Learning and Leveraging World Models in Visual Representation Learning | Link |
| 206 | Improving LLM Code Generation with Grammar Augmentation | Link |
| 207 | The Hidden Attention of Mamba Models | Link |
| 208 | Training-Free Pretrained Model Merging | Link |
| 209 | Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures | Link |
| 210 | The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning | Link |
| 211 | Evolution Transformer: In-Context Evolutionary Optimization | Link |
| 212 | Enhancing Vision-Language Pre-training with Rich Supervisions | Link |
| 213 | Scaling Rectified Flow Transformers for High-Resolution Image Synthesis | Link |
| 214 | Design2Code: How Far Are We From Automating Front-End Engineering? | Link |
| 215 | ShortGPT: Layers in Large Language Models are More Redundant Than You Expect | Link |
| 216 | Backtracing: Retrieving the Cause of the Query | Link |
| 217 | Learning to Decode Collaboratively with Multiple Language Models | Link |
| 218 | SaulLM-7B: A pioneering Large Language Model for Law | Link |
| 219 | Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning | Link |
| 220 | 3D Diffusion Policy | Link |
| 221 | MedMamba: Vision Mamba for Medical Image Classification | Link |
| 222 | GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection | Link |
| 223 | Stop Regressing: Training Value Functions via Classification for Scalable Deep RL | Link |
| 224 | How Far Are We from Intelligent Visual Deductive Reasoning? | Link |
| 225 | Common 7B Language Models Already Possess Strong Math Capabilities | Link |
| 226 | Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of Context | Link |
| 227 | Is Cosine-Similarity of Embeddings Really About Similarity? | Link |
| 228 | LLM4Decompile: Decompiling Binary Code with Large Language Models | Link |
| 229 | Algorithmic Progress in Language Models | Link |
| 230 | Stealing Part of a Production Language Model | Link |
| 231 | Chronos: Learning the Language of Time Series | Link |
| 232 | Simple and Scalable Strategies to Continually Pre-train Large Language Models | Link |
| 233 | Language Models Scale Reliably With Over-Training and on Downstream Tasks | Link |
| 234 | BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences | Link |
| 235 | LocalMamba: Visual State Space Model with Windowed Selective Scan | Link |
| 236 | GiT: Towards Generalist Vision Transformer through Universal Language Interface | Link |
| 237 | MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training | Link |
| 238 | RAFT: Adapting Language Model to Domain Specific RAG | Link |
| 239 | TnT-LLM: Text Mining at Scale with Large Language Models | Link |
| 240 | Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression | Link |
| 241 | PERL: Parameter Efficient Reinforcement Learning from Human Feedback | Link |
| 242 | RewardBench: Evaluating Reward Models for Language Modeling | Link |
| 243 | LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models | Link |
| 244 | RakutenAI-7B: Extending Large Language Models for Japanese | Link |
| 245 | SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time Series | Link |
| 246 | Can Large Language Models Explore In-Context? | Link |
| 247 | LLM2LLM: Boosting LLMs with Novel Iterative Data Enhancement | Link |
| 248 | LLM Agent Operating System | Link |
| 249 | The Unreasonable Ineffectiveness of the Deeper Layers | Link |
| 250 | BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 251 | ViTAR: Vision Transformer with Any Resolution | Link |
| 252 | Long-form Factuality in Large Language Models | Link |
| 253 | Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models | Link |
| 254 | LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning | Link |
| 255 | Mechanistic Design and Scaling of Hybrid Architectures | Link |
| 256 | MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions | Link |
| 257 | Model Stock: All We Need Is Just a Few Fine-Tuned Models | Link |
| 258 | Do Language Models Plan Ahead for Future Tokens? | Link |
| 259 | Bigger is not Always Better: Scaling Properties of Latent Diffusion Models | Link |
| 260 | The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability Analysis | Link |
| 261 | Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models | Link |
| 262 | Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language Models | Link |
| 263 | Long-context LLMs Struggle with Long In-context Learning | Link |
| 264 | Emergent Abilities in Reduced-Scale Generative Language Models | Link |
| 265 | Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks | Link |
| 266 | On the Scalability of Diffusion-based Text-to-Image Generation | Link |
| 267 | BAdam: A Memory Efficient Full Parameter Training Method for Large Language Models | Link |
| 268 | Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion Models | Link |
| 269 | Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences | Link |
| 270 | Training LLMs over Neurally Compressed Text | Link |
| 271 | CantTalkAboutThis: Aligning Language Models to Stay on Topic in Dialogues | Link |
| 272 | ReFT: Representation Finetuning for Language Models | Link |
| 273 | Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data | Link |
| 274 | Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation | Link |
| 275 | AutoCodeRover: Autonomous Program Improvement | Link |
| 276 | Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence | Link |
| 277 | CodecLM: Aligning Language Models with Tailored Synthetic Data | Link |
| 278 | MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies | Link |
| 279 | Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models | Link |
| 280 | LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders | Link |
| 281 | Adapting LLaMA Decoder to Vision Transformer | Link |
| 282 | Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention | Link |
| 283 | LLoCO: Learning Long Contexts Offline | Link |
| 284 | JetMoE: Reaching Llama2 Performance with 0.1M Dollars | Link |
| 285 | Best Practices and Lessons Learned on Synthetic Data for Language Models | Link |
| 286 | Rho-1: Not All Tokens Are What You Need | Link |
| 287 | Pre-training Small Base LMs with Fewer Tokens | Link |
| 288 | Dataset Reset Policy Optimization for RLHF | Link |
| 289 | LLM In-Context Recall is Prompt Dependent | Link |
| 290 | State Space Model for New-Generation Network Alternative to Transformers: A Survey | Link |
| 291 | Chinchilla Scaling: A Replication Attempt | Link |
| 292 | Learn Your Reference Model for Real Good Alignment | Link |
| 293 | Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study | Link |
| 294 | Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies | Link |
| 295 | How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMsโ Internal Prior | Link |
| 296 | A Survey on Retrieval-Augmented Text Generation for Large Language Models | Link |
| 297 | When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes | Link |
| 298 | Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing | Link |
| 299 | OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data | Link |
| 300 | The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 301 | How Good Are Low-bit Quantized LLaMA3 Models? An Empirical Study | Link |
| 302 | Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone | Link |
| 303 | OpenELM: An Efficient Language Model Family with Open-source Training and Inference Framework | Link |
| 304 | A Survey on Self-Evolution of Large Language Models | Link |
| 305 | Multi-Head Mixture-of-Experts | Link |
| 306 | NExT: Teaching Large Language Models to Reason about Code Execution | Link |
| 307 | Graph Machine Learning in the Era of Large Language Models (LLMs) | Link |
| 308 | Retrieval Head Mechanistically Explains Long-Context Factuality | Link |
| 309 | Layer Skip: Enabling Early Exit Inference and Self-Speculative Decoding | Link |
| 310 | Make Your LLM Fully Utilize the Context | Link |
| 311 | LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report | Link |
| 312 | Better & Faster Large Language Models via Multi-token Prediction | Link |
| 313 | RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language Processing | Link |
| 314 | A Primer on the Inner Workings of Transformer-based Language Models | Link |
| 315 | When to Retrieve: Teaching LLMs to Utilize Information Retrieval Effectively | Link |
| 316 | KAN: KolmogorovโArnold Networks | Link |
| 317 | LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives | Link |
| 318 | Searching for Best Practices in Retrieval-Augmented Generation | Link |
| 319 | Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models | Link |
| 320 | Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion | Link |
| 321 | Eliminating Position Bias of Language Models: A Mechanistic Approach | Link |
| 322 | JMInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention | Link |
| 323 | TokenPacker: Efficient Visual Projector for Multimodal LLM | Link |
| 324 | Reasoning in Large Language Models: A Geometric Perspective | Link |
| 325 | RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs | Link |
| 326 | AgentInstruct: Toward Generative Teaching with Agentic Flows | Link |
| 327 | HEMM: Holistic Evaluation of Multimodal Foundation Models | Link |
| 328 | Mixture of A Million Experts | Link |
| 329 | Learning to (Learn at Test Time): RNNs with Expressive Hidden States | Link |
| 330 | Vision Language Models Are Blind | Link |
| 331 | Self-Recognition in Language Models | Link |
| 332 | Inference Performance Optimization for Large Language Models on CPUs | Link |
| 333 | Gradient Boosting Reinforcement Learning | Link |
| 334 | FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision | Link |
| 335 | SpreadsheetLLM: Encoding Spreadsheets for Large Language Models | Link |
| 336 | New Desiderata for Direct Preference Optimization | Link |
| 337 | Context Embeddings for Efficient Answer Generation in RAG | Link |
| 338 | Qwen2 Technical Report | Link |
| 339 | The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism | Link |
| 340 | From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank Gradients | Link |
| 341 | GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache Compression | Link |
| 342 | Scaling Diffusion Transformers to 16 Billion Parameters | Link |
| 343 | NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window? | Link |
| 344 | Patch-Level Training for Large Language Models | Link |
| 345 | LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models | Link |
| 346 | A Survey of Prompt Engineering Methods in Large Language Models for Different NLP Tasks | Link |
| 347 | Spectra: A Comprehensive Study of Ternary, Quantized, and FP16 Language Models | Link |
| 348 | Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation | Link |
| 349 | Weak-to-Strong Reasoning | Link |
| 350 | Understanding Reference Policies in Direct Preference Optimization | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 351 | Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies | Link |
| 352 | BOND: Aligning LLMs with Best-of-N Distillation | Link |
| 353 | Compact Language Models via Pruning and Knowledge Distillation | Link |
| 354 | LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference | Link |
| 355 | Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training | Link |
| 356 | DDK: Distilling Domain Knowledge for Efficient Large Language Models | Link |
| 357 | Generation Constraint Scaling Can Mitigate Hallucination | Link |
| 358 | Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach | Link |
| 359 | Course-Correction: Safety Alignment Using Synthetic Preferences | Link |
| 360 | Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data? | Link |
| 361 | Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge | Link |
| 362 | Improving Retrieval Augmented Language Model with Self-Reasoning | Link |
| 363 | Apple Intelligence Foundation Language Models | Link |
| 364 | ThinK: Thinner Key Cache by Query-Driven Pruning | Link |
| 365 | The Llama 3 Herd of Models | Link |
| 366 | Gemma 2: Improving Open Language Models at a Practical Size | Link |
| 367 | SAM 2: Segment Anything in Images and Videos | Link |
| 368 | POA: Pre-training Once for Models of All Sizes | Link |
| 369 | RAGEval: Scenario Specific RAG Evaluation Dataset Generation Framework | Link |
| 370 | A Survey of Mamba | Link |
| 371 | MiniCPM-V: A GPT-4V Level MLLM on Your Phone | Link |
| 372 | RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation | Link |
| 373 | Self-Taught Evaluators | Link |
| 374 | BioMamba: A Pre-trained Biomedical Language Representation Model Leveraging Mamba | Link |
| 375 | EXAONE 3.0 7.8B Instruction Tuned Language Model | Link |
| 376 | 1.5-Pints Technical Report: Pretraining in Days, Not Months โ Your Language Model Thrives on Quality Data | Link |
| 377 | Conversational Prompt Engineering | Link |
| 378 | Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP | Link |
| 379 | The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Link |
| 380 | Hermes 3 Technical Report | Link |
| 381 | Customizing Language Models with Instance-wise LoRA for Sequential Recommendation | Link |
| 382 | Enhancing Robustness in Large Language Models: Prompting for Mitigating the Impact of Irrelevant Information | Link |
| 383 | To Code, or Not To Code? Exploring Impact of Code in Pre-training | Link |
| 384 | LLM Pruning and Distillation in Practice: The Minitron Approach | Link |
| 385 | Jamba-1.5: Hybrid Transformer-Mamba Models at Scale | Link |
| 386 | Controllable Text Generation for Large Language Models: A Survey | Link |
| 387 | Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time | Link |
| 388 | A Practitionerโs Guide to Continual Multimodal Pretraining | Link |
| 389 | Building and Better Understanding Vision-Language Models: Insights and Future Directions | Link |
| 390 | CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation | Link |
| 391 | The Mamba in the Llama: Distilling and Accelerating Hybrid Models | Link |
| 392 | ReMamba: Equip Mamba with Effective Long-Sequence Modeling | Link |
| 393 | Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling | Link |
| 394 | LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models | Link |
| 395 | Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis | Link |
| 396 | X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models | Link |
| 397 | Free Process Rewards without Process Labels | Link |
| 398 | Scaling Image Tokenizers with Grouped Spherical Quantization | Link |
| 399 | RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models | Link |
| 400 | Perception Tokens Enhance Visual Reasoning in Multimodal Language Models | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 401 | Evaluating Language Models as Synthetic Data Generators | Link |
| 402 | Best-of-N Jailbreaking | Link |
| 403 | PaliGemma 2: A Family of Versatile VLMs for Transfer | Link |
| 404 | VisionZip: Longer is Better but Not Necessary in Vision Language Models | Link |
| 405 | Evaluating and Aligning CodeLLMs on Human Preference | Link |
| 406 | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale | Link |
| 407 | Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling | Link |
| 408 | LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods | Link |
| 409 | Does RLHF Scale? Exploring the Impacts From Data, Model, and Method | Link |
| 410 | Unraveling the Complexity of Memory in RL Agents: An Approach for Classification and Evaluation | Link |
| 411 | Training Large Language Models to Reason in a Continuous Latent Space | Link |
| 412 | AutoReason: Automatic Few-Shot Reasoning Decomposition | Link |
| 413 | Large Concept Models: Language Modeling in a Sentence Representation Space | Link |
| 414 | Phi-4 Technical Report | Link |
| 415 | Byte Latent Transformer: Patches Scale Better Than Tokens | Link |
| 416 | SCBench: A KV Cache-Centric Analysis of Long-Context Methods | Link |
| 417 | Cultural Evolution of Cooperation among LLM Agents | Link |
| 418 | DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding | Link |
| 419 | No More Adam: Learning Rate Scaling at Initialization is All You Need | Link |
| 420 | Precise Length Control in Large Language Models | Link |
| 421 | The Open Source Advantage in Large Language Models (LLMs) | Link |
| 422 | A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges | Link |
| 423 | Are Your LLMs Capable of Stable Reasoning? | Link |
| 424 | LLM Post-Training Recipes: Improving Reasoning in LLMs | Link |
| 425 | Hansel: Output Length Controlling Framework for Large Language Models | Link |
| 426 | Mind Your Theory: Theory of Mind Goes Deeper Than Reasoning | Link |
| 427 | Alignment Faking in Large Language Models | Link |
| 428 | SCOPE: Optimizing Key-Value Cache Compression in Long-Context Generation | Link |
| 429 | LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-Context Multitasks | Link |
| 430 | Offline Reinforcement Learning for LLM Multi-Step Reasoning | Link |
| 431 | Mulberry: Empowering MLLM with O1-like Reasoning and Reflection via Collective Monte Carlo Tree Search | Link |
| 432 | Titans: Learning to Memorize at Test Time | Link |
| 433 | Addition is All You Need for Energy-efficient Language Models | Link |
| 434 | Quantifying Generalization Complexity for Large Language Models | Link |
| 435 | When a Language Model is Optimized for Reasoning, Does It Still Show Embers of Autoregression? | Link |
| 436 | Were RNNs All We Needed? | Link |
| 437 | Selective Attention Improves Transformer | Link |
| 438 | LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations | Link |
| 439 | LLaVA-Critic: Learning to Evaluate Multimodal Models | Link |
| 440 | Differential Transformer | Link |
| 441 | GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models | Link |
| 442 | ARIA: An Open Multimodal Native Mixture-of-Experts Model | Link |
| 443 | O1 Replication Journey: A Strategic Progress Report โ Part 1 | Link |
| 444 | Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG | Link |
| 445 | From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction Tuning | Link |
| 446 | KV Prediction for Improved Time to First Token | Link |
| 447 | Baichuan-Omni Technical Report | Link |
| 448 | MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models | Link |
| 449 | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models | Link |
| 450 | AFlow: Automating Agentic Workflow Generation | Link |
| -------------- | -------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 451 | Toward General Instruction-Following Alignment for Retrieval-Augmented Generation | Link |
| 452 | Pre-training Distillation for Large Language Models: A Design Space Exploration | Link |
| 453 | MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models | Link |
| 454 | Scalable Ranked Preference Optimization for Text-to-Image Generation | Link |
| 455 | Scaling Diffusion Language Models via Adaptation from Autoregressive Models | Link |
| 456 | Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback | Link |
| 457 | Counting Ability of Large Language Models and Impact of Tokenization | Link |
| 458 | A Survey of Small Language Models | Link |
| 459 | Accelerating Direct Preference Optimization with Prefix Sharing | Link |
| 460 | Mind Your Step (by Step): Chain-of-Thought Can Reduce Performance on Tasks Where Thinking Makes Humans Worse | Link |
| 461 | LongReward: Improving Long-context Large Language Models with AI Feedback | Link |
| 462 | ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference | Link |
| 463 | Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial Applications | Link |
| 464 | CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation | Link |
| 465 | What Happened in LLMs Layers When Trained for Fast vs. Slow Thinking: A Gradient Perspective | Link |
| 466 | GPT or BERT: Why Not Both? | Link |
| 467 | Language Models Can Self-Lengthen to Generate Long Texts | Link |
| 468 | OLMoE: Open Mixture-of-Experts Language Models | Link |
| 469 | In Defense of RAG in the Era of Long-Context Language Models | Link |
| 470 | Attention Heads of Large Language Models: A Survey | Link |
| 471 | LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA | Link |
| 472 | How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality Data | Link |
| 473 | Theory, Analysis, and Best Practices for Sigmoid Self-Attention | Link |
| 474 | LLaMA-Omni: Seamless Speech Interaction with Large Language Models | Link |
| 475 | What is the Role of Small Models in the LLM Era: A Survey | Link |
| 476 | Policy Filtration in RLHF to Fine-Tune LLM for Code Generation | Link |
| 477 | RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval | Link |
| 478 | Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement | Link |
| 479 | Qwen2.5-Coder Technical Report | Link |
| 480 | Instruction Following without Instruction Tuning | Link |
| 481 | Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis | Link |
| 482 | The Perfect Blend: Redefining RLHF with Mixture of Judges | Link |
| 483 | Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations | Link |
| 484 | Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation | Link |
| 485 | Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models | Link |
| 486 | Sample-Efficient Alignment for LLMs | Link |
| 487 | A Comprehensive Survey of Small Language Models in the Era of Large Language Models | Link |
| 488 | โGive Me BF16 or Give Me Deathโ? Accuracy-Performance Trade-Offs in LLM Quantization | Link |
| 489 | Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation | Link |
| 490 | HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems | Link |
| 491 | Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data Contamination | Link |
| 492 | Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding | Link |
| 493 | Number Cookbook: Number Understanding of Language Models and How to Improve It | Link |
| 494 | Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models | Link |
| 495 | BitNet a4.8: 4-bit Activations for 1-bit LLMs | Link |
| 496 | Scaling Laws for Precision | Link |
| 497 | Energy Efficient Protein Language Models | Link |
| 498 | Balancing Pipeline Parallelism with Vocabulary Parallelism | Link |
| 499 | Toward Optimal Search and Retrieval for RAG | Link |
| 500 | Large Language Models Can Self-Improve in Long-context Reasoning | Link |
| -------------- | ----------------------------------------------------------------------------------------------------------------------- | ------------------------------------- |
| 501 | Stronger Models are NOT Stronger Teachers for Instruction Tuning | Link |
| 502 | Direct Preference Optimization Using Sparse Feature-Level Constraints | Link |
| 503 | Cut Your Losses in Large-Vocabulary Language Models | Link |
| 504 | Does Prompt Formatting Have Any Impact on LLM Performance? | Link |
| 505 | SymDPO: Boosting In-Context Learning of Large Multimodal Models | Link |
| 506 | SageAttention2 Technical Report | Link |
| 507 | Bi-Mamba: Towards Accurate 1-Bit State Space Models | Link |
| 508 | RedPajama: An Open Dataset for Training Large Language Models | Link |
| 509 | Hymba: A Hybrid-head Architecture for Small Language Models | Link |
| 510 | Loss-to-Loss Prediction: Scaling Laws for All Datasets | Link |
| 511 | When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training | Link |
| 512 | Multimodal Autoregressive Pre-training of Large Vision Encoders | Link |
| 513 | Natural Language Reinforcement Learning | Link |
| 514 | Large Multi-modal Models Can Interpret Features in Large Multi-modal Models | Link |
| 515 | TรLU 3: Pushing Frontiers in Open Language Model Post-Training | Link |
| 516 | MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs | Link |
| 517 | LLMs Do Not Think Step-by-step In Implicit Reasoning | Link |
| 518 | O1 Replication Journey โ Part 2 | Link |
| 519 | Star Attention: Efficient LLM Inference over Long Sequences | Link |
| 520 | Low-Bit Quantization Favors Undertrained LLMs | Link |
| 521 | Rethinking Token Reduction in MLLMs | Link |
| 522 | Reverse Thinking Makes LLMs Stronger Reasoners | Link |
| 523 | Critical Tokens Matter | Link |
| 524 | Foundations of Large Language Models | Link |
| 525 | A Survey of Research in Large Language Models for Electronic Design Automation | Link |
| 526 | Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models | Link |
| 527 | Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' Questions | Link |
| 528 | The Future of AI: Exploring the Potential of Large Concept Models | Link |
| 529 | Investigating Numerical Translation with Large Language Models | Link |
| 530 | CALM: Curiosity-Driven Auditing for Large Language Models | Link |
๐ Explore cutting-edge research โจ and stay updated! ๐ง ๐
| ๐ข S.No. | ๐ Paper Title | ๐ Link |
|---|---|---|
| 531 | DeepSeek-V3 Technical Report | Link |
| 532 | DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning | Link |
| 533 | Kimi k1.5: Scaling Reinforcement Learning with LLMs | Link |
| 534 | Reasoning Language Models: A Blueprint | Link |
| 535 | Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought | Link |
| 536 | The Lessons of Developing Process Reward Models in Mathematical Reasoning | Link |
| 537 | LIMO: Less is More for Reasoning | Link |
| 538 | Demystifying Long Chain-of-Thought Reasoning in LLMs | Link |
| 539 | Competitive Programming with Large Reasoning Models | Link |
| 540 | LLMs Can Easily Learn to Reason from Demonstrations: Structure, Not Content, Is What Matters | Link |
| 541 | Training Language Models to Reason Efficiently | Link |
| 542 | Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning | Link |
| 543 | SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution | Link |
| 544 | On the Emergence of Thinking in LLMs: Searching for the Right Intuition | Link |
| 545 | Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning | Link |
| 546 | Teaching Language Models to Critique via Reinforcement Learning | Link |
| 547 | A Review of DeepSeek Models' Key Innovative Techniques | Link |
| 548 | R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning | Link |
| 549 | Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning | Link |
| 550 | Understanding R1-Zero-Like Training: A Critical Perspective | Link |
| 551 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning | Link |
| 552 | Open-Reasoner-Zero: An Open Source Approach to Scaling Up RL on the Base Model | Link |
| 553 | Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't | Link |
| 554 | Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in LLMs | Link |
| 555 | JudgeLRM: Large Reasoning Models as a Judge | Link |
| 556 | Concise Reasoning via Reinforcement Learning | Link |
| 557 | Absolute Zero: Reinforced Self-play Reasoning with Zero Data | Link |
| 558 | Qwen3 Technical Report | Link |
| 559 | MiMo: Unlocking the Reasoning Potential of Language Models โ From Pretraining to Posttraining | Link |
| 560 | Llama-Nemotron: Efficient Reasoning Models | Link |
| 561 | INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning | Link |
| 562 | AdaptThink: Reasoning Models Can Learn When to Think | Link |
| 563 | Thinkless: LLM Learns When to Think | Link |
| 564 | General-Reasoner: Advancing LLM Reasoning Across All Domains | Link |
| 565 | Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents | Link |
| 566 | Reinforcement Pre-Training | Link |
| 567 | Magistral | Link |
| 568 | AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery | Link |
| 569 | SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training | Link |
| 570 | Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond The Base Model? | Link |
| 571 | From System 1 to System 2: A Survey of Reasoning Large Language Models | Link |
| 572 | Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning LLMs | Link |
| 573 | Multi-Agent Collaboration Mechanisms: A Survey of LLMs | Link |
| 574 | Search-o1: Agentic Search-Enhanced Large Reasoning Models | Link |
| 575 | Reasoning Models Can Be Effective Without Thinking | Link |
| 576 | RM-R1: Reward Modeling as Reasoning | Link |
| 577 | QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning | Link |
| 578 | Enigmata: Scaling Logical Reasoning in LLMs with Synthetic Verifiable Puzzles | Link |
| 579 | ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in LLMs | Link |
| 580 | Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective RL for LLM Reasoning | Link |
| 581 | Spurious Rewards: Rethinking Training Signals in RLVR | Link |
| 582 | Tina: Tiny Reasoning Models via LoRA | Link |
| 583 | Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math | Link |
| 584 | VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning | Link |
| 585 | Gemini Robotics: Bringing AI Into the Physical World | Link |
| 586 | Gemini 2.0: The Era of Multimodal Agentic AI | Link |
| 587 | RARE: Retrieval-Augmented Reasoning Modeling | Link |
| 588 | Learning from Failures in Multi-Attempt Reinforcement Learning | Link |
| 589 | R1-VL: Learning to Reason with Multimodal LLMs via Step-wise Group Relative Policy Optimization | Link |
| 590 | Diffusion-Based Language Models: A Survey | Link |
| 591 | FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching | Link |
| 592 | DAPO: An Open-Source LLM Reinforcement Learning System at Scale | Link |
| 593 | Scaling Laws for Inference-Time Compute | Link |
| 594 | LLM Post-Training: A Deep Dive into Reasoning Large Language Models | Link |
| 595 | Long-VITA: Scaling Large Vision-Language Models for Long Video Understanding | Link |
| 596 | Wan: Open and Advanced Large-Scale Video Generative Models | Link |
| 597 | Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models | Link |
| 598 | Seed1.5-VL Technical Report | Link |
| 599 | A Survey on LLM-based Agents: Recent Advances and New Frontiers | Link |
| 600 | RLVR Is Not RL: Revisiting Reinforcement Learning for LLMs | Link |
| 601 | RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning | Link |
| 602 | Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning | Link |
| 603 | LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL | Link |
| 604 | Reinforcement Learning for Reasoning in Large Language Models with One Training Example | Link |
| 605 | Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model Capability | Link |
| 606 | The First Few Tokens Are All You Need: Unsupervised Prefix Fine-Tuning for Reasoning Models | Link |
| 607 | Learning to Reason without External Rewards | Link |
| 608 | Genius: A Generalizable and Purely Unsupervised Self-Training Framework for Advanced Reasoning | Link |
| 609 | Reinforcement Learning Teachers of Test Time Scaling | Link |
| 610 | Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening | Link |
7 commits