itsual/Notable-LLM-Research-Papers

Curated list of research papers published in 2024 related to Large Language Models (LLM)

50

7 commits

updated Apr 14, 2026

See the code

README

๐Ÿ“š Notable-LLM-Research-Papers ๐ŸŽ“โœจ

Hello, curious reader! ๐Ÿ‘‹ Are you someone who loves exploring the cutting-edge world of AI and large language models (LLMs)? Or maybe you're a researcher, developer, or simply an enthusiast wondering about how these incredible technologies are evolving? Either way, you're in the right place. This repository is your gateway to 600+ groundbreaking research papers that capture the exciting developments in AI (in recent months)

๐Ÿš€ Whatโ€™s in This Repository?

These papers cover a wide spectrum of topics, ranging from improving model efficiency to aligning AI with human values, and even diving into the magic of multimodal models. The field of AI is advancing faster than ever, and these works showcase the brilliant ideas and innovations shaping the future.

The papers broadly fall under the following categories:


1. Model Architecture and Efficiency ๐Ÿ—๏ธโšก

Papers that introduce innovative architectures, explore scaling laws, or improve the efficiency of LLMs. These works aim to make AI faster, cheaper, and better.


2. Preference Optimization and RLHF (Reinforcement Learning with Human Feedback) ๐Ÿค๐Ÿ’ก

Research exploring how to align AI with human preferences, ensuring outputs are not just accurate but also ethical, safe, and aligned with societal values.


3. Multimodal Models ๐Ÿ–ผ๏ธ๐ŸŽ™๏ธโœ๏ธ

Ever wondered how AI can see, read, and understand all at once? These papers focus on models that combine multiple types of data, like images, text, and audio, to build richer, more versatile systems.


4. Long-Context Learning and Inference ๐Ÿง ๐ŸŒ€

Tackling the challenges of long-context understanding, these papers discuss how to extend the memory of LLMs, making them capable of reasoning across longer documents or conversations.


5. Reasoning and Knowledge ๐Ÿ”๐Ÿง 

From enabling AI to solve complex puzzles to improving its reasoning capabilities, these papers explore how LLMs "think" and manage the vast knowledge they've been trained on.


6. Quantization and Compression ๐Ÿ“๐Ÿ—œ๏ธ

Smaller, faster, and more efficientโ€”these papers delve into techniques like quantization and compression to make AI models more practical for real-world applications.


7. Evaluation and Benchmarks ๐Ÿ“Š๐Ÿ“ˆ

A great model needs great evaluation. These papers propose new benchmarks and methodologies to assess the capabilities of AI systems more rigorously.


8. Instruction Tuning and Alignment ๐Ÿ“๐ŸŽฏ

Want your AI to follow your instructions perfectly? These papers refine how we align LLMs to understand and execute instructions across various tasks.


9. Surveys and Meta-Analysis ๐Ÿ”ฌ๐Ÿ“š

Meta-level research analyzing trends, methods, and applications in AI. Perfect for getting a birdโ€™s-eye view of where the field is heading.


10. Applications and Tooling ๐Ÿ”ง๐ŸŒ

Papers showcasing how LLMs are applied in diverse domains, from healthcare to coding and everything in between. These highlight the transformative power of AI in the real world.


๐ŸŒŸ Why Should You Care?

These papers represent the state of the art in AI research. Whether you're here to stay updated, seek inspiration, or find practical tools, this collection is a treasure trove for AI enthusiasts. Dive in, explore, and let your curiosity guide you! ๐ŸŒโœจ


Ready to start? ๐ŸŽ‰ Browse the list and explore the exciting breakthroughs shaping the future of AI! ๐Ÿš€


๐Ÿ”ข S.No.๐Ÿ“ Paper Title๐Ÿ”— Link
1LLM Maybe LongLM: Self-Extend LLM Context Window Without TuningLink
2Knowledge Fusion of Large Language ModelsLink
3A Comprehensive Study of Knowledge Editing for Large Language ModelsLink
4DiffusionGPT: LLM-Driven Text-to-Image Generation SystemLink
5Tuning Language Models by ProxyLink
6An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language ModelsLink
7Soaring from 4K to 400K: Extending LLMโ€™s Context with Activation BeaconLink
8VMamba: Visual State Space ModelLink
9LLaMA Beyond English: An Empirical Study on Language Capability TransferLink
10DeepSeek LLM: Scaling Open-Source Language Models with LongtermismLink
11Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLMLink
12LLaMA Pro: Progressive LLaMA with Block ExpansionLink
13RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationLink
14Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextLink
15Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingLink
16WARM: On the Benefits of Weight Averaged Reward ModelsLink
17SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token DetectionLink
18A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityLink
19Mixtral of ExpertsLink
20MoE-Mamba: Efficient Selective State Space Models with Mixture of ExpertsLink
21Code Generation with AlphaCodium: From Prompt Engineering to Flow EngineeringLink
22EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyLink
23KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationLink
24A Closer Look at AUROC and AUPRC under Class ImbalanceLink
25Transformers are Multi-State RNNsLink
26LLM Augmented LLMs: Expanding Capabilities through CompositionLink
27Rethinking Patch Dependence for Masked AutoencodersLink
28Astraios: Parameter-Efficient Instruction Tuning Code Large Language ModelsLink
29Pix2gestalt: Amodal Segmentation by Synthesizing WholesLink
30RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on AgricultureLink
31An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language ModelsLink
32Knowledge Fusion of Large Language ModelsLink
33Scalable Pre-training of Large Autoregressive Image ModelsLink
34SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning CapabilitiesLink
35Multimodal Pathway: Improve Transformers with Irrelevant Data from Other ModalitiesLink
36MambaByte: Token-free Selective State Space ModelLink
37ReFT: Reasoning with Reinforced Fine-TuningLink
38Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsLink
39Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingLink
40Denoising Vision TransformersLink
41Self-Rewarding Language ModelsLink
42LoRA+: Efficient Low Rank Adaptation of Large ModelsLink
43MobileVLM V2: Faster and Stronger Baseline for Vision Language ModelLink
44Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?Link
45ODIN: Disentangled Reward Mitigates Hacking in RLHFLink
46Genie: Generative Interactive EnvironmentsLink
47A Phase Transition Between Positional and Semantic Learning in a Solvable Model of Dot-Product AttentionLink
48Neural Network DiffusionLink
49More Agents Is All You NeedLink
50Scaling Laws for Downstream Task Performance of Large Language ModelsLink
-----------------------------------------------------------------------------------------------------------------------------------------------------------------------
51Repeat After Me: Transformers are Better than State Space Models at CopyingLink
52Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language ModelsLink
53AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingLink
54BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of DataLink
55LongAgent: Scaling Language Models to 128k Context through Multi-Agent CollaborationLink
56LongRoPE: Extending LLM Context Window Beyond 2 Million TokensLink
57Policy Improvement using Language Feedback ModelsLink
58DoRA: Weight-Decomposed Low-Rank AdaptationLink
59FindingEmo: An Image Dataset for Emotion Recognition in the WildLink
60TinyLLaVA: A Framework of Small-scale Large Multimodal ModelsLink
61Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation ModelsLink
62When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodLink
63Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via BootstrappingLink
64Suppressing Pink Elephants with Direct Principle FeedbackLink
65The Era of 1-bit LLMs: All Large Language Models are in 1.58 BitsLink
66LiPO: Listwise Preference Optimization through Learning-to-RankLink
67Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelLink
68The Boundary of Neural Network Trainability is FractalLink
69Recovering the Pre-Fine-Tuning Weights of Generative ModelsLink
70Scaling Laws for Fine-Grained Mixture of ExpertsLink
71Direct Language Model Alignment from Online AI FeedbackLink
72CARTE: Pretraining and Transfer for Tabular LearningLink
73Grandmaster-Level Chess Without SearchLink
74Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsLink
75Reformatted AlignmentLink
76OLMo: Accelerating the Science of Language ModelsLink
77FinTral: A Family of GPT-4 Level Multimodal Financial Large Language ModelsLink
78Mixtures of Experts Unlock Parameter Scaling for Deep RLLink
79Generative Representational Instruction TuningLink
80World Model on Million-Length Video And Language With RingAttentionLink
81Efficient Exploration for LLMsLink
82YOLOv9: Learning What You Want to Learn Using Programmable Gradient InformationLink
83Towards Cross-Tokenizer Distillation: The Universal Logit Distillation Loss for LLMsLink
84Mixtures of Experts Unlock Parameter Scaling for Deep RLLink
85Sora Generates Videos with Stunning Geometrical ConsistencyLink
86BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of DataLink
87Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsLink
88Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language ModelsLink
89DoRA: Weight-Decomposed Low-Rank AdaptationLink
90The Era of 1-bit LLMs: All Large Language Models are in 1.58 BitsLink
91Efficient Exploration for LLMsLink
92Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingLink
93You Only Cache Once: Decoder-Decoder Architectures for Language ModelsLink
94gzip Predicts Data-dependent Scaling LawsLink
95Self-Play Preference Optimization for Language Model AlignmentLink
96PHUDGE: Phi-3 as Scalable JudgeLink
97What Matters When Building Vision-Language Models?Link
98Towards Modular LLMs by Building and Reusing a Library of LoRAsLink
99Contextual Position Encoding: Learning to Count Whatโ€™s ImportantLink
100RLHF Workflow: From Reward Modeling to Online RLHFLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
101Attention as an RNNLink
102AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityLink
103Instruction Tuning With Loss Over InstructionsLink
104LoRA Learns Less and Forgets LessLink
105Trans-LoRA: Towards Data-free Transferable Parameter Efficient FinetuningLink
106VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsLink
107MoRA: High-Rank Updating for Parameter-Efficient Fine-TuningLink
108LLaMA-NAS: Efficient Neural Architecture Search for Large Language ModelsLink
109SimPO: Simple Preference Optimization with a Reference-Free RewardLink
110The Road Less ScheduledLink
111Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsLink
112Is Flash Attention Stable?Link
113Value Augmented Sampling for Language Model Alignment and PersonalizationLink
114Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Link
115vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionLink
116A Careful Examination of Large Language Model Performance on Grade School ArithmeticLink
117Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsLink
118DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelLink
119Dense Connector for MLLMsLink
120xLSTM: Extended Long Short-Term MemoryLink
121SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationLink
122Xmodel-VLM: A Simple Baseline for Multimodal Vision Language ModelLink
123Chameleon: Mixed-Modal Early-Fusion Foundation ModelsLink
124Is Bigger Edit Batch Size Always Better? An Empirical Study on Model Editing with Llama-3Link
125AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityLink
126The Prompt Report: A Systematic Survey of Prompting TechniquesLink
127Creativity Has Left the Chat: The Price of Debiasing Language ModelsLink
128Show, Donโ€™t Tell: Aligning Language Models with Demonstrated FeedbackLink
129WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildLink
130Scalable MatMul-free Language ModelingLink
131Never Miss A Beat: An Efficient Recipe for Context Window Extension of Large Language ModelsLink
132Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven ApproachLink
133Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language ModelsLink
134MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer DecodingLink
135Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLink
136An Image is Worth 32 Tokens for Reconstruction and GenerationLink
137Block Transformer: Global-to-Local Language Modeling for Fast InferenceLink
1383D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less HallucinationLink
139Transformers Need Glasses! Information Over-Squashing in Language TasksLink
140The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsLink
141BERTs are Generative In-Context LearnersLink
142An Empirical Study of Mamba-based Language ModelsLink
143Discovering Preference Optimization Algorithms with and for Large Language ModelsLink
144Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingLink
145Autoregressive Model Beats Diffusion: Llama for Scalable Image GenerationLink
146Are We Done with MMLU?Link
147OLoRA: Orthonormal Low-Rank Adaptation of Large Language ModelsLink
148Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each StepLink
149Husky: A Unified, Open-Source Language Agent for Multi-Step ReasoningLink
150Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated ParametersLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
151Simple and Effective Masked Diffusion Language ModelsLink
152The Prompt Report: A Systematic Survey of Prompting TechniquesLink
153Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsLink
154Be like a Goldfish, Donโ€™t Memorize! Mitigating Memorization in Generative LLMsLink
155TextGrad: Automatic โ€œDifferentiationโ€ via TextLink
156Large Language Models Must Be Taught to Know What They Donโ€™t KnowLink
157What If We Recaption Billions of Web Images with LLaMA-3?Link
158Discovering Preference Optimization Algorithms with and for Large Language ModelsLink
159An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual PixelsLink
160Buffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsLink
161CRAG โ€“ Comprehensive RAG BenchmarkLink
162Margin-aware Preference Optimization for Aligning Diffusion Models Without ReferenceLink
163Mixture-of-Agents Enhances Large Language Model CapabilitiesLink
164Large Language Model Unlearning via Embedding-Corrupted PromptsLink
165Bootstrapping Language Models with DPO Implicit RewardsLink
166THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response GenerationLink
167Task Me AnythingLink
168Nemotron-4 340B Technical ReportLink
169Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-JudgesLink
170How Do Large Language Models Acquire Factual Knowledge During Pretraining?Link
171mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsLink
172Unveiling Encoder-Free Vision-Language ModelsLink
173HARE: HumAn pRiors, a key to small language model EfficiencyLink
174Measuring Memorization in RLHF for Code CompletionLink
175DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceLink
176Iterative Length-Regularized Direct Preference Optimization: Improving 7B Language Models to GPT-4 LevelLink
177From RAGs to Rich Parameters: Probing How Language Models Utilize External Knowledge Over Parametric InformationLink
178DataComp-LM: In Search of the Next Generation of Training Sets for Language ModelsLink
179Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?Link
180Instruction Pre-Training: Language Models are Supervised Multitask LearnersLink
181Can LLMs Learn by Teaching? A Preliminary StudyLink
182A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG SystemsLink
183LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMsLink
184MoA: Mixture of Sparse Attention for Automatic Large Language Model CompressionLink
185Efficient Continual Pre-training by Mitigating the Stability GapLink
186Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range TransformersLink
187WARP: On the Benefits of Weight Averaged Rewarded PoliciesLink
188Adam-mini: Use Fewer Learning Rates To Gain MoreLink
189The FineWeb Datasets: Decanting the Web for the Finest Text Data at ScaleLink
190LongIns: A Challenging Long-context Instruction-based Exam for LLMsLink
191Following Length Constraints in InstructionsLink
192A Closer Look into Mixture-of-Experts in Large Language ModelsLink
193RouteLLM: Learning to Route LLMs with Preference DataLink
194Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMsLink
195Dataset Size Recovery from LoRA WeightsLink
196From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic DataLink
197Changing Answer Order Can Decrease MMLU AccuracyLink
198Direct Preference Knowledge Distillation for Large Language ModelsLink
199LLM Critics Help Catch LLM BugsLink
200Scaling Synthetic Data Creation with 1,000,000,000 PersonasLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
201Tokenization Falling Short: The Curse of TokenizationLink
202Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsLink
203Bootstrapping Language Models with DPO Implicit RewardsLink
204Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-JudgesLink
205Learning and Leveraging World Models in Visual Representation LearningLink
206Improving LLM Code Generation with Grammar AugmentationLink
207The Hidden Attention of Mamba ModelsLink
208Training-Free Pretrained Model MergingLink
209Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesLink
210The WMDP Benchmark: Measuring and Reducing Malicious Use With UnlearningLink
211Evolution Transformer: In-Context Evolutionary OptimizationLink
212Enhancing Vision-Language Pre-training with Rich SupervisionsLink
213Scaling Rectified Flow Transformers for High-Resolution Image SynthesisLink
214Design2Code: How Far Are We From Automating Front-End Engineering?Link
215ShortGPT: Layers in Large Language Models are More Redundant Than You ExpectLink
216Backtracing: Retrieving the Cause of the QueryLink
217Learning to Decode Collaboratively with Multiple Language ModelsLink
218SaulLM-7B: A pioneering Large Language Model for LawLink
219Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal ReasoningLink
2203D Diffusion PolicyLink
221MedMamba: Vision Mamba for Medical Image ClassificationLink
222GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionLink
223Stop Regressing: Training Value Functions via Classification for Scalable Deep RLLink
224How Far Are We from Intelligent Visual Deductive Reasoning?Link
225Common 7B Language Models Already Possess Strong Math CapabilitiesLink
226Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of ContextLink
227Is Cosine-Similarity of Embeddings Really About Similarity?Link
228LLM4Decompile: Decompiling Binary Code with Large Language ModelsLink
229Algorithmic Progress in Language ModelsLink
230Stealing Part of a Production Language ModelLink
231Chronos: Learning the Language of Time SeriesLink
232Simple and Scalable Strategies to Continually Pre-train Large Language ModelsLink
233Language Models Scale Reliably With Over-Training and on Downstream TasksLink
234BurstAttention: An Efficient Distributed Attention Framework for Extremely Long SequencesLink
235LocalMamba: Visual State Space Model with Windowed Selective ScanLink
236GiT: Towards Generalist Vision Transformer through Universal Language InterfaceLink
237MM1: Methods, Analysis & Insights from Multimodal LLM Pre-trainingLink
238RAFT: Adapting Language Model to Domain Specific RAGLink
239TnT-LLM: Text Mining at Scale with Large Language ModelsLink
240Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionLink
241PERL: Parameter Efficient Reinforcement Learning from Human FeedbackLink
242RewardBench: Evaluating Reward Models for Language ModelingLink
243LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language ModelsLink
244RakutenAI-7B: Extending Large Language Models for JapaneseLink
245SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time SeriesLink
246Can Large Language Models Explore In-Context?Link
247LLM2LLM: Boosting LLMs with Novel Iterative Data EnhancementLink
248LLM Agent Operating SystemLink
249The Unreasonable Ineffectiveness of the Deeper LayersLink
250BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical TextLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
251ViTAR: Vision Transformer with Any ResolutionLink
252Long-form Factuality in Large Language ModelsLink
253Mini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsLink
254LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-TuningLink
255Mechanistic Design and Scaling of Hybrid ArchitecturesLink
256MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsLink
257Model Stock: All We Need Is Just a Few Fine-Tuned ModelsLink
258Do Language Models Plan Ahead for Future Tokens?Link
259Bigger is not Always Better: Scaling Properties of Latent Diffusion ModelsLink
260The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability AnalysisLink
261Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion ModelsLink
262Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language ModelsLink
263Long-context LLMs Struggle with Long In-context LearningLink
264Emergent Abilities in Reduced-Scale Generative Language ModelsLink
265Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksLink
266On the Scalability of Diffusion-based Text-to-Image GenerationLink
267BAdam: A Memory Efficient Full Parameter Training Method for Large Language ModelsLink
268Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion ModelsLink
269Direct Nash Optimization: Teaching Language Models to Self-Improve with General PreferencesLink
270Training LLMs over Neurally Compressed TextLink
271CantTalkAboutThis: Aligning Language Models to Stay on Topic in DialoguesLink
272ReFT: Representation Finetuning for Language ModelsLink
273Verifiable by Design: Aligning Language Models to Quote from Pre-Training DataLink
274Sigma: Siamese Mamba Network for Multi-Modal Semantic SegmentationLink
275AutoCodeRover: Autonomous Program ImprovementLink
276Eagle and Finch: RWKV with Matrix-Valued States and Dynamic RecurrenceLink
277CodecLM: Aligning Language Models with Tailored Synthetic DataLink
278MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training StrategiesLink
279Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language ModelsLink
280LLM2Vec: Large Language Models Are Secretly Powerful Text EncodersLink
281Adapting LLaMA Decoder to Vision TransformerLink
282Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attentionLink
283LLoCO: Learning Long Contexts OfflineLink
284JetMoE: Reaching Llama2 Performance with 0.1M DollarsLink
285Best Practices and Lessons Learned on Synthetic Data for Language ModelsLink
286Rho-1: Not All Tokens Are What You NeedLink
287Pre-training Small Base LMs with Fewer TokensLink
288Dataset Reset Policy Optimization for RLHFLink
289LLM In-Context Recall is Prompt DependentLink
290State Space Model for New-Generation Network Alternative to Transformers: A SurveyLink
291Chinchilla Scaling: A Replication AttemptLink
292Learn Your Reference Model for Real Good AlignmentLink
293Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyLink
294Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training StrategiesLink
295How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMsโ€™ Internal PriorLink
296A Survey on Retrieval-Augmented Text Generation for Large Language ModelsLink
297When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many ClassesLink
298Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingLink
299OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction DataLink
300The Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
301How Good Are Low-bit Quantized LLaMA3 Models? An Empirical StudyLink
302Phi-3 Technical Report: A Highly Capable Language Model Locally on Your PhoneLink
303OpenELM: An Efficient Language Model Family with Open-source Training and Inference FrameworkLink
304A Survey on Self-Evolution of Large Language ModelsLink
305Multi-Head Mixture-of-ExpertsLink
306NExT: Teaching Large Language Models to Reason about Code ExecutionLink
307Graph Machine Learning in the Era of Large Language Models (LLMs)Link
308Retrieval Head Mechanistically Explains Long-Context FactualityLink
309Layer Skip: Enabling Early Exit Inference and Self-Speculative DecodingLink
310Make Your LLM Fully Utilize the ContextLink
311LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical ReportLink
312Better & Faster Large Language Models via Multi-token PredictionLink
313RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language ProcessingLink
314A Primer on the Inner Workings of Transformer-based Language ModelsLink
315When to Retrieve: Teaching LLMs to Utilize Information Retrieval EffectivelyLink
316KAN: Kolmogorovโ€“Arnold NetworksLink
317LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable ObjectivesLink
318Searching for Best Practices in Retrieval-Augmented GenerationLink
319Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsLink
320Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionLink
321Eliminating Position Bias of Language Models: A Mechanistic ApproachLink
322JMInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionLink
323TokenPacker: Efficient Visual Projector for Multimodal LLMLink
324Reasoning in Large Language Models: A Geometric PerspectiveLink
325RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsLink
326AgentInstruct: Toward Generative Teaching with Agentic FlowsLink
327HEMM: Holistic Evaluation of Multimodal Foundation ModelsLink
328Mixture of A Million ExpertsLink
329Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesLink
330Vision Language Models Are BlindLink
331Self-Recognition in Language ModelsLink
332Inference Performance Optimization for Large Language Models on CPUsLink
333Gradient Boosting Reinforcement LearningLink
334FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionLink
335SpreadsheetLLM: Encoding Spreadsheets for Large Language ModelsLink
336New Desiderata for Direct Preference OptimizationLink
337Context Embeddings for Efficient Answer Generation in RAGLink
338Qwen2 Technical ReportLink
339The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-DeterminismLink
340From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank GradientsLink
341GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache CompressionLink
342Scaling Diffusion Transformers to 16 Billion ParametersLink
343NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window?Link
344Patch-Level Training for Large Language ModelsLink
345LMMs-Eval: Reality Check on the Evaluation of Large Multimodal ModelsLink
346A Survey of Prompt Engineering Methods in Large Language Models for Different NLP TasksLink
347Spectra: A Comprehensive Study of Ternary, Quantized, and FP16 Language ModelsLink
348Attention Overflow: Language Model Input Blur during Long-Context Missing Items RecommendationLink
349Weak-to-Strong ReasoningLink
350Understanding Reference Policies in Direct Preference OptimizationLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
351Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesLink
352BOND: Aligning LLMs with Best-of-N DistillationLink
353Compact Language Models via Pruning and Knowledge DistillationLink
354LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM InferenceLink
355Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences TrainingLink
356DDK: Distilling Domain Knowledge for Efficient Large Language ModelsLink
357Generation Constraint Scaling Can Mitigate HallucinationLink
358Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid ApproachLink
359Course-Correction: Safety Alignment Using Synthetic PreferencesLink
360Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?Link
361Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeLink
362Improving Retrieval Augmented Language Model with Self-ReasoningLink
363Apple Intelligence Foundation Language ModelsLink
364ThinK: Thinner Key Cache by Query-Driven PruningLink
365The Llama 3 Herd of ModelsLink
366Gemma 2: Improving Open Language Models at a Practical SizeLink
367SAM 2: Segment Anything in Images and VideosLink
368POA: Pre-training Once for Models of All SizesLink
369RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkLink
370A Survey of MambaLink
371MiniCPM-V: A GPT-4V Level MLLM on Your PhoneLink
372RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented GenerationLink
373Self-Taught EvaluatorsLink
374BioMamba: A Pre-trained Biomedical Language Representation Model Leveraging MambaLink
375EXAONE 3.0 7.8B Instruction Tuned Language ModelLink
3761.5-Pints Technical Report: Pretraining in Days, Not Months โ€“ Your Language Model Thrives on Quality DataLink
377Conversational Prompt EngineeringLink
378Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLPLink
379The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoveryLink
380Hermes 3 Technical ReportLink
381Customizing Language Models with Instance-wise LoRA for Sequential RecommendationLink
382Enhancing Robustness in Large Language Models: Prompting for Mitigating the Impact of Irrelevant InformationLink
383To Code, or Not To Code? Exploring Impact of Code in Pre-trainingLink
384LLM Pruning and Distillation in Practice: The Minitron ApproachLink
385Jamba-1.5: Hybrid Transformer-Mamba Models at ScaleLink
386Controllable Text Generation for Large Language Models: A SurveyLink
387Multi-Layer Transformers Gradient Can be Approximated in Almost Linear TimeLink
388A Practitionerโ€™s Guide to Continual Multimodal PretrainingLink
389Building and Better Understanding Vision-Language Models: Insights and Future DirectionsLink
390CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting MitigationLink
391The Mamba in the Llama: Distilling and Accelerating Hybrid ModelsLink
392ReMamba: Equip Mamba with Effective Long-Sequence ModelingLink
393Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal SamplingLink
394LongRecipe: Recipe for Efficient Long Context Generalization in Large Language ModelsLink
395Switti: Designing Scale-Wise Transformers for Text-to-Image SynthesisLink
396X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation ModelsLink
397Free Process Rewards without Process LabelsLink
398Scaling Image Tokenizers with Grouped Spherical QuantizationLink
399RARE: Retrieval-Augmented Reasoning Enhancement for Large Language ModelsLink
400Perception Tokens Enhance Visual Reasoning in Multimodal Language ModelsLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
401Evaluating Language Models as Synthetic Data GeneratorsLink
402Best-of-N JailbreakingLink
403PaliGemma 2: A Family of Versatile VLMs for TransferLink
404VisionZip: Longer is Better but Not Necessary in Vision Language ModelsLink
405Evaluating and Aligning CodeLLMs on Human PreferenceLink
406MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at ScaleLink
407Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time ScalingLink
408LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation MethodsLink
409Does RLHF Scale? Exploring the Impacts From Data, Model, and MethodLink
410Unraveling the Complexity of Memory in RL Agents: An Approach for Classification and EvaluationLink
411Training Large Language Models to Reason in a Continuous Latent SpaceLink
412AutoReason: Automatic Few-Shot Reasoning DecompositionLink
413Large Concept Models: Language Modeling in a Sentence Representation SpaceLink
414Phi-4 Technical ReportLink
415Byte Latent Transformer: Patches Scale Better Than TokensLink
416SCBench: A KV Cache-Centric Analysis of Long-Context MethodsLink
417Cultural Evolution of Cooperation among LLM AgentsLink
418DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal UnderstandingLink
419No More Adam: Learning Rate Scaling at Initialization is All You NeedLink
420Precise Length Control in Large Language ModelsLink
421The Open Source Advantage in Large Language Models (LLMs)Link
422A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & ChallengesLink
423Are Your LLMs Capable of Stable Reasoning?Link
424LLM Post-Training Recipes: Improving Reasoning in LLMsLink
425Hansel: Output Length Controlling Framework for Large Language ModelsLink
426Mind Your Theory: Theory of Mind Goes Deeper Than ReasoningLink
427Alignment Faking in Large Language ModelsLink
428SCOPE: Optimizing Key-Value Cache Compression in Long-Context GenerationLink
429LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-Context MultitasksLink
430Offline Reinforcement Learning for LLM Multi-Step ReasoningLink
431Mulberry: Empowering MLLM with O1-like Reasoning and Reflection via Collective Monte Carlo Tree SearchLink
432Titans: Learning to Memorize at Test TimeLink
433Addition is All You Need for Energy-efficient Language ModelsLink
434Quantifying Generalization Complexity for Large Language ModelsLink
435When a Language Model is Optimized for Reasoning, Does It Still Show Embers of Autoregression?Link
436Were RNNs All We Needed?Link
437Selective Attention Improves TransformerLink
438LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsLink
439LLaVA-Critic: Learning to Evaluate Multimodal ModelsLink
440Differential TransformerLink
441GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsLink
442ARIA: An Open Multimodal Native Mixture-of-Experts ModelLink
443O1 Replication Journey: A Strategic Progress Report โ€“ Part 1Link
444Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAGLink
445From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction TuningLink
446KV Prediction for Improved Time to First TokenLink
447Baichuan-Omni Technical ReportLink
448MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsLink
449LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal ModelsLink
450AFlow: Automating Agentic Workflow GenerationLink
-----------------------------------------------------------------------------------------------------------------------------------------------------------------------
451Toward General Instruction-Following Alignment for Retrieval-Augmented GenerationLink
452Pre-training Distillation for Large Language Models: A Design Space ExplorationLink
453MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language ModelsLink
454Scalable Ranked Preference Optimization for Text-to-Image GenerationLink
455Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsLink
456Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackLink
457Counting Ability of Large Language Models and Impact of TokenizationLink
458A Survey of Small Language ModelsLink
459Accelerating Direct Preference Optimization with Prefix SharingLink
460Mind Your Step (by Step): Chain-of-Thought Can Reduce Performance on Tasks Where Thinking Makes Humans WorseLink
461LongReward: Improving Long-context Large Language Models with AI FeedbackLink
462ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM InferenceLink
463Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial ApplicationsLink
464CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation GenerationLink
465What Happened in LLMs Layers When Trained for Fast vs. Slow Thinking: A Gradient PerspectiveLink
466GPT or BERT: Why Not Both?Link
467Language Models Can Self-Lengthen to Generate Long TextsLink
468OLMoE: Open Mixture-of-Experts Language ModelsLink
469In Defense of RAG in the Era of Long-Context Language ModelsLink
470Attention Heads of Large Language Models: A SurveyLink
471LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QALink
472How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality DataLink
473Theory, Analysis, and Best Practices for Sigmoid Self-AttentionLink
474LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsLink
475What is the Role of Small Models in the LLM Era: A SurveyLink
476Policy Filtration in RLHF to Fine-Tune LLM for Code GenerationLink
477RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalLink
478Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-ImprovementLink
479Qwen2.5-Coder Technical ReportLink
480Instruction Following without Instruction TuningLink
481Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical AnalysisLink
482The Perfect Blend: Redefining RLHF with Mixture of JudgesLink
483Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsLink
484Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage AdaptationLink
485Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language ModelsLink
486Sample-Efficient Alignment for LLMsLink
487A Comprehensive Survey of Small Language Models in the Era of Large Language ModelsLink
488โ€œGive Me BF16 or Give Me Deathโ€? Accuracy-Performance Trade-Offs in LLM QuantizationLink
489Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test GenerationLink
490HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsLink
491Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data ContaminationLink
492Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-RewardingLink
493Number Cookbook: Number Understanding of Language Models and How to Improve ItLink
494Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation ModelsLink
495BitNet a4.8: 4-bit Activations for 1-bit LLMsLink
496Scaling Laws for PrecisionLink
497Energy Efficient Protein Language ModelsLink
498Balancing Pipeline Parallelism with Vocabulary ParallelismLink
499Toward Optimal Search and Retrieval for RAGLink
500Large Language Models Can Self-Improve in Long-context ReasoningLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
501Stronger Models are NOT Stronger Teachers for Instruction TuningLink
502Direct Preference Optimization Using Sparse Feature-Level ConstraintsLink
503Cut Your Losses in Large-Vocabulary Language ModelsLink
504Does Prompt Formatting Have Any Impact on LLM Performance?Link
505SymDPO: Boosting In-Context Learning of Large Multimodal ModelsLink
506SageAttention2 Technical ReportLink
507Bi-Mamba: Towards Accurate 1-Bit State Space ModelsLink
508RedPajama: An Open Dataset for Training Large Language ModelsLink
509Hymba: A Hybrid-head Architecture for Small Language ModelsLink
510Loss-to-Loss Prediction: Scaling Laws for All DatasetsLink
511When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context TrainingLink
512Multimodal Autoregressive Pre-training of Large Vision EncodersLink
513Natural Language Reinforcement LearningLink
514Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsLink
515TรœLU 3: Pushing Frontiers in Open Language Model Post-TrainingLink
516MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMsLink
517LLMs Do Not Think Step-by-step In Implicit ReasoningLink
518O1 Replication Journey โ€“ Part 2Link
519Star Attention: Efficient LLM Inference over Long SequencesLink
520Low-Bit Quantization Favors Undertrained LLMsLink
521Rethinking Token Reduction in MLLMsLink
522Reverse Thinking Makes LLMs Stronger ReasonersLink
523Critical Tokens MatterLink
524Foundations of Large Language ModelsLink
525A Survey of Research in Large Language Models for Electronic Design AutomationLink
526Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language ModelsLink
527Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' QuestionsLink
528The Future of AI: Exploring the Potential of Large Concept ModelsLink
529Investigating Numerical Translation with Large Language ModelsLink
530CALM: Curiosity-Driven Auditing for Large Language ModelsLink

๐Ÿ“Œ Explore cutting-edge research โœจ and stay updated! ๐Ÿง ๐ŸŒŸ


๐Ÿ†• 2025โ€“2026 Notable Additions

๐Ÿ”ข S.No.๐Ÿ“ Paper Title๐Ÿ”— Link
531DeepSeek-V3 Technical ReportLink
532DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningLink
533Kimi k1.5: Scaling Reinforcement Learning with LLMsLink
534Reasoning Language Models: A BlueprintLink
535Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-ThoughtLink
536The Lessons of Developing Process Reward Models in Mathematical ReasoningLink
537LIMO: Less is More for ReasoningLink
538Demystifying Long Chain-of-Thought Reasoning in LLMsLink
539Competitive Programming with Large Reasoning ModelsLink
540LLMs Can Easily Learn to Reason from Demonstrations: Structure, Not Content, Is What MattersLink
541Training Language Models to Reason EfficientlyLink
542Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement LearningLink
543SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionLink
544On the Emergence of Thinking in LLMs: Searching for the Right IntuitionLink
545Exploring the Limit of Outcome Reward for Learning Mathematical ReasoningLink
546Teaching Language Models to Critique via Reinforcement LearningLink
547A Review of DeepSeek Models' Key Innovative TechniquesLink
548R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement LearningLink
549Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningLink
550Understanding R1-Zero-Like Training: A Critical PerspectiveLink
551ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningLink
552Open-Reasoner-Zero: An Open Source Approach to Scaling Up RL on the Base ModelLink
553Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn'tLink
554Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in LLMsLink
555JudgeLRM: Large Reasoning Models as a JudgeLink
556Concise Reasoning via Reinforcement LearningLink
557Absolute Zero: Reinforced Self-play Reasoning with Zero DataLink
558Qwen3 Technical ReportLink
559MiMo: Unlocking the Reasoning Potential of Language Models โ€” From Pretraining to PosttrainingLink
560Llama-Nemotron: Efficient Reasoning ModelsLink
561INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement LearningLink
562AdaptThink: Reasoning Models Can Learn When to ThinkLink
563Thinkless: LLM Learns When to ThinkLink
564General-Reasoner: Advancing LLM Reasoning Across All DomainsLink
565Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsLink
566Reinforcement Pre-TrainingLink
567MagistralLink
568AlphaEvolve: A Coding Agent for Scientific and Algorithmic DiscoveryLink
569SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingLink
570Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond The Base Model?Link
571From System 1 to System 2: A Survey of Reasoning Large Language ModelsLink
572Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning LLMsLink
573Multi-Agent Collaboration Mechanisms: A Survey of LLMsLink
574Search-o1: Agentic Search-Enhanced Large Reasoning ModelsLink
575Reasoning Models Can Be Effective Without ThinkingLink
576RM-R1: Reward Modeling as ReasoningLink
577QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement LearningLink
578Enigmata: Scaling Logical Reasoning in LLMs with Synthetic Verifiable PuzzlesLink
579ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in LLMsLink
580Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective RL for LLM ReasoningLink
581Spurious Rewards: Rethinking Training Signals in RLVRLink
582Tina: Tiny Reasoning Models via LoRALink
583Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in MathLink
584VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningLink
585Gemini Robotics: Bringing AI Into the Physical WorldLink
586Gemini 2.0: The Era of Multimodal Agentic AILink
587RARE: Retrieval-Augmented Reasoning ModelingLink
588Learning from Failures in Multi-Attempt Reinforcement LearningLink
589R1-VL: Learning to Reason with Multimodal LLMs via Step-wise Group Relative Policy OptimizationLink
590Diffusion-Based Language Models: A SurveyLink
591FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingLink
592DAPO: An Open-Source LLM Reinforcement Learning System at ScaleLink
593Scaling Laws for Inference-Time ComputeLink
594LLM Post-Training: A Deep Dive into Reasoning Large Language ModelsLink
595Long-VITA: Scaling Large Vision-Language Models for Long Video UnderstandingLink
596Wan: Open and Advanced Large-Scale Video Generative ModelsLink
597Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning ModelsLink
598Seed1.5-VL Technical ReportLink
599A Survey on LLM-based Agents: Recent Advances and New FrontiersLink
600RLVR Is Not RL: Revisiting Reinforcement Learning for LLMsLink
601RL Tango: Reinforcing Generator and Verifier Together for Language ReasoningLink
602Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement LearningLink
603LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RLLink
604Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleLink
605Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model CapabilityLink
606The First Few Tokens Are All You Need: Unsupervised Prefix Fine-Tuning for Reasoning ModelsLink
607Learning to Reason without External RewardsLink
608Genius: A Generalizable and Purely Unsupervised Self-Training Framework for Advanced ReasoningLink
609Reinforcement Learning Teachers of Test Time ScalingLink
610Rewarding the Unlikely: Lifting GRPO Beyond Distribution SharpeningLink

Contributors

itsual

7 commits

itsual/Notable-LLM-Research-Papers

Curated list of research papers published in 2024 related to Large Language Models (LLM)

50

7 commits

updated Apr 14, 2026

See the code

README

๐Ÿ“š Notable-LLM-Research-Papers ๐ŸŽ“โœจ

Hello, curious reader! ๐Ÿ‘‹ Are you someone who loves exploring the cutting-edge world of AI and large language models (LLMs)? Or maybe you're a researcher, developer, or simply an enthusiast wondering about how these incredible technologies are evolving? Either way, you're in the right place. This repository is your gateway to 600+ groundbreaking research papers that capture the exciting developments in AI (in recent months)

๐Ÿš€ Whatโ€™s in This Repository?

These papers cover a wide spectrum of topics, ranging from improving model efficiency to aligning AI with human values, and even diving into the magic of multimodal models. The field of AI is advancing faster than ever, and these works showcase the brilliant ideas and innovations shaping the future.

The papers broadly fall under the following categories:


1. Model Architecture and Efficiency ๐Ÿ—๏ธโšก

Papers that introduce innovative architectures, explore scaling laws, or improve the efficiency of LLMs. These works aim to make AI faster, cheaper, and better.


2. Preference Optimization and RLHF (Reinforcement Learning with Human Feedback) ๐Ÿค๐Ÿ’ก

Research exploring how to align AI with human preferences, ensuring outputs are not just accurate but also ethical, safe, and aligned with societal values.


3. Multimodal Models ๐Ÿ–ผ๏ธ๐ŸŽ™๏ธโœ๏ธ

Ever wondered how AI can see, read, and understand all at once? These papers focus on models that combine multiple types of data, like images, text, and audio, to build richer, more versatile systems.


4. Long-Context Learning and Inference ๐Ÿง ๐ŸŒ€

Tackling the challenges of long-context understanding, these papers discuss how to extend the memory of LLMs, making them capable of reasoning across longer documents or conversations.


5. Reasoning and Knowledge ๐Ÿ”๐Ÿง 

From enabling AI to solve complex puzzles to improving its reasoning capabilities, these papers explore how LLMs "think" and manage the vast knowledge they've been trained on.


6. Quantization and Compression ๐Ÿ“๐Ÿ—œ๏ธ

Smaller, faster, and more efficientโ€”these papers delve into techniques like quantization and compression to make AI models more practical for real-world applications.


7. Evaluation and Benchmarks ๐Ÿ“Š๐Ÿ“ˆ

A great model needs great evaluation. These papers propose new benchmarks and methodologies to assess the capabilities of AI systems more rigorously.


8. Instruction Tuning and Alignment ๐Ÿ“๐ŸŽฏ

Want your AI to follow your instructions perfectly? These papers refine how we align LLMs to understand and execute instructions across various tasks.


9. Surveys and Meta-Analysis ๐Ÿ”ฌ๐Ÿ“š

Meta-level research analyzing trends, methods, and applications in AI. Perfect for getting a birdโ€™s-eye view of where the field is heading.


10. Applications and Tooling ๐Ÿ”ง๐ŸŒ

Papers showcasing how LLMs are applied in diverse domains, from healthcare to coding and everything in between. These highlight the transformative power of AI in the real world.


๐ŸŒŸ Why Should You Care?

These papers represent the state of the art in AI research. Whether you're here to stay updated, seek inspiration, or find practical tools, this collection is a treasure trove for AI enthusiasts. Dive in, explore, and let your curiosity guide you! ๐ŸŒโœจ


Ready to start? ๐ŸŽ‰ Browse the list and explore the exciting breakthroughs shaping the future of AI! ๐Ÿš€


๐Ÿ”ข S.No.๐Ÿ“ Paper Title๐Ÿ”— Link
1LLM Maybe LongLM: Self-Extend LLM Context Window Without TuningLink
2Knowledge Fusion of Large Language ModelsLink
3A Comprehensive Study of Knowledge Editing for Large Language ModelsLink
4DiffusionGPT: LLM-Driven Text-to-Image Generation SystemLink
5Tuning Language Models by ProxyLink
6An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language ModelsLink
7Soaring from 4K to 400K: Extending LLMโ€™s Context with Activation BeaconLink
8VMamba: Visual State Space ModelLink
9LLaMA Beyond English: An Empirical Study on Language Capability TransferLink
10DeepSeek LLM: Scaling Open-Source Language Models with LongtermismLink
11Blending Is All You Need: Cheaper, Better Alternative to Trillion-Parameters LLMLink
12LLaMA Pro: Progressive LLaMA with Block ExpansionLink
13RoSA: Accurate Parameter-Efficient Fine-Tuning via Robust AdaptationLink
14Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextLink
15Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingLink
16WARM: On the Benefits of Weight Averaged Reward ModelsLink
17SpacTor-T5: Pre-training T5 Models with Span Corruption and Replaced Token DetectionLink
18A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and ToxicityLink
19Mixtral of ExpertsLink
20MoE-Mamba: Efficient Selective State Space Models with Mixture of ExpertsLink
21Code Generation with AlphaCodium: From Prompt Engineering to Flow EngineeringLink
22EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyLink
23KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationLink
24A Closer Look at AUROC and AUPRC under Class ImbalanceLink
25Transformers are Multi-State RNNsLink
26LLM Augmented LLMs: Expanding Capabilities through CompositionLink
27Rethinking Patch Dependence for Masked AutoencodersLink
28Astraios: Parameter-Efficient Instruction Tuning Code Large Language ModelsLink
29Pix2gestalt: Amodal Segmentation by Synthesizing WholesLink
30RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on AgricultureLink
31An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language ModelsLink
32Knowledge Fusion of Large Language ModelsLink
33Scalable Pre-training of Large Autoregressive Image ModelsLink
34SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning CapabilitiesLink
35Multimodal Pathway: Improve Transformers with Irrelevant Data from Other ModalitiesLink
36MambaByte: Token-free Selective State Space ModelLink
37ReFT: Reasoning with Reinforced Fine-TuningLink
38Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsLink
39Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingLink
40Denoising Vision TransformersLink
41Self-Rewarding Language ModelsLink
42LoRA+: Efficient Low Rank Adaptation of Large ModelsLink
43MobileVLM V2: Faster and Stronger Baseline for Vision Language ModelLink
44Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?Link
45ODIN: Disentangled Reward Mitigates Hacking in RLHFLink
46Genie: Generative Interactive EnvironmentsLink
47A Phase Transition Between Positional and Semantic Learning in a Solvable Model of Dot-Product AttentionLink
48Neural Network DiffusionLink
49More Agents Is All You NeedLink
50Scaling Laws for Downstream Task Performance of Large Language ModelsLink
-----------------------------------------------------------------------------------------------------------------------------------------------------------------------
51Repeat After Me: Transformers are Better than State Space Models at CopyingLink
52Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language ModelsLink
53AnyGPT: Unified Multimodal LLM with Discrete Sequence ModelingLink
54BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of DataLink
55LongAgent: Scaling Language Models to 128k Context through Multi-Agent CollaborationLink
56LongRoPE: Extending LLM Context Window Beyond 2 Million TokensLink
57Policy Improvement using Language Feedback ModelsLink
58DoRA: Weight-Decomposed Low-Rank AdaptationLink
59FindingEmo: An Image Dataset for Emotion Recognition in the WildLink
60TinyLLaVA: A Framework of Small-scale Large Multimodal ModelsLink
61Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation ModelsLink
62When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning MethodLink
63Step-On-Feet Tuning: Scaling Self-Alignment of LLMs via BootstrappingLink
64Suppressing Pink Elephants with Direct Principle FeedbackLink
65The Era of 1-bit LLMs: All Large Language Models are in 1.58 BitsLink
66LiPO: Listwise Preference Optimization through Learning-to-RankLink
67Aya Model: An Instruction Finetuned Open-Access Multilingual Language ModelLink
68The Boundary of Neural Network Trainability is FractalLink
69Recovering the Pre-Fine-Tuning Weights of Generative ModelsLink
70Scaling Laws for Fine-Grained Mixture of ExpertsLink
71Direct Language Model Alignment from Online AI FeedbackLink
72CARTE: Pretraining and Transfer for Tabular LearningLink
73Grandmaster-Level Chess Without SearchLink
74Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsLink
75Reformatted AlignmentLink
76OLMo: Accelerating the Science of Language ModelsLink
77FinTral: A Family of GPT-4 Level Multimodal Financial Large Language ModelsLink
78Mixtures of Experts Unlock Parameter Scaling for Deep RLLink
79Generative Representational Instruction TuningLink
80World Model on Million-Length Video And Language With RingAttentionLink
81Efficient Exploration for LLMsLink
82YOLOv9: Learning What You Want to Learn Using Programmable Gradient InformationLink
83Towards Cross-Tokenizer Distillation: The Universal Logit Distillation Loss for LLMsLink
84Mixtures of Experts Unlock Parameter Scaling for Deep RLLink
85Sora Generates Videos with Stunning Geometrical ConsistencyLink
86BASE TTS: Lessons From Building a Billion-Parameter Text-to-Speech Model on 100K Hours of DataLink
87Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMsLink
88Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language ModelsLink
89DoRA: Weight-Decomposed Low-Rank AdaptationLink
90The Era of 1-bit LLMs: All Large Language Models are in 1.58 BitsLink
91Efficient Exploration for LLMsLink
92Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingLink
93You Only Cache Once: Decoder-Decoder Architectures for Language ModelsLink
94gzip Predicts Data-dependent Scaling LawsLink
95Self-Play Preference Optimization for Language Model AlignmentLink
96PHUDGE: Phi-3 as Scalable JudgeLink
97What Matters When Building Vision-Language Models?Link
98Towards Modular LLMs by Building and Reusing a Library of LoRAsLink
99Contextual Position Encoding: Learning to Count Whatโ€™s ImportantLink
100RLHF Workflow: From Reward Modeling to Online RLHFLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
101Attention as an RNNLink
102AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityLink
103Instruction Tuning With Loss Over InstructionsLink
104LoRA Learns Less and Forgets LessLink
105Trans-LoRA: Towards Data-free Transferable Parameter Efficient FinetuningLink
106VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsLink
107MoRA: High-Rank Updating for Parameter-Efficient Fine-TuningLink
108LLaMA-NAS: Efficient Neural Architecture Search for Large Language ModelsLink
109SimPO: Simple Preference Optimization with a Reference-Free RewardLink
110The Road Less ScheduledLink
111Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsLink
112Is Flash Attention Stable?Link
113Value Augmented Sampling for Language Model Alignment and PersonalizationLink
114Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Link
115vAttention: Dynamic Memory Management for Serving LLMs without PagedAttentionLink
116A Careful Examination of Large Language Model Performance on Grade School ArithmeticLink
117Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language ModelsLink
118DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language ModelLink
119Dense Connector for MLLMsLink
120xLSTM: Extended Long Short-Term MemoryLink
121SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationLink
122Xmodel-VLM: A Simple Baseline for Multimodal Vision Language ModelLink
123Chameleon: Mixed-Modal Early-Fusion Foundation ModelsLink
124Is Bigger Edit Batch Size Always Better? An Empirical Study on Model Editing with Llama-3Link
125AlignGPT: Multi-modal Large Language Models with Adaptive Alignment CapabilityLink
126The Prompt Report: A Systematic Survey of Prompting TechniquesLink
127Creativity Has Left the Chat: The Price of Debiasing Language ModelsLink
128Show, Donโ€™t Tell: Aligning Language Models with Demonstrated FeedbackLink
129WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the WildLink
130Scalable MatMul-free Language ModelingLink
131Never Miss A Beat: An Efficient Recipe for Context Window Extension of Large Language ModelsLink
132Boosting Large-scale Parallel Training Efficiency with C4: A Communication-Driven ApproachLink
133Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language ModelsLink
134MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer DecodingLink
135Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLink
136An Image is Worth 32 Tokens for Reconstruction and GenerationLink
137Block Transformer: Global-to-Local Language Modeling for Fast InferenceLink
1383D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less HallucinationLink
139Transformers Need Glasses! Information Over-Squashing in Language TasksLink
140The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsLink
141BERTs are Generative In-Context LearnersLink
142An Empirical Study of Mamba-based Language ModelsLink
143Discovering Preference Optimization Algorithms with and for Large Language ModelsLink
144Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with NothingLink
145Autoregressive Model Beats Diffusion: Llama for Scalable Image GenerationLink
146Are We Done with MMLU?Link
147OLoRA: Orthonormal Low-Rank Adaptation of Large Language ModelsLink
148Step-aware Preference Optimization: Aligning Preference with Denoising Performance at Each StepLink
149Husky: A Unified, Open-Source Language Agent for Multi-Step ReasoningLink
150Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated ParametersLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
151Simple and Effective Masked Diffusion Language ModelsLink
152The Prompt Report: A Systematic Survey of Prompting TechniquesLink
153Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsLink
154Be like a Goldfish, Donโ€™t Memorize! Mitigating Memorization in Generative LLMsLink
155TextGrad: Automatic โ€œDifferentiationโ€ via TextLink
156Large Language Models Must Be Taught to Know What They Donโ€™t KnowLink
157What If We Recaption Billions of Web Images with LLaMA-3?Link
158Discovering Preference Optimization Algorithms with and for Large Language ModelsLink
159An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual PixelsLink
160Buffer of Thoughts: Thought-Augmented Reasoning with Large Language ModelsLink
161CRAG โ€“ Comprehensive RAG BenchmarkLink
162Margin-aware Preference Optimization for Aligning Diffusion Models Without ReferenceLink
163Mixture-of-Agents Enhances Large Language Model CapabilitiesLink
164Large Language Model Unlearning via Embedding-Corrupted PromptsLink
165Bootstrapping Language Models with DPO Implicit RewardsLink
166THEANINE: Revisiting Memory Management in Long-term Conversations with Timeline-augmented Response GenerationLink
167Task Me AnythingLink
168Nemotron-4 340B Technical ReportLink
169Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-JudgesLink
170How Do Large Language Models Acquire Factual Knowledge During Pretraining?Link
171mDPO: Conditional Preference Optimization for Multimodal Large Language ModelsLink
172Unveiling Encoder-Free Vision-Language ModelsLink
173HARE: HumAn pRiors, a key to small language model EfficiencyLink
174Measuring Memorization in RLHF for Code CompletionLink
175DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code IntelligenceLink
176Iterative Length-Regularized Direct Preference Optimization: Improving 7B Language Models to GPT-4 LevelLink
177From RAGs to Rich Parameters: Probing How Language Models Utilize External Knowledge Over Parametric InformationLink
178DataComp-LM: In Search of the Next Generation of Training Sets for Language ModelsLink
179Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?Link
180Instruction Pre-Training: Language Models are Supervised Multitask LearnersLink
181Can LLMs Learn by Teaching? A Preliminary StudyLink
182A Tale of Trust and Accuracy: Base vs. Instruct LLMs in RAG SystemsLink
183LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMsLink
184MoA: Mixture of Sparse Attention for Automatic Large Language Model CompressionLink
185Efficient Continual Pre-training by Mitigating the Stability GapLink
186Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range TransformersLink
187WARP: On the Benefits of Weight Averaged Rewarded PoliciesLink
188Adam-mini: Use Fewer Learning Rates To Gain MoreLink
189The FineWeb Datasets: Decanting the Web for the Finest Text Data at ScaleLink
190LongIns: A Challenging Long-context Instruction-based Exam for LLMsLink
191Following Length Constraints in InstructionsLink
192A Closer Look into Mixture-of-Experts in Large Language ModelsLink
193RouteLLM: Learning to Route LLMs with Preference DataLink
194Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMsLink
195Dataset Size Recovery from LoRA WeightsLink
196From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic DataLink
197Changing Answer Order Can Decrease MMLU AccuracyLink
198Direct Preference Knowledge Distillation for Large Language ModelsLink
199LLM Critics Help Catch LLM BugsLink
200Scaling Synthetic Data Creation with 1,000,000,000 PersonasLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
201Tokenization Falling Short: The Curse of TokenizationLink
202Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMsLink
203Bootstrapping Language Models with DPO Implicit RewardsLink
204Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-JudgesLink
205Learning and Leveraging World Models in Visual Representation LearningLink
206Improving LLM Code Generation with Grammar AugmentationLink
207The Hidden Attention of Mamba ModelsLink
208Training-Free Pretrained Model MergingLink
209Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like ArchitecturesLink
210The WMDP Benchmark: Measuring and Reducing Malicious Use With UnlearningLink
211Evolution Transformer: In-Context Evolutionary OptimizationLink
212Enhancing Vision-Language Pre-training with Rich SupervisionsLink
213Scaling Rectified Flow Transformers for High-Resolution Image SynthesisLink
214Design2Code: How Far Are We From Automating Front-End Engineering?Link
215ShortGPT: Layers in Large Language Models are More Redundant Than You ExpectLink
216Backtracing: Retrieving the Cause of the QueryLink
217Learning to Decode Collaboratively with Multiple Language ModelsLink
218SaulLM-7B: A pioneering Large Language Model for LawLink
219Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal ReasoningLink
2203D Diffusion PolicyLink
221MedMamba: Vision Mamba for Medical Image ClassificationLink
222GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionLink
223Stop Regressing: Training Value Functions via Classification for Scalable Deep RLLink
224How Far Are We from Intelligent Visual Deductive Reasoning?Link
225Common 7B Language Models Already Possess Strong Math CapabilitiesLink
226Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of ContextLink
227Is Cosine-Similarity of Embeddings Really About Similarity?Link
228LLM4Decompile: Decompiling Binary Code with Large Language ModelsLink
229Algorithmic Progress in Language ModelsLink
230Stealing Part of a Production Language ModelLink
231Chronos: Learning the Language of Time SeriesLink
232Simple and Scalable Strategies to Continually Pre-train Large Language ModelsLink
233Language Models Scale Reliably With Over-Training and on Downstream TasksLink
234BurstAttention: An Efficient Distributed Attention Framework for Extremely Long SequencesLink
235LocalMamba: Visual State Space Model with Windowed Selective ScanLink
236GiT: Towards Generalist Vision Transformer through Universal Language InterfaceLink
237MM1: Methods, Analysis & Insights from Multimodal LLM Pre-trainingLink
238RAFT: Adapting Language Model to Domain Specific RAGLink
239TnT-LLM: Text Mining at Scale with Large Language ModelsLink
240Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionLink
241PERL: Parameter Efficient Reinforcement Learning from Human FeedbackLink
242RewardBench: Evaluating Reward Models for Language ModelingLink
243LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language ModelsLink
244RakutenAI-7B: Extending Large Language Models for JapaneseLink
245SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time SeriesLink
246Can Large Language Models Explore In-Context?Link
247LLM2LLM: Boosting LLMs with Novel Iterative Data EnhancementLink
248LLM Agent Operating SystemLink
249The Unreasonable Ineffectiveness of the Deeper LayersLink
250BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical TextLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
251ViTAR: Vision Transformer with Any ResolutionLink
252Long-form Factuality in Large Language ModelsLink
253Mini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsLink
254LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-TuningLink
255Mechanistic Design and Scaling of Hybrid ArchitecturesLink
256MagicLens: Self-Supervised Image Retrieval with Open-Ended InstructionsLink
257Model Stock: All We Need Is Just a Few Fine-Tuned ModelsLink
258Do Language Models Plan Ahead for Future Tokens?Link
259Bigger is not Always Better: Scaling Properties of Latent Diffusion ModelsLink
260The Fine Line: Navigating Large Language Model Pretraining with Down-streaming Capability AnalysisLink
261Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion ModelsLink
262Mixture-of-Depths: Dynamically Allocating Compute in Transformer-Based Language ModelsLink
263Long-context LLMs Struggle with Long In-context LearningLink
264Emergent Abilities in Reduced-Scale Generative Language ModelsLink
265Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksLink
266On the Scalability of Diffusion-based Text-to-Image GenerationLink
267BAdam: A Memory Efficient Full Parameter Training Method for Large Language ModelsLink
268Cross-Attention Makes Inference Cumbersome in Text-to-Image Diffusion ModelsLink
269Direct Nash Optimization: Teaching Language Models to Self-Improve with General PreferencesLink
270Training LLMs over Neurally Compressed TextLink
271CantTalkAboutThis: Aligning Language Models to Stay on Topic in DialoguesLink
272ReFT: Representation Finetuning for Language ModelsLink
273Verifiable by Design: Aligning Language Models to Quote from Pre-Training DataLink
274Sigma: Siamese Mamba Network for Multi-Modal Semantic SegmentationLink
275AutoCodeRover: Autonomous Program ImprovementLink
276Eagle and Finch: RWKV with Matrix-Valued States and Dynamic RecurrenceLink
277CodecLM: Aligning Language Models with Tailored Synthetic DataLink
278MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training StrategiesLink
279Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language ModelsLink
280LLM2Vec: Large Language Models Are Secretly Powerful Text EncodersLink
281Adapting LLaMA Decoder to Vision TransformerLink
282Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attentionLink
283LLoCO: Learning Long Contexts OfflineLink
284JetMoE: Reaching Llama2 Performance with 0.1M DollarsLink
285Best Practices and Lessons Learned on Synthetic Data for Language ModelsLink
286Rho-1: Not All Tokens Are What You NeedLink
287Pre-training Small Base LMs with Fewer TokensLink
288Dataset Reset Policy Optimization for RLHFLink
289LLM In-Context Recall is Prompt DependentLink
290State Space Model for New-Generation Network Alternative to Transformers: A SurveyLink
291Chinchilla Scaling: A Replication AttemptLink
292Learn Your Reference Model for Real Good AlignmentLink
293Is DPO Superior to PPO for LLM Alignment? A Comprehensive StudyLink
294Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training StrategiesLink
295How Faithful Are RAG Models? Quantifying the Tug-of-War Between RAG and LLMsโ€™ Internal PriorLink
296A Survey on Retrieval-Augmented Text Generation for Large Language ModelsLink
297When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many ClassesLink
298Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingLink
299OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction DataLink
300The Instruction Hierarchy: Training LLMs to Prioritize Privileged InstructionsLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
301How Good Are Low-bit Quantized LLaMA3 Models? An Empirical StudyLink
302Phi-3 Technical Report: A Highly Capable Language Model Locally on Your PhoneLink
303OpenELM: An Efficient Language Model Family with Open-source Training and Inference FrameworkLink
304A Survey on Self-Evolution of Large Language ModelsLink
305Multi-Head Mixture-of-ExpertsLink
306NExT: Teaching Large Language Models to Reason about Code ExecutionLink
307Graph Machine Learning in the Era of Large Language Models (LLMs)Link
308Retrieval Head Mechanistically Explains Long-Context FactualityLink
309Layer Skip: Enabling Early Exit Inference and Self-Speculative DecodingLink
310Make Your LLM Fully Utilize the ContextLink
311LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical ReportLink
312Better & Faster Large Language Models via Multi-token PredictionLink
313RAG and RAU: A Survey on Retrieval-Augmented Language Model in Natural Language ProcessingLink
314A Primer on the Inner Workings of Transformer-based Language ModelsLink
315When to Retrieve: Teaching LLMs to Utilize Information Retrieval EffectivelyLink
316KAN: Kolmogorovโ€“Arnold NetworksLink
317LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable ObjectivesLink
318Searching for Best Practices in Retrieval-Augmented GenerationLink
319Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsLink
320Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionLink
321Eliminating Position Bias of Language Models: A Mechanistic ApproachLink
322JMInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionLink
323TokenPacker: Efficient Visual Projector for Multimodal LLMLink
324Reasoning in Large Language Models: A Geometric PerspectiveLink
325RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMsLink
326AgentInstruct: Toward Generative Teaching with Agentic FlowsLink
327HEMM: Holistic Evaluation of Multimodal Foundation ModelsLink
328Mixture of A Million ExpertsLink
329Learning to (Learn at Test Time): RNNs with Expressive Hidden StatesLink
330Vision Language Models Are BlindLink
331Self-Recognition in Language ModelsLink
332Inference Performance Optimization for Large Language Models on CPUsLink
333Gradient Boosting Reinforcement LearningLink
334FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionLink
335SpreadsheetLLM: Encoding Spreadsheets for Large Language ModelsLink
336New Desiderata for Direct Preference OptimizationLink
337Context Embeddings for Efficient Answer Generation in RAGLink
338Qwen2 Technical ReportLink
339The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-DeterminismLink
340From GaLore to WeLore: How Low-Rank Weights Non-uniformly Emerge from Low-Rank GradientsLink
341GoldFinch: High Performance RWKV/Transformer Hybrid with Linear Pre-Fill and Extreme KV-Cache CompressionLink
342Scaling Diffusion Transformers to 16 Billion ParametersLink
343NeedleBench: Can LLMs Do Retrieval and Reasoning in 1 Million Context Window?Link
344Patch-Level Training for Large Language ModelsLink
345LMMs-Eval: Reality Check on the Evaluation of Large Multimodal ModelsLink
346A Survey of Prompt Engineering Methods in Large Language Models for Different NLP TasksLink
347Spectra: A Comprehensive Study of Ternary, Quantized, and FP16 Language ModelsLink
348Attention Overflow: Language Model Input Blur during Long-Context Missing Items RecommendationLink
349Weak-to-Strong ReasoningLink
350Understanding Reference Policies in Direct Preference OptimizationLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
351Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesLink
352BOND: Aligning LLMs with Best-of-N DistillationLink
353Compact Language Models via Pruning and Knowledge DistillationLink
354LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM InferenceLink
355Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences TrainingLink
356DDK: Distilling Domain Knowledge for Efficient Large Language ModelsLink
357Generation Constraint Scaling Can Mitigate HallucinationLink
358Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid ApproachLink
359Course-Correction: Safety Alignment Using Synthetic PreferencesLink
360Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?Link
361Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-JudgeLink
362Improving Retrieval Augmented Language Model with Self-ReasoningLink
363Apple Intelligence Foundation Language ModelsLink
364ThinK: Thinner Key Cache by Query-Driven PruningLink
365The Llama 3 Herd of ModelsLink
366Gemma 2: Improving Open Language Models at a Practical SizeLink
367SAM 2: Segment Anything in Images and VideosLink
368POA: Pre-training Once for Models of All SizesLink
369RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkLink
370A Survey of MambaLink
371MiniCPM-V: A GPT-4V Level MLLM on Your PhoneLink
372RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented GenerationLink
373Self-Taught EvaluatorsLink
374BioMamba: A Pre-trained Biomedical Language Representation Model Leveraging MambaLink
375EXAONE 3.0 7.8B Instruction Tuned Language ModelLink
3761.5-Pints Technical Report: Pretraining in Days, Not Months โ€“ Your Language Model Thrives on Quality DataLink
377Conversational Prompt EngineeringLink
378Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLPLink
379The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoveryLink
380Hermes 3 Technical ReportLink
381Customizing Language Models with Instance-wise LoRA for Sequential RecommendationLink
382Enhancing Robustness in Large Language Models: Prompting for Mitigating the Impact of Irrelevant InformationLink
383To Code, or Not To Code? Exploring Impact of Code in Pre-trainingLink
384LLM Pruning and Distillation in Practice: The Minitron ApproachLink
385Jamba-1.5: Hybrid Transformer-Mamba Models at ScaleLink
386Controllable Text Generation for Large Language Models: A SurveyLink
387Multi-Layer Transformers Gradient Can be Approximated in Almost Linear TimeLink
388A Practitionerโ€™s Guide to Continual Multimodal PretrainingLink
389Building and Better Understanding Vision-Language Models: Insights and Future DirectionsLink
390CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting MitigationLink
391The Mamba in the Llama: Distilling and Accelerating Hybrid ModelsLink
392ReMamba: Equip Mamba with Effective Long-Sequence ModelingLink
393Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal SamplingLink
394LongRecipe: Recipe for Efficient Long Context Generalization in Large Language ModelsLink
395Switti: Designing Scale-Wise Transformers for Text-to-Image SynthesisLink
396X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation ModelsLink
397Free Process Rewards without Process LabelsLink
398Scaling Image Tokenizers with Grouped Spherical QuantizationLink
399RARE: Retrieval-Augmented Reasoning Enhancement for Large Language ModelsLink
400Perception Tokens Enhance Visual Reasoning in Multimodal Language ModelsLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
401Evaluating Language Models as Synthetic Data GeneratorsLink
402Best-of-N JailbreakingLink
403PaliGemma 2: A Family of Versatile VLMs for TransferLink
404VisionZip: Longer is Better but Not Necessary in Vision Language ModelsLink
405Evaluating and Aligning CodeLLMs on Human PreferenceLink
406MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at ScaleLink
407Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time ScalingLink
408LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation MethodsLink
409Does RLHF Scale? Exploring the Impacts From Data, Model, and MethodLink
410Unraveling the Complexity of Memory in RL Agents: An Approach for Classification and EvaluationLink
411Training Large Language Models to Reason in a Continuous Latent SpaceLink
412AutoReason: Automatic Few-Shot Reasoning DecompositionLink
413Large Concept Models: Language Modeling in a Sentence Representation SpaceLink
414Phi-4 Technical ReportLink
415Byte Latent Transformer: Patches Scale Better Than TokensLink
416SCBench: A KV Cache-Centric Analysis of Long-Context MethodsLink
417Cultural Evolution of Cooperation among LLM AgentsLink
418DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal UnderstandingLink
419No More Adam: Learning Rate Scaling at Initialization is All You NeedLink
420Precise Length Control in Large Language ModelsLink
421The Open Source Advantage in Large Language Models (LLMs)Link
422A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & ChallengesLink
423Are Your LLMs Capable of Stable Reasoning?Link
424LLM Post-Training Recipes: Improving Reasoning in LLMsLink
425Hansel: Output Length Controlling Framework for Large Language ModelsLink
426Mind Your Theory: Theory of Mind Goes Deeper Than ReasoningLink
427Alignment Faking in Large Language ModelsLink
428SCOPE: Optimizing Key-Value Cache Compression in Long-Context GenerationLink
429LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-Context MultitasksLink
430Offline Reinforcement Learning for LLM Multi-Step ReasoningLink
431Mulberry: Empowering MLLM with O1-like Reasoning and Reflection via Collective Monte Carlo Tree SearchLink
432Titans: Learning to Memorize at Test TimeLink
433Addition is All You Need for Energy-efficient Language ModelsLink
434Quantifying Generalization Complexity for Large Language ModelsLink
435When a Language Model is Optimized for Reasoning, Does It Still Show Embers of Autoregression?Link
436Were RNNs All We Needed?Link
437Selective Attention Improves TransformerLink
438LLMs Know More Than They Show: On the Intrinsic Representation of LLM HallucinationsLink
439LLaVA-Critic: Learning to Evaluate Multimodal ModelsLink
440Differential TransformerLink
441GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsLink
442ARIA: An Open Multimodal Native Mixture-of-Experts ModelLink
443O1 Replication Journey: A Strategic Progress Report โ€“ Part 1Link
444Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAGLink
445From Generalist to Specialist: Adapting Vision Language Models via Task-Specific Visual Instruction TuningLink
446KV Prediction for Improved Time to First TokenLink
447Baichuan-Omni Technical ReportLink
448MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language ModelsLink
449LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal ModelsLink
450AFlow: Automating Agentic Workflow GenerationLink
-----------------------------------------------------------------------------------------------------------------------------------------------------------------------
451Toward General Instruction-Following Alignment for Retrieval-Augmented GenerationLink
452Pre-training Distillation for Large Language Models: A Design Space ExplorationLink
453MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language ModelsLink
454Scalable Ranked Preference Optimization for Text-to-Image GenerationLink
455Scaling Diffusion Language Models via Adaptation from Autoregressive ModelsLink
456Hybrid Preferences: Learning to Route Instances for Human vs. AI FeedbackLink
457Counting Ability of Large Language Models and Impact of TokenizationLink
458A Survey of Small Language ModelsLink
459Accelerating Direct Preference Optimization with Prefix SharingLink
460Mind Your Step (by Step): Chain-of-Thought Can Reduce Performance on Tasks Where Thinking Makes Humans WorseLink
461LongReward: Improving Long-context Large Language Models with AI FeedbackLink
462ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM InferenceLink
463Beyond Text: Optimizing RAG with Multimodal Inputs for Industrial ApplicationsLink
464CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation GenerationLink
465What Happened in LLMs Layers When Trained for Fast vs. Slow Thinking: A Gradient PerspectiveLink
466GPT or BERT: Why Not Both?Link
467Language Models Can Self-Lengthen to Generate Long TextsLink
468OLMoE: Open Mixture-of-Experts Language ModelsLink
469In Defense of RAG in the Era of Long-Context Language ModelsLink
470Attention Heads of Large Language Models: A SurveyLink
471LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QALink
472How Do Your Code LLMs Perform? Empowering Code Instruction Tuning with High-Quality DataLink
473Theory, Analysis, and Best Practices for Sigmoid Self-AttentionLink
474LLaMA-Omni: Seamless Speech Interaction with Large Language ModelsLink
475What is the Role of Small Models in the LLM Era: A SurveyLink
476Policy Filtration in RLHF to Fine-Tune LLM for Code GenerationLink
477RetrievalAttention: Accelerating Long-Context LLM Inference via Vector RetrievalLink
478Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-ImprovementLink
479Qwen2.5-Coder Technical ReportLink
480Instruction Following without Instruction TuningLink
481Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical AnalysisLink
482The Perfect Blend: Redefining RLHF with Mixture of JudgesLink
483Adding Error Bars to Evals: A Statistical Approach to Language Model EvaluationsLink
484Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage AdaptationLink
485Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language ModelsLink
486Sample-Efficient Alignment for LLMsLink
487A Comprehensive Survey of Small Language Models in the Era of Large Language ModelsLink
488โ€œGive Me BF16 or Give Me Deathโ€? Accuracy-Performance Trade-Offs in LLM QuantizationLink
489Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test GenerationLink
490HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsLink
491Both Text and Images Leaked! A Systematic Analysis of Multimodal LLM Data ContaminationLink
492Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-RewardingLink
493Number Cookbook: Number Understanding of Language Models and How to Improve ItLink
494Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation ModelsLink
495BitNet a4.8: 4-bit Activations for 1-bit LLMsLink
496Scaling Laws for PrecisionLink
497Energy Efficient Protein Language ModelsLink
498Balancing Pipeline Parallelism with Vocabulary ParallelismLink
499Toward Optimal Search and Retrieval for RAGLink
500Large Language Models Can Self-Improve in Long-context ReasoningLink
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------
501Stronger Models are NOT Stronger Teachers for Instruction TuningLink
502Direct Preference Optimization Using Sparse Feature-Level ConstraintsLink
503Cut Your Losses in Large-Vocabulary Language ModelsLink
504Does Prompt Formatting Have Any Impact on LLM Performance?Link
505SymDPO: Boosting In-Context Learning of Large Multimodal ModelsLink
506SageAttention2 Technical ReportLink
507Bi-Mamba: Towards Accurate 1-Bit State Space ModelsLink
508RedPajama: An Open Dataset for Training Large Language ModelsLink
509Hymba: A Hybrid-head Architecture for Small Language ModelsLink
510Loss-to-Loss Prediction: Scaling Laws for All DatasetsLink
511When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context TrainingLink
512Multimodal Autoregressive Pre-training of Large Vision EncodersLink
513Natural Language Reinforcement LearningLink
514Large Multi-modal Models Can Interpret Features in Large Multi-modal ModelsLink
515TรœLU 3: Pushing Frontiers in Open Language Model Post-TrainingLink
516MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMsLink
517LLMs Do Not Think Step-by-step In Implicit ReasoningLink
518O1 Replication Journey โ€“ Part 2Link
519Star Attention: Efficient LLM Inference over Long SequencesLink
520Low-Bit Quantization Favors Undertrained LLMsLink
521Rethinking Token Reduction in MLLMsLink
522Reverse Thinking Makes LLMs Stronger ReasonersLink
523Critical Tokens MatterLink
524Foundations of Large Language ModelsLink
525A Survey of Research in Large Language Models for Electronic Design AutomationLink
526Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language ModelsLink
527Large Language Models, Knowledge Graphs and Search Engines: A Crossroads for Answering Users' QuestionsLink
528The Future of AI: Exploring the Potential of Large Concept ModelsLink
529Investigating Numerical Translation with Large Language ModelsLink
530CALM: Curiosity-Driven Auditing for Large Language ModelsLink

๐Ÿ“Œ Explore cutting-edge research โœจ and stay updated! ๐Ÿง ๐ŸŒŸ


๐Ÿ†• 2025โ€“2026 Notable Additions

๐Ÿ”ข S.No.๐Ÿ“ Paper Title๐Ÿ”— Link
531DeepSeek-V3 Technical ReportLink
532DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningLink
533Kimi k1.5: Scaling Reinforcement Learning with LLMsLink
534Reasoning Language Models: A BlueprintLink
535Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-ThoughtLink
536The Lessons of Developing Process Reward Models in Mathematical ReasoningLink
537LIMO: Less is More for ReasoningLink
538Demystifying Long Chain-of-Thought Reasoning in LLMsLink
539Competitive Programming with Large Reasoning ModelsLink
540LLMs Can Easily Learn to Reason from Demonstrations: Structure, Not Content, Is What MattersLink
541Training Language Models to Reason EfficientlyLink
542Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement LearningLink
543SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software EvolutionLink
544On the Emergence of Thinking in LLMs: Searching for the Right IntuitionLink
545Exploring the Limit of Outcome Reward for Learning Mathematical ReasoningLink
546Teaching Language Models to Critique via Reinforcement LearningLink
547A Review of DeepSeek Models' Key Innovative TechniquesLink
548R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement LearningLink
549Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement LearningLink
550Understanding R1-Zero-Like Training: A Critical PerspectiveLink
551ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningLink
552Open-Reasoner-Zero: An Open Source Approach to Scaling Up RL on the Base ModelLink
553Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn'tLink
554Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in LLMsLink
555JudgeLRM: Large Reasoning Models as a JudgeLink
556Concise Reasoning via Reinforcement LearningLink
557Absolute Zero: Reinforced Self-play Reasoning with Zero DataLink
558Qwen3 Technical ReportLink
559MiMo: Unlocking the Reasoning Potential of Language Models โ€” From Pretraining to PosttrainingLink
560Llama-Nemotron: Efficient Reasoning ModelsLink
561INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement LearningLink
562AdaptThink: Reasoning Models Can Learn When to ThinkLink
563Thinkless: LLM Learns When to ThinkLink
564General-Reasoner: Advancing LLM Reasoning Across All DomainsLink
565Darwin Godel Machine: Open-Ended Evolution of Self-Improving AgentsLink
566Reinforcement Pre-TrainingLink
567MagistralLink
568AlphaEvolve: A Coding Agent for Scientific and Algorithmic DiscoveryLink
569SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingLink
570Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond The Base Model?Link
571From System 1 to System 2: A Survey of Reasoning Large Language ModelsLink
572Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning LLMsLink
573Multi-Agent Collaboration Mechanisms: A Survey of LLMsLink
574Search-o1: Agentic Search-Enhanced Large Reasoning ModelsLink
575Reasoning Models Can Be Effective Without ThinkingLink
576RM-R1: Reward Modeling as ReasoningLink
577QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement LearningLink
578Enigmata: Scaling Logical Reasoning in LLMs with Synthetic Verifiable PuzzlesLink
579ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in LLMsLink
580Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective RL for LLM ReasoningLink
581Spurious Rewards: Rethinking Training Signals in RLVRLink
582Tina: Tiny Reasoning Models via LoRALink
583Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in MathLink
584VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningLink
585Gemini Robotics: Bringing AI Into the Physical WorldLink
586Gemini 2.0: The Era of Multimodal Agentic AILink
587RARE: Retrieval-Augmented Reasoning ModelingLink
588Learning from Failures in Multi-Attempt Reinforcement LearningLink
589R1-VL: Learning to Reason with Multimodal LLMs via Step-wise Group Relative Policy OptimizationLink
590Diffusion-Based Language Models: A SurveyLink
591FlowAR: Scale-wise Autoregressive Image Generation Meets Flow MatchingLink
592DAPO: An Open-Source LLM Reinforcement Learning System at ScaleLink
593Scaling Laws for Inference-Time ComputeLink
594LLM Post-Training: A Deep Dive into Reasoning Large Language ModelsLink
595Long-VITA: Scaling Large Vision-Language Models for Long Video UnderstandingLink
596Wan: Open and Advanced Large-Scale Video Generative ModelsLink
597Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning ModelsLink
598Seed1.5-VL Technical ReportLink
599A Survey on LLM-based Agents: Recent Advances and New FrontiersLink
600RLVR Is Not RL: Revisiting Reinforcement Learning for LLMsLink
601RL Tango: Reinforcing Generator and Verifier Together for Language ReasoningLink
602Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement LearningLink
603LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RLLink
604Reinforcement Learning for Reasoning in Large Language Models with One Training ExampleLink
605Leveraging Reasoning Model Answers to Enhance Non-Reasoning Model CapabilityLink
606The First Few Tokens Are All You Need: Unsupervised Prefix Fine-Tuning for Reasoning ModelsLink
607Learning to Reason without External RewardsLink
608Genius: A Generalizable and Purely Unsupervised Self-Training Framework for Advanced ReasoningLink
609Reinforcement Learning Teachers of Test Time ScalingLink
610Rewarding the Unlikely: Lifting GRPO Beyond Distribution SharpeningLink

Contributors

itsual

7 commits