anubhavshrimal/Machine-Learning-Research-Papers

A list of research papers in the domain of machine learning, deep learning and related fields.

216

27 commits

updated Mar 3, 2024

See the code

README

Machine-Learning-Research-Papers

A list of research papers in the domain of machine learning, deep learning and related fields.

I have curated a list of research papers that I come across and read. I'll keep on updating the list of papers and their summary as I read them every week.

How to read a Research Paper

Professor Andrew Ng gave some awesome tips on how to read a research paper. I have summarised the tips in this PDF.

Table of Contents

The list of papers can be viewed based on differentiating criteria's such as (Conference venue, Year Published, Topic Covered, Authors, etc.).

The following filtered formats are available to view paper's list:

All Papers

Paper NameStatusTopicCategoryYearConferenceAuthorSummaryLink
0ZF Net (Visualizing and Understanding Convolutional Networks)ReadCNNs, CV , ImageVisualization2014ECCVMatthew D. Zeiler, Rob FergusVisualize CNN Filters / Kernels using De-Convolutions on CNN filter activations.link
1Inception-v1 (Going Deeper With Convolutions)ReadCNNs, CV , ImageArchitecture2015CVPRChristian Szegedy, Wei LiuPropose the use of 1x1 conv operations to reduce the number of parameters in a deep and wide CNNlink
2ResNet (Deep Residual Learning for Image Recognition)ReadCNNs, CV , ImageArchitecture2016CVPRKaiming He, Xiangyu ZhangIntroduces Residual or Skip Connections to allow increase in the depth of a DNNlink
3MobileNet (Efficient Convolutional Neural Networks for Mobile Vision Applications)PendingCNNs, CV , ImageArchitecture, Optimization-No. of params2017arXivAndrew G. Howard, Menglong Zhulink
4Evaluation of neural network architectures for embedded systemsReadCNNs, CV , ImageComparison2017IEEE ISCASAdam Paszke, Alfredo Canziani, Eugenio CulurcielloCompare CNN classification architectures on accuracy, memory footprint, parameters, operations count, inference time and power consumption.link
5SqueezeNetReadCNNs, CV , ImageArchitecture, Optimization-No. of params2016arXivForrest N. Iandola, Song HanExplores model compression by using 1x1 convolutions called fire modules.link
6Pruning Filters for Efficient ConvNetsPendingCNNs, CV , ImageOptimization-No. of params2017arXivAsim Kadav, Hao Lilink
7Attention is All you NeedReadAttention, Text , TransformersArchitecture2017NIPSAshish Vaswani, Illia Polosukhin, Noam Shazeer, Łukasz KaiserTalks about Transformer architecture which brings SOTA performance for different tasks in NLPlink
8GPT-2 (Language Models are Unsupervised Multitask Learners)PendingAttention, Text , Transformers2019Alec Radford, Dario Amodei, Ilya Sutskever, Jeffrey Wulink
9BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingReadAttention, Text , TransformersEmbeddings2018NAACLJacob Devlin, Kenton Lee, Kristina Toutanova, Ming-Wei ChangBERT is an extension to Transformer based architecture which introduces a masked word pretraining and next sentence prediction task to pretrain the model for a wide variety of tasks.link
10SAGAN: Self-Attention Generative Adversarial NetworksPendingAttention, GANs, ImageArchitecture2018arXivAugustus Odena, Dimitris Metaxas, Han Zhang, Ian Goodfellowlink
11Single Headed Attention RNN: Stop Thinking With Your HeadPendingAttention, LSTMs, TextOptimization-No. of params2019arXivStephen Meritylink
12Reformer: The Efficient TransformerReadAttention, Text , TransformersArchitecture, Optimization-Memory, Optimization-No. of params2020arXivAnselm Levskaya, Lukasz Kaiser, Nikita KitaevOvercome time and memory complexity of Transformers by bucketing Query, Keys and using Reversible residual connections.link
13A 2019 guide to Human Pose Estimation with Deep LearningPendingCV , Pose EstimationComparison2019BlogSudharshan Chandra Babulink
14A Simple yet Effective Baseline for 3D Human Pose EstimationPendingCV , Pose Estimation2017ICCVJames J. Little, Javier Romero, Julieta Martinez, Rayat Hossainlink
15Bag of Tricks for Image Classification with Convolutional Neural NetworksReadCV , ImageOptimizations, Tips & Tricks2018arXivTong He, Zhi ZhangShows a dozen tricks (mixup, label smoothing, etc.) to improve CNN accuracy and training time.link
16Class-Balanced Loss Based on Effective Number of SamplesPendingLoss FunctionTips & Tricks2019CVPRMenglin Jia, Yin Cuilink
17Self-Normalizing Neural NetworksPendingActivation Function, TabularOptimizations, Tips & Tricks2017NIPSAndreas Mayr, Günter Klambauer, Thomas Unterthinerlink
18A Comprehensive Guide on Activation FunctionsThis weekActivation Function2020BlogYgor Rebouças Serpalink
19Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNetReadingCNNs, CV , Image2019arXivMatthias Bethge, Wieland Brendellink
20Breaking neural networks with adversarial attacksPendingCNNs, ImageAdversarial2019BlogAnant Jainlink
21The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural NetworksReadNN Initialization, NNsOptimization-No. of params, Tips & Tricks2019ICLRJonathan Frankle, Michael CarbinLottery ticket hypothesis: dense, randomly-initialized, feed-forward networks contain subnetworks (winning tickets) that—when trained in isolation— reach test accuracy comparable to the original network in a similar number of iterations.link
22All you need is a good initPendingNN InitializationTips & Tricks2015arXivDmytro Mishkin, Jiri Mataslink
23Pix2Pix: Image-to-Image Translation with Conditional Adversarial NetsReadGANs, Image2017CVPRAlexei A. Efros, Jun-Yan Zhu, Phillip Isola, Tinghui ZhouImage to image translation using Conditional GANs and dataset of image pairs from one domain to another.link
24CycleGAN: Unpaired Image-To-Image Translation Using Cycle-Consistent Adversarial NetworksPendingGANs, ImageArchitecture2017ICCVAlexei A. Efros, Jun-Yan Zhu, Phillip Isola, Taesung Parklink
25Language-Agnostic BERT Sentence EmbeddingReadAttention, Siamese Network, Text , TransformersEmbeddings2020arXivFangxiaoyu Feng, Yinfei YangA BERT model with multilingual sentence embeddings learned over 112 languages and Zero-shot learning over unseen languages.link
26Phrase-Based & Neural Unsupervised Machine TranslationPendingNMT, Text , TransformersUnsupervised2018arXivAlexis Conneau, Guillaume Lample, Ludovic Denoyer, Marc'Aurelio Ranzato, Myle Ottlink
27Unsupervised Machine Translation Using Monolingual Corpora OnlyPendingGANs, NMT, Text , TransformersUnsupervised2017arXivAlexis Conneau, Guillaume Lample, Ludovic Denoyer, Marc'Aurelio Ranzato, Myle Ottlink
28Cross-lingual Language Model PretrainingPendingNMT, Text , TransformersUnsupervised2019arXivAlexis Conneau, Guillaume Lamplelink
29Word2Vec: Efficient Estimation of Word Representations in Vector SpacePendingTextEmbeddings, Tips & Tricks2013arXivGreg Corrado, Jeffrey Dean, Kai Chen, Tomas Mikolovlink
30Capsule Networks: Dynamic Routing Between CapsulesPendingCV , ImageArchitecture2017arXivGeoffrey E Hinton, Nicholas Frosst, Sara Sabourlink
31Graph Neural Network: Relational inductive biases, deep learning, and graph networksPendingGraphNNArchitecture2018arXivJessica B. Hamrick, Oriol Vinyals, Peter W. Battaglialink
32Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsPendingCNNs, Image2020arXivAri S. Morcos, David J. Schwab, Jonathan Franklelink
33Arbitrary Style Transfer in Real-Time With Adaptive Instance NormalizationPendingCNNs, Image2017ICCVSerge Belongie, Xun Huanglink
34How Does Batch Normalization Help Optimization?PendingNNs, NormalizationOptimizations2018arXivAleksander Madry, Andrew Ilyas, Dimitris Tsipras, Shibani Santurkarlink
35WGAN: Wasserstein GANPendingGANs, Loss Function2017arXivLéon Bottou, Martin Arjovsky, Soumith Chintalalink
36Group NormalizationPendingNNs, NormalizationOptimizations2018arXivKaiming He, Yuxin Wulink
37Spectral Normalization for GANsPendingGANs, NormalizationOptimizations2018arXivMasanori Koyama, Takeru Miyato, Toshiki Kataoka, Yuichi Yoshidalink
38One-shot Text Field Labeling using Attention and Belief Propagation for Structure Information ExtractionPendingImage , Text2020arXivJun Huang, Mengli Cheng, Minghui Qiu, Wei Lin, Xing Shilink
39Perceptual Losses for Real-Time Style Transfer and Super-ResolutionPendingLoss Function, NNs2016ECCVAlexandre Alahi, Justin Johnson, Li Fei-Feilink
40Topological Loss: Beyond the Pixel-Wise Loss for Topology-Aware DelineationPendingImage , Loss Function, Segmentation2018CVPRAgata Mosinska, Mateusz Koziński, Pablo Márquez-Neila, Pascal Fualink
41Understanding Loss Functions in Computer VisionPendingCV , GANs, Image , Loss FunctionComparison, Tips & Tricks2020BlogSowmya Yellapragadalink
42NADAM: Incorporating Nesterov Momentum into AdamPendingNNs, OptimizersComparison2016Timothy Dozatlink
43Deep Double Descent: Where Bigger Models and More Data HurtPendingNNs2019arXivBoaz Barak, Gal Kaplun, Ilya Sutskever, Preetum Nakkiran, Tristan Yang, Yamini Bansallink
44StyleGAN: A Style-Based Generator Architecture for Generative Adversarial NetworksPendingGANs, Image2019CVPRSamuli Laine, Tero Karras, Timo Ailalink
45Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?PendingGANs, Image2019ICCVPeter Wonka, Rameen Abdal, Yipeng Qinlink
46Improved Techniques for Training GANsPendingGANs, ImageSemi-Supervised2016NIPSAlec Radford, Ian Goodfellow, Tim Salimans, Vicki Cheung, Wojciech Zaremba, Xi Chenlink
47AnimeGAN: Towards the Automatic Anime Characters Creation with Generative Adversarial NetworksPendingGANs, Image2017NIPSJiakai Zhang, Minjun Li, Yanghua Jinlink
48Progressive Growing of GANs for Improved Quality, Stability, and VariationPendingGANs, ImageTips & Tricks2018ICLRJaakko Lehtinen, Samuli Laine, Tero Karras, Timo Ailalink
49BEGAN: Boundary Equilibrium Generative Adversarial NetworksPendingGANs, Image2017arXivDavid Berthelot, Luke Metz, Thomas Schummlink
50Adam: A Method for Stochastic OptimizationPendingNNs, Optimizers2015ICLRDiederik P. Kingma, Jimmy Balink
51StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image TranslationPendingGANs, Image2018CVPRJaegul Choo, Jung-Woo Ha, Minje Choi, Munyoung Kim, Sunghun Kim, Yunjey Choilink
52IMLE-GAN: Inclusive GAN: Improving Data and Minority Coverage in Generative ModelsPendingGANs2020arXivJitendra Malik, Ke Li, Larry Davis, Mario Fritz, Ning Yu, Peng Zhoulink
53Few-Shot Learning with Localization in Realistic SettingsPendingCNNs, ImageFew-shot-learning2019CVPRBharath Hariharan, Davis Wertheimerlink
54Revisiting Pose-Normalization for Fine-Grained Few-Shot RecognitionPendingCNNs, ImageFew-shot-learning2020CVPRBharath Hariharan, Davis Wertheimer, Luming Tanglink
55ATOMIC: An Atlas of Machine Commonsense for If-Then ReasoningPendingAGI, Dataset, Text2019AAAIMaarten Sap, Noah A. Smith, Ronan Le Bras, Yejin Choilink
56COMET: Commonsense Transformers for Automatic Knowledge Graph ConstructionPendingAGI, Text , Transformers2019ACLAntoine Bosselut, Hannah Rashkin, Yejin Choilink
57VisualCOMET: Reasoning about the Dynamic Context of a Still ImagePendingAGI, Dataset, Image , Text , Transformers2020ECCVAli Farhadi, Chandra Bhagavatula, Jae Sung Park, Yejin Choilink
58Occupancy Anticipation for Efficient Exploration and NavigationPendingCNNs, ImageReinforcement-Learning2020ECCVKristen Grauman, Santhosh K. Ramakrishnan, Ziad Al-Halahlink
59T5: Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerReadAttention, Text , Transformers2020JMLRColin Raffel, Noam Shazeer, Peter J. Liu, Wei Liu, Yanqi ZhouPresents a Text-to-Text transformer model with multi-task learning capabilities, simultaneously solving problems such as machine translation, document summarization, question answering, and classification tasks.link
60GPT-f: Generative Language Modeling for Automated Theorem ProvingPendingAttention, Transformers2020arXivIlya Sutskever, Stanislas Polulink
61Vision Transformer: An Image is Worth 16x16 Words: Transformers for Image Recognition at ScalePendingAttention, Image , Transformers2021ICLRAlexey Dosovitskiy, Jakob Uszkoreit, Lucas Beyer, Neil Houlsbylink
62MuZero: Mastering Go, chess, shogi and Atari without rulesPendingReinforcement-Learning2020NatureDavid Silver, Demis Hassabis, Ioannis Antonoglou, Julian Schrittwieselink
63Deconstructing Lottery Tickets: Zeros, Signs, and the SupermaskReadNN Initialization, NNsComparison, Optimization-No. of params, Tips & Tricks2019NeurIPSHattie Zhou, Janice Lan, Jason Yosinski, Rosanne LiuFollow up on Lottery Ticket Hypothesis exploring the effects of different Masking criteria as well as Mask-1 and Mask-0 actions.link
64DALL·E: Creating Images from TextPendingImage , Text , Transformers2021BlogAditya Ramesh, Gabriel Goh, Ilya Sutskever, Mikhail Pavlov, Scott Graylink
65CLIP: Connecting Text and ImagesPendingImage , Text , TransformersMultimodal, Pre-Training2021arXivAlec Radford, Ilya Sutskever, Jong Wook Kimlink
66Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded SupervisionThis weekImage , Text , TransformersMultimodal2020EMNLPHao Tan, Mohit Bansallink
67SpanBERT: Improving Pre-training by Representing and Predicting SpansReadQuestion-Answering, Text , TransformersPre-Training2020TACLDanqi Chen, Mandar JoshiA different pre-training strategy for BERT model to improve performance for Question Answering task.link
68Learning to Extract Attribute Value from Product via Question Answering: A Multi-task ApproachReadQuestion-Answering, Text , TransformersZero-shot-learning2020KDDLi Yang, Qifan WangQuestion Answering BERT model used to extract attributes from products. Introduce further No Answer loss and distillation to promote zero shot learning.link
69TransGAN: Two Transformers Can Make One Strong GANPendingGANs, Image , TransformersArchitecture2021arXivShiyu Chang, Yifan Jiang, Zhangyang Wanglink
70Interpreting Deep Learning Models in Natural Language Processing: A ReviewPendingTextComparison, Visualization2021arXivDiyi Yang, Xiaofei Sunlink
71Symbolic Knowledge Distillation: from General Language Models to Commonsense ModelsPendingDataset, Text , TransformersOptimizations, Tips & Tricks2021arXivChandra Bhagavatula, Jack Hessel, Peter West, Yejin Choilink
72Chain of Thought Prompting Elicits Reasoning in Large Language ModelsPendingQuestion-Answering, Text , Transformers2022arXivDenny Zhou, Jason Wei, Xuezhi Wanglink
73Transforming Sequence Tagging Into A Seq2Seq TaskPendingGenerative, TextComparison, Tips & Tricks2022arXivIftekhar Naim, Karthik Raman, Krishna Srinivasanlink
74Large Language Models are Zero-Shot ReasonersPendingGenerative, Question-Answering, TextTips & Tricks, Zero-shot-learning2022arXivTakeshi Kojima, Yusuke Iwasawalink
75Flan-T5: Scaling Instruction-Finetuned Language ModelsPendingGenerative, Text , TransformersArchitecture, Pre-Training2022arXivHyung Won Chung, Le Houlink
76Decoding a Neural Retriever’s Latent Space for Query SuggestionPendingTextEmbeddings, Latent space2022arXivChristian Buck, Leonard Adolphs, Michelle Chen Huebscherlink
77Training Compute-Optimal Large Language ModelsPendingLarge-Language-Models, TransformersArchitecture, Optimization-No. of params, Pre-Training, Tips & Tricks2022arXivJordan Hoffmann, Laurent Sifre, Oriol Vinyals, Sebastian Borgeaudlink
78VL-T5: Unifying Vision-and-Language Tasks via Text GenerationReadCNNs, CV , Generative, Image , Large-Language-Models, Question-Answering, Text , TransformersArchitecture, Embeddings, Multimodal, Pre-Training2021arXivHao Tan, Jaemin Cho, Jie Le, Mohit BansalUnifying two modalities (image and text) together in a single transformer model to solve multiple tasks in a single architecture using text prefixes similar to T5.link
79Scaling Instruction-Finetuned Language Models (FLAN)PendingGenerative, Large-Language-Models, Question-Answering, Text , TransformersInstruction-Finetuning2022arXivHyung Won Chung, Jason Wei, Jeffrey Dean, Le Hou, Quoc V. Le, Shayne Longprehttps://arxiv.org/abs/2210.11416 introduces FLAN (Fine-tuned LAnguage Net), an instruction finetuning method, and presents the results of its application. The study demonstrates that by fine-tuning the 540B PaLM model on 1836 tasks while incorporating Chain-of-Thought Reasoning data, FLAN achieves improvements in generalization, human usability, and zero-shot reasoning over the base model. The paper also provides detailed information on how each these aspects was evaluated.link
80ReAct: Synergizing Reasoning and Acting in Language ModelsPendingGenerative, Large-Language-Models, TextOptimizations, Tips & Tricks2023ICLRDian Yu, Izhak Shafran, Jeffrey Zhao, Karthik Narasimhan, Nan Du, Shunyu Yao, Yuan CaoThis paper introduces ReAct, a novel approach that leverages Large Language Models (LLMs) to interleave reasoning traces and task-specific actions. ReAct outperforms existing methods on various language and decision-making tasks, addressing issues like hallucination, error propagation, and improving human interpretability and trustworthiness.link
81Training language models to follow instructions with human feedbackPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning, Reinforcement-Learning, Semi-Supervised2022arXivCarroll L. Wainwright, Diogo Almeida, Jan Leike, Jeff Wu, Long Ouyang, Pamela Mishkin, Paul Christiano, Ryan Lowe, Xu JiangThis paper presents InstructGPT, a model fine-tuned with human feedback to better align with user intent across various tasks. Despite having significantly fewer parameters than larger models, InstructGPT outperforms them in human evaluations, demonstrating improved truthfulness, reduced toxicity, and minimal performance regressions on public NLP datasets, highlighting the potential of fine-tuning with human feedback for enhancing language model alignment with human intent.link
82Constitutional AI: Harmlessness from AI FeedbackPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning, Reinforcement-Learning, Unsupervised2022arXivJared Kaplan, Yuntao BaThe paper introduces Constitutional AI, a method for training a safe AI assistant without human-labeled data on harmful outputs. It combines supervised learning and reinforcement learning phases, enabling the AI to engage with harmful queries by explaining its objections, thus improving control, transparency, and human-judged performance with minimal human oversight.link
83Self-Alignment with Instruction BacktranslationPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning2023arXivJason Weston, Mike Lewis, Ping Yu, Xian LiThe paper introduces a scalable method called "instruction backtranslation" to create a high-quality instruction-following language model. This method involves self-augmentation and self-curation of training examples generated from web documents, resulting in a model that outperforms others in its category without relying on distillation data, showcasing its effective self-alignment capability.link
84Table-GPT: Table-tuned GPT for Diverse Table TasksPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning2023arXivlink
85Code Generation with AlphaCodium: From Prompt Engineering to Flow EngineeringPendingLarge-Language-ModelsPrompting, Tips & Tricks2024arXivDedy Kredo, Itamar Friedman, Tal RidnikThis paper introduces AlphaCodium, a novel test-based, multi-stage, code-oriented iterative approach for improving the performance of Language Model Models (LLMs) on code generation tasks.link
86Large Language Models for Data Annotation: A SurveyThis weekDataset, Generative, Large-Language-ModelsPrompting, Tips & Tricks2024arXivAlimohammad Beigi, Zhen Tanlink
awesome-dl
awesome-list
conference-paper
deep-learning
machine-learning
research-paper

Contributors

anubhavshrimal

27 commits

anubhavshrimal/Machine-Learning-Research-Papers

A list of research papers in the domain of machine learning, deep learning and related fields.

216

27 commits

updated Mar 3, 2024

See the code

README

Machine-Learning-Research-Papers

A list of research papers in the domain of machine learning, deep learning and related fields.

I have curated a list of research papers that I come across and read. I'll keep on updating the list of papers and their summary as I read them every week.

How to read a Research Paper

Professor Andrew Ng gave some awesome tips on how to read a research paper. I have summarised the tips in this PDF.

Table of Contents

The list of papers can be viewed based on differentiating criteria's such as (Conference venue, Year Published, Topic Covered, Authors, etc.).

The following filtered formats are available to view paper's list:

All Papers

Paper NameStatusTopicCategoryYearConferenceAuthorSummaryLink
0ZF Net (Visualizing and Understanding Convolutional Networks)ReadCNNs, CV , ImageVisualization2014ECCVMatthew D. Zeiler, Rob FergusVisualize CNN Filters / Kernels using De-Convolutions on CNN filter activations.link
1Inception-v1 (Going Deeper With Convolutions)ReadCNNs, CV , ImageArchitecture2015CVPRChristian Szegedy, Wei LiuPropose the use of 1x1 conv operations to reduce the number of parameters in a deep and wide CNNlink
2ResNet (Deep Residual Learning for Image Recognition)ReadCNNs, CV , ImageArchitecture2016CVPRKaiming He, Xiangyu ZhangIntroduces Residual or Skip Connections to allow increase in the depth of a DNNlink
3MobileNet (Efficient Convolutional Neural Networks for Mobile Vision Applications)PendingCNNs, CV , ImageArchitecture, Optimization-No. of params2017arXivAndrew G. Howard, Menglong Zhulink
4Evaluation of neural network architectures for embedded systemsReadCNNs, CV , ImageComparison2017IEEE ISCASAdam Paszke, Alfredo Canziani, Eugenio CulurcielloCompare CNN classification architectures on accuracy, memory footprint, parameters, operations count, inference time and power consumption.link
5SqueezeNetReadCNNs, CV , ImageArchitecture, Optimization-No. of params2016arXivForrest N. Iandola, Song HanExplores model compression by using 1x1 convolutions called fire modules.link
6Pruning Filters for Efficient ConvNetsPendingCNNs, CV , ImageOptimization-No. of params2017arXivAsim Kadav, Hao Lilink
7Attention is All you NeedReadAttention, Text , TransformersArchitecture2017NIPSAshish Vaswani, Illia Polosukhin, Noam Shazeer, Łukasz KaiserTalks about Transformer architecture which brings SOTA performance for different tasks in NLPlink
8GPT-2 (Language Models are Unsupervised Multitask Learners)PendingAttention, Text , Transformers2019Alec Radford, Dario Amodei, Ilya Sutskever, Jeffrey Wulink
9BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingReadAttention, Text , TransformersEmbeddings2018NAACLJacob Devlin, Kenton Lee, Kristina Toutanova, Ming-Wei ChangBERT is an extension to Transformer based architecture which introduces a masked word pretraining and next sentence prediction task to pretrain the model for a wide variety of tasks.link
10SAGAN: Self-Attention Generative Adversarial NetworksPendingAttention, GANs, ImageArchitecture2018arXivAugustus Odena, Dimitris Metaxas, Han Zhang, Ian Goodfellowlink
11Single Headed Attention RNN: Stop Thinking With Your HeadPendingAttention, LSTMs, TextOptimization-No. of params2019arXivStephen Meritylink
12Reformer: The Efficient TransformerReadAttention, Text , TransformersArchitecture, Optimization-Memory, Optimization-No. of params2020arXivAnselm Levskaya, Lukasz Kaiser, Nikita KitaevOvercome time and memory complexity of Transformers by bucketing Query, Keys and using Reversible residual connections.link
13A 2019 guide to Human Pose Estimation with Deep LearningPendingCV , Pose EstimationComparison2019BlogSudharshan Chandra Babulink
14A Simple yet Effective Baseline for 3D Human Pose EstimationPendingCV , Pose Estimation2017ICCVJames J. Little, Javier Romero, Julieta Martinez, Rayat Hossainlink
15Bag of Tricks for Image Classification with Convolutional Neural NetworksReadCV , ImageOptimizations, Tips & Tricks2018arXivTong He, Zhi ZhangShows a dozen tricks (mixup, label smoothing, etc.) to improve CNN accuracy and training time.link
16Class-Balanced Loss Based on Effective Number of SamplesPendingLoss FunctionTips & Tricks2019CVPRMenglin Jia, Yin Cuilink
17Self-Normalizing Neural NetworksPendingActivation Function, TabularOptimizations, Tips & Tricks2017NIPSAndreas Mayr, Günter Klambauer, Thomas Unterthinerlink
18A Comprehensive Guide on Activation FunctionsThis weekActivation Function2020BlogYgor Rebouças Serpalink
19Approximating CNNs with Bag-of-local-Features models works surprisingly well on ImageNetReadingCNNs, CV , Image2019arXivMatthias Bethge, Wieland Brendellink
20Breaking neural networks with adversarial attacksPendingCNNs, ImageAdversarial2019BlogAnant Jainlink
21The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural NetworksReadNN Initialization, NNsOptimization-No. of params, Tips & Tricks2019ICLRJonathan Frankle, Michael CarbinLottery ticket hypothesis: dense, randomly-initialized, feed-forward networks contain subnetworks (winning tickets) that—when trained in isolation— reach test accuracy comparable to the original network in a similar number of iterations.link
22All you need is a good initPendingNN InitializationTips & Tricks2015arXivDmytro Mishkin, Jiri Mataslink
23Pix2Pix: Image-to-Image Translation with Conditional Adversarial NetsReadGANs, Image2017CVPRAlexei A. Efros, Jun-Yan Zhu, Phillip Isola, Tinghui ZhouImage to image translation using Conditional GANs and dataset of image pairs from one domain to another.link
24CycleGAN: Unpaired Image-To-Image Translation Using Cycle-Consistent Adversarial NetworksPendingGANs, ImageArchitecture2017ICCVAlexei A. Efros, Jun-Yan Zhu, Phillip Isola, Taesung Parklink
25Language-Agnostic BERT Sentence EmbeddingReadAttention, Siamese Network, Text , TransformersEmbeddings2020arXivFangxiaoyu Feng, Yinfei YangA BERT model with multilingual sentence embeddings learned over 112 languages and Zero-shot learning over unseen languages.link
26Phrase-Based & Neural Unsupervised Machine TranslationPendingNMT, Text , TransformersUnsupervised2018arXivAlexis Conneau, Guillaume Lample, Ludovic Denoyer, Marc'Aurelio Ranzato, Myle Ottlink
27Unsupervised Machine Translation Using Monolingual Corpora OnlyPendingGANs, NMT, Text , TransformersUnsupervised2017arXivAlexis Conneau, Guillaume Lample, Ludovic Denoyer, Marc'Aurelio Ranzato, Myle Ottlink
28Cross-lingual Language Model PretrainingPendingNMT, Text , TransformersUnsupervised2019arXivAlexis Conneau, Guillaume Lamplelink
29Word2Vec: Efficient Estimation of Word Representations in Vector SpacePendingTextEmbeddings, Tips & Tricks2013arXivGreg Corrado, Jeffrey Dean, Kai Chen, Tomas Mikolovlink
30Capsule Networks: Dynamic Routing Between CapsulesPendingCV , ImageArchitecture2017arXivGeoffrey E Hinton, Nicholas Frosst, Sara Sabourlink
31Graph Neural Network: Relational inductive biases, deep learning, and graph networksPendingGraphNNArchitecture2018arXivJessica B. Hamrick, Oriol Vinyals, Peter W. Battaglialink
32Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNsPendingCNNs, Image2020arXivAri S. Morcos, David J. Schwab, Jonathan Franklelink
33Arbitrary Style Transfer in Real-Time With Adaptive Instance NormalizationPendingCNNs, Image2017ICCVSerge Belongie, Xun Huanglink
34How Does Batch Normalization Help Optimization?PendingNNs, NormalizationOptimizations2018arXivAleksander Madry, Andrew Ilyas, Dimitris Tsipras, Shibani Santurkarlink
35WGAN: Wasserstein GANPendingGANs, Loss Function2017arXivLéon Bottou, Martin Arjovsky, Soumith Chintalalink
36Group NormalizationPendingNNs, NormalizationOptimizations2018arXivKaiming He, Yuxin Wulink
37Spectral Normalization for GANsPendingGANs, NormalizationOptimizations2018arXivMasanori Koyama, Takeru Miyato, Toshiki Kataoka, Yuichi Yoshidalink
38One-shot Text Field Labeling using Attention and Belief Propagation for Structure Information ExtractionPendingImage , Text2020arXivJun Huang, Mengli Cheng, Minghui Qiu, Wei Lin, Xing Shilink
39Perceptual Losses for Real-Time Style Transfer and Super-ResolutionPendingLoss Function, NNs2016ECCVAlexandre Alahi, Justin Johnson, Li Fei-Feilink
40Topological Loss: Beyond the Pixel-Wise Loss for Topology-Aware DelineationPendingImage , Loss Function, Segmentation2018CVPRAgata Mosinska, Mateusz Koziński, Pablo Márquez-Neila, Pascal Fualink
41Understanding Loss Functions in Computer VisionPendingCV , GANs, Image , Loss FunctionComparison, Tips & Tricks2020BlogSowmya Yellapragadalink
42NADAM: Incorporating Nesterov Momentum into AdamPendingNNs, OptimizersComparison2016Timothy Dozatlink
43Deep Double Descent: Where Bigger Models and More Data HurtPendingNNs2019arXivBoaz Barak, Gal Kaplun, Ilya Sutskever, Preetum Nakkiran, Tristan Yang, Yamini Bansallink
44StyleGAN: A Style-Based Generator Architecture for Generative Adversarial NetworksPendingGANs, Image2019CVPRSamuli Laine, Tero Karras, Timo Ailalink
45Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space?PendingGANs, Image2019ICCVPeter Wonka, Rameen Abdal, Yipeng Qinlink
46Improved Techniques for Training GANsPendingGANs, ImageSemi-Supervised2016NIPSAlec Radford, Ian Goodfellow, Tim Salimans, Vicki Cheung, Wojciech Zaremba, Xi Chenlink
47AnimeGAN: Towards the Automatic Anime Characters Creation with Generative Adversarial NetworksPendingGANs, Image2017NIPSJiakai Zhang, Minjun Li, Yanghua Jinlink
48Progressive Growing of GANs for Improved Quality, Stability, and VariationPendingGANs, ImageTips & Tricks2018ICLRJaakko Lehtinen, Samuli Laine, Tero Karras, Timo Ailalink
49BEGAN: Boundary Equilibrium Generative Adversarial NetworksPendingGANs, Image2017arXivDavid Berthelot, Luke Metz, Thomas Schummlink
50Adam: A Method for Stochastic OptimizationPendingNNs, Optimizers2015ICLRDiederik P. Kingma, Jimmy Balink
51StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image TranslationPendingGANs, Image2018CVPRJaegul Choo, Jung-Woo Ha, Minje Choi, Munyoung Kim, Sunghun Kim, Yunjey Choilink
52IMLE-GAN: Inclusive GAN: Improving Data and Minority Coverage in Generative ModelsPendingGANs2020arXivJitendra Malik, Ke Li, Larry Davis, Mario Fritz, Ning Yu, Peng Zhoulink
53Few-Shot Learning with Localization in Realistic SettingsPendingCNNs, ImageFew-shot-learning2019CVPRBharath Hariharan, Davis Wertheimerlink
54Revisiting Pose-Normalization for Fine-Grained Few-Shot RecognitionPendingCNNs, ImageFew-shot-learning2020CVPRBharath Hariharan, Davis Wertheimer, Luming Tanglink
55ATOMIC: An Atlas of Machine Commonsense for If-Then ReasoningPendingAGI, Dataset, Text2019AAAIMaarten Sap, Noah A. Smith, Ronan Le Bras, Yejin Choilink
56COMET: Commonsense Transformers for Automatic Knowledge Graph ConstructionPendingAGI, Text , Transformers2019ACLAntoine Bosselut, Hannah Rashkin, Yejin Choilink
57VisualCOMET: Reasoning about the Dynamic Context of a Still ImagePendingAGI, Dataset, Image , Text , Transformers2020ECCVAli Farhadi, Chandra Bhagavatula, Jae Sung Park, Yejin Choilink
58Occupancy Anticipation for Efficient Exploration and NavigationPendingCNNs, ImageReinforcement-Learning2020ECCVKristen Grauman, Santhosh K. Ramakrishnan, Ziad Al-Halahlink
59T5: Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerReadAttention, Text , Transformers2020JMLRColin Raffel, Noam Shazeer, Peter J. Liu, Wei Liu, Yanqi ZhouPresents a Text-to-Text transformer model with multi-task learning capabilities, simultaneously solving problems such as machine translation, document summarization, question answering, and classification tasks.link
60GPT-f: Generative Language Modeling for Automated Theorem ProvingPendingAttention, Transformers2020arXivIlya Sutskever, Stanislas Polulink
61Vision Transformer: An Image is Worth 16x16 Words: Transformers for Image Recognition at ScalePendingAttention, Image , Transformers2021ICLRAlexey Dosovitskiy, Jakob Uszkoreit, Lucas Beyer, Neil Houlsbylink
62MuZero: Mastering Go, chess, shogi and Atari without rulesPendingReinforcement-Learning2020NatureDavid Silver, Demis Hassabis, Ioannis Antonoglou, Julian Schrittwieselink
63Deconstructing Lottery Tickets: Zeros, Signs, and the SupermaskReadNN Initialization, NNsComparison, Optimization-No. of params, Tips & Tricks2019NeurIPSHattie Zhou, Janice Lan, Jason Yosinski, Rosanne LiuFollow up on Lottery Ticket Hypothesis exploring the effects of different Masking criteria as well as Mask-1 and Mask-0 actions.link
64DALL·E: Creating Images from TextPendingImage , Text , Transformers2021BlogAditya Ramesh, Gabriel Goh, Ilya Sutskever, Mikhail Pavlov, Scott Graylink
65CLIP: Connecting Text and ImagesPendingImage , Text , TransformersMultimodal, Pre-Training2021arXivAlec Radford, Ilya Sutskever, Jong Wook Kimlink
66Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded SupervisionThis weekImage , Text , TransformersMultimodal2020EMNLPHao Tan, Mohit Bansallink
67SpanBERT: Improving Pre-training by Representing and Predicting SpansReadQuestion-Answering, Text , TransformersPre-Training2020TACLDanqi Chen, Mandar JoshiA different pre-training strategy for BERT model to improve performance for Question Answering task.link
68Learning to Extract Attribute Value from Product via Question Answering: A Multi-task ApproachReadQuestion-Answering, Text , TransformersZero-shot-learning2020KDDLi Yang, Qifan WangQuestion Answering BERT model used to extract attributes from products. Introduce further No Answer loss and distillation to promote zero shot learning.link
69TransGAN: Two Transformers Can Make One Strong GANPendingGANs, Image , TransformersArchitecture2021arXivShiyu Chang, Yifan Jiang, Zhangyang Wanglink
70Interpreting Deep Learning Models in Natural Language Processing: A ReviewPendingTextComparison, Visualization2021arXivDiyi Yang, Xiaofei Sunlink
71Symbolic Knowledge Distillation: from General Language Models to Commonsense ModelsPendingDataset, Text , TransformersOptimizations, Tips & Tricks2021arXivChandra Bhagavatula, Jack Hessel, Peter West, Yejin Choilink
72Chain of Thought Prompting Elicits Reasoning in Large Language ModelsPendingQuestion-Answering, Text , Transformers2022arXivDenny Zhou, Jason Wei, Xuezhi Wanglink
73Transforming Sequence Tagging Into A Seq2Seq TaskPendingGenerative, TextComparison, Tips & Tricks2022arXivIftekhar Naim, Karthik Raman, Krishna Srinivasanlink
74Large Language Models are Zero-Shot ReasonersPendingGenerative, Question-Answering, TextTips & Tricks, Zero-shot-learning2022arXivTakeshi Kojima, Yusuke Iwasawalink
75Flan-T5: Scaling Instruction-Finetuned Language ModelsPendingGenerative, Text , TransformersArchitecture, Pre-Training2022arXivHyung Won Chung, Le Houlink
76Decoding a Neural Retriever’s Latent Space for Query SuggestionPendingTextEmbeddings, Latent space2022arXivChristian Buck, Leonard Adolphs, Michelle Chen Huebscherlink
77Training Compute-Optimal Large Language ModelsPendingLarge-Language-Models, TransformersArchitecture, Optimization-No. of params, Pre-Training, Tips & Tricks2022arXivJordan Hoffmann, Laurent Sifre, Oriol Vinyals, Sebastian Borgeaudlink
78VL-T5: Unifying Vision-and-Language Tasks via Text GenerationReadCNNs, CV , Generative, Image , Large-Language-Models, Question-Answering, Text , TransformersArchitecture, Embeddings, Multimodal, Pre-Training2021arXivHao Tan, Jaemin Cho, Jie Le, Mohit BansalUnifying two modalities (image and text) together in a single transformer model to solve multiple tasks in a single architecture using text prefixes similar to T5.link
79Scaling Instruction-Finetuned Language Models (FLAN)PendingGenerative, Large-Language-Models, Question-Answering, Text , TransformersInstruction-Finetuning2022arXivHyung Won Chung, Jason Wei, Jeffrey Dean, Le Hou, Quoc V. Le, Shayne Longprehttps://arxiv.org/abs/2210.11416 introduces FLAN (Fine-tuned LAnguage Net), an instruction finetuning method, and presents the results of its application. The study demonstrates that by fine-tuning the 540B PaLM model on 1836 tasks while incorporating Chain-of-Thought Reasoning data, FLAN achieves improvements in generalization, human usability, and zero-shot reasoning over the base model. The paper also provides detailed information on how each these aspects was evaluated.link
80ReAct: Synergizing Reasoning and Acting in Language ModelsPendingGenerative, Large-Language-Models, TextOptimizations, Tips & Tricks2023ICLRDian Yu, Izhak Shafran, Jeffrey Zhao, Karthik Narasimhan, Nan Du, Shunyu Yao, Yuan CaoThis paper introduces ReAct, a novel approach that leverages Large Language Models (LLMs) to interleave reasoning traces and task-specific actions. ReAct outperforms existing methods on various language and decision-making tasks, addressing issues like hallucination, error propagation, and improving human interpretability and trustworthiness.link
81Training language models to follow instructions with human feedbackPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning, Reinforcement-Learning, Semi-Supervised2022arXivCarroll L. Wainwright, Diogo Almeida, Jan Leike, Jeff Wu, Long Ouyang, Pamela Mishkin, Paul Christiano, Ryan Lowe, Xu JiangThis paper presents InstructGPT, a model fine-tuned with human feedback to better align with user intent across various tasks. Despite having significantly fewer parameters than larger models, InstructGPT outperforms them in human evaluations, demonstrating improved truthfulness, reduced toxicity, and minimal performance regressions on public NLP datasets, highlighting the potential of fine-tuning with human feedback for enhancing language model alignment with human intent.link
82Constitutional AI: Harmlessness from AI FeedbackPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning, Reinforcement-Learning, Unsupervised2022arXivJared Kaplan, Yuntao BaThe paper introduces Constitutional AI, a method for training a safe AI assistant without human-labeled data on harmful outputs. It combines supervised learning and reinforcement learning phases, enabling the AI to engage with harmful queries by explaining its objections, thus improving control, transparency, and human-judged performance with minimal human oversight.link
83Self-Alignment with Instruction BacktranslationPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning2023arXivJason Weston, Mike Lewis, Ping Yu, Xian LiThe paper introduces a scalable method called "instruction backtranslation" to create a high-quality instruction-following language model. This method involves self-augmentation and self-curation of training examples generated from web documents, resulting in a model that outperforms others in its category without relying on distillation data, showcasing its effective self-alignment capability.link
84Table-GPT: Table-tuned GPT for Diverse Table TasksPendingGenerative, Large-Language-Models, Training MethodInstruction-Finetuning2023arXivlink
85Code Generation with AlphaCodium: From Prompt Engineering to Flow EngineeringPendingLarge-Language-ModelsPrompting, Tips & Tricks2024arXivDedy Kredo, Itamar Friedman, Tal RidnikThis paper introduces AlphaCodium, a novel test-based, multi-stage, code-oriented iterative approach for improving the performance of Language Model Models (LLMs) on code generation tasks.link
86Large Language Models for Data Annotation: A SurveyThis weekDataset, Generative, Large-Language-ModelsPrompting, Tips & Tricks2024arXivAlimohammad Beigi, Zhen Tanlink
awesome-dl
awesome-list
conference-paper
deep-learning
machine-learning
research-paper

Contributors

anubhavshrimal

27 commits