pmiller10/cambridge-ai

58

283 commits

updated Aug 26, 2024

See the code

README

SIPB Deep Learning Group

The schedule of readings for the SIPB/Cambridge AI Deep Learning Group If you have any papers you'd like to discuss, please either make a pull request, or send an email to the group and we'll add it. Papers with implementations available are strongly preferred.

Suggested Papers:

Schedule:

DatePaperImplementation
8.26.24The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoverySakanaAI/AI-Scientist
8.19.24Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetSonyResearch/micro_diffusion (pending)
8.12.24Scaling and evaluating sparse autoencodersopenai/sparse_autoencoder
4.10.24Score-Based Generative Modeling through Stochastic Differential Equations
3.28.24Generative Modeling by Estimating Gradients of the Data Distribution
3.21.24Humanoid Locomotion as Next Token Prediction
3.14.24TIES-Merging: Resolving Interference When Merging Modelsprateeky2806/ties-merging
2.8.24Merging Models with Fisher-Weighted Averagingarcee-ai/mergekit
2.1.24Averaging Weights Leads to Wider Optima and Better Generalization
1.18.24Hyena Hierarchy: Towards Larger Convolutional Language Models
1.04.24Mamba: Linear-Time Sequence Modeling with Selective State Spacesstate-spaces/mamba
12.07.23Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
11.30.233D Gaussian Splatting for Real-Time Radiance Field Renderinggraphdeco-inria/gaussian-splatting
11.16.23LILO: Learning Interpretable Libraries by Compressing and Documenting Codegabegrand/lilo
11.09.23Human-like systematic generalization through a meta-learning neural networkbrendenlake/MLC and brendenlake/MLC-ML
9.28.23Retrieval-Augmented Generation for Knowledge-Intensive NLP Taskshuggingface/transformers/examples/research_projects/rag
9.14.23Gradient-based Adversarial Attacks against Text Transformersfacebookresearch/text-adversarial-attack
8.10.23Reflexion: Language Agents with Verbal Reinforcement Learningnoahshinn024/reflexion
6.15.23RWKV: Reinventing RNNs for the Transformer EraBlinkDL/RWKV-LM
5.18.23Toy Models of Superposition
5.11.23LoRA: Low-Rank Adaptation of Large Language Modelstloen/alpaca-lora and huggingface/blog/lora
5.04.23Efficiently Modeling Long Sequences with Structured State SpacesHazyResearch/state-spaces
4.06.23Generating Sequences by Learning to Self-Correct
3.30.23The Capacity for Moral Self-Correction in Large Language Models
3.23.23LLaMA: Open and Efficient Foundation Language Modelsfacebookresearch/llama and huggingface/llama
3.16.23Language Is Not All You Need: Aligning Perception with Language Models
3.02.23Guiding Pretraining in Reinforcement Learning with Large Language Models
2.23.23Toolformer: Language Models Can Teach Themselves to Use Tools
2.16.23What learning algorithm is in-context learning? Investigations with linear modelsekinakyurek/incontext
2.09.23Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationTutorial
1.26.23Mastering Diverse Domains through World Models
1.12.23The Forward-Forward Algorithm: Some Preliminary Investigations
12.08.22Training language models to follow instructions with human feedback
9.22.22Git Re-Basin: Merging Models modulo Permutation Symmetries
9.08.22Transformers are Sample-Efficient World Models
8.25.22A Path Towards Autonomous Machine Intelligence
8.18.22Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfermicrosoft/mup
7.14.22Learning Iterative Reasoning through Energy Minimizationyilundu/irem_code_release
6.16.22Sharpness-Aware Minimization for Efficiently Improving Generalizationgoogle-research/sam
5.26.22Neural Tangent Kernel: Convergence and Generalization in Neural Networks
4.28.22A Modern Self-Referential Weight Matrix That Learns to Modify ItselfIDSIA/modern-srwm
4.14.22Hierarchical Perceiver
3.24.22Dual Diffusion Implicit Bridges for Image-to-Image Translation
3.10.22Understanding Generalization through Visualizationswronnyhuang/gen-viz
2.17.22Divide and Contrast: Self-supervised Learning from Uncurated Data
2.10.22Investigating Human Priors for Playing Video Gamesrach0012/humanRL_prior_games
1.27.22data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Languagepytorch/data2vec
1.20.22Consistent Video Depth Estimationfacebookresearch/consistent_depth
1.13.22Masked Autoencoders Are Scalable Vision Learners
12.02.21Training Verifiers to Solve Math Word Problems
11.18.21(StyleGan3) Alias-Free Generative Adversarial NetworksNVlabs/stylegan3
11.04.21Do Vision Transformers See Like Convolutional Neural Networks?
10.21.21CoBERL: Contrastive BERT for Reinforcement Learning
10.14.21WarpedGANSpace: Finding non-linear RBF paths in GAN latent spacechi0tzp/WarpedGANSpace
10.06.21RAFT: Recurrent All-Pairs Field Transforms for Optical Flowprinceton-vl/RAFT
9.16.21Bootstrapped Meta-Learning
9.09.21Program Synthesis with Large Language Models
8.19.21Perceiver IO: A General Architecture for Structured Inputs & Outputsdeepmind/perceiver
8.12.21Reward is enough
8.05.21Learning Compositional Rules via Neural Program Synthesismtensor/rulesynthesis
6.24.21Thinking Like Transformers
6.17.21Equilibrium Propagation: Bridging the Gap between Energy-Based Models and Backpropagation
6.10.21Unsupervised Learning by Competing Hidden Units
5.27.21Pay Attention to MLPs
5.20.21Memory Based Trajectory-conditioned Policies for Learning from Sparse Rewards
5.13.21Emerging Properties in Self-Supervised Vision Transformers
5.06.21Implicit Neural Representations with Periodic Activation Functionsvsitzmann/siren
4.29.21How to represent part-whole hierarchies in a neural networklucidrains/glom-pytorch RedRyan111/GLOM ArneBinder/GlomImpl
4.15.21Perceiver: General Perception with Iterative Attention
4.01.21Synthetic Returns for Long-Term Credit Assignment
3.25.21The Pitfalls of Simplicity Bias in Neural Networks
3.18.21Bootstrap your own latent: A new approach to self-supervised Learning
3.11.21Meta Learning Backpropagation And Improving It
3.04.21Taming Transformers for High-Resolution Image SynthesisCompVis/taming-transformers
2.18.21Pre-training without Natural Imageshirokatsukataoka16/FractalDB-Pretrained-ResNet-PyTorch
2.11.21Revisiting Locally Supervised Learning: an Alternative to End-to-end Trainingblackfeather-wang/InfoPro-Pytorch
2.04.21Neural Power Units
1.28.21Representation Learning via Invariant Causal Mechanisms
1.21.21γ-Models: Generative Temporal Difference Learning for Infinite-Horizon PredictionJannerM/gamma-models
1.14.21Improving Generalisation for Temporal Difference Learning: The Successor Representation
12.17.20Learning Associative Inference Using Fast Weight Memory
Hopfield Networks cycle ends
12.10.20Hopfield Networks is All You Needml-jku/hopfield-layers
12.03.20On a model of associative memory with huge storage capacity
11.19.20Dense Associative Memory for Pattern Recognition
11.12.20Neural Networks and Physical Systems with Emergent Collective Computational Abilities (= "the Hopfield Networks paper")
Hopfield Networks cycle of papers - from the original paper on Hopfield networks to "Hopfield Networks is All You Need"
11.05.20Training Generative Adversarial Networks with Limited DataNVlabs/stylegan2-ada
10.29.20Memories from patterns: Attractor and integrator networks in the brain
10.15.20Entities as Experts: Sparse Memory Access with Entity Supervision
10.08.20A Primer in BERTology: What we know about how BERT works
10.01.20It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learnerstimoschick/pet
9.24.20End-to-End Object Detection with Transformersfacebookresearch/detr
9.17.20Gated Linear Networks
7.23.20A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
7.02.20DreamCoder: Building interpretable hierarchical knowledge representations with wake-sleep Bayesian program learningellisk42/ec
6.18.20SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solverlocuslab/SATNet
6.4.20Adaptive Attention Span in Transformers
5.28.20Complexity control by gradient descent in deep networks
5.21.20What Can Learned Intrinsic Rewards Capture?
5.14.20COMET: Commonsense Transformers for Automatic Knowledge Graph Construction
5.7.20Write, Execute, Assess: Program Synthesis With a REPLflxsosa/ProgramSearch
4.23.20Graph Representations for Higher-Order Logic and Theorem Proving
4.16.20Mathematical Reasoning in Latent Space
4.9.20MEMO: A Deep Network for Flexible Combination of Episodic Memories
4.2.20Creating High Resolution Images with a Latent Adversarial Generator
3.26.20Invertible Residual Networks
3.5.20Value-driven Hindsight Modelling
2.27.20Analyzing and Improving the Image Quality of StyleGAN
2.13.20Axiomatic Attribution for Deep Networks
2.6.20Automated curricula through setter-solver interactions
1.30.20Protein structure prediction ...deepmind
1.23.20Putting An End to End-to-End: Gradient-Isolated Learning of Representations
1.16.20Normalizing Flows: An Introduction and Review of Current Methods
12.19.19Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
12.5.19On the Measure of Intelligence
11.21.19Understanding the Neural Tangent Kernelrajatvd
11.14.19XLNet: Generalized Autoregressive Pretraining for Language Understanding
11.7.19Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
10.31.19Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
10.24.19N-BEATS: Neural basis expansion analysis for interpretable time series forecasting
10.17.19Unsupervised Doodling and Painting with Improved SPIRAL
10.10.19Adversarial Robustness as a Prior for Learned RepresentationsMadryLab
10.3.19Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
9.26.19Image Transformer
9.19.19Generating Diverse High-Fidelity Images with VQ-VAE-2
9.12.19Neural Discrete Representation Learning
9.5.19Neural Text Generation with Unlikelihood Training
8.29.19Learning Representations by Maximizing Mutual Information Across Views
breakswitch from Tuesdays to Thursdays after the break
6.11.19BERT Rediscovers the Classical NLP Pipeline
6.4.19Semantic Visual Localization
5.28.19AlgoNet: C^∞ Smooth Algorithmic Neural Networks
5.14.19Unsupervised Data Augmentation for Consistency Training
4.30.19Augmented Neural ODEs
4.9.19Wasserstein Dependency Measure for Representation Learning
4.2.19Leveraging Knowledge Bases in LSTMs for Improving Machine Reading
3.26.19Meta Particle Flow for Sequential Bayesian Inference
3.19.19A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms
3.12.19The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
2.26.19Language Models are Unsupervised Multitask Learnersopenai
2.19.19Learning to Understand Goal Specifications by Modelling Reward
1.29.19GamePad: A Learning Environment for Theorem Proving
1.15.19Matrix capsules with EM routing
12.4.18Optimizing Agent Behavior over Long Time Scales by Transporting Value
11.27.18Embedding Logical Queries on Knowledge Graphswilliamleif
11.20.18Large-Scale Study of Curiosity-Driven Learningopenai
11.13.18Sparse Attentive Backtracking: Temporal Credit Assignment Through Remindingnke001
11.6.18Generalizing Hamiltonian Monte Carlo with Neural Networksbrain-research
10.23.18A Conceptual Introduction to Hamiltonian Monte Carlo
10.16.18MaskGAN: Better Text Generation via Filling in the ...
10.9.18Large Scale GAN Training for High Fidelity Natural Image Synthesis
10.2.18Improving Variational Inference with Inverse Autoregressive Flow
9.25.18Artificial Intelligence - The Revolution Hasn’t Happened Yet
9.18.18Learning deep representations by mutual information estimation and maximization
9.11.18The Variational Homoencoder: Learning to learn high capacity generative models from few examplesinsperatum
9.4.18Towards Conceptual Compressiongeosada
8.28.18Vector-based navigation using grid-like representations in artificial agentsdeepmind
break in maintaining this file; filled on April 10, 2020
--------------------------------
8.21.18Universal Transformerstensorflow
8.14.18Neural Arithmetic Logic Unitsgautam1858
8.7.18Neural Scene Representation and Rendering
7.31.18Measuring Abstract Reasoning in Neural Networks
6.26.18Improving Language Understanding by Generative Pre-Trainingopenai
6.19.18Associative Compression Networks for Representation Learning
6.12.18On Characterizing the Capacity of Neural Networks using Algebraic Topology
6.5.18Causal Effect Inference with Deep Latent-Variable ModelsAMLab
5.29.18ML beyond Curve Fitting
5.22.18Synthesizing Programs for Images using Reinforced Adversarial Learning
5.15.18TensorFlow Overviewr1.8
5.8.18Compositional Attention Networks for Machine Reasoningstanfordnlp
4.24.18The Annotated Transformer
4.3.18How Developers Iterate on Machine Learning Workflows
3.27.18Faster R-CNN: Towards Real-Time Object,Detection with Region Proposal Networks
3.20.18Attention Is All You Needtensor2tensor
3.6.18Generating Wikipedia by Summarizing Long Sequenceswikisum, per this gist
2.27.18AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial NetworksStackGAN-v2
2.20.18Information DropoutInformationDropout, official implementation
2.13.18Nested LSTMsNested-LSTM
2.6.18Deep vs. Shallow Networks: An Approximation Theory Perspective
1.30.18The Case for Learned Index Structures
1.23.18Visualizing The Loss Landscape Of Neural Nets
1.16.18Go for a Walk and Arrive at the Answer, RelNet: End-to-End Modeling of Entities & Relations
1.9.18Intro to Coq
12.12.17Chains of Reasoning over Entities, Relations, and Text using Recurrent Neural Networks(ChainsofReasoning)
12.5.17Stochastic Neural Networks for Hierarchical Reinforcement Learningsnn4hrl
11.28.17Emergent Complexity via Multi-Agent Competition (blog post)multiagent-competition
11.14.17Mastering the game of Go without human knowledge
11.7.17Meta-Learning with Memory-Augmented Neural Networksntm-meta-learning
10.24.17Poincaré Embeddings for Learning Hierarchical Representationspoincare_embeddings
10.17.17What does Attention in Neural Machine Translation Pay Attention to?
10.10.17Zero-Shot Learning Through Cross-Modal Transferzslearning
9.26.17Variational Boosting: Iteratively Refining Posterior Approximationsvboost
9.19.17Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networkscbfinn
9.12.17Neuroscience-inspired AI
9.5.17Recurrent Dropout Without Memory Lossrnn_cell_mulint_modern.py
8.29.17Deep Transfer Learning with Joint Adaptation Networksjmmd.{cpp,hpp}
8.22.17Designing Neural Network Architectures using Reinforcement Learningmetaqnn
8.15.17Phased LSTM: Accelerating Recurrent Network Training for Long or Event-based Sequencesplstm
8.8.17Hyper Networksotoro blog
8.1.17Full-Capacity Unitary Recurrent Neural Networkscomplex_RNN, urnn
7.25.17Decoupled Neural Interfaces using Synthetic Gradients & follow-updni.pytorch
7.18.17A simple neural network module for relational reasoningrelation-network
7.11.17Speaker diarization using deep neural network embeddings
6.20.17Neural Episodic ControlPFCM
6.13.17Lie-Access Neural Turing Machinesharvardnlp
6.6.17Artistic style transfer for videosartistic video
5.30.17High-Dimensional Continuous Control Using Generalized Advantage Estimationmodular_rl
5.23.17Emergence of Grounded Compositional Language in Multi-Agent Populations
5.16.17Trust Region Policy Optimizationmodular_rl
5.9.17Improved Training of Wasserstein GANscode
5.4.17Using Fast Weights to Attend to the Recent Past
4.25.17Strategic Attentive Writer for Learning Macro-Actions
4.18.17Massive Exploration of Neural Machine Translation Architectures
4.4.17End to End Learning for Self-Driving Cars
3.28.17Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning
3.21.17Image-to-Image Translation with Conditional Adversarial Networks
3.7.17Neural Programmer Interpreters
2.14.17Wasserstein GAN
2.7.17Towards Principled Methods for Training GANs
1.31.17Mastering the Game of Go with Deep Networks
1.24.17Understanding Deep Learning Requires Rethinking Generalization
1.17.17Neural Semantic Encoders
12.21.16StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
12.14.16Key-Value Memory Networks for Directly Reading Documents
12.7.16InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets

Contributors

anhinga

174 commits

pmiller10

57 commits

coventry

32 commits

jalexvig

7 commits

pmiller10/cambridge-ai

58

283 commits

updated Aug 26, 2024

See the code

README

SIPB Deep Learning Group

The schedule of readings for the SIPB/Cambridge AI Deep Learning Group If you have any papers you'd like to discuss, please either make a pull request, or send an email to the group and we'll add it. Papers with implementations available are strongly preferred.

Suggested Papers:

Schedule:

DatePaperImplementation
8.26.24The AI Scientist: Towards Fully Automated Open-Ended Scientific DiscoverySakanaAI/AI-Scientist
8.19.24Stretching Each Dollar: Diffusion Training from Scratch on a Micro-BudgetSonyResearch/micro_diffusion (pending)
8.12.24Scaling and evaluating sparse autoencodersopenai/sparse_autoencoder
4.10.24Score-Based Generative Modeling through Stochastic Differential Equations
3.28.24Generative Modeling by Estimating Gradients of the Data Distribution
3.21.24Humanoid Locomotion as Next Token Prediction
3.14.24TIES-Merging: Resolving Interference When Merging Modelsprateeky2806/ties-merging
2.8.24Merging Models with Fisher-Weighted Averagingarcee-ai/mergekit
2.1.24Averaging Weights Leads to Wider Optima and Better Generalization
1.18.24Hyena Hierarchy: Towards Larger Convolutional Language Models
1.04.24Mamba: Linear-Time Sequence Modeling with Selective State Spacesstate-spaces/mamba
12.07.23Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
11.30.233D Gaussian Splatting for Real-Time Radiance Field Renderinggraphdeco-inria/gaussian-splatting
11.16.23LILO: Learning Interpretable Libraries by Compressing and Documenting Codegabegrand/lilo
11.09.23Human-like systematic generalization through a meta-learning neural networkbrendenlake/MLC and brendenlake/MLC-ML
9.28.23Retrieval-Augmented Generation for Knowledge-Intensive NLP Taskshuggingface/transformers/examples/research_projects/rag
9.14.23Gradient-based Adversarial Attacks against Text Transformersfacebookresearch/text-adversarial-attack
8.10.23Reflexion: Language Agents with Verbal Reinforcement Learningnoahshinn024/reflexion
6.15.23RWKV: Reinventing RNNs for the Transformer EraBlinkDL/RWKV-LM
5.18.23Toy Models of Superposition
5.11.23LoRA: Low-Rank Adaptation of Large Language Modelstloen/alpaca-lora and huggingface/blog/lora
5.04.23Efficiently Modeling Long Sequences with Structured State SpacesHazyResearch/state-spaces
4.06.23Generating Sequences by Learning to Self-Correct
3.30.23The Capacity for Moral Self-Correction in Large Language Models
3.23.23LLaMA: Open and Efficient Foundation Language Modelsfacebookresearch/llama and huggingface/llama
3.16.23Language Is Not All You Need: Aligning Perception with Language Models
3.02.23Guiding Pretraining in Reinforcement Learning with Large Language Models
2.23.23Toolformer: Language Models Can Teach Themselves to Use Tools
2.16.23What learning algorithm is in-context learning? Investigations with linear modelsekinakyurek/incontext
2.09.23Flow Network based Generative Models for Non-Iterative Diverse Candidate GenerationTutorial
1.26.23Mastering Diverse Domains through World Models
1.12.23The Forward-Forward Algorithm: Some Preliminary Investigations
12.08.22Training language models to follow instructions with human feedback
9.22.22Git Re-Basin: Merging Models modulo Permutation Symmetries
9.08.22Transformers are Sample-Efficient World Models
8.25.22A Path Towards Autonomous Machine Intelligence
8.18.22Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfermicrosoft/mup
7.14.22Learning Iterative Reasoning through Energy Minimizationyilundu/irem_code_release
6.16.22Sharpness-Aware Minimization for Efficiently Improving Generalizationgoogle-research/sam
5.26.22Neural Tangent Kernel: Convergence and Generalization in Neural Networks
4.28.22A Modern Self-Referential Weight Matrix That Learns to Modify ItselfIDSIA/modern-srwm
4.14.22Hierarchical Perceiver
3.24.22Dual Diffusion Implicit Bridges for Image-to-Image Translation
3.10.22Understanding Generalization through Visualizationswronnyhuang/gen-viz
2.17.22Divide and Contrast: Self-supervised Learning from Uncurated Data
2.10.22Investigating Human Priors for Playing Video Gamesrach0012/humanRL_prior_games
1.27.22data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Languagepytorch/data2vec
1.20.22Consistent Video Depth Estimationfacebookresearch/consistent_depth
1.13.22Masked Autoencoders Are Scalable Vision Learners
12.02.21Training Verifiers to Solve Math Word Problems
11.18.21(StyleGan3) Alias-Free Generative Adversarial NetworksNVlabs/stylegan3
11.04.21Do Vision Transformers See Like Convolutional Neural Networks?
10.21.21CoBERL: Contrastive BERT for Reinforcement Learning
10.14.21WarpedGANSpace: Finding non-linear RBF paths in GAN latent spacechi0tzp/WarpedGANSpace
10.06.21RAFT: Recurrent All-Pairs Field Transforms for Optical Flowprinceton-vl/RAFT
9.16.21Bootstrapped Meta-Learning
9.09.21Program Synthesis with Large Language Models
8.19.21Perceiver IO: A General Architecture for Structured Inputs & Outputsdeepmind/perceiver
8.12.21Reward is enough
8.05.21Learning Compositional Rules via Neural Program Synthesismtensor/rulesynthesis
6.24.21Thinking Like Transformers
6.17.21Equilibrium Propagation: Bridging the Gap between Energy-Based Models and Backpropagation
6.10.21Unsupervised Learning by Competing Hidden Units
5.27.21Pay Attention to MLPs
5.20.21Memory Based Trajectory-conditioned Policies for Learning from Sparse Rewards
5.13.21Emerging Properties in Self-Supervised Vision Transformers
5.06.21Implicit Neural Representations with Periodic Activation Functionsvsitzmann/siren
4.29.21How to represent part-whole hierarchies in a neural networklucidrains/glom-pytorch RedRyan111/GLOM ArneBinder/GlomImpl
4.15.21Perceiver: General Perception with Iterative Attention
4.01.21Synthetic Returns for Long-Term Credit Assignment
3.25.21The Pitfalls of Simplicity Bias in Neural Networks
3.18.21Bootstrap your own latent: A new approach to self-supervised Learning
3.11.21Meta Learning Backpropagation And Improving It
3.04.21Taming Transformers for High-Resolution Image SynthesisCompVis/taming-transformers
2.18.21Pre-training without Natural Imageshirokatsukataoka16/FractalDB-Pretrained-ResNet-PyTorch
2.11.21Revisiting Locally Supervised Learning: an Alternative to End-to-end Trainingblackfeather-wang/InfoPro-Pytorch
2.04.21Neural Power Units
1.28.21Representation Learning via Invariant Causal Mechanisms
1.21.21γ-Models: Generative Temporal Difference Learning for Infinite-Horizon PredictionJannerM/gamma-models
1.14.21Improving Generalisation for Temporal Difference Learning: The Successor Representation
12.17.20Learning Associative Inference Using Fast Weight Memory
Hopfield Networks cycle ends
12.10.20Hopfield Networks is All You Needml-jku/hopfield-layers
12.03.20On a model of associative memory with huge storage capacity
11.19.20Dense Associative Memory for Pattern Recognition
11.12.20Neural Networks and Physical Systems with Emergent Collective Computational Abilities (= "the Hopfield Networks paper")
Hopfield Networks cycle of papers - from the original paper on Hopfield networks to "Hopfield Networks is All You Need"
11.05.20Training Generative Adversarial Networks with Limited DataNVlabs/stylegan2-ada
10.29.20Memories from patterns: Attractor and integrator networks in the brain
10.15.20Entities as Experts: Sparse Memory Access with Entity Supervision
10.08.20A Primer in BERTology: What we know about how BERT works
10.01.20It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learnerstimoschick/pet
9.24.20End-to-End Object Detection with Transformersfacebookresearch/detr
9.17.20Gated Linear Networks
7.23.20A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
7.02.20DreamCoder: Building interpretable hierarchical knowledge representations with wake-sleep Bayesian program learningellisk42/ec
6.18.20SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solverlocuslab/SATNet
6.4.20Adaptive Attention Span in Transformers
5.28.20Complexity control by gradient descent in deep networks
5.21.20What Can Learned Intrinsic Rewards Capture?
5.14.20COMET: Commonsense Transformers for Automatic Knowledge Graph Construction
5.7.20Write, Execute, Assess: Program Synthesis With a REPLflxsosa/ProgramSearch
4.23.20Graph Representations for Higher-Order Logic and Theorem Proving
4.16.20Mathematical Reasoning in Latent Space
4.9.20MEMO: A Deep Network for Flexible Combination of Episodic Memories
4.2.20Creating High Resolution Images with a Latent Adversarial Generator
3.26.20Invertible Residual Networks
3.5.20Value-driven Hindsight Modelling
2.27.20Analyzing and Improving the Image Quality of StyleGAN
2.13.20Axiomatic Attribution for Deep Networks
2.6.20Automated curricula through setter-solver interactions
1.30.20Protein structure prediction ...deepmind
1.23.20Putting An End to End-to-End: Gradient-Isolated Learning of Representations
1.16.20Normalizing Flows: An Introduction and Review of Current Methods
12.19.19Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
12.5.19On the Measure of Intelligence
11.21.19Understanding the Neural Tangent Kernelrajatvd
11.14.19XLNet: Generalized Autoregressive Pretraining for Language Understanding
11.7.19Learning to Predict Without Looking Ahead: World Models Without Forward Prediction
10.31.19Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
10.24.19N-BEATS: Neural basis expansion analysis for interpretable time series forecasting
10.17.19Unsupervised Doodling and Painting with Improved SPIRAL
10.10.19Adversarial Robustness as a Prior for Learned RepresentationsMadryLab
10.3.19Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
9.26.19Image Transformer
9.19.19Generating Diverse High-Fidelity Images with VQ-VAE-2
9.12.19Neural Discrete Representation Learning
9.5.19Neural Text Generation with Unlikelihood Training
8.29.19Learning Representations by Maximizing Mutual Information Across Views
breakswitch from Tuesdays to Thursdays after the break
6.11.19BERT Rediscovers the Classical NLP Pipeline
6.4.19Semantic Visual Localization
5.28.19AlgoNet: C^∞ Smooth Algorithmic Neural Networks
5.14.19Unsupervised Data Augmentation for Consistency Training
4.30.19Augmented Neural ODEs
4.9.19Wasserstein Dependency Measure for Representation Learning
4.2.19Leveraging Knowledge Bases in LSTMs for Improving Machine Reading
3.26.19Meta Particle Flow for Sequential Bayesian Inference
3.19.19A Meta-Transfer Objective for Learning to Disentangle Causal Mechanisms
3.12.19The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
2.26.19Language Models are Unsupervised Multitask Learnersopenai
2.19.19Learning to Understand Goal Specifications by Modelling Reward
1.29.19GamePad: A Learning Environment for Theorem Proving
1.15.19Matrix capsules with EM routing
12.4.18Optimizing Agent Behavior over Long Time Scales by Transporting Value
11.27.18Embedding Logical Queries on Knowledge Graphswilliamleif
11.20.18Large-Scale Study of Curiosity-Driven Learningopenai
11.13.18Sparse Attentive Backtracking: Temporal Credit Assignment Through Remindingnke001
11.6.18Generalizing Hamiltonian Monte Carlo with Neural Networksbrain-research
10.23.18A Conceptual Introduction to Hamiltonian Monte Carlo
10.16.18MaskGAN: Better Text Generation via Filling in the ...
10.9.18Large Scale GAN Training for High Fidelity Natural Image Synthesis
10.2.18Improving Variational Inference with Inverse Autoregressive Flow
9.25.18Artificial Intelligence - The Revolution Hasn’t Happened Yet
9.18.18Learning deep representations by mutual information estimation and maximization
9.11.18The Variational Homoencoder: Learning to learn high capacity generative models from few examplesinsperatum
9.4.18Towards Conceptual Compressiongeosada
8.28.18Vector-based navigation using grid-like representations in artificial agentsdeepmind
break in maintaining this file; filled on April 10, 2020
--------------------------------
8.21.18Universal Transformerstensorflow
8.14.18Neural Arithmetic Logic Unitsgautam1858
8.7.18Neural Scene Representation and Rendering
7.31.18Measuring Abstract Reasoning in Neural Networks
6.26.18Improving Language Understanding by Generative Pre-Trainingopenai
6.19.18Associative Compression Networks for Representation Learning
6.12.18On Characterizing the Capacity of Neural Networks using Algebraic Topology
6.5.18Causal Effect Inference with Deep Latent-Variable ModelsAMLab
5.29.18ML beyond Curve Fitting
5.22.18Synthesizing Programs for Images using Reinforced Adversarial Learning
5.15.18TensorFlow Overviewr1.8
5.8.18Compositional Attention Networks for Machine Reasoningstanfordnlp
4.24.18The Annotated Transformer
4.3.18How Developers Iterate on Machine Learning Workflows
3.27.18Faster R-CNN: Towards Real-Time Object,Detection with Region Proposal Networks
3.20.18Attention Is All You Needtensor2tensor
3.6.18Generating Wikipedia by Summarizing Long Sequenceswikisum, per this gist
2.27.18AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial NetworksStackGAN-v2
2.20.18Information DropoutInformationDropout, official implementation
2.13.18Nested LSTMsNested-LSTM
2.6.18Deep vs. Shallow Networks: An Approximation Theory Perspective
1.30.18The Case for Learned Index Structures
1.23.18Visualizing The Loss Landscape Of Neural Nets
1.16.18Go for a Walk and Arrive at the Answer, RelNet: End-to-End Modeling of Entities & Relations
1.9.18Intro to Coq
12.12.17Chains of Reasoning over Entities, Relations, and Text using Recurrent Neural Networks(ChainsofReasoning)
12.5.17Stochastic Neural Networks for Hierarchical Reinforcement Learningsnn4hrl
11.28.17Emergent Complexity via Multi-Agent Competition (blog post)multiagent-competition
11.14.17Mastering the game of Go without human knowledge
11.7.17Meta-Learning with Memory-Augmented Neural Networksntm-meta-learning
10.24.17Poincaré Embeddings for Learning Hierarchical Representationspoincare_embeddings
10.17.17What does Attention in Neural Machine Translation Pay Attention to?
10.10.17Zero-Shot Learning Through Cross-Modal Transferzslearning
9.26.17Variational Boosting: Iteratively Refining Posterior Approximationsvboost
9.19.17Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networkscbfinn
9.12.17Neuroscience-inspired AI
9.5.17Recurrent Dropout Without Memory Lossrnn_cell_mulint_modern.py
8.29.17Deep Transfer Learning with Joint Adaptation Networksjmmd.{cpp,hpp}
8.22.17Designing Neural Network Architectures using Reinforcement Learningmetaqnn
8.15.17Phased LSTM: Accelerating Recurrent Network Training for Long or Event-based Sequencesplstm
8.8.17Hyper Networksotoro blog
8.1.17Full-Capacity Unitary Recurrent Neural Networkscomplex_RNN, urnn
7.25.17Decoupled Neural Interfaces using Synthetic Gradients & follow-updni.pytorch
7.18.17A simple neural network module for relational reasoningrelation-network
7.11.17Speaker diarization using deep neural network embeddings
6.20.17Neural Episodic ControlPFCM
6.13.17Lie-Access Neural Turing Machinesharvardnlp
6.6.17Artistic style transfer for videosartistic video
5.30.17High-Dimensional Continuous Control Using Generalized Advantage Estimationmodular_rl
5.23.17Emergence of Grounded Compositional Language in Multi-Agent Populations
5.16.17Trust Region Policy Optimizationmodular_rl
5.9.17Improved Training of Wasserstein GANscode
5.4.17Using Fast Weights to Attend to the Recent Past
4.25.17Strategic Attentive Writer for Learning Macro-Actions
4.18.17Massive Exploration of Neural Machine Translation Architectures
4.4.17End to End Learning for Self-Driving Cars
3.28.17Variational Information Maximisation for Intrinsically Motivated Reinforcement Learning
3.21.17Image-to-Image Translation with Conditional Adversarial Networks
3.7.17Neural Programmer Interpreters
2.14.17Wasserstein GAN
2.7.17Towards Principled Methods for Training GANs
1.31.17Mastering the Game of Go with Deep Networks
1.24.17Understanding Deep Learning Requires Rethinking Generalization
1.17.17Neural Semantic Encoders
12.21.16StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks
12.14.16Key-Value Memory Networks for Directly Reading Documents
12.7.16InfoGAN: Interpretable Representation Learning by Information Maximizing Generative Adversarial Nets

Contributors

anhinga

174 commits

pmiller10

57 commits

coventry

32 commits

jalexvig

7 commits