MSWEIMZ/video-cnn-interpretability

视频 CNN 可解释性相关论文自动搜索与整理

Python

2

179 commits

updated Sep 23, 2026

See the code

README

English | 中文

📚 Video CNN/XAI Research Hub

Automated paper curation for video deep learning & explainability research

papers core strongly_related arXiv Semantic Scholar last_update license


Quick Navigation · 🏆 Influential · 🔥 Trending · 📄 Core · 📎 Strongly Related · 🏷️ Topics · 📈 Trends · 🖥️ Dashboard · 📋 Full List


📊 Overview

MetricCount
📚 Total Papers1143
🔥 Core Papers509
📎 Strongly Related634
🆕 New This Month151
📡 arXiv822
🔬 Semantic Scholar316
🔗 CrossRef Enriched16
✍️ Manual5
⏰ Last Updated2026-09-23 05:47:15

🏆 Top 5 Most Influential

This list highlights long-term impact; see Trending for recent work.


YearTitleSummaryCitationsScore
2024VideoMamba: State Space Model for Efficient Video UnderstandAddressing the dual challenges of local redundancy and global dependencies in vi5984.7
2024LongVU: Spatiotemporal Adaptive Compression for Long Video-LMultimodal Large Language Models (MLLMs) have shown promising progress in unders3544.7
2024Benchmarking Micro-Action Recognition: Dataset, Methods, andMicro-action is an imperceptible non-verbal behaviour characterised by low-inten1384.0
2025Video deepfake detection using a hybrid CNN-LSTM-TransformerThe proliferation of deepfake technology poses significant challenges due to its615.0
2025StreamForest: Efficient Online Video Understanding with PersMultimodal Large Language Models (MLLMs) have recently achieved remarkable progr574.7

🔥 Latest Core Papers

YearTitleSummaryAuthorScore
2026WA-JEPA: Rethinking the Video JEPA Paradigm for World-ActionVideo Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemXinlin Wang, Yujiao Xiang+5.8
2026Thinking Beyond Videos: Unifying Video Reasoning and Deep ReOpen-world video understanding often requires a model to locate sparse visual evWenqi Liu, Shijie Ma+5.8
2026VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Native visual reasoning treats visual generation as the medium of reasoning itseJunxiang Xu, Ruisi Wang+5.8
2026MyoMechanix: Biomechanically-Grounded Compositional Skilled Existing action quality assessment (AQA) datasets and methods rely primarily onHao Yin, Paritosh Parmar+5.8
2026Addressable Memory for Video World ModelsWe study visual persistence in interactive video world models. These models relyXindi Wu, Sven Elflein+5.7
2026MoTE: Mixture of Task Experts for Multi-Task Video UnderstanProcedural video-language models must solve heterogeneous tasks from the same viMuhammad Asad Ali, Umar Khan+5.7
2026Searching Videos as Trees: Self-Correcting Agents for GroundGrounded long-video question answering (Grounded LVQA) requires answering a quesCe Zhang, Ziyang Wang+5.5
2026GROVE: Growing and Reasoning over Temporally Stratified MemoA wearable assistant should both answer questions about its visual history and rSitong Gong, Caixin Kang+5.5
2026Video-DeepResearch: Towards the Next-Generation Multimodal DWe introduce Video-DeepResearch (Video-DR), extending multimodal agents from staZhen Fang, Yu Zeng+5.5
2026HelloWorld: Enabling Socially Interactive Characters in VideDespite the remarkable recent progress of video world models, social interactionLiangyang Ouyang, Ruicong Liu+5.5
2026Towards Expert-level Medical AI for Real-time Video ConsultaAudio-visual interaction is the standard for patient-physician consultations, enMahvish Nagda, Jihyeon Lee+5.5
2026X-LMC: Cross-View Spatiotemporal Collateral Circulation ScorDigital subtraction angiography (DSA) is the reference standard for leptomeningeMaedeh Hafezi Moghadas, Hakim Baazaoui+5.5
2026LeVJEPA: Efficient & Scalable Video Pretraining without the Video carries the temporal structure of the physical world, yet learning represeLukas Kuhn, Lucas Maes+5.5
2026Video-Based Palm-Vein Authentication under Challenging CondiPalm-vein biometrics are increasingly used for secure, contactless authenticatioXiaofeng Yan, Kechen Liu+5.5
2026MARS: What Retrieval Signals Are Hidden in Multimodal Large Text-video retrieval requires representations that can distinguish videos with sUicheol Jung, Juyoung Hong+5.5
2026Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Multimodal Large Language Models have demonstrated impressive video understandinZhaoyang Wei, Zipeng Wang+5.5
2026QuantWM: Temporally Consistent 2-Bit KV Cache Quantization fKV cache memory has become a major deployment bottleneck for video generation anJiaqi Zhao, Xiaobin Hu+5.5
2026Visual Representation Matters: Exploiting Temporal DifferencVideo-to-audio (V2A) generation extends image-to-audio generation (I2A) by introZehua Chen, Junyou Wang+5.4
2026LAION-BVD: A 10-Million-Hour Open Video Dataset for MultimodWe present LAION-BVD, a large-scale open video dataset for multimodal learning,Andreas Hochlehnert, Marianna Nezhurina+5.4
2026Post-Training VLMs for Video Mistake DetectionHuman mistakes are inevitable when following instructions, yet they can lead toFederico Spurio, Olga Zatsarynna+5.4
YearTitleSummaryAuthorScore
2026SM4RT: Learning Structured Motion Geometry for 4D ReconstrucGeometry Foundation Models (GFMs) have substantially advanced monocular 3D reconShing Ho J. Lin, Wenzhao Zheng+3.9
2026From Local Payoffs to Global Instabilities: A Spectral CartoWe develop a motif-based framework for spatiotemporal chaos in spatial evolutionOzgur Aydogmus3.9
2026From Passive Video to Editable Experience: Physically GroundThe key bottleneck in embodied AI is not model architecture but data. Although bJia Luo3.9
2026QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for QFault-tolerant quantum computing (FTQC) relies on quantum error correction to suRan Miao, Rui Luo+3.9
2026Faster-WAM: Do World Action Models Need Deep Action Modules?World Action Models (WAMs) couple robot action prediction with video world modelLiheng Ma, Rui Heng Yang+3.9
2026Context-Aware Mixture of Domain Experts for Bodily ExpressioThe same body posture can convey entirely different emotions depending on its suMohammad Mahdi Dehshibi, David Masip3.9
2026Identity-Faithful Audio-Visual Target Speaker Extraction witAudio-visual target speaker extraction should return the speaker indicated by thPeijun Yang, Zhan Jin+3.9
2026Multimodal Spatiotemporal Atmospheric Data Assimilation withData assimilation (DA) uses Bayesian inference to update the state of a numericaDibyajyoti Chakraborty, Romit Maulik3.9
2026BendTwin: Robust Dense-to-Sparse Physical Reconstruction witReconstructing objects with mechanical properties from video observations enableYixiong Jing, Qi Wang+3.9
2026Geometry-Aware Camera Localization for BronchoscopyCamera localization in bronchoscopy remains a challenging problem due to stringeLumin Chen, Qingyao Tian+3.9

📅 2026 (647 papers)
TagTitleSummaryAuthorScore
🔥WA-JEPA: Rethinking the Video JEPA Paradigm for WoVideo Joint Embedding Predictive Architecture (V-JEPA) learns powerfulXinlin Wang, Yujiao Xiang+5.8
🔥Thinking Beyond Videos: Unifying Video Reasoning aOpen-world video understanding often requires a model to locate sparseWenqi Liu, Shijie Ma+5.8
🔥VBVR-Pro: A Scalable and Verifiable Suite for NatiNative visual reasoning treats visual generation as the medium of reasJunxiang Xu, Ruisi Wang+5.8
🔥MyoMechanix: Biomechanically-Grounded CompositionaExisting action quality assessment (AQA) datasets and methods rely priHao Yin, Paritosh Parmar+5.8
🔥Addressable Memory for Video World ModelsWe study visual persistence in interactive video world models. These mXindi Wu, Sven Elflein+5.7
🔥MoTE: Mixture of Task Experts for Multi-Task VideoProcedural video-language models must solve heterogeneous tasks from tMuhammad Asad Ali, Umar Khan+5.7
🔥Searching Videos as Trees: Self-Correcting Agents Grounded long-video question answering (Grounded LVQA) requires answerCe Zhang, Ziyang Wang+5.5
🔥GROVE: Growing and Reasoning over Temporally StratA wearable assistant should both answer questions about its visual hisSitong Gong, Caixin Kang+5.5
🔥Video-DeepResearch: Towards the Next-Generation MuWe introduce Video-DeepResearch (Video-DR), extending multimodal agentZhen Fang, Yu Zeng+5.5
🔥HelloWorld: Enabling Socially Interactive CharacteDespite the remarkable recent progress of video world models, social iLiangyang Ouyang, Ruicong Liu+5.5
🔥Towards Expert-level Medical AI for Real-time VideAudio-visual interaction is the standard for patient-physician consultMahvish Nagda, Jihyeon Lee+5.5
🔥X-LMC: Cross-View Spatiotemporal Collateral CirculDigital subtraction angiography (DSA) is the reference standard for leMaedeh Hafezi Moghadas, Hakim Baazaoui+5.5

Showing 12 of 647 papers. See ALL_PAPERS.md for all entries.

📅 2025 (107 papers)
TagTitleSummaryAuthorScore
🔥Enhancing Video Understanding: Deep Neural NetworkIt's no secret that video has become the primary way we share informatAmir Hosein Fadaei, Mohammad-Reza A. Dehaqani5.8
🔥Fine tuning 3D Convolutional Networks for enhancedThe study of Human Activity Recognition (HAR) has attracted considerabAbir Frad, Hend Basly+5.4
🔥A Hybrid 3D CNNs Transformer Architecture for VideVideo-Based Human Action Recognition (HAR) remains challenging due toEngin Seven, Eylem Yücel Demirel5.4
🔥Video-CoT: A Comprehensive Dataset for SpatiotempoVideo content comprehension is essential for various applications, ranShuyi Zhang, Xiaoshuai Hao+5.2
🔥V-STaR: Benchmarking Video-LLMs on Video Spatio-TeHuman processes video reasoning in a sequential spatio-temporal reasonZixu Cheng, Jian Hu+5.2
🔥Video deepfake detection using a hybrid CNN-LSTM-TThe proliferation of deepfake technology poses significant challengesG. Petmezas, Vazgken Vanian+5.0
🔥TinyLLaVA-Video: Towards Smaller LMMs for Video UnVideo behavior recognition and scene understanding are fundamental tasXingjian Zhang, Xi Weng+4.9
🔥Harnessing Synthetic Preference Data for EnhancingWhile Video Large Language Models (Video-LLMs) have demonstrated remarSameep Vani, Shreyas Jena+4.9
🔥AceVFI: A Comprehensive Survey of Advances in VideVideo Frame Interpolation (VFI) is a core low-level vision task that sDahyeon Kye, Changhyun Roh+4.9
🔥How Much 3D Do Video Foundation Models Encode?Videos are continuous 2D projections of 3D worlds. After training on lZixuan Huang, Xiang Li+4.9
🔥A Novel 3D Convolutional Neural Network-Based DeepAccurate analysis of medical videos remains a major challenge in deepM. K. Dhar, Mou Deb+4.9
🔥RepAttn3D: Re-parameterizing 3D attention with spaThe technique of structural re-parameterization has been widely adopteXiusheng Lu, Lechao Cheng+4.8

Showing 12 of 107 papers. See ALL_PAPERS.md for all entries.

📅 2024 (103 papers)
TagTitleSummaryAuthorScore
🔥InternVideo2: Scaling Foundation Models for MultimWe introduce InternVideo2, a new family of video foundation models (ViYi Wang, Kunchang Li+5.2
🔥Various frameworks for integrating image and videoHuman action recognition has been identified as an important researchShaimaa Yosry, Lamiaa A. Elrefaei+4.8
🔥Video-based Exercise Classification and Activated This paper introduces a simple yet effective strategy for exercise claManvik Pasula, Pramit Saha4.7
🔥LongVU: Spatiotemporal Adaptive Compression for LoMultimodal Large Language Models (MLLMs) have shown promising progressXiaoqian Shen, Yunyang Xiong+4.7
🔥VideoMamba: State Space Model for Efficient Video Addressing the dual challenges of local redundancy and global dependenKunchang Li, Xinhao Li+4.7
🔥Isolated Video-Based Sign Language Recognition UsiSign language is a complex language that uses hand gestures, body moveDiksha Kumari, Radhey Shyam Anand4.7
🔥Can VLMs be used on videos for action recognition?Recent advancements have introduced multiple vision-language models (VHarsh Lunia4.4
🔥Prompting Video-Language Foundation Models with DoVideo Question Answering (VideoQA) represents a crucial intersection bTing Yu, Kunhao Fu+4.4
🔥Relevance-guided Audio Visual Fusion for Video SalAudio data, often synchronized with video frames, plays a crucial roleLi Yu, Xuanzhe Sun+4.4
🔥Automated diagnosis of respiratory diseases from lAn automated computerized approach can aid radiologists in the early dArefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan+4.3
🔥Facial Expression Recognition in Video Using 3D-CNThe focus of research work presented in this paper on improving perforSathisha G, C. K. Subbaraya+4.3
🔥Interpretability in Video-based Human Action RecogInterpretability plays a vital role in understanding complex deep learJorge Garcia-Torres Fernandez4.3

Showing 12 of 103 papers. See ALL_PAPERS.md for all entries.

📅 2023 (57 papers)
TagTitleSummaryAuthorScore
🔥Hierarchical Spatiotemporal Feature Fusion NetworkCurrent video saliency prediction methods have made great progress relYunzuo Zhang, Tian Zhang+5.7
🔥Video-FocalNets: Spatio-Temporal Focal Modulation Recent video recognition models utilize Transformer models for long-raSyed Talal Wasim, Muhammad Uzair Khattak+5.5
🔥Deep Neural Networks in Video Human Action RecogniCurrently, video behavior recognition is one of the most foundationalZihan Wang, Yang Yang+5.5
🔥Video Understanding with Large Language Models: A With the burgeoning growth of online video platforms and the escalatinYolo Y. Tang, Jing Bi+5.2
🔥Audio-visual Saliency for Omnidirectional VideosVisual saliency prediction for omnidirectional videos (ODVs) has shownYuxin Zhu, Xilei Zhu+5.2
🔥Understanding Video Transformers for Segmentation:Video segmentation encompasses a wide range of categories of problem fRezaul Karim, Richard P. Wildes4.8
🔥VMC: Video Motion Customization using Temporal AttText-to-video diffusion models have advanced video generation significHyeonho Jeong, Geon Yeong Park+4.7
🔥A Video Is Worth 4096 Tokens: Verbalize Videos To Multimedia content, such as advertisements and story videos, exhibit aAanisha Bhattacharya, Yaman K Singla+4.6
🔥A dynamic gesture recognition method based on R(2+Efficient spatial-temporal feature extraction from input video streamsYupeng Huo, Jie Shen+4.6
🔥Spatio-Temporal Features based Human Action Recogn—Recognition of human intention is crucial and challenging due to subtSaifuddin Saif, E. Wollega+4.5
🔥AMS-Net: Modeling Adaptive Multi-Granularity SpatiEffective spatio-temporal modeling as a core of video representation lQilong Wang, Qiyao Hu+4.5
🔥Video Traffic Analysis for Real-Time Emotion RecogSince the outbreak of the COVID-19 crisis, the transition to remote edAyoub Sassi, W. Jaafar+4.5

Showing 12 of 57 papers. See ALL_PAPERS.md for all entries.

📅 2022 (51 papers)
TagTitleSummaryAuthorScore
🔥Large-scale Robustness Analysis of Video Action ReWe have seen a great progress in video action recognition in recent yeMadeline Chantry Schiappa, Naman Biyani+5.5
🔥VRT: A Video Restoration TransformerVideo restoration (e. g. , video super-resolution) aims to restore higJingyun Liang, Jiezhang Cao+5.5
🔥3D Convolutional with Attention for Action RecogniHuman action recognition is one of the challenging tasks in computer vLabina Shrestha, Shikha Dubey+5.0
🔥UniFormerV2: Spatiotemporal Learning by Arming ImaLearning discriminative spatiotemporal representation is the key problKunchang Li, Yali Wang+5.0
🔥Video Human Action Recognition Algorithm Based on The traditional action recognition algorithm based on manual feature eYu Wang, Jiaxi Sun4.8
🔥No-Reference Video Quality Assessment Using Multi-With the constantly growing popularity of video-based services and appD. Varga4.8
🔥Action Recognition Using Action Sequences OptimizaEffective extraction and representation of action information are critXin Xiong, Weidong Min+4.5
🔥Enhancing Deformable Convolution based Video FrameThis paper presents a new deformable convolution-based video frame intDuolikun Danier, Fan Zhang+4.4
🔥Skeleton Graph-Neural-Network-Based Human Action RHuman action recognition has been applied in many fields, such as videMiao Feng, Jean Meunier4.4
🔥Two-stream fusion model using 3D-CNN and 2D-CNN viHand gestures are useful tools for many applications in the human-compDebajit Sarma, V. Kavyasree+4.3
🔥Sign Language Recognition Based on R(2+1)D With SpPrevious work utilized three-dimensional (3-D) convolutional neural neXiangzu Han, Fei Lu+4.3
🔥Video Visual Relation Detection via 3D ConvolutionVideo visual relation detection, which aims to detect the visual relatMingcheng Qu, Jianxun Cui+4.3

Showing 12 of 51 papers. See ALL_PAPERS.md for all entries.

📅 2021 (49 papers)
TagTitleSummaryAuthorScore
🔥Spatiotemporal Dilated Convolution with Uncertain In this paper, we propose a novel SpatioTemporal convolutional Dense NYu-Jen Ma, Hong-Han Shuai+6.3
🔥Action Transformer: A Self-Attention Model for ShoDeep neural networks based purely on attention have been successful acVittorio Mazzia, Simone Angarano+5.7
🔥Temporal-attentive Covariance Pooling Networks forFor video recognition task, a global representation summarizing the whZilin Gao, Qilong Wang+5.7
🔥Is Space-Time Attention All You Need for Video UndWe present a convolution-free approach to video classification built eGedas Bertasius, Heng Wang+5.6
🔥TAda! Temporally-Adaptive Convolutions for Video USpatial convolutions are widely used in numerous deep video models. ItZiyuan Huang, Shiwei Zhang+5.2
🔥Efficient Action Recognition with Introducing R(2+The mainstream methods in video action recognition includes 3D convoluHao Jin, Jianming Yang+5.2
🔥Token Shift Transformer for Video ClassificationTransformer achieves remarkable successes in understanding 1 and 2-dimHao Zhang, Y. Hao+5.0
🔥Improved CNN-based Learning of Interpolation FilteThe versatility of recent machine learning approaches makes them idealLuka Murn, Saverio Blasi+4.9
🔥SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-KiThis report presents the technical details of our submission to the EPSwathikiran Sudhakaran, Adrian Bulat+4.7
🔥"Knights": First Place Submission for VIPriors21 AThis technical report presents our approach "Knights" to solve the actIshan Dave, Naman Biyani+4.4
🔥Revisiting Video Saliency Prediction in the Deep LPredicting where people look in static scenes, a. k. a visual saliencyWenguan Wang, Jianbing Shen+4.3
🔥Recent Advances in Video Action Recognition with 3SUMMARY The performance of video action recognition has improved signiKensho Hara4.3

Showing 12 of 49 papers. See ALL_PAPERS.md for all entries.

📅 2020 (41 papers)
TagTitleSummaryAuthorScore
🔥Deep Analysis of CNN-based Spatio-temporal RepreseIn recent years, a number of approaches based on 2D or 3D convolutionaChun-Fu Chen, Rameswar Panda+5.8
🔥Unified Image and Video Saliency ModelingVisual saliency modeling for images and videos is treated as two indepRichard Droste, Jianbo Jiao+5.7
🔥TAM: Temporal Adaptive Module for Video RecognitioVideo data is with complex temporal dynamics due to various factors suZhaoyang Liu, Limin Wang+5.7
🔥Dissected 3D CNNs: Temporal Skip Connections for EConvolutional Neural Networks with 3D kernels (3D-CNNs) currently achiOkan Köpüklü, Stefan Hörmann+5.5
🔥TEA: Temporal Excitation and Aggregation for ActioTemporal modeling is key for action recognition in videos. It normallyYan Li, Bin Ji+5.5
🔥RANP: Resource Aware Neuron Pruning at InitializatAlthough 3D Convolutional Neural Networks (CNNs) are essential for mosZhiwei Xu, Thalaiyasingam Ajanthan+5.2
🔥Developing Motion Code Embedding for Action RecognIn this work, we propose a motion embedding strategy known as motion cMaxat Alibayev, David Paulius+5.2
🔥Would Mega-scale Datasets Further Enhance SpatioteHow can we collect and use a video dataset to further improve spatioteHirokatsu Kataoka, Tenga Wakamiya+5.0
🔥Learnable Sampling 3D Convolution for Video EnhancA key challenge in video enhancement and action recognition is to fuseShuyang Gu, Jianmin Bao+5.0
🔥Challenge report:VIPriors Action Recognition ChallThis paper is a brief report to our submission to the VIPriors ActionZhipeng Luo, Dawei Xu+4.7
🔥Res3ATN -- Deep 3D Residual Attention Network for Hand gesture recognition is a strenuous task to solve in videos. In thNaina Dhingra, Andreas Kunz4.6
🔥Toward Accurate Person-level Action Recognition inDetecting and recognizing human action in videos with crowded scenes iLi Yuan, Yichen Zhou+4.4

Showing 12 of 41 papers. See ALL_PAPERS.md for all entries.

📅 2019 (27 papers)
TagTitleSummaryAuthorScore
🔥Spatio-Temporal FAST 3D Convolutions for Human ActEffective processing of video input is essential for the recognition oAlexandros Stergiou, Ronald Poppe5.8
🔥A review of Convolutional-Neural-Network-based actAbstract Video action recognition is widely applied in video indexing,Guangle Yao, Tao Lei+5.3
🔥Spatiotemporal distilled dense-connectivity networAbstract Two-stream convolutional neural networks show great promise fWangli Hao, Zhaoxiang Zhang5.3
🔥Resource Efficient 3D Convolutional Neural NetworkRecently, convolutional neural networks with 3D kernels (3D CNNs) haveOkan Köpüklü, Neslihan Kose+5.0
🔥Image and Video Compression with Neural Networks: In recent years, the image and video coding technologies have advancedSiwei Ma, Xinfeng Zhang+4.9
🔥Improving Action Recognition with the Graph-NeuralRecent human action recognition methods mainly model a two-stream or 3Wu Luo, Chongyang Zhang+4.5
🔥Multi-teacher Knowledge Distillation for CompresseRecently, convolutional neural networks (CNNs) have seen great progresMeng-Chieh Wu, C. Chiu+4.2
🔥Motion Sickness Prediction in Stereoscopic Videos In this paper, we propose a three-dimensional (3D) convolutional neuraTae Min Lee, Jong-Chul Yoon+4.2
🔥Predicting 3D Human Dynamics from VideoGiven a video of a person in action, we can easily guess the 3D futureJason Y. Zhang, Panna Felsen+4.1
🔥FBK-HUPBA Submission to the EPIC-Kitchens 2019 ActIn this report we describe the technical details of our submission toSwathikiran Sudhakaran, Sergio Escalera+4.1
📎Explainable Deep Learning for Video Recognition TaThe popularity of Deep Learning for real-world applications is ever-grLiam Hiley, Alun Preece+3.9
📎Deep 3D Convolutional Neural Network for AutomatedComputer Aided Diagnosis has emerged as an indispensible technique forSumita Mishra, Naresh Kumar Chaudhary+3.9

Showing 12 of 27 papers. See ALL_PAPERS.md for all entries.

📅 2018 (28 papers)
TagTitleSummaryAuthorScore
🔥Interpretable Spatio-temporal Attention for Video Inspired by the observation that humans are able to process videos effLili Meng, Bo Zhao+7.4
🔥Revisiting Video Saliency: A Large-scale BenchmarkIn this work, we contribute to video saliency research in two ways. FiWenguan Wang, Jianbing Shen+5.3
🔥Review of Visual Saliency Detection with ComprehenVisual saliency detection model simulates the human visual system to pRunmin Cong, Jianjun Lei+5.0
🔥Reduced-Gate Convolutional LSTM Using Predictive CSpatiotemporal sequence prediction is an important problem in deep leaNelly Elsayed, Anthony S. Maida+4.9
🔥Recurrent Convolutions for Causal 3D CNNsRecently, three dimensional (3D) convolutional neural networks (CNNs)Gurkirt Singh, Fabio Cuzzolin4.8
🔥Non-local NetVLAD Encoding for Video ClassificatioThis paper describes our solution for the 2$^\text{nd}$ YouTube-8M vidYongyi Tang, Xing Zhang+4.6
🔥Morph: Flexible Acceleration for 3D CNN-Based VideThe past several years have seen both an explosion in the use of ConvoKartik Hegde, R. Agrawal+4.5
🔥ECO: Efficient Convolutional Network for Online ViThe state of the art in video understanding suffers from two problems:Mohammadreza Zolfaghari, Kamaljeet Singh+4.4
🔥Non-Local Video Denoising by CNNNon-local patch based methods were until recently state-of-the-art forAxel Davy, Thibaud Ehret+4.4
🔥Recurrence to the Rescue: Towards Causal SpatiotemRecently, three dimensional (3D) convolutional neural networks (CNNs)Gurkirt Singh, Fabio Cuzzolin4.3
🔥Benchmark 3D eye-tracking dataset for visual salieVisual Attention Models (VAMs) predict the location of an image or vidAmin Banitalebi-Dehkordi, Eleni Nasiopoulos+4.2
🔥SlowFast Networks for Video RecognitionWe present SlowFast networks for video recognition. Our model involvesChristoph Feichtenhofer, Haoqi Fan+4.2

Showing 12 of 28 papers. See ALL_PAPERS.md for all entries.

📅 2017 (19 papers)
TagTitleSummaryAuthorScore
🔥The Monkeytyping Solution to the YouTube-8M Video This article describes the final solution of team monkeytyping, who fiHe-Da Wang, Teng Zhang+5.2
🔥A Closer Look at Spatiotemporal Convolutions for AIn this paper we discuss several forms of spatiotemporal convolutionsDu Tran, Heng Wang+5.1
🔥Predicting Video Saliency with Object-to-Motion CNOver the past few years, deep neural networks (DNNs) have exhibited grLai Jiang, Mai Xu+5.0
🔥Hierarchical Deep Recurrent Architecture for VideoThis paper introduces the system we developed for the Youtube-8M VideoLuming Tang, Boyang Deng+4.9
🔥Graph-Theoretic Spatiotemporal Context Modeling foAs an important and challenging problem in computer vision, video saliLina Wei, Fangfang Wang+4.8
🔥Video Classification With CNNs: Using The Codec AsWe investigate video classification via a two-stream convolutional neuAaron Chadha, Alhabib Abbas+4.7
🔥Facial Expression Recognition Using Enhanced Deep Deep Neural Networks (DNNs) have shown to outperform traditional methoBehzad Hasani, Mohammad H. Mahoor4.4
🔥Temporal Relational Reasoning in VideosTemporal relational reasoning, the ability to link meaningful transforBolei Zhou, A. Andonian+4.2
🔥Two-Stream 3D Convolutional Neural Network for SkeIt remains a challenge to efficiently extract spatialtemporal informatHong Liu, Juanhui Tu+4.2
🔥Spatio-Temporal Facial Expression Recognition UsinAutomated Facial Expression Recognition (FER) has been a challenging tBehzad Hasani, Mohammad H. Mahoor4.1
📎A Brief Survey of Deep Reinforcement LearningDeep reinforcement learning is poised to revolutionise the field of AIKai Arulkumaran, Marc Peter Deisenroth+3.9
📎Rethinking Spatiotemporal Feature Learning For VidRethinking Spatiotemporal Feature Learning For Video UnderstandingSaining Xie, Chen Sun+3.9

Showing 12 of 19 papers. See ALL_PAPERS.md for all entries.

📅 2016 (6 papers)
TagTitleSummaryAuthorScore
🔥Convolutional Two-Stream Network Fusion for Video Recent applications of Convolutional Neural Networks (ConvNets) for huChristoph Feichtenhofer, A. Pinz+5.0
🔥Large-Scale Shape Retrieval with Sparse 3D ConvoluIn this paper we present results of performance evaluation of S3DCNN -Alexandr Notchenko, Ermek Kapushev+4.4
📎Deep Learning for Saliency Prediction in Natural VThe purpose of this paper is the detection of salient areas in naturalSouad Chaabouni, Jenny Benois-Pineau+3.4
📎Grad-CAM: Visual Explanations from Deep Networks vWe propose a technique for producing ‘visual explanations’ for decisioRamprasaath R. Selvaraju, Abhishek Das+3.3
📎A Deep Learning Approach for Joint Video Frame andReinforcement learning is concerned with identifying reward-maximizingFelix Leibfried, Nate Kushman+2.6
📎Transfer learning with deep networks for saliency Transfer learning with deep networks for saliency prediction in naturaS. Chaabouni, J. Benois-Pineau+2.6
📅 2015 (6 papers)
TagTitleSummaryAuthorScore
🔥C3D: Generic Features for Video AnalysisWe propose a simple yet effective approach for spatiotemporal featureDu Tran, Lubomir Bourdev+5.7
🔥Intra-and-Inter-Constraint-based Video EnhancementVideo enhancement plays an important role in various video applicationYuanzhe Chen, Weiyao Lin+4.1
📎Activity Recognition Using A Combination of CategoThis paper presents a novel approach for automatic recognition of humaWeiyao Lin, Ming-Ting Sun+3.8
📎A new network-based algorithm for human activity rIn this paper, a new network-transmission-based (NTB) algorithm is proWeiyao Lin, Yuanzhe Chen+3.8
📎Group Event Detection with a Varying Number of GroThis paper presents a novel approach for automatic recognition of grouWeiyao Lin, Ming-Ting Sun+2.9
📎VoxNet: A 3D Convolutional Neural Network for realVoxNet: A 3D Convolutional Neural Network for real-time object recogniD. Maturana, S. Scherer2.8
📅 2014 (2 papers)
TagTitleSummaryAuthorScore
🔥Visualizing and Understanding Convolutional NetworLarge Convolutional Network models have recently demonstrated impressiMatthew D. Zeiler, Rob Fergus5.4
📎Deep Inside Convolutional Networks: Visualising ImThis paper addresses the visualisation of image classification models,Karen Simonyan, Andrea Vedaldi+3.7

📄 Full paper list: ALL_PAPERS.md

🏗️ Architecture

┌─────────────────────────────────────────────────────────────┐
│                    search_config.json                       │
│              (24 queries × 3 depth layers)                  │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                     Collector                               │
│         arXiv API  +  Semantic Scholar API                  │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│              Normalizer + CrossRef Enrichment               │
│        ID/version normalization, dedup, citations           │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                      Scorer                                 │
│    Keyword Match + Citations + Venue + Survey Bonus          │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                      Storage                                │
│   papers/index.jsonl + papers/quarantine.jsonl              │
└───────────────────────────┬─────────────────────────────────┘
                            │
          ┌─────────────────┼─────────────────┐
          ▼                 ▼                 ▼
    ┌──────────┐     ┌──────────┐     ┌──────────┐
    │ README   │     │ Feishu   │     │ Dashboard │
    │ (Display)│     │ (Notify) │     │  (HTML)   │
    └──────────┘     └──────────┘     └──────────┘

✨ Features

🎯 Smart Search📊 Data Enhancement🌐 Multi-Source🔔 Auto Notify
Daily arXiv searchCrossRef enrichment for new papersarXiv + Semantic ScholarFeishu Webhook
24 layered queriesCN/EN summary generationTitle dedup + ID normSuccess/Failure alerts
Score-based filteringTopic clusteringCitation + Venue boostGitHub Actions

⚙️ Auto Update

GitHub Actions aggregates arXiv and Semantic Scholar daily, then normalizes, deduplicates, scores, and enriches accepted papers with CrossRef.

📄 License

For academic research use only

Contributors

MSWEIMZ

2 commits

MSWEIMZ/video-cnn-interpretability

视频 CNN 可解释性相关论文自动搜索与整理

Python

2

179 commits

updated Sep 23, 2026

See the code

README

English | 中文

📚 Video CNN/XAI Research Hub

Automated paper curation for video deep learning & explainability research

papers core strongly_related arXiv Semantic Scholar last_update license


Quick Navigation · 🏆 Influential · 🔥 Trending · 📄 Core · 📎 Strongly Related · 🏷️ Topics · 📈 Trends · 🖥️ Dashboard · 📋 Full List


📊 Overview

MetricCount
📚 Total Papers1143
🔥 Core Papers509
📎 Strongly Related634
🆕 New This Month151
📡 arXiv822
🔬 Semantic Scholar316
🔗 CrossRef Enriched16
✍️ Manual5
⏰ Last Updated2026-09-23 05:47:15

🏆 Top 5 Most Influential

This list highlights long-term impact; see Trending for recent work.


YearTitleSummaryCitationsScore
2024VideoMamba: State Space Model for Efficient Video UnderstandAddressing the dual challenges of local redundancy and global dependencies in vi5984.7
2024LongVU: Spatiotemporal Adaptive Compression for Long Video-LMultimodal Large Language Models (MLLMs) have shown promising progress in unders3544.7
2024Benchmarking Micro-Action Recognition: Dataset, Methods, andMicro-action is an imperceptible non-verbal behaviour characterised by low-inten1384.0
2025Video deepfake detection using a hybrid CNN-LSTM-TransformerThe proliferation of deepfake technology poses significant challenges due to its615.0
2025StreamForest: Efficient Online Video Understanding with PersMultimodal Large Language Models (MLLMs) have recently achieved remarkable progr574.7

🔥 Latest Core Papers

YearTitleSummaryAuthorScore
2026WA-JEPA: Rethinking the Video JEPA Paradigm for World-ActionVideo Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemXinlin Wang, Yujiao Xiang+5.8
2026Thinking Beyond Videos: Unifying Video Reasoning and Deep ReOpen-world video understanding often requires a model to locate sparse visual evWenqi Liu, Shijie Ma+5.8
2026VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Native visual reasoning treats visual generation as the medium of reasoning itseJunxiang Xu, Ruisi Wang+5.8
2026MyoMechanix: Biomechanically-Grounded Compositional Skilled Existing action quality assessment (AQA) datasets and methods rely primarily onHao Yin, Paritosh Parmar+5.8
2026Addressable Memory for Video World ModelsWe study visual persistence in interactive video world models. These models relyXindi Wu, Sven Elflein+5.7
2026MoTE: Mixture of Task Experts for Multi-Task Video UnderstanProcedural video-language models must solve heterogeneous tasks from the same viMuhammad Asad Ali, Umar Khan+5.7
2026Searching Videos as Trees: Self-Correcting Agents for GroundGrounded long-video question answering (Grounded LVQA) requires answering a quesCe Zhang, Ziyang Wang+5.5
2026GROVE: Growing and Reasoning over Temporally Stratified MemoA wearable assistant should both answer questions about its visual history and rSitong Gong, Caixin Kang+5.5
2026Video-DeepResearch: Towards the Next-Generation Multimodal DWe introduce Video-DeepResearch (Video-DR), extending multimodal agents from staZhen Fang, Yu Zeng+5.5
2026HelloWorld: Enabling Socially Interactive Characters in VideDespite the remarkable recent progress of video world models, social interactionLiangyang Ouyang, Ruicong Liu+5.5
2026Towards Expert-level Medical AI for Real-time Video ConsultaAudio-visual interaction is the standard for patient-physician consultations, enMahvish Nagda, Jihyeon Lee+5.5
2026X-LMC: Cross-View Spatiotemporal Collateral Circulation ScorDigital subtraction angiography (DSA) is the reference standard for leptomeningeMaedeh Hafezi Moghadas, Hakim Baazaoui+5.5
2026LeVJEPA: Efficient & Scalable Video Pretraining without the Video carries the temporal structure of the physical world, yet learning represeLukas Kuhn, Lucas Maes+5.5
2026Video-Based Palm-Vein Authentication under Challenging CondiPalm-vein biometrics are increasingly used for secure, contactless authenticatioXiaofeng Yan, Kechen Liu+5.5
2026MARS: What Retrieval Signals Are Hidden in Multimodal Large Text-video retrieval requires representations that can distinguish videos with sUicheol Jung, Juyoung Hong+5.5
2026Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Multimodal Large Language Models have demonstrated impressive video understandinZhaoyang Wei, Zipeng Wang+5.5
2026QuantWM: Temporally Consistent 2-Bit KV Cache Quantization fKV cache memory has become a major deployment bottleneck for video generation anJiaqi Zhao, Xiaobin Hu+5.5
2026Visual Representation Matters: Exploiting Temporal DifferencVideo-to-audio (V2A) generation extends image-to-audio generation (I2A) by introZehua Chen, Junyou Wang+5.4
2026LAION-BVD: A 10-Million-Hour Open Video Dataset for MultimodWe present LAION-BVD, a large-scale open video dataset for multimodal learning,Andreas Hochlehnert, Marianna Nezhurina+5.4
2026Post-Training VLMs for Video Mistake DetectionHuman mistakes are inevitable when following instructions, yet they can lead toFederico Spurio, Olga Zatsarynna+5.4
YearTitleSummaryAuthorScore
2026SM4RT: Learning Structured Motion Geometry for 4D ReconstrucGeometry Foundation Models (GFMs) have substantially advanced monocular 3D reconShing Ho J. Lin, Wenzhao Zheng+3.9
2026From Local Payoffs to Global Instabilities: A Spectral CartoWe develop a motif-based framework for spatiotemporal chaos in spatial evolutionOzgur Aydogmus3.9
2026From Passive Video to Editable Experience: Physically GroundThe key bottleneck in embodied AI is not model architecture but data. Although bJia Luo3.9
2026QAdapt: A Noise-Adaptive Neural Pre-Decoding Framework for QFault-tolerant quantum computing (FTQC) relies on quantum error correction to suRan Miao, Rui Luo+3.9
2026Faster-WAM: Do World Action Models Need Deep Action Modules?World Action Models (WAMs) couple robot action prediction with video world modelLiheng Ma, Rui Heng Yang+3.9
2026Context-Aware Mixture of Domain Experts for Bodily ExpressioThe same body posture can convey entirely different emotions depending on its suMohammad Mahdi Dehshibi, David Masip3.9
2026Identity-Faithful Audio-Visual Target Speaker Extraction witAudio-visual target speaker extraction should return the speaker indicated by thPeijun Yang, Zhan Jin+3.9
2026Multimodal Spatiotemporal Atmospheric Data Assimilation withData assimilation (DA) uses Bayesian inference to update the state of a numericaDibyajyoti Chakraborty, Romit Maulik3.9
2026BendTwin: Robust Dense-to-Sparse Physical Reconstruction witReconstructing objects with mechanical properties from video observations enableYixiong Jing, Qi Wang+3.9
2026Geometry-Aware Camera Localization for BronchoscopyCamera localization in bronchoscopy remains a challenging problem due to stringeLumin Chen, Qingyao Tian+3.9

📅 2026 (647 papers)
TagTitleSummaryAuthorScore
🔥WA-JEPA: Rethinking the Video JEPA Paradigm for WoVideo Joint Embedding Predictive Architecture (V-JEPA) learns powerfulXinlin Wang, Yujiao Xiang+5.8
🔥Thinking Beyond Videos: Unifying Video Reasoning aOpen-world video understanding often requires a model to locate sparseWenqi Liu, Shijie Ma+5.8
🔥VBVR-Pro: A Scalable and Verifiable Suite for NatiNative visual reasoning treats visual generation as the medium of reasJunxiang Xu, Ruisi Wang+5.8
🔥MyoMechanix: Biomechanically-Grounded CompositionaExisting action quality assessment (AQA) datasets and methods rely priHao Yin, Paritosh Parmar+5.8
🔥Addressable Memory for Video World ModelsWe study visual persistence in interactive video world models. These mXindi Wu, Sven Elflein+5.7
🔥MoTE: Mixture of Task Experts for Multi-Task VideoProcedural video-language models must solve heterogeneous tasks from tMuhammad Asad Ali, Umar Khan+5.7
🔥Searching Videos as Trees: Self-Correcting Agents Grounded long-video question answering (Grounded LVQA) requires answerCe Zhang, Ziyang Wang+5.5
🔥GROVE: Growing and Reasoning over Temporally StratA wearable assistant should both answer questions about its visual hisSitong Gong, Caixin Kang+5.5
🔥Video-DeepResearch: Towards the Next-Generation MuWe introduce Video-DeepResearch (Video-DR), extending multimodal agentZhen Fang, Yu Zeng+5.5
🔥HelloWorld: Enabling Socially Interactive CharacteDespite the remarkable recent progress of video world models, social iLiangyang Ouyang, Ruicong Liu+5.5
🔥Towards Expert-level Medical AI for Real-time VideAudio-visual interaction is the standard for patient-physician consultMahvish Nagda, Jihyeon Lee+5.5
🔥X-LMC: Cross-View Spatiotemporal Collateral CirculDigital subtraction angiography (DSA) is the reference standard for leMaedeh Hafezi Moghadas, Hakim Baazaoui+5.5

Showing 12 of 647 papers. See ALL_PAPERS.md for all entries.

📅 2025 (107 papers)
TagTitleSummaryAuthorScore
🔥Enhancing Video Understanding: Deep Neural NetworkIt's no secret that video has become the primary way we share informatAmir Hosein Fadaei, Mohammad-Reza A. Dehaqani5.8
🔥Fine tuning 3D Convolutional Networks for enhancedThe study of Human Activity Recognition (HAR) has attracted considerabAbir Frad, Hend Basly+5.4
🔥A Hybrid 3D CNNs Transformer Architecture for VideVideo-Based Human Action Recognition (HAR) remains challenging due toEngin Seven, Eylem Yücel Demirel5.4
🔥Video-CoT: A Comprehensive Dataset for SpatiotempoVideo content comprehension is essential for various applications, ranShuyi Zhang, Xiaoshuai Hao+5.2
🔥V-STaR: Benchmarking Video-LLMs on Video Spatio-TeHuman processes video reasoning in a sequential spatio-temporal reasonZixu Cheng, Jian Hu+5.2
🔥Video deepfake detection using a hybrid CNN-LSTM-TThe proliferation of deepfake technology poses significant challengesG. Petmezas, Vazgken Vanian+5.0
🔥TinyLLaVA-Video: Towards Smaller LMMs for Video UnVideo behavior recognition and scene understanding are fundamental tasXingjian Zhang, Xi Weng+4.9
🔥Harnessing Synthetic Preference Data for EnhancingWhile Video Large Language Models (Video-LLMs) have demonstrated remarSameep Vani, Shreyas Jena+4.9
🔥AceVFI: A Comprehensive Survey of Advances in VideVideo Frame Interpolation (VFI) is a core low-level vision task that sDahyeon Kye, Changhyun Roh+4.9
🔥How Much 3D Do Video Foundation Models Encode?Videos are continuous 2D projections of 3D worlds. After training on lZixuan Huang, Xiang Li+4.9
🔥A Novel 3D Convolutional Neural Network-Based DeepAccurate analysis of medical videos remains a major challenge in deepM. K. Dhar, Mou Deb+4.9
🔥RepAttn3D: Re-parameterizing 3D attention with spaThe technique of structural re-parameterization has been widely adopteXiusheng Lu, Lechao Cheng+4.8

Showing 12 of 107 papers. See ALL_PAPERS.md for all entries.

📅 2024 (103 papers)
TagTitleSummaryAuthorScore
🔥InternVideo2: Scaling Foundation Models for MultimWe introduce InternVideo2, a new family of video foundation models (ViYi Wang, Kunchang Li+5.2
🔥Various frameworks for integrating image and videoHuman action recognition has been identified as an important researchShaimaa Yosry, Lamiaa A. Elrefaei+4.8
🔥Video-based Exercise Classification and Activated This paper introduces a simple yet effective strategy for exercise claManvik Pasula, Pramit Saha4.7
🔥LongVU: Spatiotemporal Adaptive Compression for LoMultimodal Large Language Models (MLLMs) have shown promising progressXiaoqian Shen, Yunyang Xiong+4.7
🔥VideoMamba: State Space Model for Efficient Video Addressing the dual challenges of local redundancy and global dependenKunchang Li, Xinhao Li+4.7
🔥Isolated Video-Based Sign Language Recognition UsiSign language is a complex language that uses hand gestures, body moveDiksha Kumari, Radhey Shyam Anand4.7
🔥Can VLMs be used on videos for action recognition?Recent advancements have introduced multiple vision-language models (VHarsh Lunia4.4
🔥Prompting Video-Language Foundation Models with DoVideo Question Answering (VideoQA) represents a crucial intersection bTing Yu, Kunhao Fu+4.4
🔥Relevance-guided Audio Visual Fusion for Video SalAudio data, often synchronized with video frames, plays a crucial roleLi Yu, Xuanzhe Sun+4.4
🔥Automated diagnosis of respiratory diseases from lAn automated computerized approach can aid radiologists in the early dArefin Ittesafun Abian, Mohaimenul Azam Khan Raiaan+4.3
🔥Facial Expression Recognition in Video Using 3D-CNThe focus of research work presented in this paper on improving perforSathisha G, C. K. Subbaraya+4.3
🔥Interpretability in Video-based Human Action RecogInterpretability plays a vital role in understanding complex deep learJorge Garcia-Torres Fernandez4.3

Showing 12 of 103 papers. See ALL_PAPERS.md for all entries.

📅 2023 (57 papers)
TagTitleSummaryAuthorScore
🔥Hierarchical Spatiotemporal Feature Fusion NetworkCurrent video saliency prediction methods have made great progress relYunzuo Zhang, Tian Zhang+5.7
🔥Video-FocalNets: Spatio-Temporal Focal Modulation Recent video recognition models utilize Transformer models for long-raSyed Talal Wasim, Muhammad Uzair Khattak+5.5
🔥Deep Neural Networks in Video Human Action RecogniCurrently, video behavior recognition is one of the most foundationalZihan Wang, Yang Yang+5.5
🔥Video Understanding with Large Language Models: A With the burgeoning growth of online video platforms and the escalatinYolo Y. Tang, Jing Bi+5.2
🔥Audio-visual Saliency for Omnidirectional VideosVisual saliency prediction for omnidirectional videos (ODVs) has shownYuxin Zhu, Xilei Zhu+5.2
🔥Understanding Video Transformers for Segmentation:Video segmentation encompasses a wide range of categories of problem fRezaul Karim, Richard P. Wildes4.8
🔥VMC: Video Motion Customization using Temporal AttText-to-video diffusion models have advanced video generation significHyeonho Jeong, Geon Yeong Park+4.7
🔥A Video Is Worth 4096 Tokens: Verbalize Videos To Multimedia content, such as advertisements and story videos, exhibit aAanisha Bhattacharya, Yaman K Singla+4.6
🔥A dynamic gesture recognition method based on R(2+Efficient spatial-temporal feature extraction from input video streamsYupeng Huo, Jie Shen+4.6
🔥Spatio-Temporal Features based Human Action Recogn—Recognition of human intention is crucial and challenging due to subtSaifuddin Saif, E. Wollega+4.5
🔥AMS-Net: Modeling Adaptive Multi-Granularity SpatiEffective spatio-temporal modeling as a core of video representation lQilong Wang, Qiyao Hu+4.5
🔥Video Traffic Analysis for Real-Time Emotion RecogSince the outbreak of the COVID-19 crisis, the transition to remote edAyoub Sassi, W. Jaafar+4.5

Showing 12 of 57 papers. See ALL_PAPERS.md for all entries.

📅 2022 (51 papers)
TagTitleSummaryAuthorScore
🔥Large-scale Robustness Analysis of Video Action ReWe have seen a great progress in video action recognition in recent yeMadeline Chantry Schiappa, Naman Biyani+5.5
🔥VRT: A Video Restoration TransformerVideo restoration (e. g. , video super-resolution) aims to restore higJingyun Liang, Jiezhang Cao+5.5
🔥3D Convolutional with Attention for Action RecogniHuman action recognition is one of the challenging tasks in computer vLabina Shrestha, Shikha Dubey+5.0
🔥UniFormerV2: Spatiotemporal Learning by Arming ImaLearning discriminative spatiotemporal representation is the key problKunchang Li, Yali Wang+5.0
🔥Video Human Action Recognition Algorithm Based on The traditional action recognition algorithm based on manual feature eYu Wang, Jiaxi Sun4.8
🔥No-Reference Video Quality Assessment Using Multi-With the constantly growing popularity of video-based services and appD. Varga4.8
🔥Action Recognition Using Action Sequences OptimizaEffective extraction and representation of action information are critXin Xiong, Weidong Min+4.5
🔥Enhancing Deformable Convolution based Video FrameThis paper presents a new deformable convolution-based video frame intDuolikun Danier, Fan Zhang+4.4
🔥Skeleton Graph-Neural-Network-Based Human Action RHuman action recognition has been applied in many fields, such as videMiao Feng, Jean Meunier4.4
🔥Two-stream fusion model using 3D-CNN and 2D-CNN viHand gestures are useful tools for many applications in the human-compDebajit Sarma, V. Kavyasree+4.3
🔥Sign Language Recognition Based on R(2+1)D With SpPrevious work utilized three-dimensional (3-D) convolutional neural neXiangzu Han, Fei Lu+4.3
🔥Video Visual Relation Detection via 3D ConvolutionVideo visual relation detection, which aims to detect the visual relatMingcheng Qu, Jianxun Cui+4.3

Showing 12 of 51 papers. See ALL_PAPERS.md for all entries.

📅 2021 (49 papers)
TagTitleSummaryAuthorScore
🔥Spatiotemporal Dilated Convolution with Uncertain In this paper, we propose a novel SpatioTemporal convolutional Dense NYu-Jen Ma, Hong-Han Shuai+6.3
🔥Action Transformer: A Self-Attention Model for ShoDeep neural networks based purely on attention have been successful acVittorio Mazzia, Simone Angarano+5.7
🔥Temporal-attentive Covariance Pooling Networks forFor video recognition task, a global representation summarizing the whZilin Gao, Qilong Wang+5.7
🔥Is Space-Time Attention All You Need for Video UndWe present a convolution-free approach to video classification built eGedas Bertasius, Heng Wang+5.6
🔥TAda! Temporally-Adaptive Convolutions for Video USpatial convolutions are widely used in numerous deep video models. ItZiyuan Huang, Shiwei Zhang+5.2
🔥Efficient Action Recognition with Introducing R(2+The mainstream methods in video action recognition includes 3D convoluHao Jin, Jianming Yang+5.2
🔥Token Shift Transformer for Video ClassificationTransformer achieves remarkable successes in understanding 1 and 2-dimHao Zhang, Y. Hao+5.0
🔥Improved CNN-based Learning of Interpolation FilteThe versatility of recent machine learning approaches makes them idealLuka Murn, Saverio Blasi+4.9
🔥SAIC_Cambridge-HuPBA-FBK Submission to the EPIC-KiThis report presents the technical details of our submission to the EPSwathikiran Sudhakaran, Adrian Bulat+4.7
🔥"Knights": First Place Submission for VIPriors21 AThis technical report presents our approach "Knights" to solve the actIshan Dave, Naman Biyani+4.4
🔥Revisiting Video Saliency Prediction in the Deep LPredicting where people look in static scenes, a. k. a visual saliencyWenguan Wang, Jianbing Shen+4.3
🔥Recent Advances in Video Action Recognition with 3SUMMARY The performance of video action recognition has improved signiKensho Hara4.3

Showing 12 of 49 papers. See ALL_PAPERS.md for all entries.

📅 2020 (41 papers)
TagTitleSummaryAuthorScore
🔥Deep Analysis of CNN-based Spatio-temporal RepreseIn recent years, a number of approaches based on 2D or 3D convolutionaChun-Fu Chen, Rameswar Panda+5.8
🔥Unified Image and Video Saliency ModelingVisual saliency modeling for images and videos is treated as two indepRichard Droste, Jianbo Jiao+5.7
🔥TAM: Temporal Adaptive Module for Video RecognitioVideo data is with complex temporal dynamics due to various factors suZhaoyang Liu, Limin Wang+5.7
🔥Dissected 3D CNNs: Temporal Skip Connections for EConvolutional Neural Networks with 3D kernels (3D-CNNs) currently achiOkan Köpüklü, Stefan Hörmann+5.5
🔥TEA: Temporal Excitation and Aggregation for ActioTemporal modeling is key for action recognition in videos. It normallyYan Li, Bin Ji+5.5
🔥RANP: Resource Aware Neuron Pruning at InitializatAlthough 3D Convolutional Neural Networks (CNNs) are essential for mosZhiwei Xu, Thalaiyasingam Ajanthan+5.2
🔥Developing Motion Code Embedding for Action RecognIn this work, we propose a motion embedding strategy known as motion cMaxat Alibayev, David Paulius+5.2
🔥Would Mega-scale Datasets Further Enhance SpatioteHow can we collect and use a video dataset to further improve spatioteHirokatsu Kataoka, Tenga Wakamiya+5.0
🔥Learnable Sampling 3D Convolution for Video EnhancA key challenge in video enhancement and action recognition is to fuseShuyang Gu, Jianmin Bao+5.0
🔥Challenge report:VIPriors Action Recognition ChallThis paper is a brief report to our submission to the VIPriors ActionZhipeng Luo, Dawei Xu+4.7
🔥Res3ATN -- Deep 3D Residual Attention Network for Hand gesture recognition is a strenuous task to solve in videos. In thNaina Dhingra, Andreas Kunz4.6
🔥Toward Accurate Person-level Action Recognition inDetecting and recognizing human action in videos with crowded scenes iLi Yuan, Yichen Zhou+4.4

Showing 12 of 41 papers. See ALL_PAPERS.md for all entries.

📅 2019 (27 papers)
TagTitleSummaryAuthorScore
🔥Spatio-Temporal FAST 3D Convolutions for Human ActEffective processing of video input is essential for the recognition oAlexandros Stergiou, Ronald Poppe5.8
🔥A review of Convolutional-Neural-Network-based actAbstract Video action recognition is widely applied in video indexing,Guangle Yao, Tao Lei+5.3
🔥Spatiotemporal distilled dense-connectivity networAbstract Two-stream convolutional neural networks show great promise fWangli Hao, Zhaoxiang Zhang5.3
🔥Resource Efficient 3D Convolutional Neural NetworkRecently, convolutional neural networks with 3D kernels (3D CNNs) haveOkan Köpüklü, Neslihan Kose+5.0
🔥Image and Video Compression with Neural Networks: In recent years, the image and video coding technologies have advancedSiwei Ma, Xinfeng Zhang+4.9
🔥Improving Action Recognition with the Graph-NeuralRecent human action recognition methods mainly model a two-stream or 3Wu Luo, Chongyang Zhang+4.5
🔥Multi-teacher Knowledge Distillation for CompresseRecently, convolutional neural networks (CNNs) have seen great progresMeng-Chieh Wu, C. Chiu+4.2
🔥Motion Sickness Prediction in Stereoscopic Videos In this paper, we propose a three-dimensional (3D) convolutional neuraTae Min Lee, Jong-Chul Yoon+4.2
🔥Predicting 3D Human Dynamics from VideoGiven a video of a person in action, we can easily guess the 3D futureJason Y. Zhang, Panna Felsen+4.1
🔥FBK-HUPBA Submission to the EPIC-Kitchens 2019 ActIn this report we describe the technical details of our submission toSwathikiran Sudhakaran, Sergio Escalera+4.1
📎Explainable Deep Learning for Video Recognition TaThe popularity of Deep Learning for real-world applications is ever-grLiam Hiley, Alun Preece+3.9
📎Deep 3D Convolutional Neural Network for AutomatedComputer Aided Diagnosis has emerged as an indispensible technique forSumita Mishra, Naresh Kumar Chaudhary+3.9

Showing 12 of 27 papers. See ALL_PAPERS.md for all entries.

📅 2018 (28 papers)
TagTitleSummaryAuthorScore
🔥Interpretable Spatio-temporal Attention for Video Inspired by the observation that humans are able to process videos effLili Meng, Bo Zhao+7.4
🔥Revisiting Video Saliency: A Large-scale BenchmarkIn this work, we contribute to video saliency research in two ways. FiWenguan Wang, Jianbing Shen+5.3
🔥Review of Visual Saliency Detection with ComprehenVisual saliency detection model simulates the human visual system to pRunmin Cong, Jianjun Lei+5.0
🔥Reduced-Gate Convolutional LSTM Using Predictive CSpatiotemporal sequence prediction is an important problem in deep leaNelly Elsayed, Anthony S. Maida+4.9
🔥Recurrent Convolutions for Causal 3D CNNsRecently, three dimensional (3D) convolutional neural networks (CNNs)Gurkirt Singh, Fabio Cuzzolin4.8
🔥Non-local NetVLAD Encoding for Video ClassificatioThis paper describes our solution for the 2$^\text{nd}$ YouTube-8M vidYongyi Tang, Xing Zhang+4.6
🔥Morph: Flexible Acceleration for 3D CNN-Based VideThe past several years have seen both an explosion in the use of ConvoKartik Hegde, R. Agrawal+4.5
🔥ECO: Efficient Convolutional Network for Online ViThe state of the art in video understanding suffers from two problems:Mohammadreza Zolfaghari, Kamaljeet Singh+4.4
🔥Non-Local Video Denoising by CNNNon-local patch based methods were until recently state-of-the-art forAxel Davy, Thibaud Ehret+4.4
🔥Recurrence to the Rescue: Towards Causal SpatiotemRecently, three dimensional (3D) convolutional neural networks (CNNs)Gurkirt Singh, Fabio Cuzzolin4.3
🔥Benchmark 3D eye-tracking dataset for visual salieVisual Attention Models (VAMs) predict the location of an image or vidAmin Banitalebi-Dehkordi, Eleni Nasiopoulos+4.2
🔥SlowFast Networks for Video RecognitionWe present SlowFast networks for video recognition. Our model involvesChristoph Feichtenhofer, Haoqi Fan+4.2

Showing 12 of 28 papers. See ALL_PAPERS.md for all entries.

📅 2017 (19 papers)
TagTitleSummaryAuthorScore
🔥The Monkeytyping Solution to the YouTube-8M Video This article describes the final solution of team monkeytyping, who fiHe-Da Wang, Teng Zhang+5.2
🔥A Closer Look at Spatiotemporal Convolutions for AIn this paper we discuss several forms of spatiotemporal convolutionsDu Tran, Heng Wang+5.1
🔥Predicting Video Saliency with Object-to-Motion CNOver the past few years, deep neural networks (DNNs) have exhibited grLai Jiang, Mai Xu+5.0
🔥Hierarchical Deep Recurrent Architecture for VideoThis paper introduces the system we developed for the Youtube-8M VideoLuming Tang, Boyang Deng+4.9
🔥Graph-Theoretic Spatiotemporal Context Modeling foAs an important and challenging problem in computer vision, video saliLina Wei, Fangfang Wang+4.8
🔥Video Classification With CNNs: Using The Codec AsWe investigate video classification via a two-stream convolutional neuAaron Chadha, Alhabib Abbas+4.7
🔥Facial Expression Recognition Using Enhanced Deep Deep Neural Networks (DNNs) have shown to outperform traditional methoBehzad Hasani, Mohammad H. Mahoor4.4
🔥Temporal Relational Reasoning in VideosTemporal relational reasoning, the ability to link meaningful transforBolei Zhou, A. Andonian+4.2
🔥Two-Stream 3D Convolutional Neural Network for SkeIt remains a challenge to efficiently extract spatialtemporal informatHong Liu, Juanhui Tu+4.2
🔥Spatio-Temporal Facial Expression Recognition UsinAutomated Facial Expression Recognition (FER) has been a challenging tBehzad Hasani, Mohammad H. Mahoor4.1
📎A Brief Survey of Deep Reinforcement LearningDeep reinforcement learning is poised to revolutionise the field of AIKai Arulkumaran, Marc Peter Deisenroth+3.9
📎Rethinking Spatiotemporal Feature Learning For VidRethinking Spatiotemporal Feature Learning For Video UnderstandingSaining Xie, Chen Sun+3.9

Showing 12 of 19 papers. See ALL_PAPERS.md for all entries.

📅 2016 (6 papers)
TagTitleSummaryAuthorScore
🔥Convolutional Two-Stream Network Fusion for Video Recent applications of Convolutional Neural Networks (ConvNets) for huChristoph Feichtenhofer, A. Pinz+5.0
🔥Large-Scale Shape Retrieval with Sparse 3D ConvoluIn this paper we present results of performance evaluation of S3DCNN -Alexandr Notchenko, Ermek Kapushev+4.4
📎Deep Learning for Saliency Prediction in Natural VThe purpose of this paper is the detection of salient areas in naturalSouad Chaabouni, Jenny Benois-Pineau+3.4
📎Grad-CAM: Visual Explanations from Deep Networks vWe propose a technique for producing ‘visual explanations’ for decisioRamprasaath R. Selvaraju, Abhishek Das+3.3
📎A Deep Learning Approach for Joint Video Frame andReinforcement learning is concerned with identifying reward-maximizingFelix Leibfried, Nate Kushman+2.6
📎Transfer learning with deep networks for saliency Transfer learning with deep networks for saliency prediction in naturaS. Chaabouni, J. Benois-Pineau+2.6
📅 2015 (6 papers)
TagTitleSummaryAuthorScore
🔥C3D: Generic Features for Video AnalysisWe propose a simple yet effective approach for spatiotemporal featureDu Tran, Lubomir Bourdev+5.7
🔥Intra-and-Inter-Constraint-based Video EnhancementVideo enhancement plays an important role in various video applicationYuanzhe Chen, Weiyao Lin+4.1
📎Activity Recognition Using A Combination of CategoThis paper presents a novel approach for automatic recognition of humaWeiyao Lin, Ming-Ting Sun+3.8
📎A new network-based algorithm for human activity rIn this paper, a new network-transmission-based (NTB) algorithm is proWeiyao Lin, Yuanzhe Chen+3.8
📎Group Event Detection with a Varying Number of GroThis paper presents a novel approach for automatic recognition of grouWeiyao Lin, Ming-Ting Sun+2.9
📎VoxNet: A 3D Convolutional Neural Network for realVoxNet: A 3D Convolutional Neural Network for real-time object recogniD. Maturana, S. Scherer2.8
📅 2014 (2 papers)
TagTitleSummaryAuthorScore
🔥Visualizing and Understanding Convolutional NetworLarge Convolutional Network models have recently demonstrated impressiMatthew D. Zeiler, Rob Fergus5.4
📎Deep Inside Convolutional Networks: Visualising ImThis paper addresses the visualisation of image classification models,Karen Simonyan, Andrea Vedaldi+3.7

📄 Full paper list: ALL_PAPERS.md

🏗️ Architecture

┌─────────────────────────────────────────────────────────────┐
│                    search_config.json                       │
│              (24 queries × 3 depth layers)                  │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                     Collector                               │
│         arXiv API  +  Semantic Scholar API                  │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│              Normalizer + CrossRef Enrichment               │
│        ID/version normalization, dedup, citations           │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                      Scorer                                 │
│    Keyword Match + Citations + Venue + Survey Bonus          │
└───────────────────────────┬─────────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────────┐
│                      Storage                                │
│   papers/index.jsonl + papers/quarantine.jsonl              │
└───────────────────────────┬─────────────────────────────────┘
                            │
          ┌─────────────────┼─────────────────┐
          ▼                 ▼                 ▼
    ┌──────────┐     ┌──────────┐     ┌──────────┐
    │ README   │     │ Feishu   │     │ Dashboard │
    │ (Display)│     │ (Notify) │     │  (HTML)   │
    └──────────┘     └──────────┘     └──────────┘

✨ Features

🎯 Smart Search📊 Data Enhancement🌐 Multi-Source🔔 Auto Notify
Daily arXiv searchCrossRef enrichment for new papersarXiv + Semantic ScholarFeishu Webhook
24 layered queriesCN/EN summary generationTitle dedup + ID normSuccess/Failure alerts
Score-based filteringTopic clusteringCitation + Venue boostGitHub Actions

⚙️ Auto Update

GitHub Actions aggregates arXiv and Semantic Scholar daily, then normalizes, deduplicates, scores, and enriches accepted papers with CrossRef.

📄 License

For academic research use only

Contributors

MSWEIMZ

2 commits

Languages

Python

85.8%

HTML

14.2%