Quantifying the Carbon Emissions of Machine Learning
arXiv159 reposarXiv:1910.09700
stable-diffusion-v1-4, timesfm-tourism-monthly, EleutherAI_pythia-6.9b-deduped__sft__tldr, SmolLM3-3B-QAT-Baseline-Q, KD-Tinker, MMed-Llama3.1-70B, roberta-large-mnli, distilbert-base-uncased-distilled-squad, stable-diffusion-v-1-4-original, mistral-europe_culture, mistral-northamerica_culture, mistral-news_l, Llama3-v2-iterative-DPO-iter3, Llama3-v2-iterative-DPO-iter1, mistral-reddit_l, mistral-reddit_c, Llama3-v2-iterative-DPO-iter2, mistral-africa_culture, mistral-asia_culture, mistral-southamerica_culture, mistral-news_c, mistral-news_r, mistral-reddit_r, SongGen_mixed_pro, SongGen_interleaving_A_V, flan-t5-xl, FlexCAD, stable-diffusion-x4-upscaler, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, coreml-stable-diffusion-2-1-base, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-2-base-palettized, coreml-stable-diffusion-2-1-base-palettized, coreml-stable-diffusion-2-base, stable-diffusion-2-1-base, higgs-audio-v2-tokenizer, QwQ-32B-Preview-Pruned, Marco-o1-7B-Pruned, dalle-mega, dalle-mini, flan-t5-large, sup-simcse-roberta-large, sup-simcse-bert-large-uncased, hana-alpha44, hana-small_MNIST-MODERNBERT, hana-alpha30, hana-alpha32, hana-alpha14_cifar10-128_TS-1000_1000e, hana-small_alpha8_TS-1000_400e, hana-alpha17, hana-small_alpha7, hana-small_MNIST-5e, hana-alpha45, hana-alpha42, hana-alpha39, hana-alpha38, hana-alpha31, hana-alpha29, hana-small_MNIST-MODERNBERT-3e, hana-small_alpha10, hana-alpha11_cifar10_TS-1000_500e, hana-alpha35, hana-alpha12_cifar10_TS-1000_500e, hana-small_MNIST-BATCHED, hana-alpha22, hana-alpha37, mistral-7b-hermes-rm-skywork, mistral-7b-ppo-hermes-v0.3, llama3-8b-final-ppo-v0.3, mistral-7b-hermes-dpo-v0.2, mistral-7b-hermes-cdpo-v0.2, mistral-7b-ppo-clean-hermes, prometheus-bgb-8x7b-v2.0, speecht5_tts, flan-t5-base, Memalpha-4B, stable-diffusion-2-1, llava-llama-3-8b-hqedit, ecot-openvla-7b-oxe, Mistral-7B-Instruct-v0.2-gad-bv4nogram3-merged, Mistral-7B-Instruct-v0.2-gad-slianogram3-merged, Mistral-7B-Instruct-v0.2-gad-cp8-merged, woodsoloadd_codeonly_rt_add2, wood_v2_sftr1, wood_v2_sftr4_filt, llama2-mha-from-scratch-ckpt-125M, stable-diffusion-inpainting, rad-dino, LangSAMP, GPT-J-6B-Skein, EmoWhisper-AnS-Small-v0.1, k2-v1, Llasa-1B-GRPO-2000, llama-2-7b-chat-it, llama-2-7b-chat-zh, Mistral-7B-Instruct-v0.2-es, Mixtral-8x7B-Instruct-nl, Llama-3-8B-Instruct-nl, Llama-3-8B-Instruct-it, Llama-3-8B-Instruct-es, Llama-3-8B-Instruct-fr, Llama-3-8B-Instruct-de, Llama-3-8B-Instruct-pt, Llama-3-8B-Instruct-ru, Llama-3-8B-Instruct-hi, llama-2-7b-chat-nl, llama-2-7b-chat-fr, llama-2-7b-chat-de, llama-2-7b-chat-tr, llama-2-7b-chat-bn, llama-2-7b-chat-es, llama-2-7b-chat-ru, Mistral-7B-Instruct-v0.2-nl, Mistral-7B-Instruct-v0.2-de, llama-2-13b-chat-nl, llama-2-13b-chat-es, llama-2-13b-chat-fr, llama-2-7b-chat-pt, llama-2-7b-chat-hi, llama-2-7b-chat-pl-polish-polski, llama-2-7b-chat-eu, llava-mlan-llama2-7b, llava-mlan-vicuna-7b, llava-mlan-v-llama2-7b, llava-mlan-v-vicuna-7b, quant_deepseekmath, flan-t5-small, Sentiment-Reasoning-Eng, t5-large-encoder-only-bf16, MalayaLLM-Paligemma-VQA-3B-Adapters, distilgpt2, untoken-v1, untoken-v2, bert-base-turkish-sentiment-analysis, LLaVA_MORE-llama_3_1-reasoning-finetuning, LLaVA_MORE-gemma_2_9b-siglip2-finetuning, LLaVA_MORE-gemma_2_9b-finetuning, ke-t5-base, bert-base, gemma-2b-quote-generation-82000, stable-diffusion-safety-checker, NatureLM-audio, CADFusion, Rationale_predictor, chatbot-qa-path, X-VLA-libero-object-peft, X-VLA-simpler-widowx-peft, X-VLA-libero-spatial-peft, X-VLA-libero-long-peft, X-VLA-libero-goal-peft, PredEx_Llama-2-7B_Pred-Exp_Instruction-Tuned, PredEx_Llama-2-7B_Pred-Exp, show-o-512x512, show-o-w-clip-vit-512x512, show-o-512x512-wo-llava-tuning, UniRL, stable-diffusion-v1.5-webnn
LLaMA: Open and Efficient Foundation Language Models
arXiv72 reposarXiv:2302.13971
baize-lora-7B, ik_llama.cpp, llama.cpp, stanford_alpaca, api-for-open-llm, OpenOrca, HuatuoGPT, PMC-LLaMA, RAIN, llama.cpp, open_llama_7b_v2, llama.cpp.qwen2.5vl, vicuna-7B-1.1-HF, GPTQ-for-LLaMa, openalpaca_7b_700bt_preview, openalpaca_3b_600bt_preview, Llama-Chinese, atomic-llama-cpp-turboquant, Llama-Chinese, Llama-X, Chinese-Vicuna, vicuna-13b-delta-v1.1, vicuna-13B-1.1-HF, vicuna-7b-delta-v1.1, alpaca-7b-reproduced, beaver-7b-v2.0-cost, beaver-7b-v1.0-cost, beaver-7b-unified-cost, beaver-7b-v3.0-cost, beaver-7b-v3.0-reward, beaver-7b-v1.0, beaver-7b-v3.0, beaver-7b-v2.0-reward, beaver-7b-v2.0, beaver-7b-v1.0-reward, beaver-7b-unified-reward, baize-lora-30B, baize-healthcare-lora-7b, baize-lora-13B, mpt-1b-redpajama-200b, mpt-1b-redpajama-200b-dolly, vicuna-7b-delta-v0, ClinicalNLP_PMCVQA, RRHF, calm2-7b-chat, calm2-7B-chat-GPTQ, llama2_chinese, JobList, llama.cpp, Chinese-Vicuna, vicuna-13b-v1.3, KnowLM, beaver-dam-7b, llmtools, open_llama_7b, GPlatty-30B, Platypus-30B, SuperPlatty-30B, llama-chat, llama.cpp, vicuna-7b-v1.3, llmtools, koalpaca, OpenOrca-KO, kwater, KOR-OpenOrca-Platypus, OpenOrca-gugugo-ko, medalpaca, llama2.zig, llama2.scala, llama2.java, llama2.cs
Llama 2: Open Foundation and Fine-Tuned Chat Models
arXiv54 reposarXiv:2307.09288
StableBeluga2, RAIN, eCeLLM, Llama-Chinese, Llama-Chinese, Chinese-Llama-2-7b, heron-chat-git-Llama-2-7b-v0, heron-preliminary-git-Llama-2-70b-v0, heron-chat-git-ELYZA-fast-7b-v0, ELYZA-japanese-Llama-2-7b-fast, Llama-2-7B-Chat-GGUF, Llama-2-70B-chat-AWQ, Llama-2-13B-Chat-GGUF, ELYZA-japanese-Llama-2-7b-fast-instruct, llama2_chinese, Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-GPTQ, speechless-llama2-13b, Speechless-Llama2-13B-GGML, speechless-llama2-dolphin-orca-platypus-13b, Speechless-Llama2-13B-GPTQ, speechless-llama2-hermes-orca-platypus-wizardlm-13b, Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-GGUF, Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-AWQ, speechless-llama2-hermes-orca-platypus-13b, speechless-llama2-luban-orca-platypus-13b, Speechless-Llama2-13B-GGUF, ELYZA-japanese-CodeLlama-7b-instruct, Llama-2-7B-Chat-GGML, Llama-2-13B-Chat-GGML, llama-2-7b-chat, Llama-2-7b-Chat-GPTQ, llama-2-13b-chat, Explore_llamav2_with_TGI, codellama-13b-chat, Platypus2-70B, Platypus2-70B-instruct, Camel-Platypus2-13B, Platypus2-13B, Platypus-70B-adapters, Camel-Platypus2-70B, Stable-Platypus2-13B, Platypus2-7B, Platypus-7B-adapters, Platypus-13B-adapters, corningQA-llama2-13b-chat, genai-llama2-infer-vertex, EMGerman-LLama-Streamlit-Chatbot, ContraDecode, KO-Platypus2-7B-ex, Llama-2-ko-7b-Chat, StableBeluga-7B, Llama-2-7B-GPTQ, Llama-2-13B-Chat-GPTQ, llama2.zig
LoRA: Low-Rank Adaptation of Large Language Models
arXiv52 reposarXiv:2106.09685
prompt-api, YiVal, Qwen, writing-assistance-apis, stanford_alpaca, MindSpeed-MM, deepseek-r1-finetune, Continual-NExT, RWKV-LM-LoRA, SanAssist, Llama-Chinese, Llama-Chinese, Chinese-Vicuna, quickllm, LoRA, h2o-llmstudio, adapt_med_seg, vigogne, BLOOM-LORA, colpali, colpali-v1.2, colpali-v1.1, colpali-v1.3, colqwen2-v1.0, colqwen2.5-v0.2, colqwen2.5-v0.1, colSmol-500M, colSmol-256M, colqwen2-v0.1, llama2_chinese, NTU_ADL_Team11_Final, Bloom-Lora, Chinese-Vicuna, llmtools, Platypus-30B, Platypus, CharForge, robopoint-v1-vicuna-v1.5-13b-lora, robopoint-v1-llama-2-13b-lora, robopoint-v1-llama-2-7b-lora, robopoint-v1-vicuna-v1.5-7b-lora, lit-llama, llmtools, MoA, Role-Playing-LLM-Megumin, medalpaca-lora-7b-8bit, medalpaca-lora-30b-8bit, medalpaca-lora-13b-8bit, level3_nlp_finalproject-nlp-12, topxgen, AXRU, ruvector
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
arXiv46 reposarXiv:2404.16821
InternVL2-8B-MPO, InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-2B, InternVL3_5-4B-HF, InternVL3_5-8B-HF, InternVL3_5-4B, InternVL-Chat-V1-5, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL-Chat-V1-2-Plus, InternVL3_5-8B, InternVL-Chat-V1-2, InternVL3_5-38B-HF, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-30B-A3B-HF, MMPR-v1.1, InternVL3_5-1B, InternVL-14B-224px, InternVL-Chat-V1-5-AWQ, InternVL2_5-1B, InternVL3_5-14B, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL-Chat-V1-1, MMPR, InternVL2_5-78B, InternVL3_5-1B-HF, InternVL3_5-241B-A28B-HF, InternVL2-1B, InternVL2_5-8B, Mini-InternVL-Chat-4B-V1-5, InternVL2_5-4B, InternVL2_5-2B, Mini-InternVL-Chat-2B-V1-5, InternVL2-8B, InternVL2-4B, InternVL2-2B, InternVL3-38B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, detikzify-v2-8b, InternVL2_5-26B, InternVL3-1B-hf
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
arXiv45 reposarXiv:2312.14238
InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-2B, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, InternVL-Chat-V1-5, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL-Chat-V1-2-Plus, InternVL3_5-8B, InternVL-Chat-V1-2, InternVL3_5-8B-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL3_5-14B-HF, InternVL3_5-30B-A3B-HF, MMPR-v1.1, InternVL2_5-78B, InternVL2-8B-MPO, InternVL-Chat-V1-5-AWQ, InternVL2_5-1B, InternVL3_5-14B, InternVL3_5-2B-HF, MMPR, InternVL-Chat-V1-1, InternVL-14B-224px, InternVL3_5-1B, InternVL3_5-1B-HF, InternVL3_5-38B, InternVL3_5-30B-A3B, InternVL3_5-241B-A28B-HF, InternVL2-1B, InternVL2_5-4B, InternVL2_5-2B, Mini-InternVL-Chat-2B-V1-5, InternVL2-8B, InternVL2-4B, InternVL2-2B, Mini-InternVL-Chat-4B-V1-5, InternVL2_5-8B, InternVL3-38B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, InternVL2_5-26B, InternVL3-1B-hf
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
arXiv45 reposarXiv:2412.05271
InternVL3_5-1B-HF, InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL-Chat-V1-5, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL-Chat-V1-2-Plus, InternVL3_5-8B, InternVL-Chat-V1-2, InternVL3_5-2B, InternVL3_5-8B-HF, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-30B-A3B-HF, InternVL3_5-4B-HF, MMPR-v1.1, InternVL3_5-1B, InternVL-14B-224px, InternVL-Chat-V1-5-AWQ, InternVL3_5-14B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL3_5-241B-A28B-HF, InternVL-Chat-V1-1, InternVL2_5-1B, InternVL2_5-78B, InternVL3_5-30B-A3B, InternVL2-1B, InternVL2_5-8B, Mini-InternVL-Chat-4B-V1-5, InternVL2_5-4B, InternVL2_5-2B, Mini-InternVL-Chat-2B-V1-5, InternVL2-8B, InternVL2-4B, InternVL2-2B, InternVL3-38B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, InternVL2_5-26B, OpenVid-1M, OpenVid, InternVL3-1B-hf
High-Resolution Image Synthesis with Latent Diffusion Models
arXiv42 reposarXiv:2112.10752
stable-diffusion, stable-diffusion-xl-base-1.0, stable-diffusion-v1-4, stable-diffusion-xl-1.0-inpainting-0.1, stable-diffusion-v-1-4-original, Diff-Pruning, visual-chatgpt-zh, stable-diffusion-x4-upscaler, PixArt-LCM-XL-2-1024-MS, PixArt-XL-2-256x256, PixArt-XL-2-1024-MS, PixArt-XL-2-512x512, Taiyi-Stable-Diffusion-1B-Chinese-EN-v0.1, Taiyi-Stable-Diffusion-1B-Chinese-v0.1, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, coreml-stable-diffusion-2-base-palettized, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-2-1-base, coreml-stable-diffusion-xl-base-with-refiner, PixArt-Sigma-XL-2-512-MS, coreml-stable-diffusion-2-1-base-palettized, coreml-stable-diffusion-xl-base-ios, ldm-celebahq-256, coreml-stable-diffusion-2-base, stable-diffusion-2-1-base, cool-japan-diffusion-for-learning-2-0, picasso-diffusion-1-1, diff-mining, stable-diffusion-2-1, stable-diffusion-xl-refiner-1.0, sdxl-vae, Stable-Diffusion-FineTuned-zh-v1, Stable-Diffusion-FineTuned-zh-v2, Stable-Diffusion-FineTuned-zh-v0, stable-diffusion-inpainting, stable-diffusion-uncrop, ldm3d-4c, Cover-Generator, stable-diffusion-v1.5-webnn
Building a Large Japanese Web Corpus for Large Language Models
arXiv42 reposarXiv:2404.17733
Swallow-70b-hf, Swallow-70b-NVE-hf, Swallow-7b-NVE-hf, Swallow-MS-7b-v0.1, Swallow-7b-NVE-instruct-hf, Llama-3-Swallow-70B-v0.1, Swallow-70b-instruct-hf, Swallow-7b-plus-hf, Swallow-70b-NVE-instruct-hf, Llama-3.1-Swallow-8B-v0.1, Swallow-7b-hf, Swallow-MX-8x7b-NVE-v0.1, Llama-3-Swallow-8B-v0.1, Gemma-2-Llama-Swallow-9b-pt-v0.1, Llama-3.1-Swallow-70B-v0.1, Swallow-7b-instruct-hf, Llama-3.1-Swallow-8B-v0.2, Llama-3.3-Swallow-70B-v0.4, Swallow-13b-hf, Swallow-13b-instruct-hf, Swallow-13b-NVE-hf, Gemma-2-Llama-Swallow-2b-pt-v0.1, Gemma-2-Llama-Swallow-27b-pt-v0.1, Qwen3-Swallow-32B-SFT-v0.2, Qwen3-Swallow-30B-A3B-CPT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2, Llama-3.1-Swallow-8B-v0.5, Qwen3-Swallow-32B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2, GPT-OSS-Swallow-20B-RL-v0.1, GPT-OSS-Swallow-120B-SFT-v0.1, Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-20B-SFT-v0.1, Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-8B-SFT-v0.2, GPT-OSS-Swallow-120B-RL-v0.1-MXFP4, Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-20B-RL-v0.1-MXFP4, Qwen3-Swallow-8B-CPT-v0.2, Qwen3-Swallow-30B-A3B-SFT-v0.2, GPT-OSS-Swallow-120B-RL-v0.1, Qwen3-Swallow-32B-CPT-v0.2
Qwen3 Technical Report
arXiv41 reposarXiv:2505.09388
Qwen3-30B-A3B, Qwen3-32B, Qwen3-235B-A22B, Qwen3-Coder-30B-A3B-Instruct, Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-4B-Thinking-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-Coder-480B-A35B-Instruct, Qwen3-Coder-30B-A3B-Instruct-FP8, Qwen3-Coder-480B-A35B-Instruct-FP8, Qwen3-0.6B-GPTQ-Int8, Qwen3-4B-AWQ, gliner-stream-pii-v1.0, qwen3-8b-base, SDAR-8B-Chat-b16, SDAR-8B-Chat-b32, SDAR-8B-Chat-b64, SDAR, SDAR-30B-A3B-Sci, SDAR-4B-Chat, SDAR-8B-Chat, SDAR-1.7B-Chat, SDAR-30B-A3B-Chat, Qwen3-Next-80B-A3B-Instruct, Qwen3-0.6B, Qwen3-Swallow-8B-SFT-v0.2, Qwen3-Swallow-32B-SFT-v0.2, Qwen3-Swallow-30B-A3B-CPT-v0.2, Medical-Qwen3-Swallow-32B, Qwen3-Swallow-30B-A3B-SFT-v0.2, Qwen3-Swallow-32B-CPT-v0.2, Qwen3-Swallow-32B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4, Medical-Qwen3-Swallow-30B-A3B, Medical-Qwen3-Swallow-8B, Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-8B-CPT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2
YaRN: Efficient Context Window Extension of Large Language Models
arXiv40 reposarXiv:2309.00071
Yarn-Mistral-7b-64k, Qwen2.5-VL, Qwen3-VL, QwQ, Qwen3-30B-A3B, Qwen2.5-72B-Instruct, Qwen2.5-32B-Instruct, Qwen3-32B, Qwen3-235B-A22B, Qwen2.5-Coder-32B-Instruct, QwQ-32B, Qwen2.5-Coder-14B-Instruct, Qwen3-Coder-Next-GGUF, Qwen2.5-VL-72B-Instruct, Llama-3-8B-Instruct-Gradient-1048k, Qwen2-VL, Qwen2.5-VL-3B-Instruct-GGUF, Qwen2-72B-Instruct, Qwen3-4B-AWQ, Qwen2.5-14B-Instruct, Llama-3-70B-Instruct-Gradient-1048k, Llama-3-70B-Instruct-Gradient-262k, Llama-3-8B-Instruct-262k, Llama-3-8B-Instruct-Gradient-4194k, Qwen2-57B-A14B-Instruct, Chinese-LLaMA-Alpaca-2, Qwen2.5-VL-7B-Instruct-GGUF, Yarn-Llama-2-13b-128k, Yarn-Llama-2-7b-128k, yarn, Yarn-Llama-2-13b-64k, Yarn-Llama-2-70b-32k, Yarn-Llama-2-7b-64k, Yarn-Solar-10b-64k, Yarn-Mistral-7b-128k, Yarn-Solar-10b-32k, Qwen3-Next-80B-A3B-Instruct, Qwen2.5-14B-Instruct-AWQ, Qwen2.5-VL-7B-Instruct-GGUF, Qwen2.5-Coder-7B-Instruct
Learning Transferable Visual Models From Natural Language Supervision
arXiv36 reposarXiv:2103.00020
CLIP, clip-vit-base-patch32, stable-diffusion-v1-4, clip-ViT-B-32, stable-diffusion-v-1-4-original, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, contrastors, Semantic-Segment-Anything, contrastors, tf-transformers, clip-italian, clip-italian, clip-image-search, BiomedCoOp, clip-vit-base-patch16, LAVIS, ml-mobileclip, stable-diffusion-inpainting, visual-spatial-reasoning, CLIP-Chinese, Multilingual-CLIP, turkish-clip, vilmedic, prismatic-vlms, BiomedVLP-CXR-BERT-specialized, vit_large_patch14_clip_224.openai, stable-diffusion-safety-checker, TiViT, ru-clip, ru-clip, vit-dog, align-text-encoders, stable-diffusion-v1.5-webnn
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
arXiv30 reposarXiv:1810.04805
bert-base-uncased, bert, bert-base-NER, bert-base-chinese, bert-base-cased, bert-base-multilingual-cased, DiagnosisCoding, EMS-Pipeline, tf-transformers, legalbert-large-1.7M-1, legalbert-large-1.7M-2, europeana-bert, dllm, MCBG, Malware_survey_my_experiments, CodeXGLUE, Vorbereitung, Bert-Implementation-From-Scratch, Word-Embeddings-Repository-for-Turkish, BertWithPretrained, BertWithPretrained, bangla-bert-base, bangla-bert, kcbert-base, KcBERT, BERTify, IndicAbusive, rudetoxifier, ASR-Knowledge-Transferring, caption
Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
arXiv30 reposarXiv:2505.02881
Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0002500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0012500, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0005000, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0002500, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0002500, Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0010000, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0012500, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0010000, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0010000, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0002500, Llama-3.3-Swallow-70B-v0.4, Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0002500, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0007500, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0010000, swallow-code, swallow-code-v2, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0012500, Llama-3.1-Swallow-8B-v0.5, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0012500, swallow-math, swallow-math-v2, swallow-code-math
On the Dangers of Stochastic Parrots
ACM28 reposACM:3442188.3445922
hacker-laws, roberta-large-mnli, distilbert-base-uncased-distilled-squad, transfo-xl-wt103, bert-base-chinese, idefics-80b-instruct, sup-simcse-roberta-large, sup-simcse-bert-large-uncased, geolm-base-toponym-recognition, cosmo-xl, idefics2-8b, legal-xlm-longformer-base, legal-xlm-roberta-large, legal-xlm-roberta-base, idefics-9b-instruct, distilgpt2, legal-english-roberta-base, legal-croatian-roberta-base, legal-swiss-roberta-large, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large, bert-base-uncased-hatexplain-rationale-two, ke-t5-base, bert-base, stable-diffusion-safety-checker, opus-mt-ru-en
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
arXiv28 reposarXiv:1910.10683
ChatLM-mini-Chinese, t5x, chronos-t5-base, c4, SwissArmyTransformer, t5-v1_1-xxl, chronos-bolt-base, chronos-bolt-tiny, chronos-bolt-mini, chronos-bolt-small, chronos-t5-tiny, chronos-t5-small, chronos-t5-mini, chronos-t5-large, google_t5-v1_1-xxl_encoderonly, Glot500, tk-instruct-11b-def-pos, Tk-Instruct, tf-transformers, turkish-bert, tk-instruct-3b-def, nano_flan_t5, text-to-text-transfer-transformer, tk-instruct-11b-def, ke-t5, level2-nlp-datacentric-nlp-03, russe_detox_2022, russe_detox_2022
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
arXiv28 reposarXiv:2306.00978
awq-embed, lorax, ScaleLLM, nunchaku, llm-awq, AutoAWQ, VILA, efficient-transformers, InternVL-Chat-V1-5-AWQ, vllm, vllm, nunchaku, vllm-turboquant, trtllm, TensorRT-LLM, smash, dgx-spark-quantization, Qwen3-VL-2B-GRACE-W4G128-AWQ, awq4nvomni, llm-awq, lmdeploy-v100, vllm-old, vllm_amd_sleep, vllm, higgs-audio-vllm, lmdeploy-dev, cosyvoice3-inference-acceleration, lmdeploy
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
arXiv28 reposarXiv:2411.10442
InternVL3_5-30B-A3B-HF, InternVL3_5-1B-HF, InternVL3-2B, xtuner, InternVL, MMPR-Tiny, MMPR-v1.2, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL3_5-8B, InternVL3_5-8B-HF, InternVL3_5-2B, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, MMPR-v1.1, InternVL2-8B-MPO, InternVL3_5-1B, InternVL3_5-14B, MMPR, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL3_5-241B-A28B-HF, InternVL3-38B, InternVL3-1B-hf
DFlash: Block Diffusion for Flash Speculative Decoding
arXiv27 reposarXiv:2602.06036
Muse-Glimmer-30B, SpecForge, dflash, Qwen3-4B-DFlash-b16, DeepSpec, tpu-spec-decode, dflash-b4d1109c, dflash_benchmark, Kimi-K2.5-DFlash, Qwen3.6-27B-DFlash, Qwen3.6-35B-A3B-DFlash, gemma-4-26B-A4B-it-DFlash, gemma-4-31B-it-DFlash, MiniMax-M2.5-DFlash, Qwen3.5-27B-DFlash, Qwen3.5-122B-A10B-DFlash, Qwen3.5-35B-A3B-DFlash, gpt-oss-20b-DFlash, gpt-oss-120b-DFlash, Qwen3-Coder-30B-A3B-DFlash, Qwen3.5-9B-DFlash, Qwen3.5-4B-DFlash, Qwen3-8B-DFlash-b16, Qwen3-Coder-Next-DFlash, LLaMA3.1-8B-Instruct-DFlash-UltraChat, dflasher, dflash_eval
Attention Is All You Need
arXiv26 reposarXiv:1706.03762
speech-emotion-recognition, atomic-agents, pi-web-agent, bert, graphify, NeMo-Agent-Toolkit-Examples, nature-skills, prompt-engineering, onnxruntime-training-examples, minGPT, document-to-podcast, dissertation-project, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, Speech-Transformer, nmt, nugie-jax-nemotron-3-nano, tiny-vllm, canary-1b, diar_sortformer_4spk-v1, transformer-tricks, indonesian-language-models, LaTeX-OCR, NAVER_AIRUSH_Grammar_Error_Correction, llama2.zig, llama2.go
ColPali: Efficient Document Retrieval with Vision Language Models
arXiv26 reposarXiv:2407.01449
vidore-benchmark, colpali, colpali, colpali-v1.2, colpali-v1.1, colpali-v1.3, tomoro-colqwen3-embed-4b, colqwen2-v1.0, colqwen2.5-v0.1, colqwen2.5-v0.2, colSmol-256M, colSmol-500M, colpali, colpali, colqwen2-v0.1, colpali_train_set, SampadaEmbed, vllm-factory, esg_reports_v2, biomedical_lectures_v2, economics_reports_v2, esg_reports_human_labeled_v2, multi-modal-rag-with-colpali, colpali_t, colpali-tutorial, Colpali-Reproducibility
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
arXiv25 reposarXiv:2406.12793
GLM-Z1-32B-0414, glm-4-9b-chat, GLM-4, glm-4-9b, glm-4-9b-chat-1m, GLM-Z1-32B-0414, glm-4-9b-chat-hf, glm-4-9b-chat-1m-hf, glm-4v-9b, glm-4v-9b, chatglm2-6b-32k, chatglm2-6b, chatglm-6b-int8, chatglm3-6b-32k, chatglm-6b, chatglm2-6b-int4, chatglm-6b-int4, GLM-Z1-9B-0414, chatglm3-6b-32k, glm-4-9b-chat-1m, chatglm3-6b-base, GLM-4, chatglm-6b-int8, chatglm-6b-int4, visualglm-6b
Code Llama: Open Foundation Models for Code
arXiv24 reposarXiv:2308.12950
CodeLlama-7b-Instruct-hf, speechless-codellama-34b-v2.0, speechless-codellama-34b-v2.0-GGUF, speechless-codellama-34b-v2.0-GPTQ, speechless-codellama-34b-v1.0, speechless-codellama-34b-v2.0-AWQ, speechless-codellama-orca-13b, speechless-codellama-airoboros-orca-platypus-13b, speechless-codellama-dolphin-orca-platypus-34b, speechless-codellama-platypus-13b, CodeLlama-7b-hf, ELYZA-japanese-CodeLlama-7b-instruct, CodeLlama-7B-Instruct-GPTQ, CodeLlama-7B-GGUF, CodeLlama-7b-Python-hf, CodeLlama-34b-Python-hf, CodeLlama-13b-hf, CodeLlama-13b-Python-hf, CodeLlama-13b-Instruct-hf, CodeLlama-34b-hf, CodeLlama-34b-Instruct-hf, CodeLlama-70b-hf, CodeLlama-70b-Python-hf, CodeLlama-70b-Instruct-hf
EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test
arXiv24 reposarXiv:2503.01840
SpecForge, EAGLE-llama2-chat-7B, EAGLE-LLaMA3-Instruct-70B, EAGLE-Qwen2-7B-Instruct, DeepSpec, EAGLE, EAGLE3-Vicuna1.3-13B, EAGLE3-LLaMA3.1-Instruct-8B, EAGLE3-LLaMA3.3-Instruct-70B, EAGLE3-DeepSeek-R1-Distill-LLaMA-8B, EAGLE-llama2-chat-13B, EAGLE-Vicuna-7B-v1.3, GLM-4.7-Flash-Eagle3, EAGLE-Vicuna-33B-v1.3, EAGLE-Vicuna-13B-v1.3, EAGLE-LLaMA3-Instruct-8B, EAGLE-llama2-chat-70B, EAGLE-mixtral-instruct-8x7B, EAGLE-Qwen2-72B-Instruct, EAGLE-LLaMA3.1-Instruct-8B, tpu-spec-decode, quantized-eagle, EAGLE, Enhanced-Eagle
Language Models are Few-Shot Learners
arXiv23 reposarXiv:2005.14165
awesome-prompt-engineering, ik_llama.cpp, llama.cpp, lm-evaluation-harness, llm.c, llm.c, mmlu, llama.cpp, promptsource, falcon-40b, falcon-7b, falcon-rw-1b, llama.cpp.qwen2.5vl, atomic-llama-cpp-turboquant, lm-evaluation-harness, opt-125m, falcon-40b-instruct, opt-13b, opt-2.7b, llama.cpp, opt-350m, relbert, llama.cpp
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
arXiv23 reposarXiv:2210.17323
deepcompressor, ChatGLM-6B, glq, ChatGLM-6B, text-generation-inference, lorax, ScaleLLM, llm-awq, gptq, PodGPT, efficient-transformers, vllm, vllm, vllm-turboquant, smash, GPTQ-for-LLaMa, awq-embed, awq4nvomni, llm-awq, vllm-old, vllm_amd_sleep, vllm, higgs-audio-vllm
MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
arXiv23 reposarXiv:2508.02669
MedVLSynther-7B-RL_5K_internvl-glm, MedVLSynther-7B-RL_5K_no-verify, MedVLThinker-3B-SFT_5K, MedVLThinker, MedVLSynther-3B-RL_1K, MedVLSynther-3B-RL_2K, MedVLSynther-3B-RL_10K, MedVLSynther-3B-RL_5K_qwen-glm, MedVLSynther-3B-RL_5K_glm-glm, MedVLSynther-3B-RL_5K, MedVLSynther-3B-RL_5K_internvl-glm, MedVLSynther-3B-RL_13K, MedVLSynther-7B-RL_2K, MedVLSynther-7B-RL_10K, MedVLSynther-7B-RL_5K_qwen-glm, MedVLSynther-7B-RL_1K, MedVLSynther-3B-RL_5K_no-verify, MedVLSynther-7B-RL_5K, MedVLSynther-3B-RL_5K_PMC-style, MedVLSynther-7B-RL_13K, MedVLSynther-7B-RL_5K_glm-glm, MedVLSynther-7B-RL_5K_PMC-style, MedVLThinker-7B-SFT_5K
Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
arXiv23 reposarXiv:2510.25867
MedVLSynther-7B-RL_5K_internvl-glm, MedVLSynther-7B-RL_5K_no-verify, MedVLSynther, MedVLSynther-3B-RL_1K, MedVLSynther-3B-RL_2K, MedVLSynther-3B-RL_10K, MedVLSynther-3B-RL_5K_qwen-glm, MedVLSynther-3B-RL_5K_glm-glm, MedVLSynther-3B-RL_5K, MedVLSynther-3B-RL_5K_internvl-glm, MedVLSynther-3B-RL_13K, MedVLSynther-7B-RL_2K, MedVLSynther-7B-RL_10K, MedVLSynther-7B-RL_5K_qwen-glm, MedVLSynther-7B-RL_1K, MedVLSynther-3B-RL_5K_no-verify, MedVLSynther-7B-RL_5K, MedVLSynther-3B-RL_5K_PMC-style, MedVLSynther-7B-RL_13K, MedVLSynther-7B-RL_5K_glm-glm, MedVLThinker-3B-SFT_5K, MedVLSynther-7B-RL_5K_PMC-style, MedVLThinker-7B-SFT_5K
Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
arXiv22 reposarXiv:2410.16261
InternVL, InternVL-Chat-V1-5, InternVL-Chat-V1-2-Plus, InternVL-Chat-V1-2, InternVL2_5-1B, InternVL-Chat-V1-1, InternVL-Chat-V1-5-AWQ, InternVL-14B-224px, InternVL2_5-78B, InternVL2-1B, InternVL2_5-8B, InternVL2_5-4B, Mini-InternVL-Chat-4B-V1-5, Mini-InternVL-Chat-2B-V1-5, InternVL2_5-2B, InternVL2-8B, InternVL2-2B, InternVL2-4B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, InternVL2_5-26B
TimeLMs: Diachronic Language Models from Twitter
arXiv21 reposarXiv:2202.03829
tweeteval, twitter-roberta-base-sentiment-latest, twitter-roberta-base-2021-124m, timelms, twitter-roberta-base-2019-90m, twitter-roberta-base-dec2020, twitter-roberta-base-jun2021, twitter-roberta-base-mar2020, twitter-roberta-base-jun2020, twitter-roberta-base-sep2020, twitter-roberta-base-mar2021, twitter-roberta-base-dec2021, twitter-roberta-base-jun2022, twitter-roberta-base-sep2021, twitter-roberta-base-mar2022, twitter-roberta-base-mar2022-15M-incr, twitter-roberta-base-jun2022-15M-incr, twitter-roberta-base-2022-154m, twitter-roberta-base-sep2022, twitter-roberta-large-2022-154m, tgbot-hate-speech
WizardLM: Empowering large pre-trained language models to follow complex instructions
arXiv21 reposarXiv:2304.12244
WizardLM_evol_instruct_70k, WizardCoder-15B-V1.0, WizardLM-70B-V1.0, EasyInstruct, WizardMath-7B-V1.1, WizardLM-70B-V1.0, WizardMath-70B-V1.0, WizardCoder-Python-13B-V1.0, WizardCoder-15B-V1.0, WizardMath-7B-V1.0, WizardLM-13B-V1.0, WizardCoder-Python-34B-V1.0, WizardLM_evol_instruct_V2_196k, WizardMath-7B-V1.1, WizardCoder-33B-V1.1, WizardLM-13B-V1.2, WizardLM_evol_instruct_70k, slm-innovator-lab, evolve-instruct, ko-instruction-dataset, Ko.WizardLM_evol_instruct_V2_196k
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
arXiv21 reposarXiv:2401.15077
EAGLE-llama2-chat-7B, EAGLE, EAGLE3-LLaMA3.3-Instruct-70B, EAGLE3-Vicuna1.3-13B, EAGLE3-LLaMA3.1-Instruct-8B, EAGLE3-DeepSeek-R1-Distill-LLaMA-8B, EAGLE-llama2-chat-13B, EAGLE-LLaMA3-Instruct-70B, EAGLE-Qwen2-7B-Instruct, EAGLE-Vicuna-7B-v1.3, EAGLE-Vicuna-33B-v1.3, EAGLE-Vicuna-13B-v1.3, EAGLE-llama2-chat-70B, EAGLE-mixtral-instruct-8x7B, EAGLE-LLaMA3-Instruct-8B, EAGLE-Qwen2-72B-Instruct, EAGLE-LLaMA3.1-Instruct-8B, quantized-eagle, EAGLE, Enhanced-Eagle, mlx-flash
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
arXiv21 reposarXiv:2402.03300
minimind, tunix, MedVLM-R1, AntAngelMed, AntAngelMed, Qwen2.5-3B-GRPO-MATH-1EPOCH, Qwen2.5-1.5B-GRPO-MATH-1EPOCH, async-grpo, xtuner, deepseek-math-7b-rl, deepseek-math-7b-base, MM-EUREKA, CPGD-7B, MedicalGPT, detikzify-v2.5-8b, SkinTokens, SkinTokens, marcello, rl-unsloth, chart-rvr-3b, chart-rvr-hard-3b
Chronos: Learning the Language of Time Series
arXiv21 reposarXiv:2403.07815
chronos-t5-base, sundial-base-128m, uni2ts, chronos-forecasting, chronos-t5-large, chronos-bolt-base, chronos-bolt-tiny, chronos-bolt-mini, chronos-2, chronos-bolt-small, chronos-t5-tiny, chronos-t5-mini, moirai-2.0-R-small, chronos-t5-small, stock_forecasting, Samay, A2TTA, moment, patchtst-fm-r1, granite-timeseries-patchtst-fm-r1, uni2ts
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
arXiv21 reposarXiv:2504.10479
InternVL3_5-30B-A3B-HF, InternVL3_5-1B-HF, InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL3_5-8B, InternVL3_5-8B-HF, InternVL3_5-2B, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, InternVL3_5-14B, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL3_5-241B-A28B-HF, InternVL3_5-1B, InternVL3-38B, InternVL3-1B-hf
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
arXiv21 reposarXiv:2508.01691
voxlect-mandarin-cantonese-dialect-whisper-small, voxlect-indic-lid-whisper-small, voxlect-german-dialect-whisper-small, voxlect-german-dialect-whisper-large-v3, voxlect-spanish-dialect-mms-lid-256, voxlect-english-dialect-mms-lid-256, voxlect-indic-lid-whisper-large-v3, voxlect-french-dialect-mms-lid-256, voxlect-thai-dialect-mms-lid-256, voxlect-spanish-dialect-whisper-small, voxlect-mandarin-cantonese-dialect-mms-lid-256, voxlect-english-dialect-whisper-large-v3, voxlect-indic-lid-mms-lid-256, voxlect-spanish-dialect-whisper-large-v3, voxlect-german-dialect-mms-lid-256, voxlect-thai-dialect-whisper-small, voxlect-thai-dialect-whisper-large-v3, voxlect-mandarin-cantonese-dialect-whisper-large-v3, voxlect-french-dialect-whisper-small, voxlect-english-dialect-whisper-small, voxlect-french-dialect-whisper-large-v3
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
arXiv20 reposarXiv:2004.12832
RAG-RLRC-LaySum, colqwen2.5-v0.2, docs-reference, bge-m3, sweet-search, colpali, colpali-v1.2, colpali-v1.1, colpali-v1.3, colqwen2-v1.0, colqwen2.5-v0.1, colSmol-500M, colSmol-256M, colqwen2-v0.1, plaidrepro, colbertv2.0, ColBERT, odqa_baseline_code, jina-colbert-v1-en, KolBERT
Efficient Memory Management for Large Language Model Serving with PagedAttention
arXiv20 reposarXiv:2309.06180
vllm, flash-attention, vllm, flash-attention, lorax, vllm, vllm, vllm, vllm-turboquant, vllm, tiny-vllm, vllm-old, vllm_amd_sleep, vllm, higgs-audio-vllm, vllm, flash-attention, flash-attention, KsanaLLM, vllm
Qwen Technical Report
arXiv20 reposarXiv:2309.16609
Qwen, Qwen-7B-Chat, Qwen-14B-Chat, Qwen1.5-0.5B, Qwen1.5-7B-Chat, Qwen1.5-7B, Qwen1.5-14B, Qwen1.5-14B-Chat, Qwen1.5-0.5B-Chat, Qwen1.5-1.8B, Qwen1.5-1.8B-Chat, Qwen-7B, Qwen-1_8B, CodeQwen1.5-7B-Chat, Qwen1.5-110B-Chat, Qwen1.5-32B-Chat, Qwen1.5-4B-Chat, Qwen1.5-72B, Qwen1.5-32B-Chat-GPTQ-Int4, Qwen1.5-4B
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
arXiv20 reposarXiv:2405.18991
EasyAnimate, EasyAnimateV5.1-7b-zh, EasyAnimateV5.1-12b-zh, EasyAnimateV5.1-7b-zh-Control, EasyAnimateV3-XL-2-InP-512x512, EasyAnimateV2-XL-2-512x512, EasyAnimateV2-XL-2-768x768, EasyAnimateV4-XL-2-InP, EasyAnimateV5-7b-zh, EasyAnimateV5.1-12b-zh-InP, EasyAnimateV5.1-7b-zh-diffusers, EasyAnimateV5.1-12b-zh-Control-Camera, EasyAnimateV5.1-7b-zh-Control-Camera, EasyAnimateV3-XL-2-InP-960x960, EasyAnimateV5.1-7b-zh-InP, EasyAnimateV3-XL-2-InP-768x768, EasyAnimateV5.1-12b-zh-Control, EasyAnimateV5-12b-zh, EasyAnimateV5-7b-zh-InP, EasyAnimateV5-12b-zh-Control
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
arXiv20 reposarXiv:2406.16858
EAGLE-llama2-chat-7B, EAGLE-Qwen2-7B-Instruct, EAGLE, EAGLE3-Vicuna1.3-13B, EAGLE3-LLaMA3.1-Instruct-8B, EAGLE3-LLaMA3.3-Instruct-70B, EAGLE3-DeepSeek-R1-Distill-LLaMA-8B, EAGLE-llama2-chat-13B, EAGLE-LLaMA3-Instruct-70B, EAGLE-Vicuna-7B-v1.3, EAGLE-Vicuna-33B-v1.3, EAGLE-Vicuna-13B-v1.3, EAGLE-llama2-chat-70B, EAGLE-mixtral-instruct-8x7B, EAGLE-LLaMA3-Instruct-8B, EAGLE-Qwen2-72B-Instruct, EAGLE-LLaMA3.1-Instruct-8B, quantized-eagle, EAGLE, Enhanced-Eagle
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
arXiv20 reposarXiv:2412.02595
NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-3-Nano-4B-GGUF, Qwen3-Swallow-32B-SFT-v0.2, Qwen3-Swallow-30B-A3B-CPT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2, Qwen3-Swallow-30B-A3B-SFT-v0.2, Qwen3-Swallow-32B-CPT-v0.2, Qwen3-Swallow-32B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2, GPT-OSS-Swallow-120B-RL-v0.1, GPT-OSS-Swallow-20B-RL-v0.1, GPT-OSS-Swallow-20B-SFT-v0.1, GPT-OSS-Swallow-120B-SFT-v0.1, Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-8B-SFT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-120B-RL-v0.1-MXFP4, Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-20B-RL-v0.1-MXFP4, Qwen3-Swallow-8B-CPT-v0.2
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
arXiv20 reposarXiv:2504.11651
BAGEL-7B-MoT-DF11, DFloat11, Bagel-DFloat11, DeepSeek-R1-Distill-Qwen-32B-DF11, DeepSeek-R1-Distill-Qwen-14B-DF11, Llama-3.1-8B-Instruct-DF11, gemma-3-12b-it-DF11, FLUX.1-Krea-dev-DF11, FLUX.1-dev-DF11, DeepSeek-R1-Distill-Llama-8B-DF11, Wan2.1-T2V-14B-Diffusers-DF11, Qwen3-32B-DF11, Qwen3-4B-DF11, Qwen3-14B-DF11, DeepSeek-R1-Distill-Qwen-7B-DF11, Phi-4-reasoning-plus-DF11, gemma-3-4b-it-DF11, Qwen3-8B-DF11, gemma-3-27b-it-DF11, BAGEL-DFloat11-Windows
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
arXiv20 reposarXiv:2508.18265
InternVL3_5-30B-A3B-HF, MMPR-Tiny, MMPR-v1.2, InternVL3_5-4B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-38B-HF, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL3_5-8B-HF, InternVL3_5-8B, InternVL3_5-2B, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, InternVL3_5-1B, InternVL3_5-14B, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-241B-A28B-HF, InternVL3_5-1B-HF, InternVL3_5-38B
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
arXiv19 reposarXiv:2006.03654
LoRA, LEXTREME, mdeberta-v3-base, DeBERTa, deberta-v3-xsmall, deberta-v2-xxlarge, deberta-large, deberta-base-mnli, deberta-v2-xlarge, deberta-v2-xxlarge-mnli, deberta-v2-xlarge-mnli, deberta-v3-small, deberta-base, deberta-xlarge, deberta-v3-base, deberta-v3-large, deberta-xlarge-mnli, deberta-large-mnli, DeBERTa_TxtClassifier
RoFormer: Enhanced Transformer with Rotary Position Embedding
arXiv19 reposarXiv:2104.09864
falcon-40b, roformer_v2_chinese_char_base, Baichuan-13B-Chat, falcon-7b, roformer_chinese_base, Baichuan-13B-Base, nomic-bert-2048, cholesky_encoder, snowflake-arctic-embed-m-long, arctic-embed, stablecode-completion-alpha-3b-4k, mesh-transformer-jax, falcon-40b-instruct, tiny-vllm, polyglot-ko-1.3b, polyglot-ko-3.8b, polyglot-ko-5.8b, llama2.zig, llama2.go
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
arXiv19 reposarXiv:2406.08464
magpie, Magpie-Reasoning-150K, Magpie-Qwen2-Air-3M-v0.1, Magpie-Qwen2.5-Pro-1M-v0.1, Magpie-Llama-3.1-Pro-DPO-100K-v0.1, Magpie-Llama-3.3-Pro-1M-v0.1, Llama-3-8B-Magpie-Align-v0.2, Magpie-Pro-DPO-100K-v0.1, Llama-3-8B-Magpie-Align-v0.3, Magpie-Air-DPO-100K-v0.1, Magpie-Llama-3.1-Pro-1M-v0.1, Llama-3-8B-Magpie-Align-v0.1, Magpie-Qwen2-Pro-1M-v0.1, Magpie-Qwen2-Pro-200K-Chinese, aurora-m2, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-70B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.2, swallow-gemma-magpie-v0.1
Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
arXiv19 reposarXiv:2504.09895
zephyr-7b-alpha-conf-sft, RefAlign, NQ-Subset-500, Llama-3.3-70B-Inst-awq_ultrafeedback_1in3, llama3-ultrafeedback-bertscore-bart-large-mnli, alpaca-7b-ref-meteor, Llama-2-13b-hf-conf-refalign, Mistral-7B-v0.1-conf-sft, Llama-3.3-70B-Inst-awq_SafeRLHF, alpaca-7b-ref-bertscore, Llama-2-7b-hf-conf-refalign, Llama-2-7b-hf-conf-sft, Mistral-7B-v0.1-conf-refalign, Llama-2-13b-hf-conf-sft, zephyr-7b-alpha-conf-refalign, Mistral-7B-Instruct-v0.2-ref-simpo, Mistral-7B-Instruct-v0.2-refalign, Llama-3-8B-Instruct-ref-simpo, Llama-3-8B-Instruct-refalign
Scaling Instruction-Finetuned Language Models
arXiv18 reposarXiv:2210.11416
flan-t5-xl, flan-t5-large, pdfai-back, flan-t5-base, optimized-parler-tts, flan-alpaca-gpt4-xl, flan-alpaca, flan-alpaca-large, flan-alpaca-base, flan-alpaca-xl, flan-alpaca-xxl, flan-gpt4all-xl, flan-sharegpt-xl, blip2-flan-t5-xl, blip2-flan-t5-xxl, flan-t5-small, flan-ul2-dolly, flan-ul2-dolly-lora
Robust Speech Recognition via Large-Scale Weak Supervision
arXiv18 reposarXiv:2212.04356
whisper, whisper-large-v3, whisper-large-v3-turbo, whisper-medium, whisper-tiny, kotoba-whisper-v2.0, kotoba-whisper-v1.0, dissertation-project, whisper-jax, whisper-tiny.en, whisper-base, whisper-large, whisper-small.en, whisper-small, whisper-medium.en, whisperspeech, final-project-level3-nlp-01, whisper-base-webnn
WizardCoder: Empowering Code Large Language Models with Evol-Instruct
arXiv18 reposarXiv:2306.08568
evalplus, WizardLM_evol_instruct_70k, WizardCoder-15B-V1.0, WizardLM-70B-V1.0, WizardMath-7B-V1.1, WizardLM-70B-V1.0, WizardMath-7B-V1.0, WizardLM-13B-V1.2, WizardCoder-Python-13B-V1.0, WizardCoder-15B-V1.0, WizardMath-70B-V1.0, WizardLM-13B-V1.0, WizardCoder-Python-34B-V1.0, WizardMath-7B-V1.1, WizardLM_evol_instruct_70k, WizardLM_evol_instruct_V2_196k, WizardCoder-33B-V1.1, Ko.WizardLM_evol_instruct_V2_196k
Improved Baselines with Visual Instruction Tuning
arXiv18 reposarXiv:2310.03744
LLaVA, minimind-v, M4-Instruct-Data, llava-v1.6-34b-hf, llava-v1.6-vicuna-13b-hf, llava-v1.6-vicuna-7b-hf, llava-v1.6-mistral-7b-hf, libra-llava-rad, llava-rad, table-llava-v1.5-13b, table-llava-v1.5-7b, table-llava-v1.5-7b-hf, llava-bench-in-the-wild, llava-1.5-665k-instructions, TriPlaneLLaVA, LLaVA-toy, multi_token, TinyLLava
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
arXiv18 reposarXiv:2404.05961
DermL2V-tmp, llm2vec, DermL2V-training, LLM2Vec-Llama-2-7b-chat-hf-mntp, LLM2Vec-Mistral-7B-Instruct-v2-unsup-simcse, LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervised, LLM2Vec-Llama-2-7b-chat-hf-mntp-supervised, LLM2Vec-Llama-2-7b-chat-hf-unsup-simcse, LLM2Vec-Sheared-LLaMA-mntp-supervised, LLM2Vec-Meta-Llama-3-8B-Instruct-mntp, LLM2Vec-Mistral-7B-Instruct-v2-mntp, LLM2Vec-Sheared-LLaMA-unsup-simcse, LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-unsup-simcse, LLM2Vec-Sheared-LLaMA-mntp, LLM2Vec-Mistral-7B-Instruct-v2-mntp-supervised, NetTAG, llm-idiosyncrasies, bach-or-bot
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
arXiv18 reposarXiv:2406.07476
VideoLLaMA2, VideoLLaMA2.1-7B-16F-Base, Multi-Source-Video-Captioning, VideoLLaMA2.1-7B-AV, VideoLLaMA2-72B, VideoLLaMA2.1-7B-16F, VideoLLaMA2-72B-Base, VideoLLaMA2-7B, VideoLLaMA2-8x7B-Base, VideoLLaMA2-8x7B, VideoLLaMA2-7B-16F, VideoLLaMA2-7B-Base, VideoLLaMA2-7B-16F-Base, VideoLLaMA3-7B-Image, VideoLLaMA3-2B-Image, AVProunRLForVideoLLaMa2, MM-PreTrain, JavisUnd-Eval
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
arXiv18 reposarXiv:2606.19100
AMALIA, amalia-vl-eval, DocVQA-PT, MME-PT, MMMU-Pro-PT, AI2D-PT, TextVQA-PT, ChartQA-PT, OCRBench-PT, InfographicVQA-PT, POPE-PT, COCO-Caption2017-PT, MMStar-PT, MATH-Vision-PT, MMMU-PT, SEED-Bench-PT, RealWorldQA-PT, EmbSpatial-Bench-PT
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
arXiv17 reposarXiv:2010.11929
vision_transformer, tf-transformers, hls-foundation-os, coyo-dataset, coyo-700m, coyo-labeled-300m, vit-l16-coyo-labeled-300m-i1k384, vit-l16-coyo-labeled-300m-i1k512, vit-l16-coyo-labeled-300m, vilmedic, vit-base-patch16-224-in21k, vit-large-patch16-224-in21k, camie-tagger-v2, nsfw_image_detection, LaTeX-OCR, vit-base-patch32-384, vit-dog
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
arXiv17 reposarXiv:2203.13474
sven_modified, sven, codegen-16B-nl, CodeGen, MOSS, moss-moon-003-sft-int4, moss-moon-003-sft-plugin-int4, moss-moon-003-base, moss-moon-003-sft-plugin-int8, moss-moon-003-sft, moss-moon-003-sft-int8, moss-moon-003-sft-plugin, MOSS, codegen-6B-mono, jaxformer, codegen-16B-mono, diff-codegen-6b-v2
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
arXiv17 reposarXiv:2205.14135
graphify, flash-attention, web-stable-diffusion, flash-attention, GPT-2, falcon-40b, falcon-rw-1b, falcon-7b, Chinese-CLIP, falcon-40b-instruct, mpt-1b-redpajama-200b, mpt-1b-redpajama-200b-dolly, nano_flan_t5, speechless-starcoder2-15b, flash-attention, flash-attention, Block-Sparse-Attention
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
arXiv17 reposarXiv:2301.12597
Make-It-3D, heron-chat-blip-ja-stablelm-base-7b-v1-llava-620k, heron-chat-blip-ja-stablelm-base-7b-v1, heron-chat-blip-ja-stablelm-base-7b-v0, LAVIS, mBLIP, mblip-mt0-xl, mblip-bloomz-7b, blip2-flan-t5-xl, vlis, blip2-flan-t5-xxl, vqazero, Zero-and-Few-Shot-Visual-Question-Answering, VLSA, visualglm-6b, VisualGLM-6B, VisualGLM-6B
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
arXiv17 reposarXiv:2306.02858
Video-LLaMA-Series, VideoLLaMA2, VideoLLaMA2.1-7B-16F-Base, Multi-Source-Video-Captioning, VideoLLaMA2.1-7B-AV, VideoLLaMA2-72B, VideoLLaMA2.1-7B-16F, VideoLLaMA2-72B-Base, VideoLLaMA2-7B, VideoLLaMA2-8x7B-Base, VideoLLaMA2-8x7B, VideoLLaMA2-7B-16F, VideoLLaMA2-7B-Base, VideoLLaMA2-7B-16F-Base, VideoLLaMA3-7B-Image, VideoLLaMA3-2B-Image, AVProunRLForVideoLLaMa2
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
arXiv17 reposarXiv:2308.09583
WizardLM_evol_instruct_70k, WizardCoder-15B-V1.0, WizardLM-70B-V1.0, WizardMath-7B-V1.1, WizardLM-70B-V1.0, WizardLM-13B-V1.2, WizardCoder-15B-V1.0, WizardMath-70B-V1.0, WizardLM-13B-V1.0, WizardCoder-Python-34B-V1.0, WizardLM_evol_instruct_V2_196k, WizardCoder-Python-13B-V1.0, WizardMath-7B-V1.1, WizardCoder-33B-V1.1, WizardMath-7B-V1.0, WizardLM_evol_instruct_70k, Ko.WizardLM_evol_instruct_V2_196k
VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text
arXiv17 reposarXiv:2406.06462
VCR-wiki-en-easy, VCR, VCR-wiki-en-easy-test-100, VCR-wiki-zh-easy-test-500, VCR-wiki-en-easy-test-500, VCR-wiki-en-hard-test-500, VCR-wiki-en-hard-test-100, VCR-wiki-zh-hard-test-100, VCR-wiki-zh-easy-test-100, VCR-wiki-zh-hard-test, VCR-wiki-en-easy-test, VCR-wiki-zh-hard-test-500, VCR-wiki-zh-easy-test, VCR-wiki-zh-hard, VCR-wiki-en-hard-test, VCR-wiki-en-hard, VCR-wiki-zh-easy
Rank1: Test-Time Compute for Reranking in Information Retrieval
arXiv17 reposarXiv:2502.18418
rank1-7b, rank1, rank1-1.5b, rank1-llama3-8b-awq, rank1-R1-MSMARCO, rank1-mistral-2501-24b, rank1-mistral-2501-24b-awq, rank1-training-data, rank1-14b-awq, rank1-Run-Files, rank1-32b-awq, rank1-7b-awq, rank1-3b, rank1-14b, rank1-0.5b, rank1-32b, rank1-llama3-8b
Perception Encoder: The best visual embeddings are not at the output of the network
arXiv17 reposarXiv:2504.13181
Muse-Glimmer-30B, perception_models, PE-Video, PE-Core-T16-384, PE-Core-S16-384, PE-Core-B16-224, PE-Core-L14-336, PE-Core-G14-448, PE-Lang-L14-448, PE-Lang-L14-448-Tiling, PE-Lang-G14-448, PE-Spatial-B16-512, PE-Spatial-L14-448, PE-Lang-G14-448-Tiling, PE-Spatial-G14-448, PE-Spatial-T16-512, PE-Spatial-S16-512
Seq vs Seq: An Open Suite of Paired Encoders and Decoders
arXiv17 reposarXiv:2507.11412
ettin-encoder-150m, ettin-encoder-400m, ettin-encoder-32m, ettin-encoder-68m, ettin-encoder-vs-decoder, ettin-decoder-1b, ettin-decoder-17m, ettin-decoder-400m, ettin-decoder-32m, ettin-encoder-17m, ettin-checkpoints, ettin-encoder-1b, ettin-decay-data, ettin-pretraining-data, ettin-decoder-68m, ettin-decoder-150m, ettin-extension-data
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
arXiv16 reposarXiv:2101.00027
GLM, falcon-40b, minipile, falcon-7b, pile-uncopyrighted, pythia-12b, pythia-1.4b, pythia-410m, LLaMA-MiLe-Loss, BiLLa, gpt-neo-2.7B, GPT-J-6B-Janeway, GPT-J-6B-Shinen, pythia-2.8b, diff-codegen-6b-v2, gpt-neo-1.3B
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
arXiv16 reposarXiv:2306.05685
synthetic_text_to_sql, mt_bench_prompts, Multi-Modality-Arena, vicuna-7B-1.1-HF, mt_bench_human_judgments, vicuna-13b-delta-v1.1, vicuna-13B-1.1-HF, vicuna-7b-delta-v1.1, vicuna-7b-delta-v0, llm-jailbreaking-defense, stablelm-zephyr-3b, vicuna-13b-v1.3, llm-router, vicuna-7b-v1.3, KoMT-Bench, KoMT-Bench
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
arXiv16 reposarXiv:2307.04657
beaver-7b-v2.0-cost, beaver-7b-unified-cost, PKU-SafeRLHF-10K, beaver-7b-v3.0-cost, beaver-7b-v3.0-reward, beaver-7b-v1.0, beaver-7b-v1.0-cost, beaver-7b-v3.0, beaver-7b-v2.0, beaver-7b-v1.0-reward, beaver-7b-unified-reward, beaver-7b-v2.0-reward, PKU-SafeRLHF-30K, beavertails, BeaverTails, BeaverTails-Evaluation
RoBERTa: A Robustly Optimized BERT Pretraining Approach
arXiv15 reposarXiv:1907.11692
roberta-large-mnli, EasyNLP, roberta-base, DiagnosisCoding, tf-transformers, legalbert-large-1.7M-1, legalbert-large-1.7M-2, LoRA, GPT_Ranker, roberta_toxicity_classifier, Reddit-Sports-Sentiment-Analysis, roberta-large, deid_roberta_i2b2, ehr_deidentification, caption
SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
arXiv15 reposarXiv:2105.15203
segformer-b0-finetuned-ade-512-512, SegFormer, segformer-b2-finetuned-ade-512-512, segformer-b0-finetuned-cityscapes-1024-1024, Semantic-Segment-Anything, segformer-b5-finetuned-cityscapes-1024-1024, segformer-b5-finetuned-ade-640-640, segformer-b2-finetuned-cityscapes-1024-1024, segformer_b2_clothes, segformer-tf-transformers, SegFormer-Training_From-Scratch_vs_HuggingFace, EdgeSeg, segformer-tf-transformers, segformer-b3-fashion, segformer_b3_clothes
Crosslingual Generalization through Multitask Finetuning
arXiv15 reposarXiv:2211.01786
awesome-totally-open-chatgpt, bloomz, mt0-base, xmtf, xP3megds, xwinograd, xP3, xP3x, bloomz-560m, bloomz-3b, mt0-xl, bloomz-1b1, bloomz-7b1-mt, bloomz-1b7, bloomz-7b1
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
arXiv15 reposarXiv:2305.18290
BELLE, tunix, zephyr-7b-alpha, zephyr-7b-beta, Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, Reinforcement-Learning-Full-Pipeline, shisa-7b-v1, alpaca_farm, MedicalGPT, stablelm-zephyr-3b, sacpo, p-sacpo, dpo-arithmo-mistral-7B, llm_factuality_tuning
arXiv:2308.07317
arXiv15 reposarXiv:2308.07317
Jellyfish-13B, Platypus2-70B, Platypus-13B-adapters, Platypus2-70B-instruct, Camel-Platypus2-13B, Stable-Platypus2-13B, Platypus2-13B, Platypus-70B-adapters, Camel-Platypus2-70B, Platypus2-7B, Platypus-7B-adapters, Platypus-30B, Platypus, KOpen-platypus, KO-Platypus2-7B-ex
StarVector: Generating Scalable Vector Graphics Code from Images and Text
arXiv15 reposarXiv:2312.11556
vtracer, star-vector, starvector-1b-im2svg, starvector-8b-im2svg, svg-stack, text2svg-stack, svg-stack-simple, svg-diagrams, svg-fonts, svg-emoji, svg-fonts-simple, svg-emoji-simple, FIGR-SVG, svg-icons, svg-icons-simple
xLAM: A Family of Large Action Models to Empower AI Agent Systems
arXiv15 reposarXiv:2409.03215
xLAM-7b-r, xLAM, xLAM-1b-fc-r, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, xLAM-2-1b-fc-r, xLAM-v0.1-r, xLAM-2-3b-fc-r-gguf, xLAM-2-32b-fc-r, xLAM-8x22b-r, xLAM-7b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, xLAM-8x7b-r, Llama-xLAM-2-8b-fc-r, xLAM-2-3b-fc-r
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
arXiv15 reposarXiv:2409.12191
Qwen2.5-VL, Qwen3-VL, Qwen2-VL-7B-Instruct, Qwen2.5-VL-72B-Instruct, Qwen2-VL, Qwen2.5-VL-3B-Instruct-GGUF, Baichuan-Omni-1.5, UGround-V1-7B, UGround-V1-2B, UGround-V1-72B, Qwen2.5-VL-7B-Instruct-GGUF, Qwen2-VL-7B-Instruct, qwen2vit600m, Qwen2.5-VL-7B-Instruct-GGUF, ChatVLA_public
Self-Instruct: Aligning Language Models with Self-Generated Instructions
arXiv14 reposarXiv:2212.10560
stanford_alpaca, camel, alpaca_eval, self-instruct-seed, COIG, COIG, Chinese-Vicuna, moss-002-sft-data, EasyInstruct, vigogne, openchat, Chinese-Vicuna, koalpaca, kwater
Adding Conditional Control to Text-to-Image Diffusion Models
arXiv14 reposarXiv:2302.05543
ControlNet, visual-chatgpt-zh, MistoLine, sd-controlnet-canny, control_v11p_sd15_openpose, controlnet-canny-sdxl-1.0, BDM1.0, ControlNet_AnimalPose, controlnet-openpose-sdxl-1.0, controlnet-scribble-sdxl-1.0, sd-controlnet-seg, control_v11f1p_sd15_depth, control_v11p_sd15_normalbae, sd-controlnet-scribble
Sigmoid Loss for Language Image Pre-Training
arXiv14 reposarXiv:2303.15343
siglip-so400m-patch14-384, siglip2-giant-opt-patch16-384, siglip2-so400m-patch16-512, siglip2-base-patch16-224, siglip2-base-patch16-naflex, open-value, siglip-base-patch16-224, ml-mobileclip, CCD, siglip2-so400m-patch16-naflex, prismatic-vlms, vit_base_patch16_siglip_512.v2_webli, siglip2-so400m-patch14-384, siglip2-so400m-patch14-224
DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
arXiv14 reposarXiv:2309.14509
VeOmni, MindSpeed-MM, Wan2.2-T2V-A14B-Diffusers, Wan2.1-VACE-14B, EasyContext, Wan2.2, Wan2.1, Wan2.2-T2V-A14B, Wan2.1-VACE-1.3B, Wan2.1-FLF2V-14B-720P, maxdiffusion, OpenDiT, Wan2.2-Lightning, Wan2.2-Lightning
Simple and Effective Masked Diffusion Language Models
arXiv14 reposarXiv:2406.07524
mdlm-owt, dllm, SDAR, dLLM-RL, FreeDave, mdlm, bd3lm, BDM, bd3lms, BD_DNA, bd3lm-owt-block_size1024-pretrain, ar-noeos-owt, diffusion-ebm, Open-dLLM
MiniCPM-V: A GPT-4V Level MLLM on Your Phone
arXiv14 reposarXiv:2408.01800
OmniLMM, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB, MiniCPM-V-4.6-Thinking-GPTQ, MiniCPM-o-4_5-gguf, MiniCPM-o-4_5-AWQ, MiniCPM-o-4_5
ActionStudio: A Lightweight Framework for Data and Training of Large Action Models
arXiv14 reposarXiv:2503.22673
xLAM-7b-r, xLAM, xLAM-1b-fc-r, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, Llama-xLAM-2-8b-fc-r, xLAM-2-1b-fc-r, xLAM-2-3b-fc-r, xLAM-2-32b-fc-r, xLAM-8x22b-r, xLAM-7b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, xLAM-8x7b-r, xLAM-2-3b-fc-r-gguf
Helios: Real Real-Time Long Video Generation Model
arXiv14 reposarXiv:2603.04379
Open-Sora-Plan, MagicTime, ChronoMagic-Bench, ConsisID, Helios, Helios-14B-RealTime, Helios-14B-RealTime-AOTI, Helios-Base, HeliosBench-Weights, Helios-Mid, Helios-Distilled, OpenS2V-Nexus, helios, tele_Gen
Measuring Massive Multitask Language Understanding
arXiv13 reposarXiv:2009.03300
XVERSE-7B, XuanYuan, mmlu, PodGPT, mmlu-redux, Baichuan2, Qwen-14B-Chat, Qwen-7B-Chat, Baichuan-13B-Base, Baichuan-13B-Chat, Qwen-1_8B, JMMLU, mmlu_ru
BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
arXiv13 reposarXiv:2104.08663
beir, SFR-Embedding-Mistral, multilingual-e5-base, multilingual-e5-large, e5-large-unsupervised, e5-small-unsupervised, e5-base-unsupervised, e5-mistral-7b-instruct, multilingual-e5-large-instruct, e5-large-v2, gte-multilingual-base, RAG, RAGatouille
SimCSE: Simple Contrastive Learning of Sentence Embeddings
arXiv13 reposarXiv:2104.08821
DermL2V-training, DermL2V-tmp, llm2vec, SimCSE, sup-simcse-roberta-large, sup-simcse-bert-large-uncased, Thai-Sentence-Vector-Benchmark, simcse-model-roberta-base-thai, simcse-model-XLMR, simcse-model-m-bert-thai-cased, simcse-model-wangchanberta, simcse-model-phayathaibert, simcse-model-distil-m-bert
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
arXiv13 reposarXiv:2108.12409
bloom, bloom-optimizer-states, falcon-rw-1b, Baichuan-13B-Base, Baichuan-13B-Chat, bloom-560m, bloom-1b7, mpt-1b-redpajama-200b, mpt-1b-redpajama-200b-dolly, bloom-7b1, bloom-1b1, bloom-3b, jina-colbert-v1-en
OCR-free Document Understanding Transformer
arXiv13 reposarXiv:2111.15664
donut, donut-base-finetuned-cord-v2, donut-base-finetuned-cord-v1, donut-base-finetuned-cord-v1-2560, donut-base-finetuned-zhtrainticket, donut-base-finetuned-rvlcdip, donut-base-finetuned-docvqa, donut-base, donut-proto, aipeaks-pipeline-workshop, donut, donut1, donut-master
Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition
arXiv13 reposarXiv:2305.05084
parakeet-rnnt-1.1b, parakeet-tdt_ctc-1.1b, parakeet-ctc-0.6b, parakeet-tdt_ctc-110m, parakeet-tdt-1.1b, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, parakeet-rnnt-0.6b, SLU_pipeline, stt_ar_fastconformer_hybrid_large_pcd_v1.0, canary-1b, diar_sortformer_4spk-v1, stt_ru_fastconformer_hybrid_large_pc
MultiLegalPile: A 689GB Multilingual Legal Corpus
arXiv13 reposarXiv:2306.02069
LEXTREME, Multi_Legal_Pile, legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, MultilingualLegalLMPretraining, legal-swiss-roberta-large, legal-english-roberta-base, legal-croatian-roberta-base, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large
Orca: Progressive Learning from Complex Explanation Traces of GPT-4
arXiv13 reposarXiv:2306.02707
StableBeluga2, OpenOrca, orca_dpo_pairs, orca_mini_7B-GPTQ, neural-chat-7b-v3-1, vigogne, Mistral-7B-OpenOrca, Jellyfish-13B, OpenOrca-KO, KOR-OpenOrca-Platypus, OpenOrca-gugugo-ko, orca_mini_3B-GGML, StableBeluga-7B
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
arXiv13 reposarXiv:2310.00426
PixArt-alpha, PixArt-alpha, PixArt-LCM-XL-2-1024-MS, PixArt-alpha, PixArt-XL-2-512x512, PixArt-LCM, PixArt-XL-2-256x256, PixArt-XL-2-1024-MS, PixArt-Sigma-XL-2-512-MS, Open-Sora-Plan-v1.2.0, pixeart, PixArt-alpha, PixArt-LCM
Safe RLHF: Safe Reinforcement Learning from Human Feedback
arXiv13 reposarXiv:2310.12773
alpaca-7b-reproduced, safe-rlhf, beaver-7b-v2.0-cost, beaver-7b-v3.0-cost, beaver-7b-v3.0-reward, beaver-7b-v1.0, beaver-7b-v1.0-cost, beaver-7b-unified-cost, beaver-7b-v3.0, beaver-7b-v2.0, beaver-7b-v1.0-reward, beaver-7b-unified-reward, beaver-7b-v2.0-reward
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
arXiv13 reposarXiv:2311.10122
LanguageBind, MoE-LLaVA, Video-LLaVA, MoE-LLaVA-Qwen-1.8B-4e, MoE-LLaVA-Phi2-2.7B-4e, MoE-LLaVA-Phi2-2.7B-4e-384, MoE-LLaVA-StableLM-1.6B-4e-384, MoE-LLaVA-StableLM-Pretrain, MoE-LLaVA-Phi2-Pretrain, MoE-LLaVA-StableLM-1.6B-4e, MoE-LLaVA-Qwen-Pretrain, MoE-LLaVA-Phi2-384-Pretrain, LLMBind
Efficient Multimodal Learning from Data-centric Perspective
arXiv13 reposarXiv:2402.11530
Bunny-v1_1-data, Bunny, Bunny-Llama-3-8B-V, Bunny-v1_0-3B-zh, Bunny-v1_0-4B-gguf, Bunny-v1_0-4B, Bunny-v1_1-4B, Bunny-v1_1-Llama-3-8B-V, Bunny-Llama-3-8B-V-gguf, Bunny-v1_0-data, Bunny-v1_0-2B-zh, bunny-phi-2-siglip-lora, Bunny-v1_0-3B
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning
arXiv13 reposarXiv:2402.15506
xLAM-7b-r, xLAM, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, xLAM-8x22b-r, xLAM-2-3b-fc-r-gguf, xLAM-2-32b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, xLAM-8x7b-r, Llama-xLAM-2-8b-fc-r, xLAM-v0.1-r, xLAM-2-1b-fc-r, xLAM-2-3b-fc-r
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
arXiv13 reposarXiv:2404.02132
ViTamin, ViTamin-XL-384px, ViTamin-L-256px, ViTamin-L-224px, ViTamin-L-336px, ViTamin-L-384px, ViTamin-L2-384px, ViTamin-L2-256px, ViTamin-XL-256px-s13B, ViTamin-L2-224px, ViTamin-L2-336px, ViTamin-XL-336px, ViTamin-XL-256px
YOLOv10: Real-Time End-to-End Object Detection
arXiv13 reposarXiv:2405.14458
yolov10, yolov10s, YOLOv10, Yolov10, yolov10n, yolov10m, yolov10b, yolov10l, yolov10x, yolov10, spectra, yolo10-train, CalEstimator
arXiv:2407.21783
arXiv13 reposarXiv:2407.21783
STRING, Llama-3.1-Swallow-70B-v0.1, Llama-3.1-Swallow-8B-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-70B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.2, Llama-3.1-Swallow-8B-v0.2, Llama-3.1-Swallow-70B-Instruct-v0.3, Llama-3.3-Swallow-70B-v0.4, Llama-3.3-Swallow-70B-Instruct-v0.4, Llama-3.1-Swallow-8B-Instruct-v0.3, Llama-3.1-Swallow-8B-Instruct-v0.5, Llama-3.1-Swallow-8B-v0.5
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
arXiv13 reposarXiv:2503.09573
ECHO, dllm, LightningRL, SDAR, dLLM-RL, FreeDave, bd3lm, BDM, bd3lms, BD_DNA, bd3lm-owt-block_size1024-pretrain, sedd-noeos-owt, ar-noeos-owt
Wan: Open and Advanced Large-Scale Video Generative Models
arXiv13 reposarXiv:2503.20314
Wan2.2-I2V-A14B-Diffusers, Wan2.2-Animate-14B, Wan2.2-TI2V-5B-Diffusers, Wan2.1-VACE-14B, Wan2.2-T2V-A14B-Diffusers, Wan2.2-S2V-14B, Wan2.2, Wan2.1, Wan2.2-T2V-A14B, Wan2.2-TI2V-5B, Wan2.2-I2V-A14B, Wan2.1-VACE-1.3B, Wan2.1-FLF2V-14B-720P
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
arXiv12 reposarXiv:1910.01108
distilbert-base-uncased, scifact, distilbert-base-uncased-distilled-squad, DistilBERT-BERT_WNLI_NER, berts, turkish-bert, distilbert-base-turkish-cased, europeana-bert, swift-coreml-transformers, DistilBERT, LEXTREME, distilgpt2
Data Bootstrapping Approaches to Improve Low Resource Abusive Language Detection for Indic Languages
arXiv12 reposarXiv:2204.12543
english-abusive-MuRIL, IndicAbusive, malayalam-codemixed-abusive-MuRIL, bengali-abusive-MuRIL, tamil-codemixed-abusive-MuRIL, kannada-codemixed-abusive-MuRIL, hindi-abusive-MuRIL, marathi-codemixed-abusive-MuRIL, urdu-codemixed-abusive-MuRIL, indic-abusive-allInOne-MuRIL, hindi-codemixed-abusive-MuRIL, urdu-abusive-MuRIL
MTEB: Massive Text Embedding Benchmark
arXiv12 reposarXiv:2210.07316
mteb, SFR-Embedding-Mistral, multilingual-e5-base, multilingual-e5-large, e5-large-unsupervised, e5-small-unsupervised, e5-base-unsupervised, e5-mistral-7b-instruct, multilingual-e5-large-instruct, e5-large-v2, mteb-1.34.14, ru_sci_bench_mteb
Segment Anything
arXiv12 reposarXiv:2304.02643
geti-instant-learn, sd-webui-inpaint-anything, sam-vit-large, Depth-Estimation, sd-webui-inpaint-anything, MobileSAM, Flow-Inference-Time-Scaling, SEED-Data-Edit-Part2-3, SEED-Data-Edit, 3D-LLM, medsam-vit-base, SAMReg
Universal and Transferable Adversarial Attacks on Aligned Language Models
arXiv12 reposarXiv:2307.15043
OBLITERATUS, JBB-Behaviors, llm-attacks, jailbreakbench, RAIN, nanoGCG, Magic_Words, certified-llm-safety, Jailbreak_LLM, GA, llm-jailbreaking-defense, LLMart
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
arXiv12 reposarXiv:2308.12966
Qwen2-VL-7B-Instruct, Qwen2.5-VL-72B-Instruct, Qwen2.5-VL-3B-Instruct-GGUF, UGround-V1-7B, UGround-V1-2B, UGround-V1-72B, Qwen2.5-VL-7B-Instruct-GGUF, Qwen-VL, Qwen2-VL-7B-Instruct, Qwen-VL-Chat, CoBSAT, Qwen2.5-VL-7B-Instruct-GGUF
Mistral 7B
arXiv12 reposarXiv:2310.06825
Mistral-7B-v0.1, Mistral-7B-Instruct-v0.2, eCeLLM, Mistral-7B-Instruct-v0.1, SFR-Embedding-Mistral, ROOT-RAG, Mixtral-8x7B-Instruct-v0.1, speechless-mistral-six-in-one-7b, speechless-mistral-dolphin-orca-platypus-samantha-7B-GGUF, speechless-mistral-dolphin-orca-platypus-samantha-7b, speechless-mistral-dolphin-orca-platypus-samantha-7B-GPTQ, speechless-mistral-dolphin-orca-platypus-samantha-7B-AWQ
CogVLM: Visual Expert for Pretrained Language Models
arXiv12 reposarXiv:2311.03079
glm-4v-9b, glm-4v-9b, cogvlm2-llama3-chat-19B, cogagent-chat-hf, cogvlm-chat-hf, cogvlm-base-490-hf, cogvlm-grounding-base-hf, cogvlm-grounding-generalist-hf, cogagent-vqa-hf, cogvlm-base-224-hf, cogagent-9b-20241220, visualglm-6b
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
arXiv12 reposarXiv:2401.15947
LanguageBind, MoE-LLaVA, Video-LLaVA, MoE-LLaVA-Phi2-2.7B-4e, MoE-LLaVA-Qwen-1.8B-4e, MoE-LLaVA-Phi2-2.7B-4e-384, MoE-LLaVA-StableLM-1.6B-4e-384, MoE-LLaVA-Phi2-Pretrain, MoE-LLaVA-StableLM-1.6B-4e, MoE-LLaVA-Phi2-384-Pretrain, MoE-LLaVA-StableLM-Pretrain, MoE-LLaVA-Qwen-Pretrain
Yi: Open Foundation Models by 01.AI
arXiv12 reposarXiv:2403.04652
Yi, Yi-1.5-9B-Chat, Yi-1.5-34B-Chat, Yi-1.5-6B-Chat, Yi-34B-Chat, Yi-1.5-6B, Yi-1.5-9B-Chat-16K, Yi-1.5-9B, Yi-1.5-9B-32K, Yi-VL-6B, Yi-VL-34B, llm_project
InternLM2 Technical Report
arXiv12 reposarXiv:2403.17297
internlm2_5-7b-chat-1m, internlm2-20b, internlm3-8b-instruct, internlm2-chat-7b, internlm2_5-7b-chat, internlm2-chat-1_8b, internlm2_5-7b, internlm2-7b, internlm2-chat-20b, Bespoke-MiniCheck-7B, internlm2-7b-reward, internlm2-20b-reward
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
arXiv12 reposarXiv:2403.18814
MGM, MGM-2B, MGM-7B, MGM-13B, MGM-8B, MGM-8x7B, MGM-34B, MGM-7B-HD, MGM-13B-HD, MGM-8B-HD, MGM-8x7B-HD, MGM-34B-HD
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
arXiv12 reposarXiv:2404.17790
Swallow-70b-hf, Swallow-70b-NVE-hf, Swallow-13b-NVE-hf, Swallow-7b-NVE-hf, Swallow-7b-NVE-instruct-hf, Swallow-70b-instruct-hf, Swallow-70b-NVE-instruct-hf, Swallow-7b-plus-hf, Swallow-7b-hf, Swallow-13b-instruct-hf, Swallow-7b-instruct-hf, Swallow-13b-hf
Qwen2 Technical Report
arXiv12 reposarXiv:2407.10671
evalplus, Qwen2.5-72B-Instruct, Qwen2.5-32B-Instruct, Qwen2.5-Coder-32B-Instruct, Qwen2.5-Coder-14B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-0.5B-Instruct, Qwen2.5-3B, Qwen2.5-14B-Instruct, Qwen2.5-14B-Instruct-AWQ, Qwen2.5-Coder-7B-Instruct, Qwen2.5-Coder-0.5B
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
arXiv12 reposarXiv:2502.14786
siglip2-giant-opt-patch16-384, siglip2-so400m-patch16-512, siglip2-base-patch16-224, siglip2-base-patch16-naflex, pd12m, siglip2-so400m-patch16-naflex, vit_base_patch16_siglip_512.v2_webli, RAE, siglip2-so400m-patch14-384, piedomains-image, TiViT, siglip2-so400m-patch14-224
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
arXiv12 reposarXiv:2503.14734
Eagle, physical-ai-studio, GR00T-N1.6-G1-PnPAppleToPlate, GR00T-N1.6-DROID, GR00T-N1.6-bridge, GR00T-N1.6-fractal, GR00T-N1.6-BEHAVIOR1k, GR00T-N1.7-SimplerEnv-Fractal, GR00T-N1.7-DROID, GR00T-N1.7-SimplerEnv-Bridge, GR00T-N1.7-LIBERO, GR00T-N1.7-3B
FG-CLIP: Fine-Grained Visual and Textual Alignment
arXiv12 reposarXiv:2505.05071
fg-clip-base, FG-CLIP, fg-clip2-large, fg-clip2-base, DCI-CN, fg-clip-large, BoxClass-CN, DOCCI-CN, fg-clip2-so400m, FineHARD, LIT-CN, vit-dog
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
arXiv12 reposarXiv:2508.06471
GLM-4.7, GLM-4.7-Flash, GLM-4.5, GLM-4.5-FP8, GLM-4.7-FP8, GLM-4.5-Air, GLM-4.6, GLM-4.6-FP8, GLM-4.5, GLM-4.5-Air-FP8, GLM-4.5-Base, GLM-4.5-Air-Base
BYOL: Bring Your Own Language Into LLMs
arXiv12 reposarXiv:2601.10804
byol, Global-MMLU-Lite, byol-nya-12b-cpt, byol-mri-1b-cpt, byol-nya-12b-merged, byol-mri-4b-cpt, byol-mri-4b-merged, byol-nya-4b-cpt, byol-nya-1b-cpt, byol-mri-12b-cpt, byol-mri-12b-merged, byol-nya-4b-merged
RLDX-1 Technical Report
arXiv12 reposarXiv:2605.03269
RLDX-1, RLDX-1-PT-IMG, RLDX-1-MT-DROID, RLDX-1-PT, RLDX-1-MT-ALLEX, RLDX-1-FT-LIBERO, RLDX-1-FT-GR1, RLDX-1-FT-SIMPLER-GOOGLE, RLDX-1-FT-ROBOCASA, RLDX-1-FT-SIMPLER-WIDOWX, RLDX-1-FT-RC365, RLDX-FineAct
LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?
arXiv12 reposarXiv:2605.08985
OmniLMM, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB, MiniCPM-V-4.6-Thinking-GPTQ, LLaVA-UHD-v4
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
arXiv11 reposarXiv:1908.10084
paraphrase-multilingual-mpnet-base-v2, SimCSE, LateOn-Code, LateOn-Code-edge, distiluse-base-multilingual-cased-v2, splade-ecommerce-esci, SauerkrautLM-Multi-Reason-ModernColBERT, ColBERT-Zero, langcache-embed-v1, langcache-embed-v2, average_word_embeddings_glove.6B.300d
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
arXiv11 reposarXiv:1909.08053
awesome-gpu-engineering, bloom, MindSpeed-MM, bloom-optimizer-states, bloom-560m, awsome-llm-papers, bloom-1b7, mesh-transformer-jax, bloom-7b1, bloom-1b1, bloom-3b
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
arXiv11 reposarXiv:2103.10360
GLM, SwissArmyTransformer, chatglm2-6b-32k, chatglm2-6b, chatglm-6b, chatglm2-6b-int4, chatglm3-6b-32k, chatglm3-6b-32k, chatglm3-6b-base, visualglm-6b, KoGLM
LexGLUE: A Benchmark Dataset for Legal Language Understanding in English
arXiv11 reposarXiv:2110.00976
legalbert-large-1.7M-1, legalbert-large-1.7M-2, legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, legal-english-roberta-base, legal-swiss-roberta-large, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large
Training language models to follow instructions with human feedback
arXiv11 reposarXiv:2203.02155
ik_llama.cpp, llama.cpp, Open-Assistant, databricks-dolly-15k, no_robots, llama.cpp, llama.cpp.qwen2.5vl, atomic-llama-cpp-turboquant, alpaca_farm, llama.cpp, llama.cpp
BigVGAN: A Universal Neural Vocoder with Large-Scale Training
arXiv11 reposarXiv:2206.04658
HierSpeechpp, bigvgan_v2_24khz_100band_256x, BigVGAN, bigvgan_v2_44khz_128band_512x, bigvgan_base_24khz_100band, bigvgan_22khz_80band, bigvgan_v2_22khz_80band_fmax8k_256x, bigvgan_v2_22khz_80band_256x, bigvgan_24khz_100band, bigvgan_v2_44khz_128band_256x, bigvgan_base_22khz_80band
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
arXiv11 reposarXiv:2301.13126
LEXTREME, legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, lextreme, legal-swiss-roberta-large, legal-english-roberta-base, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large
DINOv2: Learning Robust Visual Features without Supervision
arXiv11 reposarXiv:2304.07193
geti-instant-learn, rf-detr, dinov2-small, dinov2-large, dinov2-giant, hashing-baseline, prismatic-vlms, RAE, CVProject, DinoBloom, TiViT
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
arXiv11 reposarXiv:2305.14233
ultrachat_200k, zephyr-7b-alpha, zephyr-7b-beta, Llama-3-8B-Instruct-Gradient-1048k, Llama-3-70B-Instruct-Gradient-262k, Llama-3-70B-Instruct-Gradient-1048k, Llama-3-8B-Instruct-262k, Llama-3-8B-Instruct-Gradient-4194k, llm-finetuning, ultrachat, Instruct-SkillMix-SDD
QLoRA: Efficient Finetuning of Quantized LLMs
arXiv11 reposarXiv:2305.14314
Qwen, qlora, text-generation-inference, smash, GPTQ-for-LLaMa, vigogne, llmtools, llama2-70b, llmtools, level3_nlp_finalproject-nlp-12, ainn-final
Unified Training of Universal Time Series Forecasting Transformers
arXiv11 reposarXiv:2402.02592
moirai-1.0-R-base, uni2ts, moirai-2.0-R-small, lotsa_data, stock_forecasting, Samay, A2TTA, moment, moirai-1.0-R-large, moirai-1.0-R-small, uni2ts
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
arXiv11 reposarXiv:2406.10721
RoboPoint, robopoint-v1-vicuna-v1.5-13b, robopoint-v1-vicuna-v1.5-13b-lora, robopoint-v1-llama-2-13b, robopoint_data, robopoint-v1-llama-2-13b-lora, robopoint-v1-llama-2-7b-lora, robopoint-v1-vicuna-v1.5-7b-lora, where2place, Robopoint_Humble, Cosmos-Reason2-32B
Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
arXiv11 reposarXiv:2503.06520
refCOCOg_2k_840, Seg-Zero, refCOCOg_9k_840, Seg-Zero-7B, VisionReasoner_multi_object_7k_840, ReasonSeg_val, ReasonSeg_test, VisionReasoner-7B, VisionReasoner, Seg-Zero, VisionReasoner
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
arXiv11 reposarXiv:2503.15621
LLaVA-MORE, LLaVA_MORE-llama_3_1-8B-pretrain, LLaVA_MORE-llama_3_1-8B-siglip-finetuning, LLaVA_MORE-llama_3_1-8B-S2-pretrain, LLaVA_MORE-llama_3_1-8B-siglip-pretrain, LLaVA_MORE-llama_3_1-8B-finetuning, LLaVA_MORE-llama_3_1-8B-S2-finetuning, LLaVA_MORE-llama_3_1-8B-S2-siglip-pretrain, LLaVA_MORE-llama_3_1-8B-S2-siglip-finetuning, LLaVA_MORE-phi_4-finetuning, LLaVA_MORE-gemma_2_9b-finetuning
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
arXiv11 reposarXiv:2505.12081
Seg-Zero, VisionReasoner_multi_object_7k_840, refCOCOg_9k_840, ReasonSeg_val, VisionReasoner_multi_object_1k_840, ReasonSeg_test, VisionReasoner-7B, VisionReasoner, TaskRouter-1.5B, VisionReasoner, Seg-Zero
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
arXiv11 reposarXiv:2507.17702
AntAngelMed, AntAngelMed, Ling-flash-base-2.0, Ling-mini-2.0, Ling-mini-base-2.0-5T, Ling-mini-base-2.0, Ling-mini-base-2.0-10T, Ling-mini-base-2.0-15T, Ling-mini-base-2.0-20T, Ling-flash-2.0, Ling-1T
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
arXiv11 reposarXiv:2509.18154
OmniLMM, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB, MiniCPM-V-4.6-Thinking-GPTQ
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
arXiv11 reposarXiv:2510.10274
X-VLA-RoboTwin2, X-VLA-Google-Robot, xvla-base, X-VLA, X-VLA-Libero, X-VLA-SoftFold, X-VLA-Calvin-ABC_D, X-VLA-Pt, X-VLA-AgiWorld-Challenge, X-VLA-WidowX, X-VLA-VLABench
A visual-language foundation model for computational pathology
Nature11 reposNature:s41591-024-02856-4
AtlasPatch, PIANO, KEEP, KEEP, VLSA, SEAL, HistAug, histaug-conch, PathPT, Histopathology_Benchmark, dpfm_factory
A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
arXiv10 reposarXiv:1910.04867
CLIP-convnext_xxlarge-laion2B-s34B-b82K-augreg-soup, CLIP-ViT-H-14-laion2B-s32B-b79K, CLIP-ViT-bigG-14-laion2B-39B-b160k, CLIP-ViT-B-32-laion2B-s34B-b79K, CLIP-ViT-L-14-laion2B-s32B-b82K, CLIP-ViT-g-14-laion2B-s12B-b42K, CLIP-ViT-B-32-roberta-base-laion2B-s12B-b32k, CLIP-ViT-H-14-frozen-xlm-roberta-large-laion5B-s13B-b90k, CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft-soup, CLIP-ViT-B-16-laion2B-s34B-b88K
TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification
arXiv10 reposarXiv:2010.12421
tweetnlp, twitter-roberta-base, twitter-roberta-base-irony, twitter-roberta-base-offensive, twitter-roberta-base-emoji, twitter-roberta-base-emotion, tweeteval, twitter-roberta-base-hate, twitter-roberta-base-sentiment, lares
Self-attention Does Not Need $O(n^2)$ Memory
arXiv10 reposarXiv:2112.05682
LLaVA, FastChat, FastChat, Libra, multilingual_mt_bench, FastChat, TriPlaneLLaVA, hermes-llava, LLaVA-toy, TinyLLava
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
arXiv10 reposarXiv:2205.11487
stable-diffusion, stable-diffusion-v1-4, stable-diffusion-v-1-4-original, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, stable-diffusion-inpainting, stable-diffusion-uncrop, stable-diffusion-v1.5-webnn
Matryoshka Representation Learning
arXiv10 reposarXiv:2205.13147
RAG-Retrieval, ogham-mcp, contrastors, nomic-embed-text-v2-moe, nomic-embed-text-v1.5, UForm, contrastors, cholesky_encoder, arctic-embed, langcache-embed-v2
Classifier-Free Diffusion Guidance
arXiv10 reposarXiv:2207.12598
stable-diffusion, stable-diffusion-v1-4, stable-diffusion-v-1-4-original, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, stable-diffusion-inpainting, smalldiffusion, stable-diffusion-v1.5-webnn
Visual Instruction Tuning
arXiv10 reposarXiv:2304.08485
LLaVA, minimind-v, MedAI-project, llava-implementation, nanoMFM, TriPlaneLLaVA, CoBSAT, LLaVA-toy, multi_token, TinyLLava
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
arXiv10 reposarXiv:2306.00890
LLaVA, llava-med-v1.5-mistral-7b, LLaVA-Med, libra-llava-med-v1.5-mistral-7b, libra-llava-rad, llava-rad, TriPlaneLLaVA, LLaVA-toy, LLaDA-MedV, TinyLLava
h2oGPT: Democratizing Large Language Models
arXiv10 reposarXiv:2306.08161
h2ogpt, fork, honogpt, OpenGPT-v1, h2ogpt, H2O_AI, h2ogpt, h2ogpt, llmbot, h2ogpt
C-Pack: Packed Resources For General Chinese Embeddings
arXiv10 reposarXiv:2309.07597
bge-small-zh-v1.5, bge-large-zh-v1.5, bge-base-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-en-icl, bge-large-en, bge-base-en, bge-small-en, mteb-1.34.14
A decoder-only foundation model for time-series forecasting
arXiv10 reposarXiv:2310.10688
timesfm, timesfm-3.0-pytorch, timesfm-1.0-200m, uni2ts, Samay, A2TTA, timesfm-1.0-200m-pytorch, moment, timesfm-2.5-200m-pytorch, uni2ts
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
arXiv10 reposarXiv:2311.09774
HuatuoGPT2-SFT-GPT4-140K, HuatuoGPT2-13B, HuatuoGPT2-Pretraining-Instruction, HuatuoGPT2-34B, HuatuoGPT2-7B, HuatuoGPT-II, HuatuoGPT2-7B-8bits, HuatuoGPT2-7B-4bits, HuatuoGPT2-34B-8bits, HuatuoGPT2-34B-4bits
Self-Play Preference Optimization for Language Model Alignment
arXiv10 reposarXiv:2405.00675
SPPO, Mistral7B-PairRM-SPPO-Iter1, Gemma-2-9B-It-SPPO-Iter1, Gemma-2-9B-It-SPPO-Iter3, Llama-3-Instruct-8B-SPPO-Iter1, Mistral7B-PairRM-SPPO-Iter2, Gemma-2-9B-It-SPPO-Iter2, Mistral7B-PairRM-SPPO-Iter3, Llama-3-Instruct-8B-SPPO-Iter3, Llama-3-Instruct-8B-SPPO-Iter2
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
arXiv10 reposarXiv:2408.06072
CogVideo, CogVideoX1.5-5B, CogVideoX-2b, CogVideoX-5b-I2V, CogVideoX1.5-5B-I2V, CogVideoX-2b, CogVideoX1.5-5B, CogVideoX-5b, CogVideoX-5b-I2V, cogvlm2-llama3-caption
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
arXiv10 reposarXiv:2409.18042
emova_speech_tokenizer_hf, EMOVA, qwen2vit600m, emova-qwen-2-5-3b-hf, emova-qwen-2-5-3b, emova-qwen-2-5-7b-hf, emova-qwen-2-5-72b, EMOVA_speech_tokenizer, emova-qwen-2-5-7b, emova-qwen-2-5-72b-hf
Open-Sora Plan: Open-Source Large Video Generation Model
arXiv10 reposarXiv:2412.00131
Open-Sora-Plan, Open-Sora-Plan-v1.3.0, MagicTime, ChronoMagic-Bench, MoE-LLaVA, Video-LLaVA, ConsisID, OpenS2V-Nexus, UniWorld-V1, UniWorld-V1-NF4
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
arXiv10 reposarXiv:2501.02790
DenseRewardRLHF-PPO, Phi-3-mini-4k-token-ppo-60k, Phi-3-mini-4k-segment-ppo-60k, meta-llama-3.1-instruct-8b-token-ppo-60k, meta-llama-3.1-instruct-8b-segment-ppo-60k, Phi-3-mini-4k-bandit-ppo-60k, meta-llama-3.1-instruct-8b-bandit-ppo-60k, rlhflow-llama-3-sft-8b-v2-segment-ppo-60k, rlhflow-llama-3-sft-8b-v2-token-ppo-60k, rlhflow-llama-3-sft-8b-v2-bandit-ppo-60k
Building Instruction-Tuning Datasets from Human-Written Instructions with Open-Weight Large Language Models
arXiv10 reposarXiv:2503.23714
Llama-3.1-Swallow-8B-Instruct-v0.5, Llama-3.1-Swallow-70B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.2, Llama-3.1-Swallow-8B-Instruct-v0.3, Llama-3.1-Swallow-70B-Instruct-v0.3, Llama-3.3-Swallow-70B-Instruct-v0.4, Gemma-2-Llama-Swallow-9b-it-v0.1, Gemma-2-Llama-Swallow-27b-it-v0.1, Gemma-2-Llama-Swallow-2b-it-v0.1
APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
arXiv10 reposarXiv:2504.03601
xLAM, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, Llama-xLAM-2-8b-fc-r, xLAM-2-1b-fc-r, xLAM-2-3b-fc-r, xLAM-2-32b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, APIGen-MT-5k, xLAM-2-3b-fc-r-gguf
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
arXiv10 reposarXiv:2510.13999
Qwen3-Coder-30B-A3B-REAP-AWQ, llm-compressor, Qwen3-Coder-REAP-25B-A3B, Qwen3-Coder-REAP-25B-A3B-AWQ, inkling-mlx, Inkling-MLX-REAP12-4bit, Inkling-Small-MLX-REAP25-4bit, Inkling-MLX-REAP25-4bit, Inkling-MLX-REAP50-4bit, turboquant-vllm
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
arXiv10 reposarXiv:2512.19687
pe-av-large, pe-av-small-16-frame, pe-av-base-16-frame, pe-av-large-16-frame, pe-av-small, pe-a-frame-small, pe-a-frame-base, pe-a-frame-large, pe-av-base, Echo-ViLD
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
arXiv10 reposarXiv:2604.27393
MiniCPM-V-4.6-Thinking-GPTQ, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB
MolmoAct2: Action Reasoning Models for Real-world Deployment
arXiv10 reposarXiv:2605.02881
MolmoAct2, MolmoAct2-Pretrain, Molmo2-ER, MolmoAct2-SO100_101, MolmoAct2-Think, MolmoAct2-BimanualYAM, MolmoAct2-Think-LIBERO, MolmoAct2-LIBERO, MolmoAct2-DROID, vla-edge
MOSS-VL Technical Report
arXiv10 reposarXiv:2608.15045
MOSS-VL-Base-0708, MOSS-VL-Base-0408, MOSS-VL, MOSS-VL-Instruct-0708-NF4, MOSS-VL-Instruct-0708-FP8, MOSS-VL-Realtime-NF4, MOSS-VL-Realtime-FP8, MOSS-VL-Realtime, MOSS-VL-Instruct-0408, MOSS-VL-Instruct-0708
Fast Transformer Decoding: One Write-Head is All You Need
arXiv9 reposarXiv:1911.02150
ChatGLM-6B, ChatGLM-6B, falcon-40b, chatglm2-6b-32k, chatglm2-6b, falcon-7b, chatglm2-6b-int4, falcon-40b-instruct, WebGLM
Revisiting Pre-Trained Models for Chinese Natural Language Processing
arXiv9 reposarXiv:2004.13922
chinese-roberta-wwm-ext, chinese-xlnet-base, chinese-macbert-base, chinese-macbert-large, chinese-bert-wwm-ext, chinese-electra-base-discriminator, pycorrector, macbert4csc-base-chinese, repo-5146-pycorrector
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding
arXiv9 reposarXiv:2009.05387
indobert-base-p2, indobert-lite-large-p1, indobert-lite-large-p2, indobert-base-p1, indobert-large-p1, indobert-lite-base-p2, indobert-lite-base-p1, indobert-large-p2, TweetSentimentIndoBERT
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
arXiv9 reposarXiv:2204.07705
instructor-embedding, COIG, COIG, FLAN, tk-instruct-11b-def-pos, Tk-Instruct, instructor-embedding, tk-instruct-3b-def, tk-instruct-11b-def
GIT: A Generative Image-to-text Transformer for Vision and Language
arXiv9 reposarXiv:2205.14100
git-large, GenerativeImage2Text, heron-chat-git-Llama-2-7b-v0, heron-preliminary-git-Llama-2-70b-v0, heron-chat-git-ja-stablelm-base-7b-v1, heron-chat-git-ja-stablelm-base-7b-v0, heron-chat-git-ELYZA-fast-7b-v0, CVProject, git-large-coco
PaLI: A Jointly-Scaled Multilingual Language-Image Model
arXiv9 reposarXiv:2209.06794
siglip-so400m-patch14-384, siglip2-giant-opt-patch16-384, siglip2-so400m-patch16-512, siglip2-base-patch16-224, siglip2-base-patch16-naflex, siglip-base-patch16-224, siglip2-so400m-patch16-naflex, siglip2-so400m-patch14-384, siglip2-so400m-patch14-224
GLM-130B: An Open Bilingual Pre-trained Model
arXiv9 reposarXiv:2210.02414
chatglm2-6b-32k, chatglm2-6b, chatglm2-6b-int4, chatglm3-6b-32k, chatglm-6b, chatglm3-6b-32k, chatglm3-6b-base, visualglm-6b, P-tuning-v2
ReAct: Synergizing Reasoning and Acting in Language Models
arXiv9 reposarXiv:2210.03629
Awesome-OpenClaw, prompt-engineering, tau-bench, Qwen-7B-Chat, Qwen-14B-Chat, ReWOO, agentodyssey, AgentTuning, hacktech24_app_testing
ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge
arXiv9 reposarXiv:2303.14070
medalpaca-13b, medalpaca-7b, doctorwithbloom, doctorwithbloom, doctorwithbloomz-7b1, doctorwithbloomz-7b1-mt, medalpaca-lora-30b-8bit, medalpaca-lora-7b-8bit, medalpaca-lora-13b-8bit
Exploring Human-Like Translation Strategy with Large Language Models
arXiv9 reposarXiv:2305.04118
llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA
Iterative Translation Refinement with Large Language Models
arXiv9 reposarXiv:2306.03856
llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA, topxgen-llama-4-scout-TEaR
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
arXiv9 reposarXiv:2306.09237
legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, legal-english-roberta-base, legal-swiss-roberta-large, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large
ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
arXiv9 reposarXiv:2310.04564
PowerInfer, ReluLLaMA-7B, ReluLLaMA-13B, prosparse-llama-2-13b, prosparse-llama-2-7b, Bamboo-base-v0_1, ReluFalcon-40B, Bamboo-DPO-v0_1, ReluLLaMA-70B
Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
arXiv9 reposarXiv:2312.02145
Marigold, marigold, marigold, marigold-depth-v1-1, marigold-normals-v1-1, marigold-iid-appearance-v1-1, marigold-iid-lighting-v1-1, marigold-depth-lcm-v1-0, prisma
Bilateral Reference for High-Resolution Dichotomous Image Segmentation
arXiv9 reposarXiv:2401.03407
BiRefNet_HR, BiRefNet, FeyNobg, BiRefNet, BiRefNet_HR-matting, BiRefNet_dynamic, BiRefNet-matting, BiRefNet_lite-2K, BiRefNet-portrait
TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
arXiv9 reposarXiv:2402.16379
llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA, topxgen-llama-4-scout-TEaR
Evolutionary Optimization of Model Merging Recipes
arXiv9 reposarXiv:2403.13187
evolutionary-model-merge, EvoLLM-JP-A-v1-7B, EvoLLM-JP-v1-10B, EvoLLM-JP-v1-7B, EvoVLM-JP-v1-7B, JA-VG-VQA-500, JA-VLM-Bench-In-the-Wild, python_moder_merge, adaptation-evolutionary-model-merge
OpenVLA: An Open-Source Vision-Language-Action Model
arXiv9 reposarXiv:2406.09246
modified_libero_rlds, openvla-7b, openvla, openvla-7b-finetuned-libero-spatial, openvla-7b-finetuned-libero-object, openvla-7b-finetuned-libero-10, openvla-7b-finetuned-libero-goal, openvla-7b-prismatic, openvla-main-dhot1
Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts
arXiv9 reposarXiv:2409.06790
llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
arXiv9 reposarXiv:2410.05160
VLM2Vec-V2.0, MMEB-train, VLM2Vec-LLaVa-Next, VLM2Vec-Full, MMEB-V2, VLM2Vec-Qwen2VL-2B, MMEB-V3, VLM2Vec-Qwen2VL-7B, MMEB-eval
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
arXiv9 reposarXiv:2410.05243
UGround, UGround-V1-72B, UGround-V1-2B, UGround-V1-7B, llava_uground, android_world_seeact_v, Mind2Web_Live_SeeAct_V, UGround, mind2web-live-seeact-v
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
arXiv9 reposarXiv:2501.02669
VLM_S2H, Eagle-X2-Llama3-8B-ConsecutiveTableReadout-Mix-160k, Eagle-X2-Llama3-8B-GridNavigation-AlignMixPlus-120k, Eagle-X2-Llama3-8B, Eagle-X2-Llama3-8B-GridNavigation-MixPlus-120k, Eagle-X2-Llama3-8B-VisualAnalogy-AlignMixPlus-120k, Eagle-X2-Llama3-8B-TableReadout-MixPlus-240k, Eagle-X2-Llama3-8B-VisualAnalogy-MixPlus-120k, Eagle-X2-Llama3-8B-TableReadout-AlignMixPlus-240k
Provence: efficient and robust context pruning for retrieval-augmented generation
arXiv9 reposarXiv:2501.16214
provence-reranker-debertav3-v1, squeez, semantic-highlight-bilingual-v1, open_provence, open-provence-reranker-v1, query-context-pruner-multilingual-Qwen3-4B, open-provence-reranker-v1-gte-modernbert-base, open-provence-reranker-large-v1, open-provence-reranker-xsmall-v1
DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
arXiv9 reposarXiv:2503.01183
DiffRhythm, DiffRhythm-base, DiffRhythm-full, DiffRhythm-vae, DiffRhythm2, Diffrhythm-linux, DiffRhythm, DiffRhythm-windows, DiffRhythm
Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation
arXiv9 reposarXiv:2503.04554
topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-BoA, llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-TEaR
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
arXiv9 reposarXiv:2504.10160
topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-BoA, llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp
Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
arXiv9 reposarXiv:2505.09358
Marigold, marigold, marigold-depth-lcm-v1-0, marigold-normals, marigold-iid, marigold-depth-v1-1, marigold-normals-v1-1, marigold-iid-appearance-v1-1, marigold-iid-lighting-v1-1
Learning to Reason without External Rewards
arXiv9 reposarXiv:2505.19590
Qwen3-14B-Intuitor-MATH-1EPOCH, Qwen2.5-1.5B-Intuitor-MATH-1EPOCH, OLMo-2-7B-SFT-Intuitor-MATH-1EPOCH, Qwen2.5-3B-Intuitor-MATH-1EPOCH, Intuitor, Qwen2.5-3B-GRPO-MATH-1EPOCH, OLMo-2-7B-SFT-GRPO-MATH-1EPOCH, Qwen3-14B-GRPO-MATH-1EPOCH, Qwen2.5-1.5B-GRPO-MATH-1EPOCH
Show-o2: Improved Native Unified Multimodal Models
arXiv9 reposarXiv:2506.15564
Show-o, show-o2-1.5B, show-o2-1.5B-w-video-und, show-o2-7B, show-o2-7B-w-video-und, show-o-512x512-wo-llava-tuning, show-o, show-o2-1.5B-HQ, show-o-w-clip-vit
gpt-oss-120b & gpt-oss-20b Model Card
arXiv9 reposarXiv:2508.10925
gpt-oss-120b, gpt-oss-20b, GPT-OSS-Swallow-120B-RL-v0.1, GPT-OSS-Swallow-20B-RL-v0.1, GPT-OSS-Swallow-20B-SFT-v0.1, GPT-OSS-Swallow-120B-SFT-v0.1, Medical-GPT-OSS-Swallow-120B, GPT-OSS-Swallow-120B-RL-v0.1-MXFP4, GPT-OSS-Swallow-20B-RL-v0.1-MXFP4
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
arXiv9 reposarXiv:2510.08668
Hulu-Med-32B, Hulu-Med-4B, Hulu-Med-7B, Hulu-Med-14B, Hulu-Med, Hulu-Med-235A22, Hulu-Med-Flash-Preview-27B, Hulu-Med-30A3, hulumed-prober
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
arXiv9 reposarXiv:2510.11919
topxgen-llama-4-scout-REFINE, llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
arXiv9 reposarXiv:2511.19418
CoVT, CoVT-Dataset, CoVT-7B-seg, CoVT-7B-seg_depth_dino_edge, CoVT-LLaVA-13B-depth, CoVT-7B-seg_depth_dino, CoVT-7B-depth, covt_data, COVT
LFM2 Technical Report
arXiv9 reposarXiv:2511.23404
LFM2.5-Audio-1.5B, LFM2.5-350M, LFM2-VL-450M-ONNX, LFM2-2.6B-Exp-ONNX, LFM2-8B-A1B-ONNX, LFM2.5-Audio-1.5B-JP, LFM2-ColBERT-350M, kani-tts-2-en, kani-tts-2-pt
Radiology Report Generation with Layer-Wise Anatomical Attention
arXiv9 reposarXiv:2512.16841
LAnA-v3, layer-wise-anatomical-attention, LAnA-Arxiv, LAnA, LAnA-v2, LAnA-MIMIC, LAnA-v5, LAnA-MIMIC-CHEXPERT, LAnA-v4
Data Science and Technology Towards AGI Part I: Tiered Data Management
arXiv9 reposarXiv:2602.09003
Ultra-FineWeb-L1, UltraData-RL-2609, UltraData-SFT-Agent-2609, UltraData-Code, MiniCPM5-2B, MiniCPM5-2B-GGUF, MiniCPM5-1B-GGUF, MiniCPM5-1B, Ultra-FineWeb
MOSS-TTS Technical Report
arXiv9 reposarXiv:2603.18090
MOSS-TTS-v1.5, MOSS-TTS-Nano, MOSS-TTS, MOSS-TTS-Local-Transformer-v1.5, MOSS-Audio-Tokenizer-Nano, MOSS-TTS-Realtime, MOSS-TTS-Norwegian-LoRA, MOSS-TTS-Nano-100M-ONNX, MOSS-Audio-Tokenizer-Nano-ONNX
MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention
arXiv9 reposarXiv:2606.07639
MOSS-VL-Base-0408, MOSS-VL-Instruct-0708-NF4, MOSS-VL-Instruct-0708-FP8, MOSS-VL-Realtime-NF4, MOSS-VL-Realtime-FP8, MOSS-VL-Realtime, MOSS-VL-Base-0708, MOSS-VL-Instruct-0408, MOSS-VL-Instruct-0708
Towards a general-purpose foundation model for computational pathology
Nature9 reposNature:s41591-024-02857-3
AtlasPatch, PIANO, TRIDENT, CPathPatchFeature, SEAL, HistAug, TridentEdited, histaug-uni, aegis
DCSpell: A Detector-Corrector Framework for Chinese Spelling Error Correction
ACM8 reposACM:3404835.3463050
macbert4mdcspell_v1, macbert4mdcspell_v2, macbert4csc_v2, macbert4mdcspell_v3, macbert4csc_v1, bert4csc_v1, Chinese-text-correction-papers, CTCResources
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
arXiv8 reposarXiv:2010.05646
MockingBird, nemo-nano-codec-22khz-1.89kbps-21.5fps, tts-hifigan-libritts-16kHz, fish-diffusion, tts-hifigan-ljspeech, dla-tts, hifi-gan, univnet
KLUE: Korean Language Understanding Evaluation
arXiv8 reposarXiv:2105.09680
roberta-large, klue, pko-t5-base, bert-base, KLUE, pko-t5, pko-t5-large, pko-t5-small
SciFive: a text-to-text transformer model for biomedical literature
arXiv8 reposarXiv:2106.03598
SciFive-large-Pubmed_PMC-MedNLI, SciFive, SciFive-base-Pubmed_PMC, SciFive-large-Pubmed_PMC, SciFive-large-Pubmed, SciFive-base-Pubmed, SciFive-base-PMC, SciFive-large-PMC
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
arXiv8 reposarXiv:2111.09543
DiagnosisCoding, LEXTREME, mdeberta-v3-base, DeBERTa, deberta-v3-small, deberta-v3-xsmall, deberta-v3-base, deberta-v3-large
OPT: Open Pre-trained Transformer Language Models
arXiv8 reposarXiv:2205.01068
safe-rlhf, opt-125m, OPT-13B-Erebus, opt-13b, opt-2.7b, OPT-6B-nerys-v2, OPT-6.7B-Erebus, opt-350m
Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence
arXiv8 reposarXiv:2209.02970
Fengshenbang-LM, Erlangshen-MegatronBert-1.3B, Taiyi-Stable-Diffusion-1B-Chinese-EN-v0.1, Taiyi-Stable-Diffusion-1B-Chinese-v0.1, Erlangshen-DeBERTa-v2-710M-Chinese, Erlangshen-DeBERTa-v2-97M-Chinese, Erlangshen-DeBERTa-v2-320M-Chinese, BDM1.0
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
arXiv8 reposarXiv:2211.06687
larger_clap_music, clap-htsat-unfused, hashing-baseline, CLAP, clap-demo, larger_clap_music_and_speech, audio-embeddings, shira_audio
Scalable Diffusion Models with Transformers
arXiv8 reposarXiv:2212.09748
DiT, nanoMFM, smalldiffusion, mdlm, bd3lms, bd3lm, BDM, BD_DNA
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
arXiv8 reposarXiv:2301.13688
OpenOrca, dolma, FLAN, Mistral-7B-OpenOrca, Jellyfish-13B, OpenOrca-KO, KOR-OpenOrca-Platypus, OpenOrca-gugugo-ko
T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
arXiv8 reposarXiv:2302.08453
t2i-adapter-canny-sdxl-1.0, T2I-Adapter, T2I-Adapter, t2i-adapter-openpose-sdxl-1.0, t2i-adapter-sketch-sdxl-1.0, t2i-adapter-lineart-sdxl-1.0, t2i-adapter-depth-zoe-sdxl-1.0, t2i-adapter-depth-midas-sdxl-1.0
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
arXiv8 reposarXiv:2306.01116
fineweb, falcon-refinedweb, falcon-40b, falcon-rw-1b, falcon-7b, falcon-40b-instruct, SEA-PILE-v1, llm_project
INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models
arXiv8 reposarXiv:2306.04757
InstructEvalImpact, flan-alpaca-gpt4-xl, flan-gpt4all-xl, flan-alpaca-large, flan-alpaca-base, flan-alpaca-xl, flan-alpaca-xxl, flan-sharegpt-xl
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
arXiv8 reposarXiv:2308.01390
Otter, OpenFlamingo-9B-vitl-mpt7b, Otter, open_flamingo, OpenFlamingo-3B-vitl-mpt1b-langinstruct, OpenFlamingo-3B-vitl-mpt1b, OpenFlamingo-4B-vitl-rpj3b, OpenFlamingo-4B-vitl-rpj3b-langinstruct
Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
arXiv8 reposarXiv:2308.09662
flan-alpaca-gpt4-xl, red-instruct, starling-7B, HarmfulQA, flan-alpaca-large, flan-alpaca-base, flan-alpaca-xl, flan-sharegpt-xl
ModelScope-Agent: Building Your Customizable Agent System with Open-source Large Language Models
arXiv8 reposarXiv:2309.00986
ms-agent, swift, modelscope-agent, ms-swift, ms-swift, vmopd, ms, dense-retention-rl
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
arXiv8 reposarXiv:2309.12284
SOLAR-10.7B-Instruct-v1.0, Arithmo2-Mistral-7B, MetaMathQA, Arithmo2-Mistral-7B-adapter, Arithmo-Mistral-7B, Arithmo-Data, GSM8K_Backward, neural-chat-7b-v3-3
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
arXiv8 reposarXiv:2310.03731
MathCoder, MathCodeInstruct-Plus, MathCodeInstruct, MathCoder-CL-7B, MathCodeInstruct, MathCoder-L-7B, MathCoder-L-13B, MathCoder-CL-34B
A Multi-Task Embedder For Retrieval Augmented LLMs
arXiv8 reposarXiv:2310.07554
bge-small-zh-v1.5, bge-large-zh-v1.5, bge-base-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-large-en, bge-base-en, bge-small-en
Skywork: A More Open Bilingual Foundation Model
arXiv8 reposarXiv:2310.19341
Skywork-13B-base, Skywork, mock_gsm8k_test, Skywork-13B-Math-8bits, Skywork-13B-Math, Skywork-13B-Base-3.1TB, ChineseDomainModelingEval, Skywork-13B-Base-8bits
Seamless: Multilingual Expressive and Streaming Speech Translation
arXiv8 reposarXiv:2312.05187
w2v-bert-2.0, seamless_communication, seamless-streaming, seamless-m4t-medium, seamless-m4t-v2-large, seamless-m4t-large, conformer-shaw, fairseq2
AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One
arXiv8 reposarXiv:2312.06709
RADIO, C-RADIOv3-L, C-RADIOv3-B, C-RADIOv4-SO400M, C-RADIOv4-H, C-RADIOv3-g, C-RADIOv3-H, RADIO
CogAgent: A Visual Language Model for GUI Agents
arXiv8 reposarXiv:2312.08914
CogVLM, CogAgent, cogagent-chat-hf, cogagent-vqa-hf, CogAgent, cogagent-9b-20241220, CogVLM, CogAgent
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
arXiv8 reposarXiv:2312.16886
MobileVLM, MobileLLaMA-2.7B-Chat, MobileLLaMA-1.4B-Chat, MobileVLM-1.7B, MobileVLM-3B, MobileLLaMA-1.4B-Base, MobileLLaMA-2.7B-Base, MobileLISA
Improving Text Embeddings with Large Language Models
arXiv8 reposarXiv:2401.00368
DermL2V-training, Anchor-Embedding, DermL2V-tmp, llm2vec, SFR-Embedding-Mistral, tinyllama-embed, e5-mistral-7b-instruct, multilingual-e5-large-instruct
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
arXiv8 reposarXiv:2401.06066
DeepSeek-Coder-V2-Instruct, DeepSeek-Coder-V2, DeepSeek-Coder-V2-Lite-Base, DeepSeek-Coder-V2-Lite-Instruct, DeepSeek-Coder-V2-Base, deepseek-moe-16b-chat, deepseek-moe-16b-base, nugie-jax-nemotron-3-nano
M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
arXiv8 reposarXiv:2402.03216
bge-m3, bge-reranker-v2-m3, OKEAN, semantic-highlight-bilingual-v1, pyterrier_dr, pyterrier_dr_jpq, gte-multilingual-base, bge-reranker-v2.5-gemma2-lightweight
MOMENT: A Family of Open Time-series Foundation Models
arXiv8 reposarXiv:2402.03885
Samay, MOMENT-1-large, moment, MOMENT-1-small, Timeseries-PILE, MOMENT-1-base, FM4Motor, TiViT
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
arXiv8 reposarXiv:2402.04249
OBLITERATUS, garak, JBB-Behaviors, HarmBench, HarmBench-Llama-2-13b-cls, HarmBench-Mistral-7b-val-cls, HarmBench-Llama-2-13b-cls-multimodal-behaviors, karma-electric-llama31-8b
Towards Building Multilingual Language Model for Medicine
arXiv8 reposarXiv:2402.13963
MMedC, MMed-Llama-3-8B, MMedLM, MMedLM2, MMedBench, MMedLM2-1.8B, MMed-Llama-3-8B-EnIns, MMedLM
M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
arXiv8 reposarXiv:2404.00578
M3D, M3D-RefSeg, M3D-LaMed-Llama-2-7B, M3D-CLIP, M3D-VQA, M3D-Seg, RAIdio-Agents, adapt_med_seg
Towards Large-Scale Training of Pathology Foundation Models
arXiv8 reposarXiv:2404.15217
towards_large_pathology_fms, vit_base_patch16_224.kaiko_ai_towards_large_pathology_fms, vit_small_patch16_224.kaiko_ai_towards_large_pathology_fms, vit_large_patch14_reg4_dinov2.kaiko_ai_towards_large_pathology_fms, vit_small_patch8_224.kaiko_ai_towards_large_pathology_fms, vit_base_patch8_224.kaiko_ai_towards_large_pathology_fms, Midnight, vit_large_patch14_reg4_224.kaiko_ai_towards_large_pathology_fms
RLHF Workflow: From Reward Modeling to Online RLHF
arXiv8 reposarXiv:2405.07863
Online-RLHF, LLaMA3-SFT-v2, LLaMA3-SFT, pair-preference-model-LLaMA3-8B, Llama3-SFT-v2.0-epoch3, Llama3-SFT-v2.0-epoch1, FsfairX-LLaMA3-RM-v0.1, LLaMA3-iterative-DPO-final
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
arXiv8 reposarXiv:2405.20797
Ovis2.5-9B, Ovis2-34B, Ovis1.6-Gemma2-9B-GPTQ-Int4, Ovis1.6-Llama3.2-3B-GPTQ-Int4, Ovis1.6-Llama3.2-3B, Ovis2-34B-GPTQ-Int4, Ovis1.6-Gemma2-9B, Ovis2.5-2B
Refusal in Language Models Is Mediated by a Single Direction
arXiv8 reposarXiv:2406.11717
ds4, OBLITERATUS, heretic, obliteratus, Heretic-Abliteration, som-refusal-directions, ds4, ds4
Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
arXiv8 reposarXiv:2406.16860
cambrian-13b, cambrian-34b, cambrian-phi3-3b, CV-Bench, cambrian-8b, Cambrian-10M, Cambrian-Alignment, phyworld
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
arXiv8 reposarXiv:2406.18518
xLAM-7b-r, xLAM-1b-fc-r, xLAM-7b-fc-r, xLAM-1b-fc-r-gguf, xLAM-v0.1-r, xLAM-8x22b-r, xLAM-8x7b-r, xLAM-7b-fc-r-gguf
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
arXiv8 reposarXiv:2406.18629
Math-Step-DPO-10K, Step-DPO, DeepSeekMath-RL-Step-DPO, Qwen2-7B-SFT-Step-DPO, Llama-3-70B-SFT-Step-DPO, Qwen2-57B-A14B-SFT-Step-DPO, Qwen1.5-32B-SFT-Step-DPO, Qwen2-72B-Instruct-Step-DPO
Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
arXiv8 reposarXiv:2409.11136
ru-promptriever, promptriever, RepLLaMA-reproduced, promptriever-mistral-v0.1-7b-v1, promptriever-llama3.1-8b-instruct-v1, promptriever-llama3.1-8b-v1, promptriever-llama2-7b-v1, ru-promptriever-qwen3-4b
PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation
arXiv8 reposarXiv:2410.01680
RADIO, C-RADIOv3-L, C-RADIOv3-B, C-RADIOv4-SO400M, C-RADIOv4-H, C-RADIOv3-g, C-RADIOv3-H, RADIO
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
arXiv8 reposarXiv:2410.10629
nunchaku, nunchaku, Sana_1600M_1024px_diffusers, SANA, Sana_600M_512px_diffusers, Sana_1600M_512px_diffusers, Sana-fork, Sana_1600M_512px_MultiLing
Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications
arXiv8 reposarXiv:2412.02732
Prithvi-EO-2.0-NYC-Pluvial, Prithvi-EO-2.0-300M-TL-Sen1Floods11, Prithvi-EO-2.0, Prithvi-EO-2.0-300M, Prithvi-EO-2.0-600M, Prithvi-EO-2.0-300M-TL, Prithvi-EO-2.0-600M-TL, Land-Change-Detection
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
arXiv8 reposarXiv:2501.12948
awesome-deepseek-prompts, DeepSeek-R1, DeepSeek-R1-0528, DeepSeek-R1-Distill-Llama-70B, DeepSeek-R1-Distill-Qwen-32B, VLAC, DeepSeek-R1-0528-Qwen3-8B, marcello
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
arXiv8 reposarXiv:2502.04128
ch-tts-llasa-rl-grpo, xcodec2, LLaSA_training, X-Codec-2.0, xcodec2, Llasa_opensource_speech_data_160k_hours_tokenized, xcodec2-hf, tts
R2MED: A Benchmark for Reasoning-Driven Medical Retrieval
arXiv8 reposarXiv:2505.14558
Bioinformatics, Biology, MedXpertQA-Exam, Medical-Sciences, MedQA-Diag, PMC-Clinical, IIYi-Clinical, PMC-Treatment
Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning
arXiv8 reposarXiv:2505.16410
Tool-Star, Tool-Star-Qwen-3B, Tool-Star-Qwen-1.5B, Tool-Star-Qwen-0.5B, Multi-Tool-RL-10K, Tool-Star-Qwen-7B, Tool-Star-SFT-54K, Tool-Star
OmniGen2: Towards Instruction-Aligned Multimodal Generation
arXiv8 reposarXiv:2506.18871
OmniGen2, OmniGen2, OmniGen2-EditScore7B, OmniContext, Macro-OmniGen2, X2I2, ComfyUI-OmniGen2, OmniGen2-EditScore7B-v1.1
GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface
arXiv8 reposarXiv:2507.18546
gliner2-base-v1, gliner2-large-v1, gliner2-multi-v1, gliner2.5-base-v1, gliner2.5-small-v1, gliner2.5-multi-v1, ArXiv-GLiNER2-Intelligence-Extractor, gliner2.5-multi-v1-onnx
RynnEC: Bringing MLLMs into Embodied World
arXiv8 reposarXiv:2508.14160
WorldVLA, RynnVLA-002, RynnVLA-001, RynnEC, RynnEC-2B, RynnEC-Bench, RynnEC-7B, RynnVLA-002
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
arXiv8 reposarXiv:2508.14706
ShizhenGPT-7B-Omni, TCM-Pretrain-Data-ShizhenGPT, ShizhenGPT-32B-VL, TCM-Instruction-Tuning-ShizhenGPT, ShizhenGPT, ShizhenGPT-7B-LLM, ShizhenGPT-7B-VL, ShizhenGPT-32B-LLM
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
arXiv8 reposarXiv:2508.20751
UniGenBench_Leaderboard, UniGenBench, UniGenBench-EvalModel-qwen-72b-v1, UniGenBench-Eval-Images, UniGenBench-EvalModel-qwen3vl-32b-v1, UniGenBench_Leaderboard_Chinese, UniGenBench_Leaderboard_English_Long, UniGenBench_Leaderboard_Chinese_Long
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding
arXiv8 reposarXiv:2509.23379
CCD, libra-v1.0-7b, libra-llava-med-v1.5-mistral-7b, libra-llava-rad, libra-maira-2, CheXpert-plus-RRG, IU-Xray-RRG, Medical-CXR-VQA
FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
arXiv8 reposarXiv:2510.10921
FG-CLIP, fg-clip2-base, DCI-CN, BoxClass-CN, DOCCI-CN, LIT-CN, fg-clip2-so400m, fg-clip2-large
Uniform Discrete Diffusion with Metric Path for Video Generation
arXiv8 reposarXiv:2510.24717
URSA, URSA-0.6B-FSQ320, URSA-0.6B-IBQ1024, URSA-1.7B-IBQ512-UDMGRPO-GenEval, URSA-1.7B-IBQ1024, URSA-1.7B-FSQ320, URSA-1.7B-IBQ512-UDMGRPO-PickScore, URSA-1.7B-IBQ512
OneThinker: All-in-one Reasoning Model for Image and Video
arXiv8 reposarXiv:2512.03043
Video-R1, OneThinker, OneThinker-8B, OneThinker-SFT-Qwen3-8B, OneThinker-eval, OneThinker-train-data, CoLT-8B, CoLT_Train_Dataset
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
arXiv8 reposarXiv:2512.22905
JavisDiT, JavisGPT, JavisGPT-v1.0-7B-Instruct, JavisGPT-v0.1-7B-Instruct, JavisInst-Omni, AV-FineTune, JavisUnd-Eval, MM-PreTrain
GLM-5: from Vibe Coding to Agentic Engineering
arXiv8 reposarXiv:2602.15763
GLM-5.3-Flash, GLM-5.3, GLM-5.3-Flash-GGUF, GLM-5.2, GLM-5, GLM-5.1, GLM-5.2-EXL3-TR3-3.0bpw, GLM-5.2-FP8
Lance: Unified Multimodal Modeling by Multi-Task Synergy
arXiv8 reposarXiv:2605.18678
lance-mlx, Lance, Lance-3B-bf16, Lance-3B-AWQ-INT4, Lance-3B-8bit, Lance-3B-Video-bf16, Lance, lance-quant
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
arXiv8 reposarXiv:2607.13125
Boogu-Image-0.1-Edit, Boogu-Image-0.1-Turbo, Boogu-Image-0.1-Base, Boogu-Image, Boogu-Image-0.1-Edit-Turbo, Boogu-Image-0.1-Edit-fp8, Boogu-Image-0.1-Base-fp8, Boogu-Image-0.1-Turbo-fp8
A whole-slide foundation model for digital pathology from real-world data
Nature8 reposNature:s41586-024-07441-w
AtlasPatch, PIANO, TRIDENT, CPathPatchFeature, TridentEdited, TITAN, aegis, dpfm_factory
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
arXiv7 reposarXiv:1910.13461
llama-2-jax, dalle-mini, lares, GPT_Ranker, ClipSumary, QA-SLM, final-project-level3-nlp-02
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
arXiv7 reposarXiv:2006.11477
speech-emotion-recognition, HierSpeechpp, dissertation-project, GigaAM, wav2vec2-base, tmh, gsoc-wav2vec2
8-bit Optimizers via Block-wise Quantization
arXiv7 reposarXiv:2110.02861
bloom, bloom-optimizer-states, bloom-560m, bloom-1b7, bloom-7b1, bloom-1b1, bloom-3b
Multitask Prompted Training Enables Zero-Shot Task Generalization
arXiv7 reposarXiv:2110.08207
FLAN, promptsource, T0pp, P3, T0, T0_3B, art
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
arXiv7 reposarXiv:2110.13900
ASR-Transcription-Router, CrisperWhisper, faster_CrisperWhisper, UniSpeech, CrisperWhisper, moshi, wavlm-large
Progressive Distillation for Fast Sampling of Diffusion Models
arXiv7 reposarXiv:2202.00512
stable-diffusion-x4-upscaler, coreml-stable-diffusion-2-1-base, coreml-stable-diffusion-2-base-palettized, coreml-stable-diffusion-2-1-base-palettized, coreml-stable-diffusion-2-base, stable-diffusion-2-1-base, stable-diffusion-2-1
No Language Left Behind: Scaling Human-Centered Machine Translation
arXiv7 reposarXiv:2207.04672
nllb-moe-54b, Glot500, WhisperLiveKit, flores200, SONAR, nmtscore, NoLanguageLeftWaiting
LAION-5B: An open large-scale dataset for training next generation image-text models
arXiv7 reposarXiv:2210.08402
CLIP-convnext_xxlarge-laion2B-s34B-b82K-augreg-soup, OpenFlamingo-9B-vitl-mpt7b, OpenFlamingo-3B-vitl-mpt1b-langinstruct, OpenFlamingo-4B-vitl-rpj3b-langinstruct, OpenFlamingo-3B-vitl-mpt1b, OpenFlamingo-4B-vitl-rpj3b, CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft-soup
Fast Inference from Transformers via Speculative Decoding
arXiv7 reposarXiv:2211.17192
LMFlow, distil-whisper, LLMSpeculativeSampling, RemoteSpeculativeDecoding, Hierarchical-Speculative-Decoding, GPTFast, NoLanguageLeftWaiting
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
arXiv7 reposarXiv:2303.05499
geti-instant-learn, LocateAnything-3B, GroundingDINO, GroundingDINO, Grounding_DINO_demo, DINO, Flow-Inference-Time-Scaling
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data
arXiv7 reposarXiv:2304.01196
falcon-40b-instruct, baize-chatbot, baize-v2-13b, baize-v2-7b, baize, vigogne, rulm
Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text
arXiv7 reposarXiv:2304.06939
OpenFlamingo-9B-vitl-mpt7b, OpenFlamingo-3B-vitl-mpt1b-langinstruct, OpenFlamingo-4B-vitl-rpj3b-langinstruct, OpenFlamingo-3B-vitl-mpt1b, OpenFlamingo-4B-vitl-rpj3b, SEED, mmc4
LIMA: Less Is More for Alignment
arXiv7 reposarXiv:2305.11206
LimaRP, llm-finetuning, InstructionGPT-4, LLaMA2-Accessory, CoT-llama2, open-korean-instructions, ko-lima
Extending Context Window of Large Language Models via Positional Interpolation
arXiv7 reposarXiv:2306.15595
chatglm2-6b-32k, Open-Sora-Plan, Open-Sora-Plan-v1.3.0, Chinese-LLaMA-Alpaca-2, LongQLoRA, Long-QLORA, article_gpt
OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
arXiv7 reposarXiv:2306.16527
idefics-80b-instruct, idefics2-8b, Idefics3-8B-Llama3, idefics-9b-instruct, OBELISC, OBELICS, OBELICS
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
arXiv7 reposarXiv:2307.01952
stable-diffusion-xl-base-1.0, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-xl-base-with-refiner, coreml-stable-diffusion-xl-base-ios, stable-diffusion-xl-refiner-1.0, cagliostro-webui, diffusion-augmentation
SA-Solver: Stochastic Adams Solver for Fast Sampling of Diffusion Models
arXiv7 reposarXiv:2309.05019
PixArt-alpha, PixArt-XL-2-256x256, PixArt-XL-2-1024-MS, PixArt-XL-2-512x512, PixArt-Sigma-XL-2-512-MS, pixeart, PixArt-alpha
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
arXiv7 reposarXiv:2309.05516
auto-round, DeepSeek-R1-int2-mixed-sym-inc, vllm, vllm-turboquant, vllm-old, vllm_amd_sleep, vllm
AnglE-optimized Text Embeddings
arXiv7 reposarXiv:2309.12871
UAE-Large-V1, AnglE, UAE-Code-Large-V1, pubmed-angle-base-en, angle-llama-13b-nli, pubmed-angle-large-en, angle-llama-7b-nli-v2
Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling
arXiv7 reposarXiv:2311.00430
kotoba-whisper-v2.0, kotoba-whisper, distil-whisper, distil-large-v2, kotoba-whisper-v1.0, distil-medium.en, distil-small.en
Nomic Embed: Training a Reproducible Long Context Text Embedder
arXiv7 reposarXiv:2402.01613
contrastors, nomic-embed-text-v1.5, nomic-embed-text-v1-ablated, nomic-embed-text-v1-unsupervised, contrastors, ColBERT-Zero, modernbert-embed-base
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
arXiv7 reposarXiv:2402.01912
parler-tts-mini-v1, parler-tts, dataspeech, parler-tts-large-v1, libritts-r-filtered-speaker-descriptions, mls-eng-speaker-descriptions, parler-tts-vietnamese-v1-stage2
World Model on Million-Length Video And Language With Blockwise RingAttention
arXiv7 reposarXiv:2402.08268
LWM, Llama-3-8B-Instruct-Gradient-1048k, Llama-3-70B-Instruct-Gradient-262k, Llama-3-8B-Instruct-262k, Llama-3-70B-Instruct-Gradient-1048k, Llama-3-8B-Instruct-Gradient-4194k, lwm
IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus
arXiv7 reposarXiv:2402.14710
OneKE, IEPile, iepie, llama2-13b-iepile-lora, OneKE, baichuan2-13b-iepile-lora, iepile
Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
arXiv7 reposarXiv:2403.03163
Design2Code, Design2Code_human_eval_pairwise, Design2Code_human_eval_reference_vs_gpt4v, Design2Code-hf, Design2Code-18B-v0, Design2Code, Design2Code-HARD
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
arXiv7 reposarXiv:2404.01258
train_video_and_instruction, LLaVA-Hound-DPO, LLaVA-Hound-DPO, test_video_and_instruction, LLaVA-Hound-SFT, MixEval-Archon, MixEval
MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators
arXiv7 reposarXiv:2404.05014
MagicTime, MagicTime, ChronoMagic, MagicTime, ChronoMagic-Bench, ConsisID, OpenS2V-Nexus
MANTIS: Interleaved Multi-Image Instruction Tuning
arXiv7 reposarXiv:2405.01483
Mantis, Mantis-8B-Idefics2, Mantis-8B-clip-llama3, Mantis-8B-siglip-llama3, Mantis-Instruct, Mantis-8B-Fuyu, SpaceMantis
Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models
arXiv7 reposarXiv:2405.05374
cholesky_encoder, snowflake-arctic-embed-m, snowflake-arctic-embed-xs, snowflake-arctic-embed-s, snowflake-arctic-embed-l, snowflake-arctic-embed-m-long, arctic-embed
USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
arXiv7 reposarXiv:2405.07719
xDiT, mochi-xdit, HunyuanVideo, HunyuanVideo, HunyuanVideo-I2V, HunyuanVideo-I2V, HunyuanWorld-Voyager
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
arXiv7 reposarXiv:2405.08748
HunyuanDiT-v1.2-Diffusers, HunyuanDiT, HunyuanDiT-v1.1, HunyuanDiT-v1.2-Diffusers-Distilled, HunyuanDiT, HunyuanDiT-v1.2, HunyuanDiT
BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
arXiv7 reposarXiv:2406.09952
CLIP_COCO, CLIP_TROHN-Img, TROHN-Text, TROHN-Img, CLIP_TROHN-Text, CLIP_Detector, CLIP_TROHN-Img_Detector
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
arXiv7 reposarXiv:2406.18522
MagicTime, ChronoMagic-Bench, ChronoMagic-Pro, ChronoMagic-ProH, ChronoMagic-Bench, ConsisID, OpenS2V-Nexus
Embedding And Clustering Your Data Can Improve Contrastive Pretraining
arXiv7 reposarXiv:2407.18887
cholesky_encoder, snowflake-arctic-embed-m, snowflake-arctic-embed-xs, snowflake-arctic-embed-s, snowflake-arctic-embed-m-long, snowflake-arctic-embed-l, arctic-embed
Gemma 2: Improving Open Language Models at a Practical Size
arXiv7 reposarXiv:2408.00118
modded-nanogpt, Gemma-2-Llama-Swallow-9b-pt-v0.1, Gemma-2-Llama-Swallow-2b-pt-v0.1, Gemma-2-Llama-Swallow-27b-it-v0.1, Gemma-2-Llama-Swallow-27b-pt-v0.1, Gemma-2-Llama-Swallow-2b-it-v0.1, Gemma-2-Llama-Swallow-9b-it-v0.1
Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining
arXiv7 reposarXiv:2408.02657
Lumina-T2X, Lumina-mGPT, Lumina-mGPT-7B-512, Lumina-mGPT-7B-768, Lumina-mGPT-7B-768-Omni, Lumina-mGPT-7B-1024, Lumina-mGPT-34B-512
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
arXiv7 reposarXiv:2409.17146
Molmo-7B-D-0924, molmo, Molmo-7B-O-0924, MolmoE-1B-0924, pixmo-docs, CoSyn-400K, CoSyn-point
Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach
arXiv7 reposarXiv:2410.03160
PusaV1_training, Pusa-VidGen, PusaV0.5_Training, Pusa-Wan2.2-V1, PusaV1, Pusa-V0.5, Mochi-Full-Finetuner
arXiv:2410.09724
arXiv7 reposarXiv:2410.09724
Reward-Calibration, mistral-7b-ppo-c-hermes, llama3-8b-crm-final-v0.1, llama3-8b-final-ppo-m-v0.3, mistral-7b-ppo-m-hermes, mistral-7b-hermes-crm-skywork, llama3-8b-final-ppo-c-v0.3
How to Evaluate Reward Models for RLHF
arXiv7 reposarXiv:2410.14872
PPE, PPE-Human-Preference-V1, PPE-MMLU-Pro-Best-of-K, PPE-MATH-Best-of-K, PPE-GPQA-Best-of-K, PPE-IFEval-Best-of-K, PPE-MBPP-Plus-Best-of-K
Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances
arXiv7 reposarXiv:2410.18775
VINE, VINE-R-Dec, VINE-R-Enc, VINE-B-Enc, VINE-B-Dec, W-Bench, WMCopier
LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch
arXiv7 reposarXiv:2411.11171
LLaMmlein, LLaMmlein_120M_prerelease, LLaMmlein_120M, LLaMmlein_7B, LLaMmlein_1B, LLaMmlein_1B_prerelease, LLaMmlein-Dataset
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models
arXiv7 reposarXiv:2412.07679
RADIO, C-RADIOv3-L, C-RADIOv3-B, C-RADIOv4-SO400M, C-RADIOv4-H, C-RADIOv3-g, C-RADIOv3-H
Multimodal Latent Language Modeling with Next-Token Diffusion
arXiv7 reposarXiv:2412.08635
VibeVoice, VibeVoice-1.5B, VibeVoice-Realtime-0.5B, VibeVoice-7B, VibeVoice, VibeVoice-Large, VibeVoice
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
arXiv7 reposarXiv:2412.13663
indic-modernBERT, ModernBERT-large, DiagnosisCoding, LightOnOCR-2-1B, ModernBERT, vllm-factory, LightOnOCR-2-1B-base
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
arXiv7 reposarXiv:2501.00574
VideoChat-Flash-Qwen2_5-7B-1M_res224, VideoChat-Flash-Qwen2-7B_res448, VideoChat-Flash-Training-Data, InternVL_2_5_HiCo_R16, VideoChat-Flash-Qwen2_5-2B_res448, VideoChat-Flash-Qwen2_5-7B_InternVideo2-1B, VideoChat-Flash-Qwen2-7B_res224
TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models
arXiv7 reposarXiv:2502.06608
TripoSG-scribble, TripoSG, TripoSG, 2D23D, tripoSG-pipeline, avera, TripoSG-fc5c409
Magma: A Foundation Model for Multimodal AI Agents
arXiv7 reposarXiv:2502.13130
Magma, Magma-Mind2Web-SoM, Magma-AITW-SoM, Magma-8B, Magma-OXE-ToM, Magma-Video-ToM, Magma-820K
LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification
arXiv7 reposarXiv:2502.17421
longspec-longchat-13b-16k, longspec-Llama-3-8B-Instruct-262k, longspec-longchat-7b-v1.5-32k, longspec-QwQ-32B-Preview, longspec-vicuna-7b-v1.5-16k, longspec-vicuna-13b-v1.5-16k, longspec-data
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
arXiv7 reposarXiv:2502.20583
LiteASR, lite-whisper-large-v3-fast, lite-whisper-large-v3-turbo, lite-whisper-large-v3-turbo-fast, lite-whisper-large-v3-acc, lite-whisper-large-v3, lite-whisper-large-v3-turbo-acc
Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
arXiv7 reposarXiv:2503.01774
difix_ref, harmonizer, Difix3D, difix, Fixer, Difix3d-3dgs-demo, difix3d
VACE: All-in-One Video Creation and Editing
arXiv7 reposarXiv:2503.07598
Wan2.1-VACE-14B, Wan2.1, Wan2.1-VACE-1.3B, VACE-LTX-Video-0.9, VACE-Wan2.1-1.3B-Preview, VACE-Benchmark, VACE-Annotators
OmniSVG: A Unified Scalable Vector Graphics Generation Model
arXiv7 reposarXiv:2504.06263
OmniSVG, MMSVG-Illustration, OmniSVG1.1_4B, OmniSVG1.1_8B, OmniSVG, OmniSVG-train, omnisvg-train
Step1X-Edit: A Practical Framework for General Image Editing
arXiv7 reposarXiv:2504.17761
GEdit-Bench, Step1X-Edit, Step1X-Edit-v1p2-preview, Step1X-Edit-v1p2, Step1X-Edit-v1p1-diffusers, Step1X-Edit, Step1X-Edit
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
arXiv7 reposarXiv:2504.19874
turbovec, mlx-vlm, turboquant_plus, rotorquant, quant.cpp, turboquant-vllm, sqlite-vector
General-Reasoner: Advancing LLM Reasoning Across All Domains
arXiv7 reposarXiv:2505.14652
General-Reasoner, General-Reasoner-Qwen2.5-14B, General-Reasoner-Qwen3-14B, WebInstruct-verified, General-Reasoner-Qwen2.5-7B, General-Reasoner-Qwen3-4B, RLPR-Evaluation
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
arXiv7 reposarXiv:2505.17012
SpaceQwen2.5-VL-3B-Instruct, SpaceOm, SpaceThinker-Qwen2.5VL-3B, SpatialScore, SpatialScore, SpatialCorpus, SpatialScore
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
arXiv7 reposarXiv:2505.22705
HiDream-E1, HiDream-E1-1, HiDream-I1-Full, HiDream-E1-Full, HiDream-I1, HiDream-I1-Dev, HiDream-I1-Fast
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
arXiv7 reposarXiv:2507.05257
Mem-alpha, memalphaljx, knowl, MemoryAgentBench, MemoryAgentBench, tmp-mem-alpha, agentic-memory
VibeVoice Technical Report
arXiv7 reposarXiv:2508.19205
VibeVoice, VibeVoice-1.5B, VibeVoice-Realtime-0.5B, VibeVoice-7B, VibeVoice, VibeVoice-Large, VibeVoice
GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
arXiv7 reposarXiv:2510.16051
GUIrilla, GUIrilla-Trees, GUIrilla-Gold, GUIrilla-See-3B, GUIrilla-See-7B, GUIrilla-Task, GUIrilla-See-0.7B
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
arXiv7 reposarXiv:2510.20487
steering-eval-awareness-public, large-finetune, steering-eval-awareness, eval-evasion, wood_v2_sftr4_filt, steering, steering-eval-awareness-public-v2
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
arXiv7 reposarXiv:2511.00088
Alpamayo-R1-10B, alpamayo, alpamayo_, alpamayo-recipes, alphamayo_VLA_test_webui_opti., alpamayo1.5, RiskWorld
MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition
arXiv7 reposarXiv:2512.07348
MICo-150K, OmniGen2-MICo, Qwen-Image-MICo, MICo-150K, BLIP3o-Next-MICo, BAGEL-MICo, Lumina-DiMOO-MICo
LLaDA2.0: Scaling Up Diffusion Language Models to 100B
arXiv7 reposarXiv:2512.15745
LLaDA2.0-mini, LLaDA2.0-flash-preview, LLaDA2.0-flash, LLaDA2.0-mini-preview, LLaDA2.0-flash-CAP, LLaDA2.0-mini-CAP, dllm
Creating a Hybrid Rule and Neural Network Based Semantic Tagger using Silver Standard Data: the PyMUSAS framework for Multilingual Semantic Annotation
arXiv7 reposarXiv:2601.09648
experimental-wsd, English-USAS-Mosaico, PyMUSAS-Neural-English-Small-BEM, PyMUSAS-Neural-English-Base-BEM, PyMUSAS-Neural-Multilingual-Small-BEM, PyMUSAS-Neural-Multilingual-Base-BEM, USAS-WSD
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
arXiv7 reposarXiv:2601.10611
molmo2, MolmoWeb-8B, MolmoWeb-Pretrained-8B, MolmoWeb-Pretrained-4B, MolmoWeb-4B, MolmoWeb-8B-Native, MolmoWeb-4B-Native
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
arXiv7 reposarXiv:2601.19325
Innovator-VL, Innovator-VL-8B-Instruct, Innovator-VL-8B-Thinking, Innovator-VL-RL-172K, Innovator-VL-Instruct-46M, PreMidTrainVL-Qwen3Dense, PreMidTrainVL
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
arXiv7 reposarXiv:2601.22709
LLaVA-1.5-7B-GRACE-W4G128, Qwen3-VL-2B-GRACE-W4G128-AWQ, Qwen3-VL-2B-GRACE-W4G128, Qwen3-VL-2B-GRACE-BF16, LLaVA-1.5-7B-GRACE-W4G128-AWQ, Qwen3-VL-2B-GRACE-W8G128, GRACE-VLM
ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation
arXiv7 reposarXiv:2602.00744
acestep-v15-xl-base, Ace-Step1.5, ACE-Step-v1.5-chinese-new-year-LoRA, acestep-v15-xl-sft, acestep-v15-xl-turbo, acestep-v15-base, acestep-v15-sft
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
arXiv7 reposarXiv:2602.10934
MOSS-TTS-Nano, MOSS-Audio-Tokenizer-v2, MOSS-Audio-Tokenizer-Nano, MOSS-Audio-Tokenizer, MOSS-TTS-Realtime, MOSS-Audio-Tokenizer-Nano-ONNX, MOSS-TTS-Nano-100M-ONNX
The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder
arXiv7 reposarXiv:2602.18487
gliner-bi-base-v2.0, gliner-linker-large-v1.0, gliner-bi-small-v2.0, gliner-bi-large-v2.0, gliner-bi-edge-v2.0, gliner-linker-base-v1.0, gliner-linker-rerank-v1.0
arXiv:2602.20903
arXiv7 reposarXiv:2602.20903
TextPecker, TextPecker-8B-InternVL3, SD3.5M-TextPecker-SQPA, Flux.1-dev-TextPecker-SQPA, TextPecker-1.5M, QwenImage-TextPecker-SQPA, TextPecker-8B-Qwen3VL
FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System
arXiv7 reposarXiv:2603.10420
FireRedLID, FireRedPunc, FireRedASR2-LLM, FireRedVAD, FireRedASR2S, FireRedVAD, FireRedASR2S
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
arXiv7 reposarXiv:2604.04913
deltatok, deltatok-kinetics, seg-head-vspw, depth-head-kitti, rgb-head-imagenet, seg-head-cityscapes, deltaworld-kinetics
Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis
arXiv7 reposarXiv:2606.25369
Joyo-Kanji-Yomi-Benchmark, joyo-kanji-yomi-benchmark, kana-whisper, JoyoKanji-Yomi-Benchmark, sarashina2.2-tts, Irodori-TTS-v4-Small, sarashina2.2-tts
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
arXiv7 reposarXiv:2607.17423
TimeLens2-4B, TimeLens2-2B, TimeLens2-8B, TimeLens2-2B-SFT, TimeLens2-4B-SFT, TimeLens2-8B-SFT, TimeLens2-93K
A visual–language foundation model for pathology image analysis using medical Twitter
Nature7 reposNature:s41591-023-02504-3
PIANO, AtlasPatch, KEEP, KEEP, VLSA, PathPT, Histopathology_Benchmark
BERTScore: Evaluating Text Generation with BERT
arXiv6 reposarXiv:1904.09675
YiVal, gec-metrics, bert_score, lares, mslr-shared-task, KoBERTScore
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
arXiv6 reposarXiv:1909.09436
CodeBERTa-language-id, CodeBERTa-small-v1, CodeRL, codet5-large, codet5-large-ntp-py, CodeXGLUE
Root Mean Square Layer Normalization
arXiv6 reposarXiv:1910.07467
idefics-80b-instruct, tiny-vllm, gemma-scope-2b-pt-transcoders, transformer-tricks, idefics-9b-instruct, llama2.zig
Unsupervised Cross-lingual Representation Learning at Scale
arXiv6 reposarXiv:1911.02116
Multilingual-MiniLM-L12-H384, xlm-roberta-base, xlm-roberta-large, XLM, tydiqa-primary-task-xlm-roberta-large, caption
YOLOv4: Optimal Speed and Accuracy of Object Detection
arXiv6 reposarXiv:2004.10934
darknet, darknetcv, tensorflow-yolov4-tflite, Complex-YOLOv4-Pytorch, darknet, pytorch-YOLOv4
Denoising Diffusion Probabilistic Models
arXiv6 reposarXiv:2006.11239
Diff-Pruning, ddpm-ema-bedroom-256, ddpm-cifar10-32, deepfake_multiLID, DDPM_vs_DDIM, smalldiffusion
Language-agnostic BERT Sentence Embedding
arXiv6 reposarXiv:2007.01852
Multilingual-CLIP, xnli_bn, squad_bn, russe_detox_2022, russe_detox_2022, MERA
Beyond English-Centric Multilingual Machine Translation
arXiv6 reposarXiv:2010.11125
m2m100_1.2B, EasyNMT, m2m100_418M, Easy-Translate, ContraDecode, m2m100-12B-avg-5-ckpt
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
arXiv6 reposarXiv:2108.01073
stable-diffusion, stable-diffusion-xl-base-1.0, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-xl-base-with-refiner, coreml-stable-diffusion-xl-base-ios, stable-diffusion-xl-refiner-1.0
MoEfication: Transformer Feed-forward Layers are Mixtures of Experts
arXiv6 reposarXiv:2110.01786
ReluLLaMA-7B, ReluLLaMA-13B, ReluLLaMA-70B, Bamboo-base-v0_1, ReluFalcon-40B, Bamboo-DPO-v0_1
Masked Autoencoders Are Scalable Vision Learners
arXiv6 reposarXiv:2111.06377
mae, hls-foundation-os, TiViT, vit-mae-large, vit-mae-huge, vit-mae-base
LiT: Zero-Shot Transfer with Locked-image text Tuning
arXiv6 reposarXiv:2111.07991
vision_transformer, CLIP_benchmark, contrastors, nomic-embed-vision-v1, nomic-embed-vision-v1.5, contrastors
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
arXiv6 reposarXiv:2112.01488
finetuning-bert-for-IR, plaidrepro, colbertv2.0, ColBERT, jina-colbert-v1-en, KolBERT
Datasheet for the Pile
arXiv6 reposarXiv:2201.07311
minipile, pythia-12b, pythia-1.4b, pythia-410m, pythia-2.8b, diff-codegen-6b-v2
From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective
arXiv6 reposarXiv:2205.04733
splade-cocondenser-ensembledistil, splade, splade-cocondenser-selfdistil, splade-ecommerce-esci, Splade_PP_en_v1, SPLADERunner
Generative Language Models for Paragraph-Level Question Generation
arXiv6 reposarXiv:2210.03992
tweetnlp, qg_tweetqa, qag_tweetqa, t5-small-tweetqa-qa, t5-base-tweetqa-qag, t5-large-squad
SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control
arXiv6 reposarXiv:2210.17432
ssd-lm, mdlm, bd3lm, bd3lms, BDM, BD_DNA
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers
arXiv6 reposarXiv:2211.01324
stable-diffusion-xl-base-1.0, paint-with-words-sd, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-xl-base-with-refiner, coreml-stable-diffusion-xl-base-ios, stable-diffusion-xl-refiner-1.0
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese
arXiv6 reposarXiv:2211.01335
Chinese-CLIP, chinese-clip-vit-huge-patch14, chinese-clip-vit-base-patch16, chinese-clip-vit-large-patch14, chinese-clip-vit-large-patch14-336px, BDM1.0
Text Embeddings by Weakly-Supervised Contrastive Pre-training
arXiv6 reposarXiv:2212.03533
ROOT-RAG, e5-large-unsupervised, e5-small-unsupervised, e5-base-unsupervised, e5-mistral-7b-instruct, e5-large-v2
Benchmarking Self-Supervised Learning on Diverse Pathology Datasets
arXiv6 reposarXiv:2212.04690
resnet50.lunit_mocov2, resnet50.lunit_swav, vit_small_patch8_224.lunit_dino, resnet50.lunit_bt, vit_small_patch16_224.lunit_dino, benchmark-ssl-pathology
SantaCoder: don't reach for the stars!
arXiv6 reposarXiv:2301.03988
sven_modified, sven, bigcode-evaluation-harness, santacoder-fim-task, MultiPL-E, MultiPL-E
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
arXiv6 reposarXiv:2303.00915
MediMeta-C, RobustMedCLIP, RobustMedCLIP, BiomedCoOp, libra-llava-rad, llava-rad
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
arXiv6 reposarXiv:2304.13705
aloha_sim_transfer_cube_human, physical-ai-studio, act_gyoza_cherry3_v2, act_gyoza_shiratama_v2, act_gyoza_pickplace_synth, act_gyoza_grape_v2
StarCoder: may the source be with you!
arXiv6 reposarXiv:2305.06161
MultiPL-E, magicoder, Magicoder-S-DS-6.7B, Magicoder-S-CL-7B, Magicoder-DS-6.7B, Magicoder-CL-7B
InstructIE: A Bilingual Instruction-based Information Extraction Dataset
arXiv6 reposarXiv:2305.11527
IEPile, InstructIE, llama2-13b-iepile-lora, baichuan2-13b-iepile-lora, EasyInstruct, KnowLM
Simple and Controllable Music Generation
arXiv6 reposarXiv:2306.05284
optimized-parler-tts, whisperspeech, musicgen-medium, musicgen-melody, musicgen-melody-large, encodec_32khz
ModelScope Text-to-Video Technical Report
arXiv6 reposarXiv:2308.06571
Text-To-Video-Finetuning, text-to-video-synthesis-colab, text-to-video-ms-1.7b, text-to-video-ms-1.7b, mcm, i2vgen-xl
A Family of Pretrained Transformer Language Models for Russian
arXiv6 reposarXiv:2309.10931
rugpt3large_based_on_gpt2, rugpt3medium_based_on_gpt2, rugpt3small_based_on_gpt2, rugpt3large_based_on_gpt2, rugpt3medium_based_on_gpt2, rugpt3small_based_on_gpt2
UltraFeedback: Boosting Language Models with Scaled AI Feedback
arXiv6 reposarXiv:2310.01377
zephyr-7b-alpha, zephyr-7b-beta, ultrafeedback_binarized, Llama-3-8B-Magpie-Align-v0.1, Llama-3-8B-Magpie-Align-v0.3, Llama-3-8B-Magpie-Align-v0.2
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment
arXiv6 reposarXiv:2310.01852
LanguageBind, LanguageBind, VIDAL-Depth-Thermal, MoE-LLaVA, Video-LLaVA, LLMBind
Ring Attention with Blockwise Transformers for Near-Infinite Context
arXiv6 reposarXiv:2310.01889
MindSpeed-MM, Wan2.1-VACE-14B, EasyContext, Wan2.1, Wan2.1-VACE-1.3B, Wan2.1-FLF2V-14B-720P
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
arXiv6 reposarXiv:2310.08491
prometheus, Feedback-Collection, prometheus-7b-v2.0, prometheus-7b-v1.0-fp16, prometheus-13b-v1.0-fp16, open-korean-instructions
DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors
arXiv6 reposarXiv:2310.12190
DynamiCrafter_1024, DynamiCrafter, DynamiCrafter_512, DynamiCrafter, DynamiCrafter_512_Interp, DynamiCrafter_pruned
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
arXiv6 reposarXiv:2310.16834
mdlm, bd3lm, bd3lms, BDM, BD_DNA, sedd-noeos-owt
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
arXiv6 reposarXiv:2311.03099
Twin-Merging, model_merging, Mario, supermario_v2, MergeLM, ComfyUI-LoRA-Optimizer
An Efficient Self-Supervised Cross-View Training For Sentence Embedding
arXiv6 reposarXiv:2311.03228
SCT-model-phayathaibert, SCT-KD-model-phayathaibert, SCT-model-XLMR, SCT-KD-model-wangchanberta, SCT-model-wangchanberta, SCT-KD-model-XLMR
Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
arXiv6 reposarXiv:2311.07919
LoRA-MCL, Qwen2-Audio-7B, Qwen-Audio-Chat, Qwen-Audio, Speech-IFEval, Qwen2-Audio-7B-Instruct
SegVol: Universal and Interactive Volumetric Medical Image Segmentation
arXiv6 reposarXiv:2311.13385
M3D-Seg, SegVol, M3D-RefSeg, SegVol, DL2-group5-med-seg, adapt_med_seg
LMDrive: Closed-Loop End-to-End Driving with Large Language Models
arXiv6 reposarXiv:2312.07488
LMDrive, LMDrive, LMDrive-vicuna-v1.5-7b-v1.0, LMDrive-llava-v1.5-7b-v1.0, LMDrive-vision-encoder-r50-v1.0, LMDrive-llama-7b-v1.0
VILA: On Pre-training for Visual Language Models
arXiv6 reposarXiv:2312.07533
llm-awq, VILA, awq-embed, awq4nvomni, llm-awq, VILA-2.7b
SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling
arXiv6 reposarXiv:2312.15166
SOLAR-10.7B-v1.0, SOLAR-10.7B-Instruct-v1.0, yarn, LDCC-SOLAR-10.7B, iDUS, corningQA-solar-10.7b-v1.0
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
arXiv6 reposarXiv:2401.12168
OpenSpaces, SpaceQwen2.5-VL-3B-Instruct, SpaceLLaVA, SpaceOm, SpaceThinker-Qwen2.5VL-3B, SpaceMantis
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
arXiv6 reposarXiv:2402.11746
red-instruct, CategoricalHarmfulQA, starling-7B, HarmfulQA, resta, CategoricalHarmfulQ
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
arXiv6 reposarXiv:2402.13516
prosparse-llama-2-13b, prosparse-llama-2-7b, prosparse-llama-2-7b-gguf, prosparse-llama-2-13b-gguf, prosparse-llama-2-13b-predictor, prosparse-llama-2-7b-predictor
Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation
arXiv6 reposarXiv:2403.08002
libra-llava-rad, llava-rad, llava-rad, chexprompt, llava-dino, Explanability_in_VLM
Layer-Condensed KV Cache for Efficient Inference of Large Language Models
arXiv6 reposarXiv:2405.10637
StepDeepResearch, LCKV, tinyllama-lckv-w10-100b, tinyllama-lckv-w2-100b, tinyllama-lckv-w10-ft-250b, tinyllama-lckv-w2-ft-100b
SimPO: Simple Preference Optimization with a Reference-Free Reward
arXiv6 reposarXiv:2405.14734
SimPO, CPO_SIMPO, Reinforcement-Learning-Full-Pipeline, Llama-3-8B-Magpie-Align-v0.1, Llama-3-8B-Magpie-Align-v0.3, Llama-3-8B-Magpie-Align-v0.2
arXiv:2405.17976
arXiv6 reposarXiv:2405.17976
Yuan2-M32, Yuan2-M32-gguf-int4, Yuan2-M32-hf-int8, Yuan2-M32-hf, Yuan2-M32-gguf, Yuan2-M32-hf-int4
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
arXiv6 reposarXiv:2406.05113
LlavaGuard, LlavaGuard-v1.2-0.5B-OV, LlavaGuard-v1.2-0.5B-OV-hf, LlavaGuard-v1.2-7B-OV, LlavaGuard-v1.2-7B-OV-hf, compagent
Depth Anything V2
arXiv6 reposarXiv:2406.09414
Depth-Anything-V2-Large, Depth-Anything-V2-Small, Depth-Anything-V2-Base, coreml-depth-anything-v2-small, svraster, Depth-Anything-V2-Small-hf
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
arXiv6 reposarXiv:2406.12845
Magpie-Pro-DPO-100K-v0.1, Magpie-Air-DPO-100K-v0.1, Magpie-Llama-3.1-Pro-DPO-100K-v0.1, Llama-3-8B-Magpie-Align-v0.1, Llama-3-8B-Magpie-Align-v0.3, Llama-3-8B-Magpie-Align-v0.2
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
arXiv6 reposarXiv:2406.17557
modded-nanogpt, fineweb, fineweb-edu, fineweb-2, Automodel, ml_filter
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
arXiv6 reposarXiv:2406.19280
HuatuoGPT-Vision-34B, HuatuoGPT-Vision-7B, PubMedVision, Medical_Multimodal_Evaluation_Data, HuatuoGPT-Vision-7B-Qwen2.5VL, MedAI-project
MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention
arXiv6 reposarXiv:2407.02490
LLMLingua, Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, RetrievalAttention, Block-Sparse-Attention
VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
arXiv6 reposarXiv:2407.11691
VLMEvalKit, Investigating_MultiEncoder_Redundancy, VLMEvalKit, ChatVLA_public, sa2va_eval, EAI_VLMEvalKit
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection
arXiv6 reposarXiv:2407.16237
origen_dataset_debug, OriGen, origen_dataset_instruction, OriGen_Fix, OriGen, origen_dataset_description
Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology
arXiv6 reposarXiv:2408.00738
PIANO, Midnight, SEAL, HistAug, histaug-virchow2, dpfm_factory
miniCTX: Neural Theorem Proving with (Long-)Contexts
arXiv6 reposarXiv:2408.03350
ntp-mathlib-instruct-st, miniCTX-v2, ntp-mathlib-st-deepseek-coder-1.3b, ntp-mathlib-context-deepseek-coder-1.3b, ntp-mathlib-instruct-ctx, ntp-mathlib
SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
arXiv6 reposarXiv:2408.05517
swift, ms-swift, ms-swift, vmopd, ms, dense-retention-rl
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
arXiv6 reposarXiv:2408.15998
Eagle-X5-34B-Plus, Eagle-X5-13B, Eagle-X4-13B-Plus, Eagle-X4-8B-Plus, Eagle-X5-13B-Chat, Eagle-X5-7B
Diabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management
arXiv6 reposarXiv:2409.13191
Diabetica, Diabetica-1.5B, Diabetica-SFT, Diabetica-o1, Diabetica-o1-SFT, Diabetica-7B
Moshi: a speech-text foundation model for real-time dialogue
arXiv6 reposarXiv:2410.00037
minimind-o, tts-1.6b-en_fr, moshi-rag, personaplex, moshi, eval-moshi
Aria: An Open Multimodal Native Mixture-of-Experts Model
arXiv6 reposarXiv:2410.05993
Aria-Chat, MMLongBench-Doc, Aria, Aria, Aria-Base-64K, Aria-Base-8K
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
arXiv6 reposarXiv:2411.15296
lmms-eval, Thyme, Investigating_MultiEncoder_Redundancy, VLMEvalKit, Video-MME, UniG2U
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
arXiv6 reposarXiv:2411.17440
MagicTime, ChronoMagic-Bench, ConsisID, ConsisID-preview, ConsisID-preview-Data, OpenS2V-Nexus
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
arXiv6 reposarXiv:2411.19378
libra-v1.0-7b, libra-llava-med-v1.5-mistral-7b, libra-llava-rad, libra-v1.0-3b, libra-maira-2, Libra
Multimodal Whole Slide Foundation Model for Pathology
arXiv6 reposarXiv:2411.19666
TRIDENT, VLSA, HistAug, TridentEdited, histaug-conch_v15, aegis
RedStone: Curating General, Code, Math, and QA Data for Large Language Models
arXiv6 reposarXiv:2412.03398
RedStone, RedStone, RedStone-QA-mcq, RedStone-Code-python, RedStone-Math, RedStone-QA-oq
LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
arXiv6 reposarXiv:2412.09262
LatentSync, LatentSync, LatentSync-1.6, LatentSync1.5-mac, LatentSync-1.5, screencastgen
CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models
arXiv6 reposarXiv:2412.13195
CoMPaSS, CoMPaSS-FLUX.1, CoMPaSS-SD2.1, CoMPaSS-SD1.4, CoMPaSS-SD1.5, CoMPaSS-FLUX.1-dev-ComfyUI
Qwen2.5 Technical Report
arXiv6 reposarXiv:2412.15115
QwQ-32B, d3LLM, Baichuan-Omni-1.5, smoltalk2, Baichuan-Audio, FLUXSynID
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
arXiv6 reposarXiv:2501.07542
visual-thinker, AlphaMaze-v0.2-1.5B, Mirage, MVoT, UniWM, GoViG
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models
arXiv6 reposarXiv:2501.14818
Eagle, Eagle2-2B, GR00T-N1.5-3B, Eagle2-1B, Eagle2-9B, llama-nemotron-embed-vl-1b-v2
Qwen2.5-1M Technical Report
arXiv6 reposarXiv:2501.15383
Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen2.5-7B-Instruct-1M, Qwen2.5-14B-Instruct-1M, Qwen3-Next-80B-A3B-Instruct
s1: Simple test-time scaling
arXiv6 reposarXiv:2501.19393
unlazy, s1, s1.1-32B, s1K-step-conditional-control-old, step-conditional-control-old, crrrocq
PolarQuant: Quantizing KV Caches with Polar Transformation
arXiv6 reposarXiv:2502.02617
turboquant_plus, quant.cpp, Qwen3.5-9B-PolarQuant-MLX-4bit, Qwen3.5-9B-PolarEngine-v4, Qwen3.5-9B-PolarQuant-Q5, polarengine-vllm
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
arXiv6 reposarXiv:2502.02737
smollm, finemath, smoltalk, SmolLM2-360M, SmolLM2-1.7B, smollm2-1.7b-instruct
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
arXiv6 reposarXiv:2502.05512
IndexTTS-2, parrots, Index-TTS, index-tts, indexTTS2, IndexTTS
Large Language Diffusion Models
arXiv6 reposarXiv:2502.09992
dLLM-cache, d3LLM, iLLaDA-8B-Instruct, LLaDA-MoE-7B-A1B-Base, dllm, SDAR
Qwen2.5-VL Technical Report
arXiv6 reposarXiv:2502.13923
Qwen2.5-VL, Qwen3-VL, Qwen2-VL, ScreenSpot-Pro-GUI-Grounding, CharXiv, Fara_Test
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
arXiv6 reposarXiv:2502.15543
ParamMute, ParamMute-8B-KTO, ParamMute-7B, ParamMute-8B-SFT, CoConflictQA, PIP-KAG
YuE: Scaling Open Foundation Models for Long-Form Music Generation
arXiv6 reposarXiv:2503.08638
YuE2-3B, YuE, SheetSage2, YuE2-Vae-legacy, YuE2-Vae, WildSongBench
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
arXiv6 reposarXiv:2503.14476
ReTool, tunix, DAPO, xtuner, MM-EUREKA, VL-Rethinker
m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
arXiv6 reposarXiv:2504.00869
m1, m1-7B-23K, m1-7B-1K, m1-32B-1K, m23k-tokenized, m1k-tokenized
Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
arXiv6 reposarXiv:2504.05599
Skywork-R1V, Skywork-R1V2-38B, Skywork-R1V-38B-AWQ, Skywork-R1V-38B, Skywork-R1V2-38B-AWQ, sky-rv1
arXiv:2504.07962
arXiv6 reposarXiv:2504.07962
GLUS, GLUS-S-partial, GLUS-A, GLUS-S, myGLUS, space_glus
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
arXiv6 reposarXiv:2504.13180
perception_models, PE-Lang-L14-448, PE-Lang-G14-448, PLM-Image-Auto, PLM-Video-Human, PLM-Video-Auto
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
arXiv6 reposarXiv:2504.19475
ViT-Prisma, sparse-autoencoder-clip-b-32-sae-vanilla-x64-layer-10-hook_mlp_out-l1-1e-05, ViT-Prisma, my_prisma, df-prisma, ViT-Prisma-fix
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
arXiv6 reposarXiv:2505.13886
Game-RL-InternVL2.5-8B, GameQA-text, Game-RL-InternVL3-8B, Game-RL-Qwen2.5-VL-7B, GameQA-140K, GameQA-5K
AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound
arXiv6 reposarXiv:2505.14142
AudSemThinker, audsem-simple, audsemthinker-qa, audsem, audsemthinker, audsemthinker-qa-grpo
Emerging Properties in Unified Multimodal Pretraining
arXiv6 reposarXiv:2505.14683
Bagel, bytedance_BAGEL-7B-MoT-INT8, Macro-Bagel, Vision-R1, ComfyUI-BAGEL, Bagel
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
arXiv6 reposarXiv:2505.20292
MagicTime, ChronoMagic-Bench, ConsisID, OpenS2V-5M, OpenS2V-Nexus, OpenS2V-Eval
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
arXiv6 reposarXiv:2505.21491
FrameINO, FrameINO_Wan2.2_5B_Stage2_MotionINO_v1.5, FrameINO_Wan2.2_5B_Stage2_MotionINO_v1.6, FrameINO_CogVideoX_Stage2_MotionINO_v1.0, FrameINO_CogVideoX_Stage1_Motion_v1.0, FrameINO_Wan2.2_5B_Stage1_Motion_v1.5
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
arXiv6 reposarXiv:2506.02095
cyclereward, CyclePrefDB-I2T, CycleReward-I2T, CycleReward-Combo, CyclePrefDB-T2I, CycleReward-T2I
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
arXiv6 reposarXiv:2506.03147
ImgEdit, UniWorld, UniWorld-V1, UniWorld-V1, UniWorld-V1, UniWorld-V1-NF4
SpatialLM: Training Large Language Models for Structured Indoor Modeling
arXiv6 reposarXiv:2506.07491
SpatialLM1.1-Qwen-0.5B, SpatialLM, SpatialLM-Dataset, SpatialLM1.1-Llama-1B, SpatialLM-TestSet, SPATIALLM
MiniCPM4: Ultra-Efficient LLMs on End Devices
arXiv6 reposarXiv:2506.07900
MiniCPM4-8B, FR-Spec, MiniCPM5-2B, MiniCPM5-2B-GGUF, MiniCPM5-1B-GGUF, MiniCPM5-1B
Step-Audio 2 Technical Report
arXiv6 reposarXiv:2507.16632
Step-Audio-2-mini, Step-Audio2, Step-Audio-2-mini-Think, StepEval-Audio-Toolcall, StepEval-Audio-Paralinguistic, Step-Audio-2-mini-Base
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
arXiv6 reposarXiv:2508.05629
swift, ms-swift, ms-swift, vmopd, ms, dense-retention-rl
Controllable Latent Space Augmentation for Digital Pathology
arXiv6 reposarXiv:2508.14588
HistAug, histaug-conch, histaug-conch_v15, histaug-virchow2, histaug-hoptimus1, histaug-uni
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
arXiv6 reposarXiv:2509.01986
DIM, DIM-T2I, DIM-4.6B-Edit, DIM-Edit, DIM-4.6B-T2I, DIM-4.6B-Edit-Stage1
Tequila: Trapping-free Ternary Quantization for Large Language Models
arXiv6 reposarXiv:2509.23809
Qwen3-1.7B_eagle3, Qwen3-4B_eagle3, Qwen3-8B_eagle3, Qwen3-14B_eagle3, Qwen3-a3B_eagle3, Qwen3-32B_eagle3
SpecExit: Accelerating Large Reasoning Model via Speculative Exit
arXiv6 reposarXiv:2509.24248
Qwen3-a3B_eagle3, Qwen3-14B_eagle3, Qwen3-1.7B_eagle3, Qwen3-4B_eagle3, Qwen3-8B_eagle3, Qwen3-32B_eagle3
Fast-dLLM v2: Efficient Block-Diffusion LLM
arXiv6 reposarXiv:2509.26328
Fast_dLLM_v2_7B, d3LLM, Fast-dLLM, dev-dllm, relay, Fast_dLLM_v2_1.5B
SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation
arXiv6 reposarXiv:2510.06303
SDAR, SDAR-4B-Chat, SDAR-8B-Chat, SDAR-30B-A3B-Sci, SDAR-1.7B-Chat, SDAR-30B-A3B-Chat
Emu3.5: Native Multimodal Models are World Learners
arXiv6 reposarXiv:2510.26583
Emu3.5-VisionTokenizer, Emu3.5, Emu3.5-Image, Emu3.5, Emu35-Comfyui-Nodes, Emu35-Image-NF4
Cambrian-S: Towards Spatial Supersensing in Video
arXiv6 reposarXiv:2511.04670
vsi-590k, Cambrian-S-3M, Cambrian-S-7B, cambrian-s-3b, cambrian-s-0.5b, cambrian-s-1.5b
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
arXiv6 reposarXiv:2601.07568
d3LLM_LLaDA, d3LLM, trajectory_data_dream_32, d3LLM_Dream_Coder, d3LLM_Dream, trajectory_data_llada_32
Qwen3-TTS Technical Report
arXiv6 reposarXiv:2601.15621
Qwen3-TTS-12Hz-1.7B-VoiceDesign, Qwen3-TTS-12Hz-1.7B-Base, Qwen3-TTS-12Hz-0.6B-CustomVoice, Qwen3-TTS-12Hz-0.6B-Base, Qwen3-TTS-Tokenizer-12Hz, qwen3-tts
vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models
arXiv6 reposarXiv:2602.02204
vllm-omni, vllm-GPT-SoVITS, cmu-15642, vllm-omni, vllm-omni-hunyuanimage3, vllm-omni-minicpmo45-npu
DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction
arXiv6 reposarXiv:2602.18589
lodochallenge_pixel_diffusion, lodochallenge_latent_diffusion, synchrotron_pixel_diffusion, lodoind_latent_diffusion, lodoind_pixel_diffusion, synchrotron_latent_diffusion
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
arXiv6 reposarXiv:2604.04911
SpatialEdit, SpatialEdit-500K, SpatialEdit-16B, SpatialEdit-Bench, JoyAI-Image-SpatialEdit-Bench, JoyAI-Image-SpatialEdit
ACL-Verbatim: hallucination-free question answering for research
arXiv6 reposarXiv:2605.21102
verbatim-rag, verbatim-rag-modern-bert-v2, verbatim-spans, acl-anthology-md, acl-verbatim-spans, acl-verbatim-modernbert
Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search
arXiv6 reposarXiv:2605.30917
v-splade, v-splade-quality, v-splade-efficient, SPLADE-mlx, v-splade-efficient-mlx, v-splade-quality-mlx
Gemma 4 Technical Report
arXiv6 reposarXiv:2607.02770
gemma-4-31B-it-qat-q4_0-unquantized, awesome-gemma, gemma-4-31B-it, gemma-4-e4b-it, gemma-4-12B-it, gemma-4-12b-it-qat-q4_0-gguf
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
arXiv6 reposarXiv:2607.04988
InternVLA-A1, InternVLA-A1.5-DOMINO, InternVLA-A1.5-base, InternVLA-A1.5-RoboTwin, InternVLA-A1.5-Libero, InternVLA-A-series
Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion
arXiv6 reposarXiv:2607.25572
vulnerability-attack-technique-classification-roberta-base, vulnerability-attack-technique-biencoder, cve-attack-mapping-paper, vulnerability-attack-technique-classification-roberta-base-llm-expanded, vulnerability-attack-techniques, vulnerability-attack-techniques-llm-scaling
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
arXiv6 reposarXiv:2608.10932
CamSFT-4B, CamInject-4B, CamSFT-8B, CamDistill-8B, CamInject-8B, CamDistill-4B
An Efficiency Study for SPLADE Models
ACM5 reposACM:3477495.3531833
splade, efficient-splade-V-large-doc, efficient-splade-VI-BT-large-doc, efficient-splade-V-large-query, efficient-splade-VI-BT-large-query
Microsoft COCO: Common Objects in Context
arXiv5 reposarXiv:1405.0312
LocateAnything-3B, mscoco-it, MSCOCO, huggingface-datasets_MSCOCO, COCO-Caption2017
Deep Residual Learning for Image Recognition
arXiv5 reposarXiv:1512.03385
yolov5, resnet50.tv_in1k, AIGI-Holmes, EasyOCR, resnet-50
SQuAD: 100,000+ Questions for Machine Comprehension of Text
arXiv5 reposarXiv:1606.05250
squad, squad_v2, squad_bn, t5-large-encoder-only-bf16, CoConflictQA
Verified Low-Level Programming Embedded in F*
arXiv5 reposarXiv:1703.00053
libcrux, everest, karamel, kremlin, libcrux
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
arXiv5 reposarXiv:1801.03924
materialgan, LatentSync, LatentSync1.5-mac, PerceptualSimilarity, WeavePrompt
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
arXiv5 reposarXiv:1908.08962
bert, tapas, bert-tiny-historic-multilingual-cased, bert-mini-historic-multilingual-cased, bert-small-historic-multilingual-cased
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
arXiv5 reposarXiv:1910.02054
ColossalAI, MindSpeed-MM, awsome-llm-papers, stablecode-completion-alpha-3b-4k, mesh-transformer-jax
GLU Variants Improve Transformer
arXiv5 reposarXiv:2002.05202
t5-v1_1-xxl, google_t5-v1_1-xxl_encoderonly, nomic-bert-2048, rwkv, llama2.zig
FLERT: Document-Level Features for Named Entity Recognition
arXiv5 reposarXiv:2011.06993
ner-german-large, ner-english-ontonotes-large, ner-spanish-large, flair, ner-dutch-large
HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection
arXiv5 reposarXiv:2012.10289
bert-base-uncased-hatexplain, HateXplain, bert-base-uncased-hatexplain-rationale-two, HateXplain, DeepLearningProject
VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation
arXiv5 reposarXiv:2101.00390
voxpopuli, voxpopuli, wav2vec2-large-100k-voxpopuli, unispeech-sat-large, wavlm-large
XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond
arXiv5 reposarXiv:2104.12250
twitter-xlm-roberta-base-sentiment, xlm-t, twitter-xlm-roberta-base, xlm-twitter-politics-sentiment, multilingual-hate-speech-robacofi
SpeechBrain: A General-Purpose Speech Toolkit
arXiv5 reposarXiv:2106.04624
lang-id-voxlingua107-ecapa, spkrec-ecapa-voxceleb, spkrec-xvect-voxceleb, spkrec-resnet-voxceleb, spkrec-ecapa-voxceleb-mel-spec
Evaluating Large Language Models Trained on Code
arXiv5 reposarXiv:2107.03374
human-eval, llm-humaneval-benchmarks, FTTT, openai_humaneval, research2
CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and Generation
arXiv5 reposarXiv:2109.05729
bart-base-chinese, CPT, bart-large-chinese, cpt-base, cpt-large
Challenges in Detoxifying Language Models
arXiv5 reposarXiv:2109.07445
fineweb, fineweb-edu, falcon-refinedweb, finepdfs, fineweb-2
Few-shot Learning with Multilingual Language Models
arXiv5 reposarXiv:2112.10668
xglm-2.9B, xglm-564M, polyglot, mGPT, mGPT
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
arXiv5 reposarXiv:2201.12086
Semantic-Segment-Anything, blip-image-captioning-base, LAVIS, blip-vqa-base, blip-vqa-capfilt-large
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
arXiv5 reposarXiv:2204.05862
hh-rlhf, artgpt2tox, RAIN, RRHF, LLM-Pref-Mark-UI
InCoder: A Generative Model for Code Infilling and Synthesis
arXiv5 reposarXiv:2204.05999
sven_modified, sven, incoder-6B, incoder, incoder-1B
Petals: Collaborative Inference and Fine-tuning of Large Models
arXiv5 reposarXiv:2209.01188
petals, subnet-llm, petals, bloombee_add_models, petals
Zero-Shot Learners for Natural Language Understanding via a Unified Multiple Choice Perspective
arXiv5 reposarXiv:2210.08590
Ziya-LLaMA-13B-Pretrain-v1, Fengshenbang-LM, Ziya-LLaMA-13B-v1, Ziya-BLIP2-14B-Visual-v1, Ziya-LLaMA-13B-v1.1
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities
arXiv5 reposarXiv:2211.06679
stable-diffusion-webui, FlagAI, AltDiffusion-m9, AltDiffusion, generative-ai
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
arXiv5 reposarXiv:2211.10438
nunchaku, nunchaku, trtllm, TensorRT-LLM, TensorRT-LLM
TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities
arXiv5 reposarXiv:2212.06385
gpt2-chinese-cluecorpussmall, t5-small-chinese-cluecorpussmall, t5-base-chinese-cluecorpussmall, Linly, Chinese-ChatLLaMA
Precise Zero-Shot Dense Retrieval without Relevance Labels
arXiv5 reposarXiv:2212.10496
prompt-engineering, docs-reference, llm-search, spring-ai-extensions, KoPrivateGPT
A Watermark for Large Language Models
arXiv5 reposarXiv:2301.10226
watermarks-remover, text-generation-inference, Adversarial-Paraphrasing, impossibility-watermark, lm-watermarking
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
arXiv5 reposarXiv:2303.16199
LLaMA-Adapter, Point-Bind_Point-LLM, LLaMA-Adapter, LLaMA2-Accessory, lit-llama
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
arXiv5 reposarXiv:2304.01373
TinyLlama, pythia-12b, pythia-1.4b, pythia-410m, pythia-2.8b
Scaling Speech Technology to 1,000+ Languages
arXiv5 reposarXiv:2305.13516
mms-300m, gujarati-vsr, fcbh-dataset-io, mms-tts-spa, mms-tts-eng
The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning
arXiv5 reposarXiv:2305.14045
KoCoT_2000, CoT-Collection, CoT-Collection, Multilingual-CoT-Collection, KoCommercial-Dataset
Segment Anything in High Quality
arXiv5 reposarXiv:2306.01567
geti-instant-learn, sd-webui-inpaint-anything, sd-webui-inpaint-anything, interior-segment-labeler, SAMReg
A Technical Report for Polyglot-Ko: Open-Source Large-Scale Korean Language Models
arXiv5 reposarXiv:2306.02254
polyglot, polyglot-ko-1.3b, polyglot-ko-3.8b, polyglot-ko-5.8b, level3_nlp_finalproject-nlp-12
RED$^{\rm FM}$: a Filtered and Multilingual Relation Extraction Dataset
arXiv5 reposarXiv:2306.09802
mrebel-large, SREDFM, REDFM, mdeberta-v3-base-triplet-critic-xnli, mrebel-large-32
BayLing: Bridging Cross-lingual Alignment and Instruction Following through Interactive Translation for Large Language Models
arXiv5 reposarXiv:2306.10968
BayLing, bayling-13b-v1.1, bayling-7b-diff, bayling-13b-diff, BayLing
Cross-Lingual Cross-Age Group Adaptation for Low-Resource Elderly Speech Emotion Recognition
arXiv5 reposarXiv:2306.14517
elderly_ser, SER-wav2vec2-large-xlsr-53-eng-zho-adults, SER-wav2vec2-large-xlsr-53-eng-zho-elderly, SER-wav2vec2-large-xlsr-53-eng-zho-all-age, YueMotion
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
arXiv5 reposarXiv:2307.08691
graphify, lorax, GPT-2, vigogne, Block-Sparse-Attention
AlpaGasus: Training A Better Alpaca with Fewer Data
arXiv5 reposarXiv:2307.08701
KoRAE, KoRAE-13b, KoRAE-13b-DPO, KoRAE_filtered_12k, original-KoRAE-13b-3ep
LISA: Reasoning Segmentation via Large Language Model
arXiv5 reposarXiv:2308.00692
LISA-Llama-3, DINO_LISA, LISA-Training, LISA, LISA
The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
arXiv5 reposarXiv:2308.01907
all-seeing, CRPE, AS-Core, AS-V2, AS-100M
Towards General Text Embeddings with Multi-stage Contrastive Learning
arXiv5 reposarXiv:2308.03281
gte-large-zh, gte-base-zh, gte-large-en-v1.5, gte-large, gte-base
Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs
arXiv5 reposarXiv:2308.09895
self-oss-instruct-sc2-exec-filter-50k, MultiPL-T-StarCoder2_15B, stack-dedup-python-testgen-starcoder-filter-v2, MultiPL-T-CodeLlama_34b, MultiPL-T-DeepSeekCoder_33b
How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection
arXiv5 reposarXiv:2308.13177
omchat-v2.0-13B-single-beta_hf, omchat, OmDet, OVDEval, OmAgent
Efficient Streaming Language Models with Attention Sinks
arXiv5 reposarXiv:2309.17453
streaming-llm, KVQuant, dlms-sinks, mlx-flash, Block-Sparse-Attention
Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
arXiv5 reposarXiv:2310.05130
L2D, AdaDetectGPT, fast-detect-gpt, fast-detect-gpt, worker-fast-detect-gpt
Fine-Tuning LLaMA for Multi-Stage Text Retrieval
arXiv5 reposarXiv:2310.08319
tinyllama-embed, repllama-v1-7b-lora-passage, rankllama-v1-7b-lora-passage, pyterrier_genrank, RepLLaMA-reproduced
Llemma: An Open Language Model For Mathematics
arXiv5 reposarXiv:2310.10631
math-lm, proof-pile-2, llemma_34b, llemma_7b, llmstep
SkyMath: Technical Report
arXiv5 reposarXiv:2310.16713
Skywork-13B-base, Skywork-13B-Math-8bits, Skywork-13B-Base-3.1TB, Skywork-13B-Base-8bits, Skywork-13B-Math
LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing
arXiv5 reposarXiv:2311.00571
LLaVA, GLIGEN, TriPlaneLLaVA, LLaVA-toy, TinyLLava
LongQLoRA: Efficient and Effective Method to Extend Context Length of Large Language Models
arXiv5 reposarXiv:2311.04879
LongQLoRA, LongQLoRA-Llama2-7b-8k, LongQLoRA-Vicuna-13b-8k, Long-QLORA, article_gpt
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
arXiv5 reposarXiv:2311.05437
LLaVA, TriPlaneLLaVA, LLaVA-toy, LLaVA-Plus-Codebase, TinyLLava
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
arXiv5 reposarXiv:2311.06242
Florence-2-base, Florence-2-base-ft, Florence-2-large, Florence-2-large-ft, interior-segment-labeler
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
arXiv5 reposarXiv:2311.08046
LanguageBind, Chat-UniVi-Instruct, Chat-UniVi, Chat-UniVi-13B, Chat-UniVi
GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer
arXiv5 reposarXiv:2311.08526
gliner_multi_pii-v1, gliner-stream-pii-v1.0, gliner_small-v2.5, gliner_small-v2.1, agentic-graphrag
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
arXiv5 reposarXiv:2311.12793
ShareGPT4V-7B, ShareGPT4V, ShareCaptioner, ShareGPT4V-13B, ShareGPT4V
Magicoder: Empowering Code Generation with OSS-Instruct
arXiv5 reposarXiv:2312.02120
evalplus, Magicoder-S-DS-6.7B, Magicoder-S-CL-7B, Magicoder-DS-6.7B, Magicoder-CL-7B
Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval
arXiv5 reposarXiv:2312.15503
bge-large-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-reranker-v2-m3, bge-reranker-v2.5-gemma2-lightweight
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
arXiv5 reposarXiv:2401.01044
Auffusion, auffusion-full, auffusion, auffusion-full-no-adapter, MMDisCo
GPT-4V(ision) is a Generalist Web Agent, if Grounded
arXiv5 reposarXiv:2401.01614
UGround-V1-7B, UGround-V1-2B, UGround-V1-72B, UGround, Multimodal-Mind2Web
Latte: Latent Diffusion Transformer for Video Generation
arXiv5 reposarXiv:2401.03048
Latte-1, Latte, Latte-0, Latte, Latte-1
Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series
arXiv5 reposarXiv:2401.03955
ttm-research-r2, granite-timeseries-ttm-r2, Samay, moment, granite-timeseries-ttm-v1
Executable Code Actions Elicit Better LLM Agents
arXiv5 reposarXiv:2402.01030
rlm, CodeActAgent-Llama-2-7b, code-act, CodeActAgent-Mistral-7b-v0.1, CodeActAgent-Mistral-7b-v0.1.q8_0.gguf
DoRA: Weight-Decomposed Low-Rank Adaptation
arXiv5 reposarXiv:2402.09353
hymba, ohara, llama3-chinese, llama3-chinese, Llama3-Chinese-Lora
Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models
arXiv5 reposarXiv:2402.14207
dspy, storm, gpt-researcher, awesome-dspy, dsp
The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
arXiv5 reposarXiv:2402.19474
all-seeing, CRPE, AS-Core, AS-V2, AS-100M
Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People
arXiv5 reposarXiv:2403.03640
ApolloCorpus, Apollo-7B, Apollo-34B, Apollo-72B, PodGPT
DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
arXiv5 reposarXiv:2403.08857
HunyuanDiT, HunyuanDiT-v1.1, HunyuanDiT, HunyuanDiT-v1.2, HunyuanDiT
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
arXiv5 reposarXiv:2403.15246
mteb-1.34.14, FollowIR, FollowIR-7B, FollowIR-train, ru-promptriever-qwen3-4b
Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models
arXiv5 reposarXiv:2404.07724
OmniGen2, OmniGen2, OmniGen2-EditScore7B, Macro-OmniGen2, OmniGen2-EditScore7B-v1.1
TAVGBench: Benchmarking Text to Audible-Video Generation
arXiv5 reposarXiv:2404.14381
JavisBench, JavisGPT, JavisInst-Omni, MM-PreTrain, AV-FineTune
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
arXiv5 reposarXiv:2404.14396
SEED-Data-Edit-Part2-3, SEED-X, SEED-X-17B, SEED-Data-Edit, SEED
Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation
arXiv5 reposarXiv:2406.02347
flash-diffusion, flash-pixart, flash-sd3, flash-sdxl, flash-sd
LADI v2: Multi-label Dataset and Classifiers for Low-Altitude Disaster Imagery
arXiv5 reposarXiv:2406.02780
ladi-overview, LADI-v2-dataset, LADI-v2-classifier-large, LADI-v2-classifier-large-reference, LADI-v2-classifier-small
Scaling and evaluating sparse autoencoders
arXiv5 reposarXiv:2406.04093
nanointerpret, sparsify, dictionary_learning, cli, notebooks
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
arXiv5 reposarXiv:2406.04325
ShareGPT4Video, ShareGPT4Video, sharegpt4video-8b, ShareCaptioner-Video, stereopilot-replica-accelerate
Multimodal Table Understanding
arXiv5 reposarXiv:2406.08100
table-llava-v1.5-13b, Table-LLaVA, MMTab, table-llava-v1.5-7b, table-llava-v1.5-7b-hf
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
arXiv5 reposarXiv:2407.04051
minimind-o, FunASR, SenseVoice, SenseVoice, FunASR
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
arXiv5 reposarXiv:2407.06135
Anole-7b-v0.1, anole, Anole-7b, UniWM, GoViG
FlashNorm: Fast Normalization for Transformers
arXiv5 reposarXiv:2407.09577
transformer-tricks, gemma-4-E2B-FlashNorm, Llama-3.1-8B-FlashNorm, gemma-4-E2B-FlashNorm-strict, Llama-3.2-1B-FlashNorm
CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models
arXiv5 reposarXiv:2407.15886
catvton-flux, catvton-unstudio-flux, CatVTON, CatVTON, CatVTON-MaskFree
SAM 2: Segment Anything in Images and Videos
arXiv5 reposarXiv:2408.00714
segment-anything, geti-instant-learn, sam2-hiera-large, sam2-hiera-tiny, sam2.1-hiera-large
OLMoE: Open Mixture-of-Experts Language Models
arXiv5 reposarXiv:2409.02060
OLMoE, OLMoE-mix-0924, OLMoE-1B-7B-0924-SFT, OLMoE-1B-7B-0924-Instruct, OLMoE-1B-7B-0125-Instruct
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
arXiv5 reposarXiv:2409.04429
vila-u, VLAC, OpenVid-1M, OpenVid, vila-u-7b-256
Block-Attention for Efficient Prefilling
arXiv5 reposarXiv:2409.15355
Block-Attention, Tulu3-Block-FT, Tulu3-SFT, Tulu3-RAG, GraphKV
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
arXiv5 reposarXiv:2409.17115
DCLM-pro, web-doc-refining-lm, math-doc-refining-lm, web-chunk-refining-lm, math-chunk-refining-lm
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
arXiv5 reposarXiv:2410.06885
F5-TTS, F5-TTS, StyleStream, StyleStream, X-Voice
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
arXiv5 reposarXiv:2410.10813
hippo-memory, Awareness-Market, ogham-mcp, Titan-Memory, post-graph-rag
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
arXiv5 reposarXiv:2410.18451
Skywork-Reward-Preference-80K-v0.2, Skywork-Reward-Llama-3.1-8B, Skywork-Reward-Gemma-2-27B, Skywork-Reward-Gemma-2-27B-v0.2, Skywork-Reward-Llama-3.1-8B-v0.2
HunyuanVideo: A Systematic Framework For Large Video Generative Models
arXiv5 reposarXiv:2412.03603
HunyuanVideo, HunyuanVideo-PromptRewrite, HunyuanVideo, HunyuanVideo-I2V, HunyuanVideo-I2V
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
arXiv5 reposarXiv:2412.18319
Mulberry-SFT, Mulberry, Mulberry_llama_11b, Mulberry_qwen2vl_7b, Mulberry_llava_8b
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
arXiv5 reposarXiv:2412.18605
Orient-Anything, Orient-Anything, OriNet, OriAnyV2_ckpt, orient-anything
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
arXiv5 reposarXiv:2412.18925
medical-o1-reasoning-SFT, HuatuoGPT-o1-70B, HuatuoGPT-o1-7B, HuatuoGPT-o1-8B, medical-o1-verifiable-problem
Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence
arXiv5 reposarXiv:2502.09927
granite-vision-3.3-2b, granite-vision-3.3-2b-GGUF, granite-vision-3.1-2b-preview, granite-vision-4.1-4b, granite-4.0-3b-vision
OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale
arXiv5 reposarXiv:2503.02240
OmniSQL, SynSQL-2.5M, OmniSQL-14B, OmniSQL-7B, OmniSQL-32B
YOLOE: Real-Time Seeing Anything
arXiv5 reposarXiv:2503.07465
yoloe, yolov10, yoloe, yoloe, yoloe
ViSpeak: Visual Instruction Feedback in Streaming Videos
arXiv5 reposarXiv:2503.12769
ViSpeak, StreamingBench, StreamingBench, ViSpeak-s2, ViSpeak-s3
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
arXiv5 reposarXiv:2503.21749
LeX-Art, LeX-Lumina, LeX-Bench, LeX-Data-10K, LeX-Enhancer-full
TerraMind: Large-Scale Generative Multimodality for Earth Observation
arXiv5 reposarXiv:2504.11171
TerraMind-1.0-base, terramind, TerraMind-1.0-large, TerraMind-1.0-tiny, TerraMind-1.0-small
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
arXiv5 reposarXiv:2504.16656
Skywork-R1V, Skywork-R1V2-38B, Skywork-R1V-38B-AWQ, Skywork-R1V2-38B-AWQ, sky-rv1
Process Reward Models That Think
arXiv5 reposarXiv:2504.16828
ThinkPRM, ThinkPRM-14B, ThinkPRM-7B, ThinkPRM-1.5B, thinkprm-1K-verification-cots
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
arXiv5 reposarXiv:2505.07608
MiMo-7B-RL, MiMo-7B-Base, MiMo-7B-RL-0530, MiMo-7B-RL-Zero, MiMo-7B-SFT
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
arXiv5 reposarXiv:2505.10557
MM-MathInstruct, Img2Code, MathCoder-VL-2B, FigCodifier, MathCoder-VL-8B
PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
arXiv5 reposarXiv:2506.05573
PartCrafter, PartCrafter-Scene, PartCrafter, modly-partcrafter-extension, Accelerator0701
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
arXiv5 reposarXiv:2506.07530
BitVLA-CoreAI, BitVLA, bitvla-bitsiglipL-224px-bf16, bitvla-bf16, bitvla-siglipL-224px-bf16
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
arXiv5 reposarXiv:2506.09965
ViLaSR, ViLaSR, ViLaSR, ViLaSR-cold-start, ViLaSR-data
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
arXiv5 reposarXiv:2506.21619
Awesome-AITools, IndexTTS-2, parrots, index-tts, ComfyUI-kaola-IndexTTS2
Kwai Keye-VL Technical Report
arXiv5 reposarXiv:2507.01949
Keye, Keye-VL-2.0-30B-A3B, Keye-VL-8B-Preview, Keye-VL-1.5-8B, keye-39e1f0b5
Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
arXiv5 reposarXiv:2507.16116
PusaV1_training, PusaV0.5_Training, Pusa-Wan2.2-V1, PusaV1, Pusa-V0.5
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
arXiv5 reposarXiv:2507.20939
ShortVid-Bench, ARC-Hunyuan-Video-7B, ARC-Hunyuan-Video-7B, ARC-Qwen-Video-7B, ARC-Qwen-Video-7B-Narrator
PixNerd: Pixel Neural Field Diffusion
arXiv5 reposarXiv:2507.23268
PixNerd-diffusers, PixNerd-diffusers, PixNerd, PixNerd-XXL-P16-T2I, PixNerd-XL-P16-C2I
Qwen-Image Technical Report
arXiv5 reposarXiv:2508.02324
Qwen-Image-2512, Qwen-Image, Qwen-Image, Macro-Qwen-Image-Edit, Qwen-Image-Flash
DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrieval
arXiv5 reposarXiv:2508.07995
Diver-Retriever-4B, Diver, Diver-Retriever-0.6B, Diver-Retriever-4B-1020, Diver-Retriever-1.7B
DINOv3
arXiv5 reposarXiv:2508.10104
geti-instant-learn, dinov3-vitl16-pretrain-lvd1689m, dinov3, compositio_nn, projet-vision
Thyme: Think Beyond Images
arXiv5 reposarXiv:2508.11630
Thyme, Thyme-SFT, Thyme-RL, Thyme-SFT, Thyme-RL
Kwai Keye-VL 1.5 Technical Report
arXiv5 reposarXiv:2509.01563
Keye, Keye-VL-2.0-30B-A3B, Keye-VL-8B-Preview, Keye-VL-1.5-8B, keye-39e1f0b5
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
arXiv5 reposarXiv:2509.06501
MiniMax-M2-BF16, MiniMax-M2.1, MiniMax-M2.1, MiniMax-M2, MiniMax-M2
EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
arXiv5 reposarXiv:2510.25628
EHR-R1, EHR-Bench, EHR-Ins-Reasoning, EHR-R1-1.7B, EHR-R1-8B
arXiv:2511.16175
arXiv5 reposarXiv:2511.16175
Mantis, mantis_libero_lerobot, Mantis-Base, Mantis, wla
Qwen3-VL Technical Report
arXiv5 reposarXiv:2511.21631
Qwen2.5-VL, Qwen3-VL, Qwen2-VL, nemotron-colembed-vl-4b-v2, RoboSpatial-Eval
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
arXiv5 reposarXiv:2512.23703
Robo-Dopamine-GRM-3B, Robo-Dopamine-GRM-2.0-8B-Preview, Robo-Dopamine-GRM-8B, Robo-Dopamine-GRM-2.0-4B-Preview, Robo-Dopamine-Bench
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
arXiv5 reposarXiv:2601.00393
NeoVerse, NeoVerse, NeoVerse, NeoVerse-archive, neoverse_new
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
arXiv5 reposarXiv:2601.14251
LightOnOCR-2-1B, LightOnOCR-2-1B-old-church-slavonic-line, LightOnOCR-2-1B-base, LightOnOCR-2-1B-Pinokio, LightOnOCR-mix-0126
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
arXiv5 reposarXiv:2602.06965
MedMO-8B-Next, MedMO-8B, MedMO-4B, MedMO-4B-Next, MedMO
Revisiting Text Ranking in Deep Research
arXiv5 reposarXiv:2602.21456
text-ranking-in-deep-research, browsecomp-plus-passage-corpus, browsecomp-plus-passage-corpus-pyserini, browsecomp-plus-indexes, browsecomp-plus-runs
MediX-R1: Open Ended Medical Reinforcement Learning
arXiv5 reposarXiv:2602.23363
medix-rl-data, MediX-R1-30B, MediX-R1-8B, MediX-R1-2B, MediX-R1
Can Vision-Language Models Solve the Shell Game?
arXiv5 reposarXiv:2603.08436
shellgame, Molmo2-SGCoT-Demo, Molmo2-SGCoT, vetbench, Molmo2-SGCoT
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
arXiv5 reposarXiv:2603.12262
VST, VST-32B, VST-3B, VST-Training-Data, VST-7B
Small Vision-Language Models are Smart Compressors for Long Video Understanding
arXiv5 reposarXiv:2604.08120
Tempo, Tempo-6B, Tempo-6B-Stage1, Tempo-6B-Stage2, Tempo-6B-Stage0
Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech
arXiv5 reposarXiv:2605.20830
Raon-OpenTTS-1B, Raon-OpenTTS, Raon-OpenTTS-0.3B, Raon-OpenTTS-Eval, Raon-OpenTTS-Pool
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
arXiv5 reposarXiv:2606.16533
kairos, Kairos3.1-4B-robot-480P, kairos-4B-robot-RoboTwin2.0, kairos-4B-robot-LIBERO-plus, kairos-sensenova
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
arXiv5 reposarXiv:2607.24743
ClinFusion-32B, ClinFusion-8B, ClinFusion, clinfusion-medical-vlm, ClinFusion-Eval-Data
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
arXiv5 reposarXiv:2609.08183
NeoHorse-1-4B, NeoHorse, NeoHorse-1-9B-GGUF, NeoHorse-1-4B-GGUF, NeoHorse-1-9B
A pathology foundation model for cancer diagnosis and prognosis prediction
Nature5 reposNature:s41586-024-07894-z
AtlasPatch, KEEP, KEEP, CPathPatchFeature, TITAN
Aligning Large Language Model with Direct Multi-Preference Optimization for Recommendation
ACM4 reposACM:3627673.3679611
LlamaFactory, LLaMA-Factory, LLaMA-Factory-personal, LLaMA-Factory-LFS
arXiv:0000.00000
arXiv4 reposarXiv:0000.00000
granite-4.1-3b, granite-3.3-8b-instruct, granite-4.1-8b, granite-3.1-1b-a400m-instruct
Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
arXiv4 reposarXiv:1603.09320
pgvector, pecos, sweet-search, usearch
SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine
arXiv4 reposarXiv:1704.05179
all-MiniLM-L6-v2, all-MiniLM-L12-v2, CoConflictQA, KoPrivateGPT
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
arXiv4 reposarXiv:1705.03551
CoConflictQA, KoPrivateGPT, GraphKV, GraphKV
Proximal Policy Optimization Algorithms
arXiv4 reposarXiv:1707.06347
minimind, tunix, Reinforcement-Learning-Full-Pipeline, cleanrl
Benchmarking Neural Network Robustness to Common Corruptions and Perturbations
arXiv4 reposarXiv:1903.12261
MediMeta-C, RobustMedCLIP, RobustMedCLIP, medmnistc-api
Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer
arXiv4 reposarXiv:1907.01341
Text2Video-Zero, svraster, MiDaS, prisma
Leveraging Pre-trained Checkpoints for Sequence Generation Tasks
arXiv4 reposarXiv:1907.12461
tf-transformers, wiki_split, bert2bert_L-24_wmt_de_en, Reddit-Sports-Sentiment-Analysis
Release Strategies and the Social Impacts of Language Models
arXiv4 reposarXiv:1908.09203
DetectLLMSegmentation, L2D, AdaDetectGPT, CodeXGLUE
NEZHA: Neural Contextualized Representation for Chinese Language Understanding
arXiv4 reposarXiv:1909.00204
nezha-cn-large, nezha-large-wwm, nezha-base-wwm, nezha-cn-base
Libri-Light: A Benchmark for ASR with Limited or No Supervision
arXiv4 reposarXiv:1912.07875
libriheavy, unispeech-sat-large, libri-light, wavlm-large
On the limits of cross-domain generalization in automated X-ray prediction
arXiv4 reposarXiv:2002.02497
covid-chestxray-dataset, torchxrayvision, densenet121-res224-chex, torchxrayvision
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
arXiv4 reposarXiv:2005.11401
gpt-researcher, ROOT-RAG, flan-ul2-dolly, flan-ul2-dolly-lora
ConvBERT: Improving BERT with Span-based Dynamic Convolution
arXiv4 reposarXiv:2008.02496
berts, turkish-bert, convbert-base-turkish-cased, europeana-bert
What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
arXiv4 reposarXiv:2009.13081
MedQA-USMLE-4-options, PodGPT, MedQA, Baichuan2
Prefix-Tuning: Optimizing Continuous Prompts for Generation
arXiv4 reposarXiv:2101.00190
FasterTransformer, Continual-NExT, SwissArmyTransformer, LoRA
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
arXiv4 reposarXiv:2101.03961
marker, fiddler, awsome-llm-papers, smolMoELM-custom
Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
arXiv4 reposarXiv:2102.05918
coyo-align-b7-base, coyo-dataset, coyo-700m, align-base
Improved Denoising Diffusion Probabilistic Models
arXiv4 reposarXiv:2102.09672
PixArt-alpha, pixeart, PixArt-alpha, smalldiffusion
GPT Understands, Too
arXiv4 reposarXiv:2103.10385
GLM, EasyNLP, Continual-NExT, SwissArmyTransformer
BookSum: A Collection of Datasets for Long-form Narrative Summarization
arXiv4 reposarXiv:2105.08209
airoboros-summarization, vid2cleantxt, bigbird-pegasus-large-K-booksum, long-t5-tglobal-xl-16384-book-summary
SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
arXiv4 reposarXiv:2109.10086
splade, splade_v2_max, splade_v2_distil, pinecone-text
Training Verifiers to Solve Math Word Problems
arXiv4 reposarXiv:2110.14168
gsm8k, llm-jepa, FTTT, RLPR-Evaluation
Swin Transformer V2: Scaling Up Capacity and Resolution
arXiv4 reposarXiv:2111.09883
CLIP-ViT-L-14-laion2B-s32B-b82K, MiDaS, swinv2-large-patch4-window12to16-192to256-22kto1k-ft, Swin-Transformer
Locating and Editing Factual Associations in GPT
arXiv4 reposarXiv:2202.05262
OBLITERATUS, obliteratus, causal_unlearn_llm, KEditVis-LLM-Editing
Competition-Level Code Generation with AlphaCode
arXiv4 reposarXiv:2203.07814
apps, code_contests, code_contests, code_contests
PLAID: An Efficient Engine for Late Interaction Retrieval
arXiv4 reposarXiv:2205.09707
fast-plaid, plaidrepro, colbertv2.0, ColBERT
Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset
arXiv4 reposarXiv:2207.00220
pile-of-law, legalbert-large-1.7M-1, distilbert-base-uncased-finetuned-eoir_privacy, legalbert-large-1.7M-2
Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
arXiv4 reposarXiv:2209.03003
Matcha-TTS, GR00T-N1.5-3B, nanoMFM, Rectified-Diffusion
Flow Matching for Generative Modeling
arXiv4 reposarXiv:2210.02747
Matcha-TTS, nanoMFM, com-304-FM-project-2026, Rectified-Diffusion
Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages
arXiv4 reposarXiv:2210.09984
miracl, miracl, miracl-corpus, gte-multilingual-base
High Fidelity Neural Audio Compression
arXiv4 reposarXiv:2210.13438
encodec, encodec_24khz, encodec_48khz, whisperspeech
Lila: A Unified Benchmark for Mathematical Reasoning
arXiv4 reposarXiv:2210.17517
Arithmo2-Mistral-7B, Arithmo-Mistral-7B, Arithmo2-Mistral-7B-adapter, Arithmo-Data
OneFormer: One Transformer to Rule Universal Image Segmentation
arXiv4 reposarXiv:2211.06220
oneformer_ade20k_dinat_large, OneFormer, oneformer_cityscapes_swin_large, Semantic-Segment-Anything
A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
arXiv4 reposarXiv:2211.14730
patchtst-fm-r1, granite-timeseries-patchtst-fm-r1, PatchTST, FM4Motor
Large Language Models Encode Clinical Knowledge
arXiv4 reposarXiv:2212.13138
Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, OLAPH, healthsearchqa
Zero-1-to-3: Zero-shot One Image to 3D Object
arXiv4 reposarXiv:2303.11328
zero123, zero123-live, zero123-weights, zero123-xl-diffusers
FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization
arXiv4 reposarXiv:2303.14189
coreml-FastViT-T8, coreml-FastViT-MA36, ml-fastvit, fast_donut_KIE
The Vault: A Comprehensive Multilingual Dataset for Advancing Code Understanding and Generation
arXiv4 reposarXiv:2305.06156
TheVault, the-vault-inline, the-vault-class, the-vault-function
TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
arXiv4 reposarXiv:2305.07759
esp32s3-distributed-ai, TinyStories, Tiny-Stories-Regional, llama2.ts
arXiv:2305.10703
arXiv4 reposarXiv:2305.10703
ReGen, news_contrastive_pretrain, wiki_contrastive_pretrain, review_contrastive_pretrain
LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus
arXiv4 reposarXiv:2305.18802
libritts-r-filtered-speaker-descriptions, libritts_r, kanade-tokenizer, libritts_r
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
arXiv4 reposarXiv:2306.00107
MERT-v1-95M, MERT-v1-330M, MERT-v0, YuE
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
arXiv4 reposarXiv:2306.00814
vocos-mel-24khz, vocos, whisperspeech, vocos-encodec-24khz
TIES-Merging: Resolving Interference When Merging Models
arXiv4 reposarXiv:2306.01708
Twin-Merging, Mario, MergeLM, ComfyUI-LoRA-Optimizer
MIMIC-IT: Multi-Modal In-Context Instruction Tuning
arXiv4 reposarXiv:2306.05425
idefics-80b-instruct, Otter, Otter, idefics-9b-instruct
High-Fidelity Audio Compression with Improved RVQGAN
arXiv4 reposarXiv:2306.06546
SpeechTokenizer, descript-audio-codec, BigVGAN, nemo-nano-codec-22khz-1.89kbps-21.5fps
Quilt-1M: One Million Image-Text Pairs for Histopathology
arXiv4 reposarXiv:2306.11207
quilt1m, QuiltNet-B-32, QuiltNet-B-16-PMB, QuiltNet-B-16
Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
arXiv4 reposarXiv:2306.14289
sd-webui-inpaint-anything, MobileSAM, sd-webui-inpaint-anything, MobileSAM
Large Multimodal Models: Notes on CVPR 2023 Tutorial
arXiv4 reposarXiv:2306.14895
LLaVA, TriPlaneLLaVA, LLaVA-toy, TinyLLava
ReLoRA: High-Rank Training Through Low-Rank Updates
arXiv4 reposarXiv:2307.05695
BigDL, ipex-llm, ipex-llm, llm_test
3D-LLM: Injecting the 3D World into Large Language Models
arXiv4 reposarXiv:2307.12981
PointLLM, ShapeLLM, MiniGPT-3D, 3D-LLM
AltDiffusion: A Multilingual Text-to-Image Diffusion Model
arXiv4 reposarXiv:2308.09991
AltDiffusion, AltDiffusion-m9, AltDiffusion-m18, AltDiffusion
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
arXiv4 reposarXiv:2308.16692
SpeechTokenizer, SpeechGPT, USLM, SpeechTokenizer
PointLLM: Empowering Large Language Models to Understand Point Clouds
arXiv4 reposarXiv:2308.16911
PointLLM, ShapeLLM, MiniGPT-3D, PointLLM
Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
arXiv4 reposarXiv:2309.00615
PointLLM, ShapeLLM, MiniGPT-3D, Point-Bind_Point-LLM
XGen-7B Technical Report
arXiv4 reposarXiv:2309.03450
xgen-7b-8k-base, xgen, xgen-7b-4k-base, xgen-7b-8k-inst
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
arXiv4 reposarXiv:2309.05653
Arithmo2-Mistral-7B, Arithmo-Mistral-7B, Arithmo2-Mistral-7B-adapter, Arithmo-Data
An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
arXiv4 reposarXiv:2309.09958
LLaVA, TriPlaneLLaVA, LLaVA-toy, TinyLLava
Multimodal Foundation Models: From Specialists to General-Purpose Assistants
arXiv4 reposarXiv:2309.10020
LLaVA, TriPlaneLLaVA, LLaVA-toy, TinyLLava
OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
arXiv4 reposarXiv:2309.11235
openchat, openchat_3.5, openchat-3.6-8b-20240522, openchat-3.5-0106
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
arXiv4 reposarXiv:2309.11998
FastChat, Nanoflow, multilingual_mt_bench, FastChat
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
arXiv4 reposarXiv:2309.14717
BigDL, ipex-llm, ipex-llm, llm_test
Finite Scalar Quantization: VQ-VAE Made Simple
arXiv4 reposarXiv:2309.15505
nemo-nano-codec-22khz-1.89kbps-21.5fps, SkinTokens, SkinTokens, whisper_pinyin
OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
arXiv4 reposarXiv:2310.06786
proof-pile-2, DeepSeek-Math, OpenWebMath, open-web-math
LCM-LoRA: A Universal Stable-Diffusion Acceleration Module
arXiv4 reposarXiv:2311.05556
TCD-SDXL-LoRA, lcm-lora-sdxl, lcm-lora-sdv1-5, Arc2Face
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
arXiv4 reposarXiv:2311.06607
Monkey, Monkey, Detailed_Caption, Monkey-Chat
Mustango: Toward Controllable Text-to-Music Generation
arXiv4 reposarXiv:2311.08355
mustango, mustango, MusicBench, mustango-pretrained
VideoCon: Robust Video-Language Alignment via Contrast Captions
arXiv4 reposarXiv:2311.10111
owl-con, videocon, videocon, videocon-model
Diffusion Model Alignment Using Direct Preference Optimization
arXiv4 reposarXiv:2311.12908
dpo-sdxl-text2image-v1, DiffusionDPO, dpo-sd1.5-text2image-v1, GRPO
LM-Cocktail: Resilient Tuning of Language Models via Model Merging
arXiv4 reposarXiv:2311.13534
bge-large-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-small-en
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
arXiv4 reposarXiv:2311.15596
EgoThink, EgoThink, embodied-eval, behaviour_subtask
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
arXiv4 reposarXiv:2311.16502
Yi-34B-Chat, MMMU, MMMU, MMMU
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance
arXiv4 reposarXiv:2311.18681
RaDialog_v2_modified, RaDialog_v2, RaDialog, RaDialog-interactive-radiology-report-generation
A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting
arXiv4 reposarXiv:2312.03594
PowerPaint, PowerPaint-v2-1, PowerPaint-v1, PowerPaint_v2
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
arXiv4 reposarXiv:2312.04746
quilt-llava.github.io, quilt1m, quilt-llava, Quilt-Llava-v1.5-7b
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
arXiv4 reposarXiv:2312.13382
dspy, PromptingTools.jl, awesome-dspy, dsp
emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
arXiv4 reposarXiv:2312.15185
emotion2vec_base, emotion2vec, emotion2vec_plus_seed, emotion2vec_plus_base
LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
arXiv4 reposarXiv:2312.17240
LISA-Llama-3, DINO_LISA, LISA, LISA
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
arXiv4 reposarXiv:2401.09417
Vim-small-midclstok, Vim, Vim-tiny-midclstok, Vim-base-midclstok
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
arXiv4 reposarXiv:2401.10891
Depth-Estimation, coreml-depth-anything-v2-small, coreml-depth-anything-small, Depth-Anything-V2-Small-hf
Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
arXiv4 reposarXiv:2401.12070
DetectLLMSegmentation, AdaDetectGPT, LAPD, L2D
AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
arXiv4 reposarXiv:2402.00769
AnimateLCM, AnimateLCM-I2V, AnimateLCM, AnimateLCM-SVD-xt
Timer: Generative Pre-trained Transformers Are Large Time Series Models
arXiv4 reposarXiv:2402.02368
UTSD, timer-base-84m, Large-Time-Series-Model, moment
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
arXiv4 reposarXiv:2402.03766
MobileVLM_V2-1.7B, MobileVLM, MobileVLM_V2-7B, MobileVLM_V2-3B
MEMORYLLM: Towards Self-Updatable Large Language Models
arXiv4 reposarXiv:2402.04624
MemoryLLM, memoryllm-8b-chat, memoryllm-8b, ExtendingMemoryLLM
Multilingual E5 Text Embeddings: A Technical Report
arXiv4 reposarXiv:2402.05672
multilingual-e5-base, multilingual-e5-large, OKEAN, multilingual-e5-large-instruct
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
arXiv4 reposarXiv:2402.12226
AnyInstruct, AnyGPT-chat, AnyGPT-base, SpeechGPT
BiMediX: Bilingual Medical Mixture of Experts LLM
arXiv4 reposarXiv:2402.13253
BiMediX, BiMediX-Bi, BiMediX-Eng, BiMediX-Ara
Repetition Improves Language Model Embeddings
arXiv4 reposarXiv:2402.15449
llm2vec, DermL2V-training, Anchor-Embedding, DermL2V-tmp
Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping
arXiv4 reposarXiv:2402.19159
TCD, TCD-SD21-base-LoRA, TCD-SDXL-LoRA, TCD-SD15-LoRA
StarCoder 2 and The Stack v2: The Next Generation
arXiv4 reposarXiv:2402.19173
evalplus, starchat2-15b-v0.1, speechless-starcoder2-15b, AMALIA-9B-0626-DPO
DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models
arXiv4 reposarXiv:2402.19481
nunchaku, xDiT, nunchaku, distrifuser-controlnet
Improving Diffusion Models for Authentic Virtual Try-on in the Wild
arXiv4 reposarXiv:2403.05139
IDM-VTON, IDM-VTON, IDM-VTON-train, modal_pipeline
Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation
arXiv4 reposarXiv:2403.12015
Nitro-1, Nitro-1-PixArt, Nitro-1-SD, AMD-Diffusion-Distillation
RULER: What's the Real Context Size of Your Long-Context Language Models?
arXiv4 reposarXiv:2404.06654
Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-Next-80B-A3B-Instruct
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
arXiv4 reposarXiv:2404.16771
ConsistentID, ConsistentID, FGID, ConsistentID
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
arXiv4 reposarXiv:2404.16994
PLLaVA, pllava-13b, pllava-34b, pllava-7b
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers
arXiv4 reposarXiv:2405.05945
Lumina-T2X, Lumina-Next-T2I, Lumina-Next-SFT, Lumina-Next-SFT-diffusers
Improved Distribution Matching Distillation for Fast Image Synthesis
arXiv4 reposarXiv:2405.14867
FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree, DMD2, DMD2, Qwen-Image-Flash
Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
arXiv4 reposarXiv:2405.17398
Vista, Vista, hf-example-vista, Vista
MidiCaps: A large-scale MIDI dataset with text captions
arXiv4 reposarXiv:2406.02255
Text2midi, text2midi, MidiCaps, MidiCaps
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
arXiv4 reposarXiv:2406.02430
seed-tts-eval, custom-seed-vc, seed-vc-test, maestro-seedvc
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
arXiv4 reposarXiv:2406.03482
QJL, turboquant_plus, rotorquant, quant.cpp
Improving Alignment and Robustness with Circuit Breakers
arXiv4 reposarXiv:2406.04313
abliterix, Mistral-7B-Instruct-RR-Abliterated, Llama-3-8B-Instruct-RR-Abliterated, circuit-breakers
MoreHopQA: More Than Multi-hop Reasoning
arXiv4 reposarXiv:2406.13397
morehopqa, morehopqa, GraphKV, GraphKV
MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions
arXiv4 reposarXiv:2406.17536
MediMeta-C, RobustMedCLIP, RobustMedCLIP, medmnistc-api
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
arXiv4 reposarXiv:2406.20076
evf-sam2, evf-sam, EVF-SAM, evf-sam2-multitask
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
arXiv4 reposarXiv:2407.03168
FacePoke_CLONE-THIS-REPO-TO-USE-IT, LivePortrait, FacePoke, FLUXSynID
EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions
arXiv4 reposarXiv:2407.08136
echomimic, EchoMimic, EchoMimic, EchoMimic
Qwen2-Audio Technical Report
arXiv4 reposarXiv:2407.10759
Qwen2-Audio-7B, Qwen2-Audio, Speech-IFEval, Qwen2-Audio-7B-Instruct
Multi-label Cluster Discrimination for Visual Representation Learning
arXiv4 reposarXiv:2407.17331
mlcd-vit-base-patch32-224, unicom, mlcd-vit-large-patch14-336, MLCD-Embodied-7B
LLaVA-OneVision: Easy Visual Task Transfer
arXiv4 reposarXiv:2408.03326
llava-onevision-qwen2-7b-si, LLaVA-OneVision-Data, llava-onevision-qwen2-7b-ov-hf, llava-onevision-qwen2-7b-ov
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
arXiv4 reposarXiv:2408.04594
Img-Diff, data-juicer, SciDataOS, data-juicer
Scalable Autoregressive Image Generation with Mamba
arXiv4 reposarXiv:2408.12245
AiM, aim-xlarge, aim-base, aim-large
Towards Evaluating and Building Versatile Large Language Models for Medicine
arXiv4 reposarXiv:2408.12547
MedS-Ins, MMedS-Llama-3-8B, MedS-Ins, MedS-Bench
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
arXiv4 reposarXiv:2408.14774
Instruct-SkillMix, Llama-3-8B-Instruct-SkillMix, Instruct-SkillMix-SDD, Instruct-SkillMix-SDA
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
arXiv4 reposarXiv:2408.16589
ASR-Transcription-Router, CrisperWhisper, faster_CrisperWhisper, CrisperWhisper
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
arXiv4 reposarXiv:2409.06656
FluidAudio, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, diar_sortformer_4spk-v1
OneEncoder: A Lightweight Framework for Progressive Alignment of Modalities
arXiv4 reposarXiv:2409.11059
OneEncoder-text-image-xray, OneEncoder-text-image, OneEncoder-text-image-audio, OneEncoder-text-image-video
Qwen2.5-Coder Technical Report
arXiv4 reposarXiv:2409.12186
Qwen2.5-Coder-32B-Instruct, Qwen2.5-Coder-14B-Instruct, Qwen2.5-Coder-7B-Instruct, Qwen2.5-Coder-0.5B
Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts
arXiv4 reposarXiv:2409.16040
TimeMoE-50M, time-moe, TimeMoE-200M, Time-300B
Emu3: Next-Token Prediction is All You Need
arXiv4 reposarXiv:2409.18869
Emu3, Emu3-Stage1, Emu3-Chat, Emu3-VisionTokenizer
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
arXiv4 reposarXiv:2410.03355
LANTERN, llamagen_drafter, llamagen2_drafter, anole_drafter
Pyramidal Flow Matching for Efficient Video Generative Modeling
arXiv4 reposarXiv:2410.05954
pyramid-flow-sd3, Pyramid-Flow, pyramid-flow-miniflux, pyramid-flow
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
arXiv4 reposarXiv:2410.08261
Meissonic, Monetico, meissonic, test-time-scaling
Teach Multimodal LLMs to Comprehend Electrocardiographic Images
arXiv4 reposarXiv:2410.19008
ECGBench, ECGInstruct, PULSE, PULSE-7B
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
arXiv4 reposarXiv:2411.02293
Hunyuan3D-2.1, Hunyuan3D-1, HY3D-Bench, Hunyuan3D-Omni
SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models
arXiv4 reposarXiv:2411.05007
nunchaku, deepcompressor, nunchaku-qwen-image, nunchaku
FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training
arXiv4 reposarXiv:2411.11927
FLAME, FLAME-ReCap-CC3M-MiniCPM-Llama3-V-2_5, FLAME-Mistral-Nemo-ViT-B-16-CC3M, FLAME-ReCap-YFCC15M-MiniCPM-Llama3-V-2_5
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
arXiv4 reposarXiv:2411.15738
AnyEdit, AnyEdit, AnySD, AnySD
Arctic-Embed 2.0: Multilingual Retrieval Without Compromise
arXiv4 reposarXiv:2412.04506
cholesky_encoder, snowflake-arctic-embed-l-v2.0, snowflake-arctic-embed-m-v2.0, arctic-embed
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
arXiv4 reposarXiv:2412.05237
MAmmoTH-VL-Instruct-12M, MAmmoTH-VL, MAmmoTH-VL-12M, MAmmoTH-VL-8B
ACT-Bench: Towards Action Controllable World Models for Autonomous Driving
arXiv4 reposarXiv:2412.05337
ACT-Bench, ACT-Estimator, Terra, ACT-Bench
Concept Bottleneck Large Language Models
arXiv4 reposarXiv:2412.07992
CB-LLMs, Concept-Bottleneck-LLM, Concept-Bottleneck-LLM, CBLLM-PubMed
Offline Reinforcement Learning for LLM Multi-Step Reasoning
arXiv4 reposarXiv:2412.16145
OREO, OREO, Qwen2.5-Math-1.5B-OREO, Qwen2.5-Math-1.5B-OREO-Value
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
arXiv4 reposarXiv:2412.18525
Understanding-Vision-Tasks, UVT-Terminological-based-Vision-Tasks, UVT-Explanatory-based-Vision-Tasks, UVT-7B-448
DeepSeek-V3 Technical Report
arXiv4 reposarXiv:2412.19437
DeepSeek-V3, DeepSeek-V3-0324, maxtext, DualPipe
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
arXiv4 reposarXiv:2501.01428
GPT4Scene-qwen2vl_full_sft_mark_32_3D_img512, GPT4Scene, GPT4Scene-All, GPT4Scene-and-VLN-R1
BEN: Using Confidence-Guided Matting for Dichotomous Image Segmentation
arXiv4 reposarXiv:2501.06230
BEN, BEN, BEN2, BEN2
FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
arXiv4 reposarXiv:2501.14350
FireRedASR2S, FireRedASR, FireRedASR-LLM-L, FireRedASR-AED-L
Sundial: A Family of Highly Capable Time Series Foundation Models
arXiv4 reposarXiv:2502.00816
sundial-base-128m, timer-base-84m, Large-Time-Series-Model, Sundial
NitiBench: A Comprehensive Study of LLM Framework Capabilities for Thai Legal Question Answering
arXiv4 reposarXiv:2502.10868
nitibench, nitibench, nitibench-ccl-human-finetuned-bge-m3, nitibench-ccl-auto-finetuned-bge-m3
AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO
arXiv4 reposarXiv:2502.14669
Maze-Reasoning-Reset-v0.1, AlphaMaze-v0.2-1.5B, Maze-Reasoning-v0.1, Maze-Reasoning-GRPO-v0.1
FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
arXiv4 reposarXiv:2502.14856
Qwen2-7B-Instruct-FR-Spec, FR-Spec, LLaMA3.2-Instruct-1B-FR-Spec, LLaMA3-Instruct-8B-FR-Spec
Mantis: Lightweight Foundation Model for Time Series Classification
arXiv4 reposarXiv:2502.15637
FM4Motor, mantis, MantisV2Experiments, TiViT
Muon is Scalable for LLM Training
arXiv4 reposarXiv:2502.16982
Moonlight-16B-A3B, Muon, Moonlight-16B-A3B-Instruct, Emerging-Optimizers
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
arXiv4 reposarXiv:2502.21257
RoboBrain, RoboBrain-LoRA-Affordance, RoboBrain-LoRA-Trajectory, RoboBrain2.5
Scaling Rich Style-Prompted Text-to-Speech Datasets
arXiv4 reposarXiv:2503.04713
paraspeechcaps, paraspeechcaps, parler-tts-mini-v1-paraspeechcaps, parler-tts-mini-v1-paraspeechcaps-only-base
TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
arXiv4 reposarXiv:2503.11509
AutomaTikZ, DeTikZify, detikzify-v2.5-8b, detikzify-v2-8b
Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
arXiv4 reposarXiv:2503.14911
MAKE, DermLIP_PanDerm-base-w-PubMed-256, Derm1M, DermLIP_ViT-B-16
RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation
arXiv4 reposarXiv:2503.18738
roboengine, roboengine-bg-diffusion, roboengine-sam, roboseg
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
arXiv4 reposarXiv:2503.18878
SAE-Reasoning, OpenThoughts-10k-DeepSeek-R1, DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenized, deepseek-r1-distill-llama-8b-lmsys-openthoughts
AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
arXiv4 reposarXiv:2503.19462
AccVideo, AccVideo, AccVideo-WanX-I2V-480P-14B, AccVideo-WanX-T2V-14B
Qwen2.5-Omni Technical Report
arXiv4 reposarXiv:2503.20215
Qwen2.5-Omni-7B-AWQ, Qwen2.5-Omni-7B-GPTQ-Int4, Qwen2.5-Omni-3B, Qwen2.5-Omni-7B
Unified Multimodal Discrete Diffusion
arXiv4 reposarXiv:2503.20853
unidisc, unidisc_hq, unidisc_non_interleaved, unidisc_interleaved
Video-R1: Reinforcing Video Reasoning in MLLMs
arXiv4 reposarXiv:2503.21776
Video-R1, Video-R1-7B, Qwen2.5-VL-7B-COT-SFT, Video-R1-eval
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
arXiv4 reposarXiv:2503.23377
JavisBench, JavisDiT, JavisDiT-v1.0-jav, JavisGPT-v0.1-7B-Instruct
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
arXiv4 reposarXiv:2504.01934
ILLUME_plus, illume_plus-qwen2_5-3b-hf, illume_plus-qwen2_5-7b-hf, dualvitok
LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
arXiv4 reposarXiv:2504.07448
LoRI, LoRI-S_safety_llama3_rank_32, LoRI-D_safety_llama3_rank_32, LoRI-D_code_llama3_rank_32
VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
arXiv4 reposarXiv:2504.08837
ViRL39K, VL-Rethinker, VL-Rethinker-72B, VL-Rethinker-7B
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
arXiv4 reposarXiv:2504.15279
VisuLogic, VisuLogic-Train, VisuLogic-Eval, VisuLogic-Train
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
arXiv4 reposarXiv:2504.16030
LiveSports-3K, LiveCC-7B-Instruct, Live-CC-5M, Live-WhisperX-526K
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
arXiv4 reposarXiv:2504.17343
TimeChat-Online, TimeChat-Online-139K, TimeChatOnline-7B, TimeChat
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
arXiv4 reposarXiv:2504.21356
Nexus-GenV2, Nexus-Gen, Nexus-GenV2-nf4-fp8, diffSynth-studio-notes
EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
arXiv4 reposarXiv:2505.04623
AVQA-R1-6K, EchoInk, EchoInk-R1-7B, OmniInstruct_V1_AVQA_R1
MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks
arXiv4 reposarXiv:2505.06152
SkinVL-PubMM, MM-Skin, SkinVL-MM, SkinVL-Pub
MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
arXiv4 reposarXiv:2505.13427
MM-EUREKA, MM-PRM, MM-PRM, MM-K12
On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?
arXiv4 reposarXiv:2505.15425
MediMeta-C, RobustMedCLIP, RobustMedCLIP, medmnistc-api
ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval
arXiv4 reposarXiv:2505.17166
esg_reports_v2, biomedical_lectures_v2, economics_reports_v2, esg_reports_human_labeled_v2
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
arXiv4 reposarXiv:2505.17589
CosyVoice, CosyVoice, CV3-Eval, FastCosyVoice
Distilling LLM Agent into Small Models with Retrieval and Code Tools
arXiv4 reposarXiv:2505.17612
agent-distillation, Qwen2.5-32B-Instruct_agent_trajectories_2k, agent_distilled_Qwen2.5-1.5B-Instruct, Qwen2.5-32B-Instruct_agent_trajectories_2k_prefix
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
arXiv4 reposarXiv:2505.22334
Multimodal-Cold-Start, Multimodal-RL-Data, Qwen2.5VL-7b-RLCS, Qwen2.5VL-3b-RLCS
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
arXiv4 reposarXiv:2505.22618
d3LLM, DIFFA, dllm, Fast-dLLM
Zero-Shot Vision Encoder Grafting via LLM Surrogates
arXiv4 reposarXiv:2505.22664
zero, zero-model-checkpoints, llava-1.5-665k-instructions, llava-1.5-665k-genqa-500k-instructions
Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material
arXiv4 reposarXiv:2506.15442
Hunyuan3D-2.1, Hunyuan3D-2.1, HY3D-Bench, Hunyuan3D-Omni
RLPR: Extrapolating RLVR to General Domains without Verifiers
arXiv4 reposarXiv:2506.18254
RLPR, RLPR-Evaluation, RLPR-Qwen2.5-7B-Base, RLPR-Train-Dataset
MindCube: Spatial Mental Modeling from Limited Views
arXiv4 reposarXiv:2506.21458
SpaceQwen2.5-VL-3B-Instruct, SpaceOm, MindCube, MindCube
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
arXiv4 reposarXiv:2506.23361
Open-OmniVCus, OmniVCus, OmniVCus-Test, OmniVCus-Train
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
arXiv4 reposarXiv:2507.04590
VLM2Vec-V2.0, MMEB-V2, MMEB-V3, TARA
Skywork-R1V3 Technical Report
arXiv4 reposarXiv:2507.06167
Skywork-R1V-38B-AWQ, Skywork-R1V3-38B, Skywork-R1V3-38B-AWQ, Skywork-R1V3-38B-GGUF
Group Sequence Policy Optimization
arXiv4 reposarXiv:2507.18071
tunix, MMPR-Tiny, DeepVision-103K, DeepVision-103K
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
arXiv4 reposarXiv:2507.19457
adk-python, dspy, sweet-search, dsp
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
arXiv4 reposarXiv:2507.21033
GPT-Image-Edit, GPT-Image-Edit-1.5M, gpt-image-edit-training, gpt-image-edit-benchmark-results
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
arXiv4 reposarXiv:2507.21802
MixGRPO, flow_grpo, MixGRPO, DanceGRPO
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
arXiv4 reposarXiv:2508.10433
We-Math2.0-Standard, We-Math, We-Math2.0, We-Math2.0-Pro
ToonOut: Fine-tuned Background-Removal for Anime Characters
arXiv4 reposarXiv:2509.06839
BiRefNet, toonout, BiRefNet, toonout
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
arXiv4 reposarXiv:2509.06888
mmBERT-small, mmBERT-base, mmBERT, mmbert-pretrain-p1-fineweb2-langs
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
arXiv4 reposarXiv:2509.06949
TraDo-8B-Instruct, TraDo-8B-Thinking, TraDo-4B-Instruct, dLLM-RL
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
arXiv4 reposarXiv:2509.15202
abliterix, Llama-3-8B-Instruct-DeepRefusal-Broken, DeepRefusal, Meta-Llama-3-8B-Instruct-DeepRefusal
Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
arXiv4 reposarXiv:2509.15279
Fleming-VL-8B, Fleming-R1-7B, Fleming-R1-32B, Fleming-VL-38B
UIPro: Unleashing Superior Interaction Capability For GUI Agents
arXiv4 reposarXiv:2509.17328
UIPro-7B_Stage2_Web, UIPro_1stage, UIPro-7B_Stage2_Mobile, UIPro-7B_Stage1
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
arXiv4 reposarXiv:2509.21144
nlp_team, nlp_personal, UniSS, UniST
LLaDA-MoE: A Sparse MoE Diffusion Language Model
arXiv4 reposarXiv:2509.24389
LLaDA-MoE-7B-A1B-Base, LLaDA-MoE-7B-A1B-Instruct, dllm, dLLM_Cache
Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models
arXiv4 reposarXiv:2509.25826
Kairos_50m, Kairos_23m, Kairos_10m, Kairos
Mem-α: Learning Memory Construction via Reinforcement Learning
arXiv4 reposarXiv:2509.25911
Mem-alpha, memalphaljx, tmp-mem-alpha, agentic-memory
ModernVBERT: Towards Smaller Visual Document Retrievers
arXiv4 reposarXiv:2510.01149
ColModernVBERT-CoreAI, modernvbert, modernvbert, modernvbert
SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision
arXiv4 reposarXiv:2510.02797
SongFormer, SongFormer, SongFormDB, SheetSage2
dInfer: An Efficient Inference Framework for Diffusion Language Models
arXiv4 reposarXiv:2510.08666
dInfer, dInfer_adaptive, dinfer_power_sampling, TIDE_DATA_COLLECTION
pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation
arXiv4 reposarXiv:2510.14974
LakonLab, pi-Qwen, pi-FLUX.2, pi-FLUX.1
Chronos-2: From Univariate to Universal Forecasting
arXiv4 reposarXiv:2510.15821
chronos-forecasting, chronos-2, chronos-2-small, chronos-2-synth
DeepSeek-OCR: Contexts Optical Compression
arXiv4 reposarXiv:2510.18234
DeepSeek-OCR, devanagari-ocr-benchmark, deepseek-ocr-encoder, DeepSeek-OCR-2
Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models
arXiv4 reposarXiv:2510.21204
mitra-classifier, mitra-classifier-pipeline, mitra-regressor, mitra-regressor-pipeline
FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
arXiv4 reposarXiv:2510.24645
AWorld, FunReason-MT, FunReason-MT, AWorld-RL
MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency
arXiv4 reposarXiv:2510.25897
miro, miro-ablations, miro, miro
MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
arXiv4 reposarXiv:2511.09611
MMaDA, MMaDA-Parallel-A, MMaDA-Parallel-M, MMaDA-Parallel
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
arXiv4 reposarXiv:2511.22715
ReAG, ReAG-Critic, ReAG-3B, ReAG-7B
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
arXiv4 reposarXiv:2512.14698
TimeLens-100K, TimeLens-8B, TimeLens-Bench, TimeLens-7B
Recursive Language Models
arXiv4 reposarXiv:2512.24601
LegalRAG, rlm, rlm-claude, memcp
ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos
arXiv4 reposarXiv:2601.05237
ObjectForesight-Data, ObjectForesight-EPIC-DiT, ObjectForesight, ObjectForesight-HOT3D-DiT
Orient Anything V2: Unifying Orientation and Rotation Understanding
arXiv4 reposarXiv:2601.05573
OriAnyV2_Train_Render, OriAnyV2_ckpt, Hunyuan3D-FLUX-Gen, OriAnyV2_Inference
Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers
arXiv4 reposarXiv:2601.10770
GPA, GPA-v1.5, GPA-v1.5-onnx-runtime, GPA
C-RADIOv4 (Tech Report)
arXiv4 reposarXiv:2601.17237
RADIO, C-RADIOv4-SO400M, C-RADIOv4-H, CRADIOv4
VIBEVOICE-ASR Technical Report
arXiv4 reposarXiv:2601.18184
VibeVoice, VibeVoice-ASR, VibeVoice-ASR-HF, VibeVoice
Qwen3-ASR Technical Report
arXiv4 reposarXiv:2601.21337
Qwen3-ASR-1.7B-hf, Qwen3-ForcedAligner-0.6B-hf, Qwen3-ASR-0.6B-hf, Qwen3-ForcedAligner-0.6B
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
arXiv4 reposarXiv:2602.00594
kanade-tokenizer, kanade-12.5hz, kanade-25hz-clean, kanade-25hz
ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation
arXiv4 reposarXiv:2602.04279
ECG-R1, ECG-R1-8B-RL, ECG-Protocol-Guided-Grounding-CoT, GEM
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
arXiv4 reposarXiv:2602.08711
TimeChat-Captioner, Timechat-OmniCaptioner-42K, TimeChat-Captioner-GRPO-7B, Timechat-OmniCaptioner-40K
MOVA: Towards Scalable and Synchronized Video-Audio Generation
arXiv4 reposarXiv:2602.08794
MOVA, MOVA-360p, MOVA-720p, MOVA_benchmark_for_arena
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
arXiv4 reposarXiv:2602.19163
JavisDiT, AV-DPO, JavisDiT-v1.0-jav, JavisGPT-v1.0-7B-Instruct
A Very Big Video Reasoning Suite
arXiv4 reposarXiv:2602.20159
VBVR-Wan2.2-diffsynth, VBVR-Wan2.1-diffsynth, VBVR-Dataset, VBVR-LTX2.3-diffsynth
DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer
arXiv4 reposarXiv:2602.24096
nurec-skills, harmonizer, DiffusionHarmonizer, Harmonizer
OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens
arXiv4 reposarXiv:2603.02138
OmniLottie, OmniLottie, MMLottieBench, MMLottie-2M
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
arXiv4 reposarXiv:2603.12201
Hy4-preview, Hy4-preview, GLM-5.2, GLM-5.2-FP8
Fast-WAM: Do World Action Models Need Test-time Future Imagination?
arXiv4 reposarXiv:2603.16666
FastWAM, fastwam, LIBERO-fastwam, robotwin2.0-fastwam
PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression
arXiv4 reposarXiv:2603.29078
Qwen3.5-9B-PolarQuant-MLX-4bit, Qwen3.5-9B-PolarEngine-v4, Qwen3.5-9B-PolarQuant-Q5, polarengine-vllm
arXiv:2603.74245
arXiv4 reposarXiv:2603.74245
Qwen3.5-9B-EOQ-v3, eoq-quantization, Qwen3.5-9B-PolarQuant-MLX-4bit, Qwen3.5-9B-PolarEngine-v4
REAM: Merging Improves Pruning of Experts in LLMs
arXiv4 reposarXiv:2604.04356
ream, Qwen3-30B-A3B-Instruct-2507-REAM, Qwen3-Next-80B-A3B-Instruct-REAM, GLM-4.5-Air-REAM
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence
arXiv4 reposarXiv:2604.07296
OpenSpatial-InternVL3-8B, OpenSpatial-InternVL2.5-8B, OpenSpatial-Qwen3-VL-8B, OpenSpatial-Qwen2.5-VL-7B
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
arXiv4 reposarXiv:2604.08516
MolmoWeb-8B, MolmoWeb-4B, MolmoWeb-8B-Native, MolmoWeb-4B-Native
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion
arXiv4 reposarXiv:2604.09450
ECHO_Base_block8, ECHO_block8, ECHO_Base_block4, ECHO_block4
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
arXiv4 reposarXiv:2604.13416
DF3DV, DI2FIX_HF, DF3DV-1K, DF3DV-1K-Fixer
Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation
arXiv4 reposarXiv:2604.18468
nurec-skills, asset-harvester, asset-harvester, asset-harvester
Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding
arXiv4 reposarXiv:2604.22245
LAT-Bench, LAT-Audio, LAT-Audio-Base, LAT-Chronicle
DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models
arXiv4 reposarXiv:2605.07210
diffretriever-dream-7b-single, diffretriever-llada-8b-single, diffretriever-dream-7b-multi-q4-p16, diffretriever-llada-8b-multi-q4-p4
SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning
arXiv4 reposarXiv:2605.09266
SwanLab, SeePhy-Pro, PhysRL, SeePhysPro
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
arXiv4 reposarXiv:2605.11061
HiDream-O1-Image, HiDream-O1-Image, HiDream-O1-Image-Dev, HiDream-O1-Image-Dev-2604
Post-Trained MoE Can Skip Half Experts via Self-Distillation
arXiv4 reposarXiv:2605.18643
ZEDA, ZEDA, ZEDA-Qwen3-30B-A3B-Dynamic, ZEDA-GLM-4.7-Flash-Dynamic
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
arXiv4 reposarXiv:2605.20342
ParaVT, ParaVT-Parquet, ParaVT-Source, ParaVT-8B
ETCHR: Editing To Clarify and Harness Reasoning
arXiv4 reposarXiv:2605.23897
ETCHR-GRPO-10K, ETCHR-FLUX.2-klein-9B, ETCHR-SFT-400K, DL3DV-2k
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
arXiv4 reposarXiv:2606.00020
ChineseErrorCorrector3-4B, ChineseErrorCorrector4-4B, ChineseErrorDetectorElectra, ChineseErrorCorrector4-4B
Unlimited OCR Works
arXiv4 reposarXiv:2606.23050
Unlimited-OCR, Unlimited-OCR, devanagari-ocr-benchmark, Unlimited-OCR-ROCm
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
arXiv4 reposarXiv:2607.24904
Mage, mage-7255fc67, Mage-VL-RTSP, Mage-VL
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
arXiv4 reposarXiv:2609.03796
LLaDA-Image, LLaDA-Image-FP8, LLaDA-Image-Turbo, LLaDA-Image-Turbo-FP8
A foundation model for clinical-grade computational pathology and rare cancers detection
Nature4 reposNature:s41591-024-03141-0
PIANO, TRIDENT, TridentEdited, aegis
DeepSpeed
ACM3 reposACM:3394486.3406703
DeepSpeed, DeepSpeed, LLMSurvey
Compositional Semantic Parsing on Semi-Structured Tables
arXiv3 reposarXiv:1508.00305
tapas-base-finetuned-wtq, WikiTableQuestions, BIPIA
Relation Classification via Recurrent Neural Network
arXiv3 reposarXiv:1508.01006
IEPile, iepie, iepile
TinyLFU: A Highly Efficient Cache Admission Policy
arXiv3 reposarXiv:1512.00727
ristretto, ristretto, BitFaster.Caching
Improved Techniques for Training GANs
arXiv3 reposarXiv:1606.03498
materialgan, VideoGPT, stylegan2-ada-pytorch
STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset
arXiv3 reposarXiv:1705.00823
STAIR-Captions, huggingface-datasets_STAIR-Captions, STAIR-captions
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
arXiv3 reposarXiv:1706.08500
materialgan, pytorch-fid, stylegan2-ada-pytorch
Progressive Growing of GANs for Improved Quality, Stability, and Variation
arXiv3 reposarXiv:1710.10196
materialgan, progressive_growing_of_gans, SkinDeep
ArcFace: Additive Angular Margin Loss for Deep Face Recognition
arXiv3 reposarXiv:1801.07698
Snap-Safe-Python, FLUXSynID, AuraFace-v1
Improving Distantly Supervised Relation Extraction using Word and Entity Based Attention
arXiv3 reposarXiv:1804.06987
IEPile, iepie, iepile
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
arXiv3 reposarXiv:1804.07461
glue, roberta-large-mnli, lares
A Simple Method for Commonsense Reasoning
arXiv3 reposarXiv:1806.02847
roberta-large-mnli, roberta-base, roberta-large
XNLI: Evaluating Cross-lingual Sentence Representations
arXiv3 reposarXiv:1809.05053
roberta-large-mnli, Multilingual-MiniLM-L12-H384, XLM
A Style-Based Generator Architecture for Generative Adversarial Networks
arXiv3 reposarXiv:1812.04948
materialgan, ffhq-dataset, stylegan2-ada-pytorch
IPRE: a Dataset for Inter-Personal Relationship Extraction
arXiv3 reposarXiv:1907.12801
IEPile, iepie, iepile
FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age
arXiv3 reposarXiv:1908.04913
clip-vit-base-patch32, clip-vit-base-patch16, vit_large_patch14_clip_224.openai
Text Summarization with Pretrained Encoders
arXiv3 reposarXiv:1908.08345
HiWestSum, ATS-islamic-organization-news, BertSum
UER: An Open-Source Toolkit for Pre-training Models
arXiv3 reposarXiv:1909.05658
t5-base-chinese-cluecorpussmall, t5-small-chinese-cluecorpussmall, gpt2-chinese-cluecorpussmall
PubMedQA: A Dataset for Biomedical Research Question Answering
arXiv3 reposarXiv:1909.06146
PodGPT, RARE, PubMedQA
PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
arXiv3 reposarXiv:1912.08777
GPT_Ranker, long-ke-t5, YoYAK
Lung and Colon Cancer Histopathological Image Dataset (LC25000)
arXiv3 reposarXiv:1912.12142
LC25000-clean, LC25000, lung_colon_image_set
CLUENER2020: Fine-grained Named Entity Recognition Dataset and Benchmark for Chinese
arXiv3 reposarXiv:2001.04351
IEPile, iepie, iepile
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
arXiv3 reposarXiv:2003.10555
tf-transformers, europeana-bert, YoYAK
Dense Passage Retrieval for Open-Domain Question Answering
arXiv3 reposarXiv:2004.04906
bge-m3, DPR, odqa_baseline_code
Longformer: The Long-Document Transformer
arXiv3 reposarXiv:2004.05150
speechless-starcoder2-15b, final-project-level3-nlp-02, YoYAK
End-to-End Object Detection with Transformers
arXiv3 reposarXiv:2005.12872
detr-resnet-50, coreml-detr-semantic-segmentation, DINO
DocVQA: A Dataset for VQA on Document Images
arXiv3 reposarXiv:2007.00398
DocVQA, CoExVQA, DocVQA
Relevance-guided Supervision for OpenQA with ColBERT
arXiv3 reposarXiv:2007.00814
plaidrepro, colbertv2.0, ColBERT
A Large-Scale Chinese Short-Text Conversation Dataset
arXiv3 reposarXiv:2008.03946
CDial-GPT_LCCC-base, CDial-GPT_LCCC-large, mmchat
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
arXiv3 reposarXiv:2009.11462
real-toxicity-prompts, real-toxicity-prompts, causal_unlearn_llm
Validating UTF-8 In Less Than One Instruction Per Byte
arXiv3 reposarXiv:2010.03090
simdjson, simdutf, fastvalidate-utf-8
arXiv:2010.05171
arXiv3 reposarXiv:2010.05171
TIL-2023, wav2vec2-conformer-rel-pos-large-960h-ft, s2t-small-librispeech-asr
mT5: A massively multilingual pre-trained text-to-text transformer
arXiv3 reposarXiv:2010.11934
mt5-base, mt5-xxl, tf-transformers
Scaled-YOLOv4: Scaling Cross Stage Partial Network
arXiv3 reposarXiv:2011.08036
darknet, darknet, nanodet
BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla
arXiv3 reposarXiv:2101.00204
xnli_bn, squad_bn, banglishbert
Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval
arXiv3 reposarXiv:2101.00436
plaidrepro, colbertv2.0, ColBERT
Number Parsing at a Gigabyte per Second
arXiv3 reposarXiv:2101.11408
ffc.h, fast_float, csFastFloat
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
arXiv3 reposarXiv:2102.03334
vilt-b32-finetuned-vqa, ViLT, visual-spatial-reasoning
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
arXiv3 reposarXiv:2102.08981
dalle-mini, cc12m-wds, conceptual-12m
WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
arXiv3 reposarXiv:2103.01913
clip-italian, clip-italian, wit
Vision Transformers for Dense Prediction
arXiv3 reposarXiv:2103.13413
Depth-Estimation, MiDaS, ldm3d-4c
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
arXiv3 reposarXiv:2103.14030
FoodSeg103-Benchmark-v1, MiDaS, Swin-Transformer
Towards Measuring Fairness in AI: the Casual Conversations Dataset
arXiv3 reposarXiv:2104.02821
parakeet-tdt_ctc-1.1b, parakeet-tdt_ctc-110m, parakeet-tdt-1.1b
AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
arXiv3 reposarXiv:2104.03603
pyannote-audio, 3D-Speaker, 3d-speaker
A Reinforcement Learning Environment For Job-Shop Scheduling
arXiv3 reposarXiv:2104.03760
JSSEnv, RL-Job-Shop-Scheduling, JSS
Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling
arXiv3 reposarXiv:2104.06967
splade, pyterrier_dr, pyterrier_dr_jpq
Sample and Computation Redistribution for Efficient Face Detection
arXiv3 reposarXiv:2105.04714
insightface, face-alignment, Snap-Safe-Python
CogView: Mastering Text-to-Image Generation via Transformers
arXiv3 reposarXiv:2105.13290
visualglm-6b, VisualGLM-6B, VisualGLM-6B
Structured Denoising Diffusion Models in Discrete State-Spaces
arXiv3 reposarXiv:2107.03006
dlms-sinks, mdlm, minimal-dlm
Per-Pixel Classification is Not All You Need for Semantic Segmentation
arXiv3 reposarXiv:2107.06278
mask2former-swin-large-coco-panoptic, maskformer-swin-small-coco, MaskFormer
Contrastive Language-Image Pre-training for the Italian Language
arXiv3 reposarXiv:2108.08688
clip-italian, clip-italian, clip-italian-demo
Mr. TyDi: A Multi-lingual Benchmark for Dense Retrieval
arXiv3 reposarXiv:2108.08787
multilingual-e5-base, multilingual-e5-large, KoPrivateGPT
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation
arXiv3 reposarXiv:2109.00859
CodeRL, codet5-large-ntp-py, codet5-large
Finetuned Language Models Are Zero-Shot Learners
arXiv3 reposarXiv:2109.01652
FLAN, promptsource, flan
TruthfulQA: Measuring How Models Mimic Human Falsehoods
arXiv3 reposarXiv:2109.07958
truthful_qa, RAIN, truthful_qa_de
TorchXRayVision: A library of chest X-ray datasets and models
arXiv3 reposarXiv:2111.00595
torchxrayvision, densenet121-res224-chex, torchxrayvision
GMFlow: Learning Optical Flow via Global Matching
arXiv3 reposarXiv:2111.13680
unimatch, gmflow, prisma
Masked-attention Mask Transformer for Universal Image Segmentation
arXiv3 reposarXiv:2112.01527
mask2former-swin-large-coco-panoptic, Mask2Former, Food_Calories
Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
arXiv3 reposarXiv:2201.02177
grokfast, grokking, grok
A ConvNet for the 2020s
arXiv3 reposarXiv:2201.03545
convnext_perceptual_loss, ConvNeXt, CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft-soup
SGPT: GPT Sentence Embeddings for Semantic Search
arXiv3 reposarXiv:2202.08904
sgpt-bloom-7b1-msmarco, sgpt, sgpt-bloom-1b7-nli
iSTFTNet: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform
arXiv3 reposarXiv:2203.02395
Kokoro-82M, kokoro-82M-onnx-opt, fish-diffusion
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
arXiv3 reposarXiv:2203.05482
Twin-Merging, Mario, MergeLM
BERTopic: Neural topic modeling with a class-based TF-IDF procedure
arXiv3 reposarXiv:2203.05794
turftopic, turkish-complaint-topic-clustering, BERTopic
ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
arXiv3 reposarXiv:2203.09509
toxigen-data, toxigen, toxigen-data
Training Compute-Optimal Large Language Models
arXiv3 reposarXiv:2203.15556
awesome-totally-open-chatgpt, llama2.c, falcon-refinedweb
PaLM: Scaling Language Modeling with Pathways
arXiv3 reposarXiv:2204.02311
idefics-80b-instruct, transformer-tricks, idefics-9b-instruct
KOBEST: Korean Balanced Evaluation of Significant Tasks
arXiv3 reposarXiv:2204.04541
polyglot-ko-1.3b, polyglot-ko-3.8b, polyglot-ko-5.8b
The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink
arXiv3 reposarXiv:2204.05149
Llama-3.1-AlternateTokenizer, Meta-Llama-3.1-8B-bnb-4bit, Meta-Llama-3.1-8B-llamafile
Hierarchical Text-Conditional Image Generation with CLIP Latents
arXiv3 reposarXiv:2204.06125
coyo-dataset, coyo-700m, BDM1.0
mGPT: Few-Shot Learners Go Multilingual
arXiv3 reposarXiv:2204.07580
mGPT, mgpt, mGPT
Visual Spatial Reasoning
arXiv3 reposarXiv:2205.00363
visual-spatial-reasoning, vsr_random, vsr_zeroshot
UL2: Unifying Language Learning Paradigms
arXiv3 reposarXiv:2205.05131
turkish-bert, bert5urk, flan-ul2
Vectorized and performance-portable Quicksort
arXiv3 reposarXiv:2205.05982
node, highway, gecko-dev
Pretraining is All You Need for Image-to-Image Translation
arXiv3 reposarXiv:2205.12952
Stable-Diffusion-FineTuned-zh-v1, Stable-Diffusion-FineTuned-zh-v2, Stable-Diffusion-FineTuned-zh-v0
DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps
arXiv3 reposarXiv:2206.00927
PixArt-alpha, pixeart, PixArt-alpha
Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation
arXiv3 reposarXiv:2206.02777
D-FINE-seg, D-FINE-seg, DINO
CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
arXiv3 reposarXiv:2207.01780
CodeRL, codet5-large-ntp-py, codet5-large
YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
arXiv3 reposarXiv:2207.02696
darknet, darknetcv, darknet
arXiv:2207.05987
arXiv3 reposarXiv:2207.05987
docprompting, tldr, docprompting-conala
Efficient Training of Language Models to Fill in the Middle
arXiv3 reposarXiv:2207.14255
Qwen2.5-Coder, Qwen3-Coder, speechless-starcoder2-15b
Twitter Topic Classification
arXiv3 reposarXiv:2209.09824
tweetnlp, tweet_topic_single, tweet_topic_multi
Human Motion Diffusion Model
arXiv3 reposarXiv:2209.14916
a-mdm-text-to-motion, a-motion-diffusion, motion-diffusion-model
Named Entity Recognition in Twitter: A Dataset and Analysis on Short-Term Temporal Shifts
arXiv3 reposarXiv:2210.03797
tweetnlp, tweetner7, tner
MedCLIP: Contrastive Learning from Unpaired Medical Images and Text
arXiv3 reposarXiv:2210.10163
MediMeta-C, RobustMedCLIP, RobustMedCLIP
ESB: A Benchmark For Multi-Domain End-to-End Speech Recognition
arXiv3 reposarXiv:2210.13352
distil-large-v2, distil-medium.en, distil-small.en
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
arXiv3 reposarXiv:2211.05100
api-for-open-llm, bloom, GlorIA
Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation
arXiv3 reposarXiv:2211.12572
diff-mining, plug-and-play, PnP-diffusion-features
InternVideo: General Video Foundation Models via Generative and Discriminative Learning
arXiv3 reposarXiv:2212.03191
InternVideo, InternVid-Full, internvideo-d2a11ea9
Editing Models with Task Arithmetic
arXiv3 reposarXiv:2212.04089
Twin-Merging, Mario, MergeLM
Reproducible scaling laws for contrastive language-image learning
arXiv3 reposarXiv:2212.07143
CLIP_benchmark, CLIP-ViT-bigG-14-laion2B-39B-b160k, TiViT
Objaverse: A Universe of Annotated 3D Objects
arXiv3 reposarXiv:2212.08051
Cap3D, objaverse, objaverse-xl
Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
arXiv3 reposarXiv:2212.09689
COIG, COIG, unnatural-instructions
SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization
arXiv3 reposarXiv:2212.10465
soda, sodaverse, cosmo-xl
AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
arXiv3 reposarXiv:2301.12503
AudioLDM-S-Full, MMDisCo, audioldm_eval
Accelerating Large Language Model Decoding with Speculative Sampling
arXiv3 reposarXiv:2302.01318
LLM-Sampling, LLMSpeculativeSampling, RemoteSpeculativeDecoding
MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation
arXiv3 reposarXiv:2302.08113
MultiDiffusion, MultiDiffusion, multidiffusion-region-based
UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers
arXiv3 reposarXiv:2303.00807
ColBERT, RAG, RAGatouille
Consistency Models
arXiv3 reposarXiv:2303.01469
TCD, TCD-SDXL-LoRA, LakonLab
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
arXiv3 reposarXiv:2303.04671
TaskMatrix, visual-chatgpt-zh, TaskMatrix
Tag2Text: Guiding Vision-Language Model via Image Tagging
arXiv3 reposarXiv:2303.05657
recognize-anything, recognize-anything-plus-model, recognize_anything_model
Erasing Concepts from Diffusion Models
arXiv3 reposarXiv:2303.07345
erasing, Erasing-Concepts-In-Diffusion, leco
TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
arXiv3 reposarXiv:2303.11897
tifa, llama2_tifa_question_generation, banana100-additional-iqa-models
Capabilities of GPT-4 on Medical Challenge Problems
arXiv3 reposarXiv:2303.13375
Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, med42
Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
arXiv3 reposarXiv:2303.13439
Text2Video-Zero, Text2Video-Zero, Text2Video-Zero
G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
arXiv3 reposarXiv:2303.16634
RLHF-Korean-Friendly-LLM, KULLM-RLHF, level3_nlp_finalproject-nlp-12
RRHF: Rank Responses to Align Language Models with Human Feedback without tears
arXiv3 reposarXiv:2304.05302
wombat-7b-gpt4-delta, RRHF, wombat-7b-delta
ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
arXiv3 reposarXiv:2304.05977
ImageReward, ImageRewardDB, banana100-additional-iqa-models
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations
arXiv3 reposarXiv:2304.06795
parakeet-tdt_ctc-1.1b, parakeet-tdt_ctc-110m, parakeet-tdt-1.1b
InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction
arXiv3 reposarXiv:2304.08085
IEPile, iepie, iepile
FindVehicle and VehicleFinder: A NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system
arXiv3 reposarXiv:2304.10893
IEPile, iepie, iepile
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
arXiv3 reposarXiv:2305.08322
Qwen-14B-Chat, Qwen-7B-Chat, Qwen-1_8B
Common Diffusion Noise Schedules and Sample Steps are Flawed
arXiv3 reposarXiv:2305.08891
OneTrainer, smalldiffusion, YetAnotherStableDiffusion
Towards Expert-Level Medical Question Answering with Large Language Models
arXiv3 reposarXiv:2305.09617
Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, med42
SoundStorm: Efficient Parallel Audio Generation
arXiv3 reposarXiv:2305.09636
Dia-1.6B-0626, dia, soundstorm-speechtokenizer
Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold
arXiv3 reposarXiv:2305.10973
DragonDiffusion, InternGPT, InternChat
Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages
arXiv3 reposarXiv:2305.12182
glot500-base, Glot500, Glot500
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
arXiv3 reposarXiv:2305.13245
tiny-vllm, speechless-starcoder2-15b, llama2.zig
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
arXiv3 reposarXiv:2305.14251
KoHalluLens, HalluLens, llm_factuality_tuning
arXiv:2305.14387
arXiv3 reposarXiv:2305.14387
alpaca_eval, alpaca_farm, JudgeBench
Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems
arXiv3 reposarXiv:2305.15017
zephyr-7b-sft-full124, zephyr-7b-sft-full124_d270, calc-x
HuatuoGPT, towards Taming Language Model to Be a Doctor
arXiv3 reposarXiv:2305.15075
HuatuoGPT2-SFT-GPT4-140K, HuatuoGPT2-Pretraining-Instruction, HuatuoGPT
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
arXiv3 reposarXiv:2306.03341
honest_llama, honest_llama2_chat_7B, 2023FallNLP
Recognize Anything: A Strong Image Tagging Model
arXiv3 reposarXiv:2306.03514
recognize-anything, recognize-anything-plus-model, recognize_anything_model
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
arXiv3 reposarXiv:2306.05087
JudgeBench, Alpaca-7B-v1, PandaLM
Fast Segment Anything
arXiv3 reposarXiv:2306.12156
sd-webui-inpaint-anything, FastSAM, sd-webui-inpaint-anything
arXiv:2306.14824
arXiv3 reposarXiv:2306.14824
vqazero, Zero-and-Few-Shot-Visual-Question-Answering, GRIT
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
arXiv3 reposarXiv:2306.17107
LLaVAR, LLaVAR_delta, LLaVAR
Provable Robust Watermarking for AI-Generated Text
arXiv3 reposarXiv:2306.17439
Adversarial-Paraphrasing, impossibility-watermark, lm-watermarking
JourneyDB: A Benchmark for Generative Image Understanding
arXiv3 reposarXiv:2307.00716
LaVi-Bridge, SEED-Data-Edit-Part2-3, SEED-Data-Edit
Flacuna: Unleashing the Problem Solving Power of Vicuna using FLAN Fine-Tuning
arXiv3 reposarXiv:2307.02053
flacuna-13b-v1.0, flacuna, flan-mini
SVIT: Scaling up Visual Instruction Tuning
arXiv3 reposarXiv:2307.04087
Bunny-v1_1-data, Bunny, Bunny-v1_0-data
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
arXiv3 reposarXiv:2307.04725
LatentSync-1.5, mcm, MMDisCo
Objaverse-XL: A Universe of 10M+ 3D Objects
arXiv3 reposarXiv:2307.05663
Cap3D, objaverse-xl, objaverse-xl
MMBench: Is Your Multi-modal Model an All-around Player?
arXiv3 reposarXiv:2307.06281
idefics-80b-instruct, MMBench, idefics-9b-instruct
Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
arXiv3 reposarXiv:2307.06304
Open-Sora-Plan-v1.2.0, Baichuan-Omni-1.5, idefics2-8b
mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs
arXiv3 reposarXiv:2307.06930
mBLIP, mblip-mt0-xl, mblip-bloomz-7b
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
arXiv3 reposarXiv:2307.13304
smash, llmtools, llmtools
Med-Flamingo: a Multimodal Medical Few-shot Learner
arXiv3 reposarXiv:2307.15189
med-flamingo, med-flamingo, med-flamingo
GEMRec: Towards Generative Model Recommendation
arXiv3 reposarXiv:2308.02205
GEMRec, GEMRec-Roster, GEMRec-PromptBook
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
arXiv3 reposarXiv:2308.03279
universal-ner, UniNER-7B-type-sup, UniNER-7B-all
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
arXiv3 reposarXiv:2308.03825
SecLists, jailbreak_llms, JailbreakRadar
OctoPack: Instruction Tuning Code Large Language Models
arXiv3 reposarXiv:2308.07124
reward-bench, bigcode-evaluation-harness, humanevalpack
BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine
arXiv3 reposarXiv:2308.09442
BioMedGPT-LM-7B, OpenBioMed, OpenBioMed_new
ChatHaruhi: Reviving Anime Character in Reality via Large Language Model
arXiv3 reposarXiv:2308.09597
Chat-Haruhi-Suzumiya, ChatHaruhi-Expand-118K, ChatHaruhi-54K-Role-Playing-Dialogue
SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
arXiv3 reposarXiv:2308.11596
dissertation-project, seamless-m4t-medium, seamless-m4t-large
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
arXiv3 reposarXiv:2309.07915
MIC_full, MIC, MMICL-Instructblip-T5-xxl
LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
arXiv3 reposarXiv:2309.15103
LaVie, LaVie, videophy
arXiv:2310.01218
arXiv3 reposarXiv:2310.01218
CoBSAT, SEED, SEED
OceanGPT: A Large Language Model for Ocean Science Tasks
arXiv3 reposarXiv:2310.02031
OceanGPT-7b, OceanBench, OceanGPT
NEFTune: Noisy Embeddings Improve Instruction Finetuning
arXiv3 reposarXiv:2310.05914
shisa-7b-v1, Arithmo, Arithmo-Mistral-7B
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
arXiv3 reposarXiv:2310.06770
opensre, oh-my-claudecode, SWEBench-verified-mini
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
arXiv3 reposarXiv:2310.08278
moment, Lag-Llama, lag-llama
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
arXiv3 reposarXiv:2310.10501
NeMo-Guardrails, Guardrails, Agent-Guard
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
arXiv3 reposarXiv:2310.11441
roborazzi, Magma-8B, hacktech24_app_testing
Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation
arXiv3 reposarXiv:2310.12371
diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, diar_sortformer_4spk-v1
AgentTuning: Enabling Generalized Agent Abilities for LLMs
arXiv3 reposarXiv:2310.12823
agentlm-7b, agentlm-13b, agentlm-70b
DPM-Solver-v3: Improved Diffusion ODE Solver with Empirical Model Statistics
arXiv3 reposarXiv:2310.13268
PixArt-alpha, pixeart, PixArt-alpha
Zephyr: Direct Distillation of LM Alignment
arXiv3 reposarXiv:2310.16944
zephyr-7b-alpha, zephyr-7b-beta, zephyr-7b-gemma-v0.1
Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents
arXiv3 reposarXiv:2310.19923
late-chunking, mteb-1.34.14, jina-colbert-v1-en
Instruction-Following Evaluation for Large Language Models
arXiv3 reposarXiv:2311.07911
starchat2-15b-v0.1, llm-action, IFEval
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
arXiv3 reposarXiv:2311.16079
Master-Thesis, guidelines, meditron
RETSim: Resilient and Efficient Text Similarity
arXiv3 reposarXiv:2311.17264
text-dedup, usearch, unisim
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
arXiv3 reposarXiv:2312.00752
Vim, mamba2-minimal, mdlm
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
arXiv3 reposarXiv:2312.11456
Online-RLHF, FsfairX-LLaMA3-RM-v0.1, LLaMA3-iterative-DPO-final
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
arXiv3 reposarXiv:2312.12450
EditPackFT, CanItEdit, CanItEdit
YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction
arXiv3 reposarXiv:2312.15548
IEPile, iepie, iepile
DB-GPT: Empowering Database Interactions with Private Large Language Models
arXiv3 reposarXiv:2312.17449
DB-GPT, DB-GPT, pathtraversal_mutation_352
Long Context Compression with Activation Beacon
arXiv3 reposarXiv:2401.03462
bge-large-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5
Mixtral of Experts
arXiv3 reposarXiv:2401.04088
Awesome-AITools, minimind, Swallow-MX-8x7b-NVE-v0.1
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
arXiv3 reposarXiv:2401.05252
PixArt-alpha, pixeart, PixArt-alpha
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
arXiv3 reposarXiv:2401.09047
VideoCrafter, videophy, MMDisCo
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
arXiv3 reposarXiv:2401.10935
ScreenSpot, SeeClick, transformer-final-proj
Mercury: A Code Efficiency Benchmark for Code Large Language Models
arXiv3 reposarXiv:2402.07844
Mercury, Mercury, Venus
CoLLaVO: Crayon Large Language and Vision mOdel
arXiv3 reposarXiv:2402.11248
CoLLaVO, CoLLaVO-7B, MoAI
Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion
arXiv3 reposarXiv:2402.12195
Brote-pretrain, Brote, Brote-IM-XXL
arXiv:2402.12354
arXiv3 reposarXiv:2402.12354
llama3-chinese, llama3-chinese, Llama3-Chinese-Lora
FinBen: A Holistic Financial Benchmark for Large Language Models
arXiv3 reposarXiv:2402.12659
PIXIU, flare-fomc, flare-finer-ord
Training-Free Long-Context Scaling of Large Language Models
arXiv3 reposarXiv:2402.17463
Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
arXiv3 reposarXiv:2402.17766
ShapeLLM, MiniGPT-3D, ReConV2
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
arXiv3 reposarXiv:2402.17811
TruthX, TruthX, Llama-2-7b-chat-TruthX
OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on
arXiv3 reposarXiv:2403.01779
MagicClothing, OOTDiffusion, OOTDiffusion
TripoSR: Fast 3D Object Reconstruction from a Single Image
arXiv3 reposarXiv:2403.02151
TripoSR, TripoSR, TriPlaneCLIP
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
arXiv3 reposarXiv:2403.03206
F5-TTS, maxdiffusion, Rectified-Diffusion
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
arXiv3 reposarXiv:2403.04132
FastChat, multilingual_mt_bench, lmsys-arena-human-preference-55k
TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
arXiv3 reposarXiv:2403.04473
Monkey, MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash
Common 7B Language Models Already Possess Strong Math Capabilities
arXiv3 reposarXiv:2403.04706
Xwin-LM, Xwin-Math-7B-V1.1, Xwin-Math-70B-V1.1
Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head
arXiv3 reposarXiv:2403.06892
OmDet, omdet-turbo-swin-tiny-hf, OmDet-Turbo_tiny_SWIN_T
MoAI: Mixture of All Intelligence for Large Language and Vision Models
arXiv3 reposarXiv:2403.07508
CoLLaVO, MoAI-7B, MoAI
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
arXiv3 reposarXiv:2403.07860
ELLA, LaVi-Bridge, ComfyUI-ELLA-wrapper
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
arXiv3 reposarXiv:2403.09629
quiet-star, quietstar-8-ahead, Adaptive_QuietSTaR
Arc2Face: A Foundation Model for ID-Consistent Human Faces
arXiv3 reposarXiv:2403.11641
FLUXSynID, Arc2Face, Arc2Face
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
arXiv3 reposarXiv:2403.12032
MVEdit, MVEdit, 3D-Adapter
arXiv:2403.13787
arXiv3 reposarXiv:2403.13787
MM-Eval, prometheus-eval, reward-bench
InstantSplat: Sparse-view Gaussian Splatting in Seconds
arXiv3 reposarXiv:2403.20309
InstantSplatPP, InstantSplat, InstantSplat
Evaluating Text-to-Visual Generation with Image-to-Text Generation
arXiv3 reposarXiv:2404.01291
t2v_metrics, clip-flant5-xxl, banana100-additional-iqa-models
CosmicMan: A Text-to-Image Foundation Model for Humans
arXiv3 reposarXiv:2404.01294
CosmicMan, CosmicMan-SD, CosmicMan-SDXL
Long-context LLMs Struggle with Long In-context Learning
arXiv3 reposarXiv:2404.02060
CAG, LongICLBench, LongICLBench
JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks
arXiv3 reposarXiv:2404.03027
JailBreakV_28K, JailBreakV_28K, JailBreakV_28K
Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
arXiv3 reposarXiv:2404.05892
VisualRWKV, rwkv-6-world, rwkv
ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
arXiv3 reposarXiv:2404.07987
ControlNet_Plus_Plus, MultiGen-20M_train, flownetpp
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
arXiv3 reposarXiv:2404.09956
tango, tango2, tango2-full
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
arXiv3 reposarXiv:2404.10774
MiniCheck, C2D-and-D2C-MiniCheck, Bespoke-MiniCheck-7B
Stepwise Alignment for Constrained Language Model Policy Optimization
arXiv3 reposarXiv:2404.11049
sacpo, sacpo, p-sacpo
BLINK: Multimodal Large Language Models Can See but Not Perceive
arXiv3 reposarXiv:2404.12390
BLINK, BLINK_Benchmark, BLINK-ja
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
arXiv3 reposarXiv:2404.14219
fineweb-edu, Phi-3.5-MoE-instruct, Phi-3.5-vision-instruct
PuLID: Pure and Lightning ID Customization via Contrastive Alignment
arXiv3 reposarXiv:2404.16022
PuLID, PuLID, FLUXSynID
OpenStreetView-5M: The Many Roads to Global Visual Geolocation
arXiv3 reposarXiv:2404.18873
osv5m, plonk, baseline
TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
arXiv3 reposarXiv:2404.19205
granite-4.0-3b-vision, tablevqabench, sa2va_eval
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
arXiv3 reposarXiv:2405.04532
nunchaku, nunchaku, deepcompressor
BiMix: A Bivariate Data Mixing Law for Language Model Pretraining
arXiv3 reposarXiv:2405.14908
data-juicer, SciDataOS, data-juicer
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
arXiv3 reposarXiv:2405.15306
AutomaTikZ, DeTikZify, detikzify-v2.5-8b
RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness
arXiv3 reposarXiv:2405.17220
vision-feedback-mix-binarized, vision-feedback-mix-binarized, OmniLMM
arXiv:2406.00770
arXiv3 reposarXiv:2406.00770
EvolKit, aurora-m2, slm-innovator-lab
Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
arXiv3 reposarXiv:2406.00977
Llama-3.1-8B-Dragonfly-Med-v2, Dragonfly, Llama-3.1-8B-Dragonfly-v2
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
arXiv3 reposarXiv:2406.01014
ms-agent, MobileAgent, modelscope-agent
LoFiT: Localized Fine-tuning on LLM Representations
arXiv3 reposarXiv:2406.01563
lo-fit, llama2_7B_base_lofit_mquake, llama2_7B_base_lofit_truthfulqa
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
arXiv3 reposarXiv:2406.01574
MMLU-Pro, MMLU-Pro, RLPR-Evaluation
UltraMedical: Building Specialized Generalists in Biomedicine
arXiv3 reposarXiv:2406.03949
Llama-3-70B-UltraMedical, Llama-3.1-8B-UltraMedical, UltraMedical-Preference
arXiv:2406.04127
arXiv3 reposarXiv:2406.04127
ZeroEval, mmlu-redux, mmlu-redux
MLVU: Benchmarking Multi-task Long Video Understanding
arXiv3 reposarXiv:2406.04264
MLVU, UniG2U, lmms-eval-mmllm
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
arXiv3 reposarXiv:2406.05761
prometheus-eval, BiGGen-Bench-Results, BiGGen-Bench
Scaling up masked audio encoder learning for general audio classification
arXiv3 reposarXiv:2406.06992
dasheng, Dasheng, hashing-baseline
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
arXiv3 reposarXiv:2406.07594
MLLMGuard, MLLMGuard, MLLMGuard
LVBench: An Extreme Long Video Understanding Benchmark
arXiv3 reposarXiv:2406.08035
LVBench, LVBench, LVBench
GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
arXiv3 reposarXiv:2406.08451
GUI-Odyssey, GUI-Odyssey, GUIOdyssey
Interpreting the Weight Space of Customized Diffusion Models
arXiv3 reposarXiv:2406.09413
weights2weights, weights2weights, weights2weights
RobustSAM: Segment Anything Robustly on Degraded Images
arXiv3 reposarXiv:2406.09627
robustsam-vit-large, robustsam-vit-huge, robustsam-vit-base
On the Impacts of Contexts on Repository-Level Code Generation
arXiv3 reposarXiv:2406.11927
RepoExec, RepoExec, RepoExec-Instruct
arXiv:2406.11939
arXiv3 reposarXiv:2406.11939
arena-hard-auto, PPE, JudgeBench
CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
arXiv3 reposarXiv:2406.13897
Step1X-3D, Hunyuan3D-Omni, UniTEX
ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
arXiv3 reposarXiv:2406.14130
ExVideo-CogVideoX-LoRA-129f-v1, ExVideo-SVD-128f-v1, diffSynth-studio-notes
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
arXiv3 reposarXiv:2406.16620
omchat-v2.0-13B-single-beta_hf, omchat, OmAgent
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
arXiv3 reposarXiv:2406.17806
MOSSBench, MOSSBench, MOSSBench
Nomic Embed Vision: Expanding the Latent Space
arXiv3 reposarXiv:2406.18587
contrastors, nomic-embed-vision-v1.5, contrastors
ProgressGym: Alignment with a Millennium of Moral Progress
arXiv3 reposarXiv:2406.20087
ProgressGym, ProgressGym-TimelessQA, ProgressGym-HistText
Scaling Synthetic Data Creation with 1,000,000,000 Personas
arXiv3 reposarXiv:2406.20094
camel, PersonaHub, aurora-m2
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
arXiv3 reposarXiv:2407.01284
We-Math, We-Math, We-Math2.0
OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
arXiv3 reposarXiv:2407.02371
OpenVid-1M, OpenVid-1M, OpenVid-1M-mapping
Crafting Large Language Models for Enhanced Interpretability
arXiv3 reposarXiv:2407.04307
CB-LLMs, ConceptBottleneck-GUI-Experiment, Concept-Bottleneck-LLM
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
arXiv3 reposarXiv:2407.08583
data-juicer, SciDataOS, data-juicer
Robotic Control via Embodied Chain-of-Thought Reasoning
arXiv3 reposarXiv:2407.08693
embodied-CoT, Adaptive-CoT-in-VLA, Fast-ECoT
Panacea: A foundation model for clinical trial search, summarization, design, and recruitment
arXiv3 reposarXiv:2407.11007
Panacea, Panacea-7B-Chat, TrialAlign
Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development
arXiv3 reposarXiv:2407.11784
data-juicer, SciDataOS, data-juicer
Sentiment Reasoning for Healthcare
arXiv3 reposarXiv:2407.21054
Sentiment-Reasoning, Sentiment-Reasoning, Sentiment-Reasoning
OmniParser for Pure Vision Based GUI Agent
arXiv3 reposarXiv:2408.00203
OmniParser, OmniParser-v2.0, OmniParser
Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid
arXiv3 reposarXiv:2408.02034
MiniMonkey, Monkey, MiniMokney
VidGen-1M: A Large-Scale Dataset for Text-to-video Generation
arXiv3 reposarXiv:2408.02629
VIDGEN-1M, VidGen, VIDGEN-v1.0
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
arXiv3 reposarXiv:2408.05147
gemma-scope-9b-pt-res, gemma-scope, gemma-scope-2b-pt-transcoders
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
arXiv3 reposarXiv:2408.07055
LongWriter-llama3.1-8b, LongWriter-6k, LongWriter-glm4-9b
Docling Technical Report
arXiv3 reposarXiv:2408.09869
docling, docling, my-docling
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
arXiv3 reposarXiv:2408.12528
Show-o, show-o-w-clip-vit, show-o
NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
arXiv3 reposarXiv:2408.13106
diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, diar_sortformer_4spk-v1
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
arXiv3 reposarXiv:2408.15556
HR-Bench, HR-Bench, sa2va_eval
Target-Driven Distillation: Consistency Distillation with Target Timestep Selection and Decoupled Guidance
arXiv3 reposarXiv:2409.01347
Target-Driven-Distillation, TDD, Target-Driven-Distillation
LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
arXiv3 reposarXiv:2409.02897
LongCite-glm4-9b, LongCite-llama3.1-8b, LongCite-45k
Synthetic continued pretraining
arXiv3 reposarXiv:2409.07431
Synthetic_Continued_Pretraining, entigraph-quality-corpus, llama-3-8b-entigraph-quality
Phikon-v2, A large and public feature extractor for biomarker prediction
arXiv3 reposarXiv:2409.09173
PIANO, phikon-v2, SEAL
Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models
arXiv3 reposarXiv:2409.12106
ValueLlama-3-8B, gpv, gpv
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
arXiv3 reposarXiv:2409.20007
DeSTA2, DeSTA2-8B-beta, Speech-IFEval
Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers
arXiv3 reposarXiv:2409.20537
HPT, hpt-base, hpt_locoman
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second
arXiv3 reposarXiv:2410.02073
DepthPro, ml-depth-pro, Depth-Estimation
Distilling an End-to-End Voice Assistant Without Instruction Training Data
arXiv3 reposarXiv:2410.02678
DiVA-llama-3-v0-8b, Speech-IFEval, canto-audio-llm
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?
arXiv3 reposarXiv:2410.03859
SWE-bench, SWE-bench, SWE-bench-fork
IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation
arXiv3 reposarXiv:2410.07171
IterComp, IterComp, RPG-DiffusionMaster
Large-Scale 3D Medical Image Pre-training with Geometric Context Priors
arXiv3 reposarXiv:2410.09890
PreCT-160K, VoCo, VoComni
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
arXiv3 reposarXiv:2410.10733
efficientvit, dc-ae-f32c32-sana-1.0-diffusers, Jazz
Large Continual Instruction Assistant
arXiv3 reposarXiv:2410.10868
Continual-NExT, CoIN, CoIN_Refined
DepthSplat: Connecting Gaussian Splatting and Depth
arXiv3 reposarXiv:2410.13862
unimatch, depthsplat, depthsplat
arXiv:2410.15522
arXiv3 reposarXiv:2410.15522
MM-Eval, prometheus-eval, m-rewardbench
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
arXiv3 reposarXiv:2410.16153
Pangea, Pangea-7B, PangeaInstruct
Scaling Diffusion Language Models via Adaptation from Autoregressive Models
arXiv3 reposarXiv:2410.17891
SDAR, dLLM-RL, FreeDave
Scaling up Masked Diffusion Models on Text
arXiv3 reposarXiv:2410.18514
SMDM, SMDMtry, minimal-dlm
arXiv:2410.21035
arXiv3 reposarXiv:2410.21035
sdtt, sdtt, SDTT-LaViDa
OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
arXiv3 reposarXiv:2410.23218
UI-TARS, SeeClick, transformer-final-proj
Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations
arXiv3 reposarXiv:2411.01661
LLambada, Llambada, LLambada
Cut Your Losses in Large-Vocabulary Language Models
arXiv3 reposarXiv:2411.09009
ml-cross-entropy, unsloth-zoo, unsloth_zoo
Golden Noise for Diffusion Models: A Learning Framework
arXiv3 reposarXiv:2411.09502
Golden-Noise-for-Diffusion-Models, GoldenNoiseModel, ComfyUI_Golden-Noise
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
arXiv3 reposarXiv:2411.10440
LLaVA-CoT-100k, LLaVA-CoT, Llama-3.2V-11B-cot
Adversarial Diffusion Compression for Real-World Image Super-Resolution
arXiv3 reposarXiv:2411.13383
OSEDiff, AdcSR, AdcSR
Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
arXiv3 reposarXiv:2411.13588
xDiT, mochi-xdit, DiTCacheAnalysis
BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models
arXiv3 reposarXiv:2411.15232
BiomedCoOp, BiomedCoOp, BiomedCoOp
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
arXiv3 reposarXiv:2411.16537
RoboSpatial-Eval, RoboSpatial-Home, RoboSpatial
GRAPE: Generalizing Robot Policy via Preference Alignment
arXiv3 reposarXiv:2411.19309
OpenVLA-7B-GRAPE-Simpler, GRAPE, OpenVLA-7B-SFT-Simpler
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
arXiv3 reposarXiv:2411.19509
ditto-talkinghead, ditto-talkinghead, LivePortrait
OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows
arXiv3 reposarXiv:2412.01169
OmniFlows, OmniFlow-v0.9, OmniFlow-v0.5
Structured 3D Latents for Scalable and Versatile 3D Generation
arXiv3 reposarXiv:2412.01506
TRELLIS-image-large, TRELLIS-image-large, c3-trellis-gradio
Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis
arXiv3 reposarXiv:2412.01819
Switti, VQVAE-Switti, Switti-AR
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
arXiv3 reposarXiv:2412.04431
Infinity, infinity, infinity
Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation
arXiv3 reposarXiv:2412.04954
Med-CXRGen-F, Med-CXRGen-I, RRG-BioNLP-ACL2024
BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
arXiv3 reposarXiv:2412.07769
BiMediX2-8B, BiMediX2-70B, BiMediX2
Learning Flow Fields in Attention for Controllable Person Image Generation
arXiv3 reposarXiv:2412.08486
Leffa, Leffa, Leffa
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
arXiv3 reposarXiv:2412.10117
CosyVoice, CosyVoice, FastCosyVoice
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
arXiv3 reposarXiv:2412.10302
deepseek-vl2, deepseek-vl2-tiny, deepseek-vl2-small
Empowering LLMs to Understand and Generate Complex Vector Graphics
arXiv3 reposarXiv:2412.11102
OmniSVG, OmniSVG-train, omnisvg-train
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
arXiv3 reposarXiv:2412.12661
medmax, medmax_eval_data, medmax_data
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
arXiv3 reposarXiv:2412.14171
VSI-Bench, embodied-eval, behaviour_subtask
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
arXiv3 reposarXiv:2412.15838
Align-DS-V, DollyTails-12K, align-anything
Text2midi: Generating Symbolic Music from Captions
arXiv3 reposarXiv:2412.16526
Text2midi, text2midi, text2midi
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
arXiv3 reposarXiv:2412.17574
data-juicer, SciDataOS, data-juicer
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
arXiv3 reposarXiv:2412.18911
TaylorSeer, ToCa, DuCa
2 OLMo 2 Furious
arXiv3 reposarXiv:2501.00656
OLMo-2-0425-1B, olmes, qwen3-8b-base
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
arXiv3 reposarXiv:2501.01108
MuQ-MuLan-large, MuQ, MuQ-large-msd-iter
Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos
arXiv3 reposarXiv:2501.04001
Sa2VA, Sa2VA, Sa2VA-Training
Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation
arXiv3 reposarXiv:2501.12202
Hunyuan3D-2.1, HY3D-Bench, Hunyuan3D-Omni
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
arXiv3 reposarXiv:2501.12326
UI-TARS, ScreenSpot-Pro-GUI-Grounding, UI-TARS-7B-DPO
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
arXiv3 reposarXiv:2501.12386
InternVideo, InternVL_2_5_HiCo_R16, internvideo-d2a11ea9
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
arXiv3 reposarXiv:2501.13106
VideoLLaMA2, VideoLLaMA3-7B-Image, VideoLLaMA3-2B-Image
Distilling foundation models for robust and efficient models in digital pathology
arXiv3 reposarXiv:2501.16239
plism-benchmark, SEAL, plism-dataset
M+: Extending MemoryLLM with Scalable Long-Term Memory
arXiv3 reposarXiv:2502.00592
MemoryLLM, mplus-8b, ExtendingMemoryLLM
Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
arXiv3 reposarXiv:2502.01051
LPO, LRM, LPO
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity
arXiv3 reposarXiv:2502.01776
nunchaku, nunchaku, Sparse-VideoGen
Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data
arXiv3 reposarXiv:2502.04380
data-juicer, SciDataOS, data-juicer
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
arXiv3 reposarXiv:2502.10248
stepvideo-t2v, Step-Video-T2V, stepvideo-t2v-turbo
Baichuan-M1: Pushing the Medical Capability of Large Language Models
arXiv3 reposarXiv:2502.12671
Baichuan-M1-14B-Instruct, Baichuan-M1-14B-Base, Med-R1
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
arXiv3 reposarXiv:2502.14846
pixmo-docs, CoSyn-point, CoSyn-400K
FeatSharp: Your Vision Model Features, Sharper
arXiv3 reposarXiv:2502.16025
RADIO, C-RADIOv4-SO400M, C-RADIOv4-H
LettuceDetect: A Hallucination Detection Framework for RAG Applications
arXiv3 reposarXiv:2502.17125
LettuceDetect, lettucedect-base-modernbert-en-v1, lettucedect-large-modernbert-en-v1
RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour
arXiv3 reposarXiv:2503.02572
RaceVLA, RaceVLA_models, RaceVLA_dataset
Unified Reward Model for Multimodal Understanding and Generation
arXiv3 reposarXiv:2503.05236
UnifiedReward-qwen-3b, VideoDPO, ShareGPTVideo-DPO
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
arXiv3 reposarXiv:2503.05639
VPBench, VideoPainter, VPData
GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
arXiv3 reposarXiv:2503.06073
ECG-Grounding, GEM, GEM
Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment
arXiv3 reposarXiv:2503.07334
ARRA, ARRA, ARRA-Adapt-MIMIC-7B
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
arXiv3 reposarXiv:2503.09499
data-juicer, SciDataOS, data-juicer
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
arXiv3 reposarXiv:2503.11576
docling, docling, my-docling
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
arXiv3 reposarXiv:2503.12165
VTON360, VTON360-THuman2.0, VTON360-MVHumanNet
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
arXiv3 reposarXiv:2503.13377
xiaomi-mimo-vl-miloco, Xiaomi-MiMo-VL-Miloco-7B, Xiaomi-MiMo-VL-Miloco-7B-GGUF
TerraTorch: The Geospatial Foundation Models Toolkit
arXiv3 reposarXiv:2503.20563
terratorch, terratorch, granite-geospatial-biomass
Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging
arXiv3 reposarXiv:2503.22236
ComfyUI-Hi3DGen, Hi3DGen, Hi3DGen
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
arXiv3 reposarXiv:2504.00824
ScholarCopilot, ScholarCopilot-Data-v1, ScholarCopilot-v1
WikiVideo: Article Generation from Multiple Videos
arXiv3 reposarXiv:2504.00939
CRAFT, wikivideo, wikivideo
MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs
arXiv3 reposarXiv:2504.00993
MedReason, MedReason-8B, Med-R1
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
arXiv3 reposarXiv:2504.01805
SpaceR, SpaceR, SpaceR-151k
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
arXiv3 reposarXiv:2504.02160
UNO, UNO-1M, UNO
Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data
arXiv3 reposarXiv:2504.02268
MeanCache, langcache-embed-v1, langcache-embed-v2
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
arXiv3 reposarXiv:2504.07164
MiniMax-M2-BF16, MiniMax-M2, MiniMax-M2
Kimi-VL Technical Report
arXiv3 reposarXiv:2504.07491
LocateAnything-3B, MoonViT-SO-400M, MMLongBench-Doc
$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
arXiv3 reposarXiv:2504.16054
physical-ai-studio, pi05_base, open-value
ReasonIR: Training Retrievers for Reasoning Tasks
arXiv3 reposarXiv:2504.20595
ReasonIR, ReasonIR-8B, reasonir-data
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
arXiv3 reposarXiv:2504.20690
ICEdit, normal-lora, ICEdit-MoE-LoRA
Llama-Nemotron: Efficient Reasoning Models
arXiv3 reposarXiv:2505.00949
Llama-Nemotron-Post-Training-Dataset, Llama-3_3-Nemotron-Super-49B-v1_5, Llama-3_3-Nemotron-Super-49B-GenRM
Benchmarking LLMs' Swarm intelligence
arXiv3 reposarXiv:2505.04364
YuLan-SwarmIntell, swarmbench, swarmbench
Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
arXiv3 reposarXiv:2505.07747
Step1X-3D, Step1X-3D, Step1X-3D-obj-data
Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
arXiv3 reposarXiv:2505.08699
granite-speech-3.3-8b, granite-speech-models, granite-speech-studio
SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
arXiv3 reposarXiv:2505.10793
TuneJury, SongEval, SongEval
Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining
arXiv3 reposarXiv:2505.11293
B3, B3_Qwen2_7B, B3_Qwen2_2B
MMaDA: Multimodal Large Diffusion Language Models
arXiv3 reposarXiv:2505.15809
MMaDA-8B-Base, MMaDA-8B-MixCoT, MMaDA
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
arXiv3 reposarXiv:2505.16915
data-juicer, SciDataOS, data-juicer
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
arXiv3 reposarXiv:2505.18875
nunchaku, nunchaku, Sparse-VideoGen
arXiv:2505.19706
arXiv3 reposarXiv:2505.19706
PathFinder-PRM, PathFinder-600K, PathFinder-PRM-7B
FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement
arXiv3 reposarXiv:2505.20192
FunReason, AWorld, AWorld-RL
ImgEdit: A Unified Image Editing Dataset and Benchmark
arXiv3 reposarXiv:2505.20275
ImgEdit, ImgEdit, ImgEdit_recap_mask
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
arXiv3 reposarXiv:2505.20325
Guided-by-Gut, DS-Qwen-7b-GG-CalibratedConfRL, DS-Qwen-1.5b-GG-CalibratedConfRL
arXiv:2505.20979
arXiv3 reposarXiv:2505.20979
MelodySim, melodySim, MelodySim
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
arXiv3 reposarXiv:2505.21960
Loopfree, loopfree-sd1.5, loopfree-sd2.1-base
AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
arXiv3 reposarXiv:2505.23716
anysplat, AnySplat, AnySplat
TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context Learning
arXiv3 reposarXiv:2505.23719
TiRex, tirex, TiRex-1.1-gifteval
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
arXiv3 reposarXiv:2505.24111
diarizen-wavlm-large-s80-md, diarizen-wavlm-base-s80-md, diarizen-wavlm-large-s80-md-v2
QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
arXiv3 reposarXiv:2506.00711
QoQ-Med-VL-32B, QoQ-Med-VL-7B, QoQ_Med
arXiv:2506.00830
arXiv3 reposarXiv:2506.00830
SkyReels-V2, SkyReels-A1, SkyReels-A2
Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
arXiv3 reposarXiv:2506.02408
CPathPatchFeature, E2E-WSI-ABMILX, CPath_Image
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
arXiv3 reposarXiv:2506.03150
IllumiCraft, Illumicraft-checkpoints, IllumiCraft
FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes
arXiv3 reposarXiv:2506.03278
FailureSensorIQ, AssetOpsBench, FailureSensorIQ
PixCell: A generative foundation model for digital histopathology images
arXiv3 reposarXiv:2506.05127
PixCell-1024, PixCell-sample-data, PixCell-256
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
arXiv3 reposarXiv:2506.07044
Lingshu-32B, Lingshu-7B, ReasonMed
SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
arXiv3 reposarXiv:2506.07966
SpaceQwen2.5-VL-3B-Instruct, SpaceOm, SpaceThinker-Qwen2.5VL-3B
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
arXiv3 reposarXiv:2506.09827
Voice-Acting-Pipeline, kani-tts-2-en, kani-tts-2-pt
Efficient Part-level 3D Object Generation via Dual Volume Packing
arXiv3 reposarXiv:2506.09980
PartPacker, PartPacker, PartPacker
3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks
arXiv3 reposarXiv:2506.11147
3D-RAD, 3D-RAD, M3D-RAD
SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
arXiv3 reposarXiv:2506.15154
SonicVerse, SonicVerse, SonicVerse
VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
arXiv3 reposarXiv:2506.17561
VLA-OS, VLA-OS, VLA-OS-Dataset
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
arXiv3 reposarXiv:2506.18623
diarizen-wavlm-large-s80-md, diarizen-wavlm-base-s80-md, diarizen-wavlm-large-s80-md-v2
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
arXiv3 reposarXiv:2506.18898
Tar-7B-v0.1, TA-Tok, Tar-1.5B
WorldVLA: Towards Autoregressive Action World Model
arXiv3 reposarXiv:2506.21539
WorldVLA, RynnVLA-001, RynnEC
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
arXiv3 reposarXiv:2507.01006
GLM-V, GLM-4.1V-9B-Thinking, MMLongBench-Doc
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
arXiv3 reposarXiv:2507.02768
DeSTA2, DeSTA2.5-Audio, DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct
Agentic-R1: Distilled Dual-Strategy Reasoning
arXiv3 reposarXiv:2507.05707
DualDistill, Agentic-R1-SD, Agentic-R1
Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
arXiv3 reposarXiv:2507.17801
Lumina-mGPT-2.0, Lumina-mGPT-2.0-Omni, Lumina-mGPT-2.0
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
arXiv3 reposarXiv:2507.18446
WhisperLiveKit, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1
DIFFA: Large Language Diffusion Models Can Listen and Understand
arXiv3 reposarXiv:2507.18452
DIFFA, DIFFA, DIFFA
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
arXiv3 reposarXiv:2507.20880
JAME, jamify, JAM-0.5
Music Arena: Live Evaluation for Text-to-Music
arXiv3 reposarXiv:2507.20900
music-arena, music-arena-dataset, TuneJury
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
arXiv3 reposarXiv:2507.21509
cli, FabricationGuard-linearprobe-qwen36-27b, ReasoningGuard-linearprobe-qwen36-27b
Trade-offs in Image Generation: How Do Different Dimensions Interact?
arXiv3 reposarXiv:2507.22100
TRIG, TRIG, TRIG
SDMatte: Grafting Diffusion Models for Interactive Matting
arXiv3 reposarXiv:2508.00443
SDMatte, SDMatte, LiteSDMatte
Marco-Voice Technical Report
arXiv3 reposarXiv:2508.02038
Marco-Voice, CSEMOTIONS, Marco-Voice
Kronos: A Foundation Model for the Language of Financial Markets
arXiv3 reposarXiv:2508.02739
Kronos, Kronos-V12, stock_forecasting
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
arXiv3 reposarXiv:2508.03542
speech2latex, Speech2Latex, GemmaApollo
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
arXiv3 reposarXiv:2508.07981
Omni-Effects, Omni-Effects, Omni-VFX
arXiv:2508.08098
arXiv3 reposarXiv:2508.08098
TBAC-UniImage, TBAC-UniImage-3B, TBAC-UniImage
TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation
arXiv3 reposarXiv:2508.08680
topxgen-llama-4-scout-and-llama-4-scout, topxgen-gemma-3-27b-and-nllb-3.3b, topxgen
V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking
arXiv3 reposarXiv:2508.13634
AWorld, AWorld-RL, V2P-7B
LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation
arXiv3 reposarXiv:2508.17858
LexSemBridge, LexSemBridge_eval, LexSemBridge_CLR_snowflake
Baichuan-M2: Scaling Medical Capability with Large Verifier System
arXiv3 reposarXiv:2509.02208
Baichuan-M2-32B, Baichuan-M2-32B, Baichuan-M2-32B-GPTQ-Int4
Sample-efficient Integration of New Modalities into Large Language Models
arXiv3 reposarXiv:2509.04606
sample-efficient-multimodality, capdels, sample-efficient-multimodality-ckpts
P3-SAM: Native 3D Part Segmentation
arXiv3 reposarXiv:2509.06784
HY3D-Bench, Hunyuan3D-Part, Hunyuan3D-Part
Continuous Audio Language Models
arXiv3 reposarXiv:2509.06926
pocket-tts, pocket-tts-korean-300m, pocket-tts-ungated
X-Part: high fidelity and structure coherent shape decomposition
arXiv3 reposarXiv:2509.08643
HY3D-Bench, Hunyuan3D-Part, Hunyuan3D-Part
Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
arXiv3 reposarXiv:2509.08753
tts-1.6b-en_fr, tts_longeval, delayed-streams-modeling
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
arXiv3 reposarXiv:2509.13031
PeBR-R1, PeBR_R1, PeBR_R1_dataset
FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning
arXiv3 reposarXiv:2509.13160
MiniMax-M2-BF16, MiniMax-M2, MiniMax-M2
SPATIALGEN: Layout-guided 3D Indoor Scene Generation
arXiv3 reposarXiv:2509.14981
FLUX.1-Wireframe-dev-lora, SpatialGen-Testset, SpatialGen-1.0
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
arXiv3 reposarXiv:2509.15212
WorldVLA, RynnVLA-002, RynnVLA-002
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
arXiv3 reposarXiv:2509.17664
SD-VLM-7B, SD-VLM, MSMU
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
arXiv3 reposarXiv:2509.23909
EditScore-Reward-Data, EditReward-Bench, EditScore-RL-Data
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
arXiv3 reposarXiv:2509.24650
VoxCPM, VoxCPM2, VoxCPM1.5
SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation
arXiv3 reposarXiv:2509.24980
SDPose-Body, SDPose-Wholebody, SDPose-OOD
arXiv:2509.26346
arXiv3 reposarXiv:2509.26346
EditReward, EditReward-Data, EditReward-Bench
MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
arXiv3 reposarXiv:2509.26642
MLA, MLA_RLBench_post, MLA_pretrain
AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
arXiv3 reposarXiv:2510.01268
DetectLLMSegmentation, L2D, AdaDetectGPT
Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
arXiv3 reposarXiv:2510.06961
esb-datasets-test-only-sorted, open-asr-leaderboard, asr-leaderboard-longform
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
arXiv3 reposarXiv:2510.08531
SpatialLadder-3B, SpatialLadder-26k, SPBench
ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
arXiv3 reposarXiv:2510.11000
IMIG-100K, ContextGen, IMIG-Source
SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
arXiv3 reposarXiv:2510.12784
SRUM_BAGEL_7B_MoT, SRUM, SRUM_6k_CompBench_Train
UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
arXiv3 reposarXiv:2510.13344
Uni-MoE, UniMoE-Audio-preview, UMOE-Scaling-Unified-Multimodal-LLMs
BADAS: Context Aware Collision Prediction Using Real-World Dashcam Data
arXiv3 reposarXiv:2510.14876
Cosmos-Sentinel, Cosmos_Sentinel, Nvidia-Cosmos-Cookoff
Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
arXiv3 reposarXiv:2510.15742
Ditto, Ditto-1M, Ditto_models
Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery
arXiv3 reposarXiv:2510.15869
Skyfall-GS-datasets, Skyfall-GS-eval, Skyfall-GS-ply
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
arXiv3 reposarXiv:2510.16872
DeepAnalyze, DataScience-Instruct-500K, DeepAnalyze-8B
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
arXiv3 reposarXiv:2510.17801
RoboBench, RoboBench-Results, RoboBench
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
arXiv3 reposarXiv:2510.18701
UniGenBench-EvalModel-qwen3vl-32b-v1, UniGenBench-EvalModel-qwen-72b-v1, UniGenBench-Eval-Images
Video-As-Prompt: Unified Semantic Control for Video Generation
arXiv3 reposarXiv:2510.20888
Video-As-Prompt, Video-As-Prompt-Wan2.1-14B, Video-As-Prompt-CogVideoX-5B
DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
arXiv3 reposarXiv:2510.22950
diffrhythm2, DiffRhythm2, DiffRhythm2
Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
arXiv3 reposarXiv:2510.25616
BlindVLA, openvla-7b-warmup-checkpoint_lora_002000, openvla_1k-dataset
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
arXiv3 reposarXiv:2511.01588
PDF-VLM2Vec, PDF-VLM2Vec-Qwen2VL-2B, PDF-VLM2Vec-Qwen2VL-7B
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
arXiv3 reposarXiv:2511.01670
SeaLLMs-Audio, SeaBench-Audio, SeaLLMs-Audio-7B
Step-Audio-EditX Technical Report
arXiv3 reposarXiv:2511.03601
Step-Audio-EditX, Step-Audio-EditX, Step-Audio-EditX-AWQ-4bit
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation
arXiv3 reposarXiv:2511.11434
weave, Weave, Bagel-weave
GroupRank: A Groupwise Paradigm for Effective and Efficient Passage Reranking with LLMs
arXiv3 reposarXiv:2511.11653
Diver, Diver-GroupRank-7B, Diver-GroupRank-32B
Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
arXiv3 reposarXiv:2511.12609
zen5, Uni-MoE, UMOE-Scaling-Unified-Multimodal-LLMs
RynnVLA-002: A Unified Vision-Language-Action and World Model
arXiv3 reposarXiv:2511.17502
WorldVLA, RynnVLA-001, RynnVLA-002
Fara-7B: An Efficient Agentic Model for Computer Use
arXiv3 reposarXiv:2511.19663
WebTailBench, Fara-7b, Fara_Test
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
arXiv3 reposarXiv:2511.20785
LongVT, LongVT-Parquet, LongVT-Source
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
arXiv3 reposarXiv:2511.21688
G2VLM-2B-MoT, G2VLM, g2vlm
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
arXiv3 reposarXiv:2511.21689
ToolOrchestra, Orchestrator-8B, ToolScale
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
arXiv3 reposarXiv:2512.02556
Hy4-preview, Hy4-preview, maxtext
Light-X: Generative 4D Video Rendering with Camera and Illumination Control
arXiv3 reposarXiv:2512.05115
Light-X-Uni, Light-X, Light-Syn
Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model
arXiv3 reposarXiv:2512.06999
QwenFeat-Vocal-Score, QwenFeat-Vocal-Score, Singing-Aesthetic-Assessment
Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs
arXiv3 reposarXiv:2512.09874
pdf-parse-bench, wikipedia-latex-formulas-319k, pdf-parse-bench
EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
arXiv3 reposarXiv:2512.11715
EditMGT, EditMGT, CrispEdit-2M
arXiv:2512.11831
arXiv3 reposarXiv:2512.11831
ExplicitShortCut, ESC-XL2, ESC-B2
Image Diffusion Preview with Consistency Solver
arXiv3 reposarXiv:2512.13592
consolver, consolver, EditReward
Bolmo: Byteifying the Next Generation of Language Models
arXiv3 reposarXiv:2512.15586
Bolmo-7B, Bolmo-1B, bolmo_mix
Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition
arXiv3 reposarXiv:2512.15603
Qwen-Image, Qwen-Image-Layered, Qwen-Image-Layered
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
arXiv3 reposarXiv:2512.16649
UltraData-RL-2609, MiniCPM5-1B-GGUF, MiniCPM5-1B
SAM Audio: Segment Anything in Audio
arXiv3 reposarXiv:2512.18099
sam-audio, sam-audio, Sam-Audio
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
arXiv3 reposarXiv:2512.24551
Open-PhyGDPO, PhyGDPO, PhyGDPO
From Failure to Mastery: Generating Hard Samples for Tool-use Agents
arXiv3 reposarXiv:2601.01498
AWorld, FunReason-MT, AWorld-RL
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
arXiv3 reposarXiv:2601.02456
InternVLA-A1-3B-RoboTwin, InternVLA-A1-3B, internvla
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
arXiv3 reposarXiv:2601.09536
Omni-Bench, Omni-R1-Zero, Omni-R1
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
arXiv3 reposarXiv:2601.10387
karma-electric-project, karma-electric-llama31-8b, drowse
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models
arXiv3 reposarXiv:2601.11329
f-actor-behavior-sd-nanocodec, f-actor, f-actor-behavior-sd-mimi
TeleStyle: Content-Preserving Style Transfer in Images and Videos
arXiv3 reposarXiv:2601.20175
TeleStyle, TeleStyleV2, TeleStyleV2-Qwen-Image-Edit-2511-bf16
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
arXiv3 reposarXiv:2602.00443
RVCBench, RVCBench, RVCBench
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
arXiv3 reposarXiv:2602.03328
GuardReasoner-Omni, GuardReasoner-Omni-7B, GuardReasoner-Omni-3B
LLaDA2.1: Speeding Up Text Diffusion via Token Editing
arXiv3 reposarXiv:2602.08676
LLaDA2.1-mini, LLaDA2.1-flash, dllm
StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors
arXiv3 reposarXiv:2602.08934
StealthRL, StealthRL-Benchmark, StealthRL
RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation
arXiv3 reposarXiv:2602.09973
RoboInter-VLM_llavaov_7B, RoboInter-VLM, RoboInter-VLM_qwenvl25_3b
ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
arXiv3 reposarXiv:2602.10113
ConsID-Gen, ConsID-Gen, ConsIDVid
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
arXiv3 reposarXiv:2602.20161
Mobile-O-SFT, Mobile-O-Pre-Train, Mobile-O-Post-Train
DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation
arXiv3 reposarXiv:2602.22839
PPTAgent, DeepPresenter-9B-GGUF, DeepPresenter-9B
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
arXiv3 reposarXiv:2602.23068
tada-1b, tada-codec, tada-3b-ml
Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design
arXiv3 reposarXiv:2603.00152
Dr-Seg, Dr_Seg, coconut
Fish Audio S2 Technical Report
arXiv3 reposarXiv:2603.08823
s2-pro, fish-speech, RVCBench
$Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation
arXiv3 reposarXiv:2603.12263
Psi0, psi0, Psi0
Overcoming the Modality Gap in Context-Aided Forecasting
arXiv3 reposarXiv:2603.12451
DoubleCast, CAF_7M, DoubleCast
Multimodal OCR: Parse Anything from Documents
arXiv3 reposarXiv:2603.13032
dots.ocr, dots.mocr, dots.mocr-svg
Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
arXiv3 reposarXiv:2603.21426
beta-kd, Cosine-Beta-KD-Instance, Cosine-Beta-KD-Task
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
arXiv3 reposarXiv:2603.26511
AMALIA-9B-0626-DPO, amalia-lm-eval, AMALIA
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
arXiv3 reposarXiv:2603.27064
ChartNet, granite-vision-4.1-4b, granite-4.0-3b-vision
An Empirical Recipe for Universal Phone Recognition
arXiv3 reposarXiv:2603.29042
PhoneticXeus, PhoneticXeus, PhoneticXeus
Memory Intelligence Agent
arXiv3 reposarXiv:2604.04503
MIA, MIA, MIA
FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
arXiv3 reposarXiv:2604.06757
FlowInOne, VisPrompt5M, VPBench
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
arXiv3 reposarXiv:2604.10905
audio-flamingo-next-hf, audio-flamingo-next-captioner-hf, audio-flamingo-next-think-hf
Introspective Diffusion Language Models
arXiv3 reposarXiv:2604.11035
I-DLM-8B, I-DLM-32B, I-DLM-8B-lora-r128
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
arXiv3 reposarXiv:2604.12928
moshi-rag, moshika-rag-pytorch-bf16, moshika-rag-candle-bf16
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
arXiv3 reposarXiv:2604.18518
UDM-GRPO, URSA-1.7B-IBQ512-UDMGRPO-GenEval, URSA-1.7B-IBQ512-UDMGRPO-PickScore
Vista4D: Video Reshooting with 4D Point Clouds
arXiv3 reposarXiv:2604.21915
Vista4D, Vista4D, Vista4D-Eval-Data
GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction
arXiv3 reposarXiv:2604.23941
GoClick, GoClick-Large, GoClick-Base
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence
arXiv3 reposarXiv:2604.24954
Eagle, Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8, Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
arXiv3 reposarXiv:2604.28123
PRISM, gemini_distill, gemini_public_mmr1
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
arXiv3 reposarXiv:2605.07982
gliguard-LLMGuardrails-300M, GLiNER2-Guardrails-PII-Multi, GLiNER2-Guardrails-PII-Multi-onnx
GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction
arXiv3 reposarXiv:2605.09973
gliner2-privacy-filter-PII-multi, GLiNER2-Guardrails-PII-Multi, GLiNER2-Guardrails-PII-Multi-onnx
Allegory of the Cave: Measurement-Grounded Vision-Language Learning
arXiv3 reposarXiv:2605.11727
PRISM-VL, PRSIMVL-LoRA-V1, MeasL-150K-V1
Asymmetric Flow Models
arXiv3 reposarXiv:2605.12964
AsymFLUX.2-klein-9B, LakonLab, AsymFLUX.2-klein
Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
arXiv3 reposarXiv:2605.14984
Sat3DGen, Sat3DGen, VIGOR_SAT3DGEN_add_skymask_DSM_satdepth
ReactiveGWM: Steering NPC in Reactive Game World Models
arXiv3 reposarXiv:2605.15256
ReactiveGWM-Models, ReactiveGWM-Datasets, stable-retro
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
arXiv3 reposarXiv:2605.19075
CRAFT, CRAFT-MAGMaR, CRAFT-WikiVideo
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning
arXiv3 reposarXiv:2605.20176
ClinSeekAgent, ClinSeek-35B-A3B, ClinSeek-Bench
GEM: Generative Supervision Helps Embodied Intelligence
arXiv3 reposarXiv:2605.28548
GEM, GEM-250K, GEM-2B
dots.tts Technical Report
arXiv3 reposarXiv:2606.07080
dots.tts-mf-2steps, dots.tts-mf-1step, dots.tts-mf-2steps-stts
Kwai Keye-VL-2.0 Technical Report
arXiv3 reposarXiv:2606.10651
Keye, Keye-VL-2.0-30B-A3B, keye-39e1f0b5
i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
arXiv3 reposarXiv:2606.11289
i1-3B, i1, i1-captions
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
arXiv3 reposarXiv:2606.12195
InternVideo, InternVideo3-8B-Instruct, InternVideo3_Dataset
Modality Forcing for Scalable Spatial Generation
arXiv3 reposarXiv:2606.13676
modality-forcing, modality_forcing, modality_forcing
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
arXiv3 reposarXiv:2606.17006
TuneJury, tunejury, release-scores
Fara-1.5: Scalable Learning Environments for Computer Use Agents
arXiv3 reposarXiv:2606.20785
fara, Fara1.5-4B, Fara1.5-9B
SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection
arXiv3 reposarXiv:2606.21138
SEED, SEED, GenText-Forensics-3rd-Place
Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents
arXiv3 reposarXiv:2607.00895
lettucedect-v2-mmbert-base, lettucedetect-code-hallucination, lettucedect-v2-qwen-2b
Representation Distribution Matching for One-Step Visual Generation
arXiv3 reposarXiv:2607.02375
RDM, flux2-klein-1step-demo, flux2-klein-1step-rdm
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation
arXiv3 reposarXiv:2607.02642
giga-world-1, Giga-World-1, Giga-World-1-Toydata
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
arXiv3 reposarXiv:2607.05147
SpecForge, kimi-k3-dspark, DeepSpec
FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis
arXiv3 reposarXiv:2607.09530
FreyaTTS, freya-tr-eval, freya-tts
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
arXiv3 reposarXiv:2607.14952
Megatron-Bridge, Megatron-Bridge, Megatron-Bridge
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
arXiv3 reposarXiv:2607.18213
swe-pruner-pro-mimo-v2-flash-head, swe-pruner-pro-training-corpus, swe-pruner-pro-qwen3-coder-next-head
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
arXiv3 reposarXiv:2607.19064
Mage, mage-7255fc67, Mage-VL-RTSP
arXiv:2607.29679
arXiv3 reposarXiv:2607.29679
GenAI-Caption-Pipeline, SP-PE-Qwen3.5-35B-A3B, Qwen-Image-SP
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
arXiv3 reposarXiv:2608.15767
tinycast, tinycast, tinycast-forecaster
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
arXiv3 reposarXiv:2608.15875
giga-brain-0, GigaBrain-0.7-SampleData, GigaBrain-0.7-3.5B-Base
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
arXiv3 reposarXiv:2608.20958
TLive-Omni, TLive-Omni-9B, TLive-Omni-4B
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
arXiv3 reposarXiv:2608.24053
WeMM-Embedding-9B, WeMM-Embedding-4B, WeMM-Embedding-2B
Detecting hallucinations in large language models using semantic entropy
Nature3 reposNature:s41586-024-07421-0
VASE, uqlm, Spnda
A vision–language foundation model for precision oncology
Nature3 reposNature:s41586-024-08378-w
PIANO, AtlasPatch, PathPT
Tanks and temples
ACM2 reposACM:3072959.3073599
awesome-mvs, InstantSplat
Microsoft recommenders
ACM2 reposACM:3298689.3346967
recommenders, Recommenders
Raha
ACM2 reposACM:3299869.3324956
Jellyfish-13B, raha
Microsoft Recommenders: Best Practices for Production-Ready Recommendation Systems
ACM2 reposACM:3366424.3382692
recommenders, Recommenders
WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning
ACM2 reposACM:3404835.3463257
test-big-dataset, wit
Automating QUIC Interoperability Testing
ACM2 reposACM:3405796.3405826
quic-interop-runner, quic-interop-runner
SyRust: automatic testing of Rust libraries with semantic-aware program synthesis
ACM2 reposACM:3453483.3454084
miri, rust
ZeRO-infinity
ACM2 reposACM:3458817.3476205
DeepSpeed, DeepSpeed
Hammer
ACM2 reposACM:3489517.3530672
chipyard, hammer
Retrieval-Based Gradient Boosting Decision Trees for Disease Risk Assessment
ACM2 reposACM:3534678.3539052
awesome-decision-tree-papers, awesome-gradient-boosting-papers
A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition Scheduling
ACM2 reposACM:3534678.3539081
awesome-fsrs, fsrs4anki
WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics
ACM2 reposACM:3544548.3581158
webui-all, webui
A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
ACM2 reposACM:3577193.3593704
DeepSpeed, DeepSpeed
Demo: ARFlow: A Framework for Simplifying AR Experimentation Workflow
ACM2 reposACM:3638550.3643617
rerun, Dalaran
Verified Extraction from Coq to OCaml
ACM2 reposACM:3656379
MetaCoq, metarocq
System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
ACM2 reposACM:3662158.3662806
DeepSpeed, DeepSpeed
Crabtree: Rust API Test Synthesis Guided by Coverage and Type
ACM2 reposACM:3689733
rust, miri
Rustlantis: Randomized Differential Testing of the Rust Compiler
ACM2 reposACM:3689780
rust, miri
Correct and Complete Type Checking and Certified Erasure for Coq , in Coq
ACM2 reposACM:3706056
MetaCoq, metarocq
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
ACM2 reposACM:3760250.3762217
DeepSpeed, DeepSpeed
Tactics for Reasoning modulo AC in Coq
arXiv2 reposarXiv:1106.4448
aac-tactics, aac-tactics
BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data
arXiv2 reposarXiv:1203.5485
awesome-bigdata, awesome-bigdata
arXiv:1209.2137
arXiv2 reposarXiv:1209.2137
JavaFastPFOR, FastPFor
Efficient Estimation of Word Representations in Vector Space
arXiv2 reposarXiv:1301.3781
Word-Embeddings-Repository-for-Turkish, ReAGent
arXiv:1401.6399
arXiv2 reposarXiv:1401.6399
JavaFastPFOR, FastPFor
A Fast, Minimal Memory, Consistent Hash Algorithm
arXiv2 reposarXiv:1406.2294
hash4j, grenier
Generative Adversarial Networks
arXiv2 reposarXiv:1406.2661
ocaml-torch, Data-Science
Neural Machine Translation by Jointly Learning to Align and Translate
arXiv2 reposarXiv:1409.0473
nl2bash, nmt
Leveraging Cloud Data to Mitigate User Experience from "Breaking Bad"
arXiv2 reposarXiv:1411.7955
breakout, breakout-ruby
arXiv:1502.01916
arXiv2 reposarXiv:1502.01916
FastPFor, JavaFastPFOR
Distilling the Knowledge in a Neural Network
arXiv2 reposarXiv:1503.02531
hasktorch, distilgpt2
arXiv:1503.07387
arXiv2 reposarXiv:1503.07387
JavaFastPFOR, FastPFor
DeepFont: Identify Your Font from An Image
arXiv2 reposarXiv:1507.03196
YuzuMarker.FontDetection, YuzuMarker.FontDetection
An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
arXiv2 reposarXiv:1507.05717
yas, EasyOCR
Unsupervised Deep Embedding for Clustering Analysis
arXiv2 reposarXiv:1511.06335
awesome-datascience, awesome-datascience
Perceptual Losses for Real-Time Style Transfer and Super-Resolution
arXiv2 reposarXiv:1603.08155
convnext_perceptual_loss, SkinDeep
Neural Language Correction with Character-Based Attention
arXiv2 reposarXiv:1603.09727
pycorrector, repo-5146-pycorrector
arXiv:1606.06031
arXiv2 reposarXiv:1606.06031
sdtt, SDTT-LaViDa
Bag of Tricks for Efficient Text Classification
arXiv2 reposarXiv:1607.01759
xlsum, fasttext-language-identification
Enriching Word Vectors with Subword Information
arXiv2 reposarXiv:1607.04606
Word-Embeddings-Repository-for-Turkish, fasttext-language-identification
Pointer Sentinel Mixture Models
arXiv2 reposarXiv:1609.07843
esp32s3-distributed-ai, wikitext
The HoTT Library: A formalization of homotopy type theory in Coq
arXiv2 reposarXiv:1610.04591
HoTT, Coq-HoTT
Axiomatic Attribution for Deep Networks
arXiv2 reposarXiv:1703.01365
shap, shap
Tacotron: Towards End-to-End Speech Synthesis
arXiv2 reposarXiv:1703.10135
Real-Time-Voice-Cloning, MockingBird
Learning Important Features Through Propagating Activation Differences
arXiv2 reposarXiv:1704.02685
shap, shap
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
arXiv2 reposarXiv:1704.05426
roberta-large-mnli, t5-large-encoder-only-bf16
Efficient Natural Language Response Suggestion for Smart Reply
arXiv2 reposarXiv:1705.00652
splade-ecommerce-esci, FinBot
ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases
arXiv2 reposarXiv:1705.02315
srt-cxr14-pooled-probe, diff-mining
SmoothGrad: removing noise by adding noise
arXiv2 reposarXiv:1706.03825
shap, shap
Rethinking Atrous Convolution for Semantic Image Segmentation
arXiv2 reposarXiv:1706.05587
EdgeSeg, spectra
GPU-acceleration for Large-scale Tree Boosting
arXiv2 reposarXiv:1706.08359
LightGBM, LightGBM
Crowdsourcing Multiple Choice Science Questions
arXiv2 reposarXiv:1707.06209
sciq, COMP4222-Course-Project
A Distributional Perspective on Reinforcement Learning
arXiv2 reposarXiv:1707.06887
open-value, cleanrl
S$^3$FD: Single Shot Scale-invariant Face Detector
arXiv2 reposarXiv:1708.05237
S3FD.pytorch, face-alignment
Squeeze-and-Excitation Networks
arXiv2 reposarXiv:1709.01507
Multi-fake-detective, Waifu2x
arXiv:1709.08990
arXiv2 reposarXiv:1709.08990
JavaFastPFOR, FastPFor
Searching for Activation Functions
arXiv2 reposarXiv:1710.05941
pegasus-x-base-synthsumm_open-16k, llama2.zig
Generalized End-to-End Loss for Speaker Verification
arXiv2 reposarXiv:1710.10467
Real-Time-Voice-Cloning, MockingBird
Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
arXiv2 reposarXiv:1712.05884
tacotron2, seq2seq_accent_conversion_model
The NarrativeQA Reading Comprehension Challenge
arXiv2 reposarXiv:1712.07040
GraphKV, GraphKV
DVQA: Understanding Data Visualizations via Question Answering
arXiv2 reposarXiv:1801.08163
DVQA_dataset, PlotQA
Efficient Neural Audio Synthesis
arXiv2 reposarXiv:1802.08435
Real-Time-Voice-Cloning, MockingBird
Shampoo: Preconditioned Stochastic Tensor Optimization
arXiv2 reposarXiv:1802.09568
modded-nanogpt, Emerging-Optimizers
xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems
arXiv2 reposarXiv:1803.05170
recommenders, Recommenders
YOLOv3: An Incremental Improvement
arXiv2 reposarXiv:1804.02767
darknet, darknetcv
MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio source separation
arXiv2 reposarXiv:1805.02410
demucs, demucs
arXiv:1805.08949
arXiv2 reposarXiv:1805.08949
docprompting-conala, conala
Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
arXiv2 reposarXiv:1806.04558
Real-Time-Voice-Cloning, MockingBird
The latest gossip on BFT consensus
arXiv2 reposarXiv:1807.04938
cometbft, tendermint
Learning to Describe Differences Between Pairs of Similar Images
arXiv2 reposarXiv:1808.10584
idefics-80b-instruct, idefics-9b-instruct
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
arXiv2 reposarXiv:1809.08887
llm-jepa, sql-create-context
HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering
arXiv2 reposarXiv:1809.09600
GraphKV, GraphKV
How Powerful are Graph Neural Networks?
arXiv2 reposarXiv:1810.00826
MCBG, Malware_survey_my_experiments
Model Cards for Model Reporting
arXiv2 reposarXiv:1810.03993
moonshine, galactica-120b
WikiHow: A Large Scale Text Summarization Dataset
arXiv2 reposarXiv:1810.09305
all-MiniLM-L6-v2, all-MiniLM-L12-v2
arXiv:1810.12368
arXiv2 reposarXiv:1810.12368
geolm-base-toponym-recognition, Pragmatic-Guide-to-Geoparsing-Evaluation
The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
arXiv2 reposarXiv:1811.00982
SEED-Data-Edit-Part2-3, SEED-Data-Edit
SAFE: Self-Attentive Function Embeddings for Binary Similarity
arXiv2 reposarXiv:1811.05296
MCBG, Malware_survey_my_experiments
Neural Abstractive Text Summarization with Sequence-to-Sequence Models
arXiv2 reposarXiv:1812.02303
pycorrector, repo-5146-pycorrector
TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents
arXiv2 reposarXiv:1901.08149
CDial-GPT_LCCC-large, CDial-GPT_LCCC-base
Parameter-Efficient Transfer Learning for NLP
arXiv2 reposarXiv:1902.00751
Continual-NExT, LoRA
Computing Extremely Accurate Quantiles Using t-Digests
arXiv2 reposarXiv:1902.04023
t-digest-c, t-digest
arXiv:1902.04043
arXiv2 reposarXiv:1902.04043
smac, smacv2
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
arXiv2 reposarXiv:1904.02882
F5-TTS, libritts
Publicly Available Clinical BERT Embeddings
arXiv2 reposarXiv:1904.03323
ehr_deidentification, deid_bert_i2b2
A Repository of Conversational Datasets
arXiv2 reposarXiv:1904.06472
all-MiniLM-L6-v2, all-MiniLM-L12-v2
Improved Precision and Recall Metric for Assessing Generative Models
arXiv2 reposarXiv:1904.06991
materialgan, stylegan2-ada-pytorch
Learning to Prove Theorems via Interacting with Proof Assistants
arXiv2 reposarXiv:1905.09381
coq-serapi, coq-serapi
GLTR: Statistical Detection and Visualization of Generated Text
arXiv2 reposarXiv:1906.04043
L2D, AdaDetectGPT
Pre-Training with Whole Word Masking for Chinese BERT
arXiv2 reposarXiv:1906.08101
chinese-roberta-wwm-ext, chinese-bert-wwm-ext
MediaPipe: A Framework for Building Perception Pipelines
arXiv2 reposarXiv:1906.08172
mediapipe, mediapipe
Generating Correctness Proofs with Neural Networks
arXiv2 reposarXiv:1907.07794
coq-serapi, coq-serapi
LVIS: A Dataset for Large Vocabulary Instance Segmentation
arXiv2 reposarXiv:1908.03195
LocateAnything-3B, lvis-api
TabNet: Attentive Interpretable Tabular Learning
arXiv2 reposarXiv:1908.07442
pytorch-frame, Trompt
PlotQA: Reasoning over Scientific Plots
arXiv2 reposarXiv:1909.00997
VisRAG-Ret-Train-In-domain-data, PlotQA
Fine-Tuning Language Models from Human Preferences
arXiv2 reposarXiv:1909.08593
Starling-LM-7B-beta, trlx
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
arXiv2 reposarXiv:1909.11942
tf-transformers, sahajBERT
Machine Learning in Python: Main developments and technology trends in data science, machine learning, and artificial intelligence
arXiv2 reposarXiv:2002.04803
cuml, cuml
A Simple Framework for Contrastive Learning of Visual Representations
arXiv2 reposarXiv:2002.05709
vilmedic, BiomedVLP-CXR-BERT-specialized
UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation
arXiv2 reposarXiv:2002.06353
UniVL-video-captioning, video-captioning
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
arXiv2 reposarXiv:2002.10957
torchserve-all-minilm-l6-v2, Multilingual-MiniLM-L12-H384
TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages
arXiv2 reposarXiv:2003.05002
squad_bn, tydiqa-primary-task-xlm-roberta-large
PathVQA: 30000+ Questions for Medical Visual Question Answering
arXiv2 reposarXiv:2003.10286
path-vqa, LLaDA-MedV
KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding
arXiv2 reposarXiv:2004.03289
kor_nli, kor_nlu
Deep Learning Models for Multilingual Hate Speech Detection
arXiv2 reposarXiv:2004.06465
dehatebert-mono-english, DE-LIMIT
Deep Generation of Coq Lemma Names Using Elaborated Terms
arXiv2 reposarXiv:2004.07761
coq-serapi, coq-serapi
Multi-Dimensional Gender Bias Classification
arXiv2 reposarXiv:2005.00614
STAIR-Captions, rebel-dataset
Conformer: Convolution-augmented Transformer for Speech Recognition
arXiv2 reposarXiv:2005.08100
GigaAM, stt_eu_conformer_ctc_large
Visual Transformers: Token-based Image Representation and Processing for Computer Vision
arXiv2 reposarXiv:2006.03677
vit-base-patch16-224-in21k, vit-large-patch16-224-in21k
Unsupervised Cross-lingual Representation Learning for Speech Recognition
arXiv2 reposarXiv:2006.13979
fairseq, wav2vec2-large-xlsr-53
Learning to Format Coq Code Using Language Models
arXiv2 reposarXiv:2006.16743
coq-serapi, coq-serapi
Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
arXiv2 reposarXiv:2007.00808
pyterrier_dr, pyterrier_dr_jpq
Big Bird: Transformers for Longer Sequences
arXiv2 reposarXiv:2007.14062
final-project-level3-nlp-02, Block-Sparse-Attention
Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
arXiv2 reposarXiv:2008.00401
EasyNMT, mbart-large-50-many-to-many-mmt
TransNet V2: An effective deep network architecture for fast shot transition detection
arXiv2 reposarXiv:2008.04838
shotplan, TransNetV2
KILT: a Benchmark for Knowledge Intensive Language Tasks
arXiv2 reposarXiv:2009.02252
instructor-embedding, instructor-embedding
Unconstrained Text Detection in Manga: a New Dataset and Baseline
arXiv2 reposarXiv:2009.04042
YuzuMarker.FontDetection, YuzuMarker.FontDetection
Large-Scale Intelligent Microservices
arXiv2 reposarXiv:2009.08044
awesome-spark, SynapseML
Autoregressive Entity Retrieval
arXiv2 reposarXiv:2010.00904
OKEAN, trusted_ke
D3Net: Densely connected multidilated DenseNet for music source separation
arXiv2 reposarXiv:2010.01733
demucs, demucs
Denoising Diffusion Implicit Models
arXiv2 reposarXiv:2010.02502
DDPM_vs_DDIM, smalldiffusion
Deformable DETR: Deformable Transformers for End-to-End Object Detection
arXiv2 reposarXiv:2010.04159
rf-detr, DINO
Distilling Dense Representations for Ranking using Tightly-Coupled Teachers
arXiv2 reposarXiv:2010.11386
pyterrier_dr, pyterrier_dr_jpq
GPUTreeShap: Massively Parallel Exact Calculation of SHAP Scores for Tree Ensembles
arXiv2 reposarXiv:2010.13972
shap, shap
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
arXiv2 reposarXiv:2010.15980
Magic_Words, imodelsX
Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps
arXiv2 reposarXiv:2011.01060
GraphKV, GraphKV
Multimodal Pretraining for Dense Video Captioning
arXiv2 reposarXiv:2011.11760
TimeChat-Online-139K, Video-Timeline-Tags-ViTT
Score-Based Generative Modeling through Stochastic Differential Equations
arXiv2 reposarXiv:2011.13456
DragonDiffusion, DDPM_vs_DDIM
The Third DIHARD Diarization Challenge
arXiv2 reposarXiv:2012.01477
pyannote-audio, hf-speaker-diarization-3.1
TabTransformer: Tabular Data Modeling Using Contextual Embeddings
arXiv2 reposarXiv:2012.06678
pytorch-frame, Trompt
Extracting Smart Contracts Tested and Verified in Coq
arXiv2 reposarXiv:2012.09138
ConCert, coq-rust-extraction
DeepHateExplainer: Explainable Hate Speech Detection in Under-resourced Bengali Language
arXiv2 reposarXiv:2012.14353
bangla-bert-base, bangla-bert
Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup
arXiv2 reposarXiv:2101.06983
pylate, langcache-embed-v2
arXiv:2102.01454
arXiv2 reposarXiv:2102.01454
sdtt, SDTT-LaViDa
PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them
arXiv2 reposarXiv:2102.07033
all-MiniLM-L6-v2, all-MiniLM-L12-v2
Zero-Shot Text-to-Image Generation
arXiv2 reposarXiv:2102.12092
DALL-E, vilmedic
Roosterize: Suggesting Lemma Names for Coq Verification Projects Using Deep Learning
arXiv2 reposarXiv:2103.01346
coq-serapi, coq-serapi
Perceiver: General Perception with Iterative Attention
arXiv2 reposarXiv:2103.03206
idefics2-8b, PathBench-MIL
Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training
arXiv2 reposarXiv:2104.01027
fairseq, wav2vec2-large-robust
Aggregated Contextual Transformations for High-Resolution Image Inpainting
arXiv2 reposarXiv:2104.01431
comic-translate, UnComicTranslate
AST: Audio Spectrogram Transformer
arXiv2 reposarXiv:2104.01778
ast-finetuned-audioset-10-10-0.4593, ast
FUDGE: Controlled Text Generation With Future Discriminators
arXiv2 reposarXiv:2104.05218
constrDecoding, naacl-2021-fudge-controlled-generation
LocalViT: Analyzing Locality in Vision Transformers
arXiv2 reposarXiv:2104.05707
TexTok-DiT, transformer_latent_diffusion
Emotion Classification in a Resource Constrained Language Using Transformer-based Approach
arXiv2 reposarXiv:2104.08613
bangla-bert-base, bangla-bert
GooAQ: Open Question Answering with Diverse Answer Types
arXiv2 reposarXiv:2104.08727
all-MiniLM-L6-v2, all-MiniLM-L12-v2
GermanQuAD and GermanDPR: Improving Non-English Question Answering and Passage Retrieval
arXiv2 reposarXiv:2104.12741
germandpr, germanquad
MLP-Mixer: An all-MLP Architecture for Vision
arXiv2 reposarXiv:2105.01601
vision_transformer, Swin-Transformer
A Large-Scale Benchmark for Food Image Segmentation
arXiv2 reposarXiv:2105.05409
FoodSeg103, FoodSeg103-Benchmark-v1
Measuring Coding Challenge Competence With APPS
arXiv2 reposarXiv:2105.09938
apps, apps
CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks
arXiv2 reposarXiv:2105.12655
Project_CodeNet, code_contests
CTSpine1K: A Large-Scale Dataset for Spinal Vertebrae Segmentation in Computed Tomography
arXiv2 reposarXiv:2105.14711
CTSpine1K, CTSpine1K
The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
arXiv2 reposarXiv:2106.03193
fleurs, flores_101
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
arXiv2 reposarXiv:2106.06909
unispeech-sat-large, wavlm-large
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
arXiv2 reposarXiv:2106.07447
tmh, hubert-large-ls960-ft
Revisiting Deep Learning Models for Tabular Data
arXiv2 reposarXiv:2106.11959
pytorch-frame, Trompt
Variational Diffusion Models
arXiv2 reposarXiv:2107.00630
TexTok-DiT, transformer_latent_diffusion
ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation
arXiv2 reposarXiv:2107.02137
ernie-3.0-base-zh, ernie-3.0-nano-zh
A Review of Bangla Natural Language Processing Tasks and the Utility of Transformer Models
arXiv2 reposarXiv:2107.03844
bangla-bert-base, bangla-bert
SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking
arXiv2 reposarXiv:2107.05720
bge-m3, splade
Deduplicating Training Data Makes Language Models Better
arXiv2 reposarXiv:2107.06499
falcon-refinedweb, KoCommercial-Dataset
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries
arXiv2 reposarXiv:2107.09609
TimeChat-Online-139K, moment_detr
Extracting functional programs from Coq, in Coq
arXiv2 reposarXiv:2108.02995
ConCert, coq-rust-extraction
MMChat: Multi-Modal Chat Dataset on Social Media
arXiv2 reposarXiv:2108.07154
MMChat, mmchat
Program Synthesis with Large Language Models
arXiv2 reposarXiv:2108.07732
FTTT, mbpp
Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models
arXiv2 reposarXiv:2108.08877
sentence-t5-base, sentence-t5-xxl
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
arXiv2 reposarXiv:2109.05014
idefics-80b-instruct, idefics-9b-instruct
Decoupling Magnitude and Phase Estimation with Deep ResUNet for Music Source Separation
arXiv2 reposarXiv:2109.05418
demucs, demucs
Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
arXiv2 reposarXiv:2109.10686
turkish-bert, bert5urk
Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
arXiv2 reposarXiv:2109.11680
fairseq, wav2vec2-xlsr-53-espeak-cv-ft
Swiss-Judgment-Prediction: A Multilingual Legal Judgment Prediction Benchmark
arXiv2 reposarXiv:2110.00806
SwissJudgementPrediction, swiss_judgment_prediction
UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
arXiv2 reposarXiv:2110.05752
UniSpeech, unispeech-sat-large
Mengzi: Towards Lightweight yet Ingenious Pre-trained Models for Chinese
arXiv2 reposarXiv:2110.06696
EasyNLP, t5-chinese-couplet
ByteTrack: Multi-Object Tracking by Associating Every Detection Box
arXiv2 reposarXiv:2110.06864
D-FINE-seg, CoreML-Models
Ego4D: Around the World in 3,000 Hours of Egocentric Video
arXiv2 reposarXiv:2110.07058
pyannote-audio, Ego4d
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing
arXiv2 reposarXiv:2110.07205
speecht5_tts, SpeechT5
Pre-training Molecular Graph Representation with 3D Geometry
arXiv2 reposarXiv:2110.07728
OpenBioMed, OpenBioMed_new
Hybrid Spectrogram and Waveform Source Separation
arXiv2 reposarXiv:2111.03600
demucs, demucs
Prune Once for All: Sparse Pre-Trained Language Models
arXiv2 reposarXiv:2111.05754
inteLearn_ML, intel-extension-for-transformers
Merging Models with Fisher-Weighted Averaging
arXiv2 reposarXiv:2111.09832
Mario, MergeLM
AVA-AVD: Audio-Visual Speaker Diarization in the Wild
arXiv2 reposarXiv:2111.14448
pyannote-audio, hf-speaker-diarization-3.1
A General Language Assistant as a Laboratory for Alignment
arXiv2 reposarXiv:2112.00861
stack-exchange-preferences, MERA
Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand
arXiv2 reposarXiv:2112.04139
instructor-embedding, instructor-embedding
FaceFormer: Speech-Driven 3D Facial Animation with Transformers
arXiv2 reposarXiv:2112.05329
FaceFormer, faceformer-emo
LongT5: Efficient Text-To-Text Transformer for Long Sequences
arXiv2 reposarXiv:2112.07916
long-t5-tglobal-xl-16384-book-summary, long-ke-t5
QuALITY: Question Answering with Long Input Texts, Yes!
arXiv2 reposarXiv:2112.08608
Synthetic_Continued_Pretraining, SoE
Unsupervised Dense Information Retrieval with Contrastive Learning
arXiv2 reposarXiv:2112.09118
contriever, contriever-msmarco
Image Segmentation Using Text and Image Prompts
arXiv2 reposarXiv:2112.10003
Semantic-Segment-Anything, FastSAM
Efficient Large Scale Language Modeling with Mixtures of Experts
arXiv2 reposarXiv:2112.10684
fairseq-dense-13B, fairseq-dense-2.7B
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
arXiv2 reposarXiv:2112.11446
falcon-refinedweb, GlorIA
Collapse by Conditioning: Training Class-conditional GANs with Limited Data
arXiv2 reposarXiv:2201.06578
wound-stylegan, transitional-cGAN
Synchromesh: Reliable code generation from pre-trained language models
arXiv2 reposarXiv:2201.11227
syncode, syncode-cypher-lark
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
arXiv2 reposarXiv:2201.11903
prompt-engineering, tree-of-thought-prompting
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
arXiv2 reposarXiv:2202.03052
Levels_image_captioning_NICE, OFA-Chinese
DiT: Self-supervised Pre-training for Document Image Transformer
arXiv2 reposarXiv:2203.02378
TexTok-DiT, transformer_latent_diffusion
Dawn of the transformer era in speech emotion recognition: closing the valence gap
arXiv2 reposarXiv:2203.07378
wav2vec2-large-robust-12-ft-emotion-msp-dim, w2v2-how-to
RELIC: Retrieving Evidence for Literary Claims
arXiv2 reposarXiv:2203.10053
RAG, RAGatouille
ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
arXiv2 reposarXiv:2203.10244
VisRAG-Ret-Train-In-domain-data, ChartQA
MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
arXiv2 reposarXiv:2203.14371
PodGPT, medmcqa
Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$
arXiv2 reposarXiv:2203.17189
xmtf, t5x
AdaFace: Quality Adaptive Margin for Face Recognition
arXiv2 reposarXiv:2204.00964
AdaFace, FLUXSynID
Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing
arXiv2 reposarXiv:2204.09817
BiomedVLP-BioViL-T, BiomedVLP-CXR-BERT-specialized
Evaluating Interpolation and Extrapolation Performance of Neural Retrieval Models
arXiv2 reposarXiv:2204.11447
RAG, RAGatouille
StyleGAN-Human: A Data-Centric Odyssey of Human Generation
arXiv2 reposarXiv:2204.11823
CosmicMan, DeepFashion-MultiModal
Flamingo: a Visual Language Model for Few-Shot Learning
arXiv2 reposarXiv:2204.14198
idefics-80b-instruct, idefics-9b-instruct
CoCa: Contrastive Captioners are Image-Text Foundation Models
arXiv2 reposarXiv:2205.01917
VLSA, VL-KE-T5
MS-Shift: An Analysis of MS MARCO Distribution Shifts on Neural Retrieval
arXiv2 reposarXiv:2205.02870
RAG, RAGatouille
arXiv:2205.09911
arXiv2 reposarXiv:2205.09911
Jellyfish-13B, fm_data_tasks
hmBERT: Historical Multilingual Language Models for Named Entity Recognition
arXiv2 reposarXiv:2205.15575
bert-base-historic-multilingual-cased, clef-hipe
Elucidating the Design Space of Diffusion-Based Generative Models
arXiv2 reposarXiv:2206.00364
playground-v2.5-1024px-aesthetic, YetAnotherStableDiffusion
No Parameter Left Behind: How Distillation and Model Size Affect Zero-Shot Retrieval
arXiv2 reposarXiv:2206.02873
monot5-3b-msmarco-10k, scaling-zero-shot-retrieval
CLAP: Learning Audio Concepts From Natural Language Supervision
arXiv2 reposarXiv:2206.04769
vocalsound, shira_audio
SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks
arXiv2 reposarXiv:2207.04632
FlexCAD, SkexGen
WISE: Whitebox Image Stylization by Example-based Learning
arXiv2 reposarXiv:2207.14606
Whitebox-Style-Transfer-Editing, wise
Evaluating Table Structure Recognition: A New Perspective
arXiv2 reposarXiv:2208.00385
granite-vision-4.1-4b, granite-4.0-3b-vision
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
arXiv2 reposarXiv:2208.07339
vigogne, textsum
Selective Annotation Makes Language Models Better Few-Shot Learners
arXiv2 reposarXiv:2209.01975
instructor-embedding, instructor-embedding
AudioLM: a Language Modeling Approach to Audio Generation
arXiv2 reposarXiv:2209.03143
bark, ultravox
An Empirical Study on Cross-X Transfer for Legal Judgment Prediction
arXiv2 reposarXiv:2209.12325
SwissJudgementPrediction, swiss_judgment_prediction
Music Source Separation with Band-split RNN
arXiv2 reposarXiv:2209.15174
demucs, demucs
SMaLL-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource Languages
arXiv2 reposarXiv:2210.11621
small100, ContraDecode
Contrastive Search Is What You Need For Neural Text Generation
arXiv2 reposarXiv:2210.14140
Adaptive-Contrastive-Search, 4th-Bookathon-The-Unbearable-Heaviness-of-GPT
Autoregressive Structured Prediction with Language Models
arXiv2 reposarXiv:2210.14698
trusted_ke, knowledge-graph-on-research-paper
QuaLA-MiniLM: a Quantized Length Adaptive MiniLM
arXiv2 reposarXiv:2210.17114
inteLearn_ML, intel-extension-for-transformers
Large Language Models Are Human-Level Prompt Engineers
arXiv2 reposarXiv:2211.01910
YiVal, PRL-Prompts-from-Reinforcement-Learning
Efficient Spatially Sparse Inference for Conditional GANs and Diffusion Models
arXiv2 reposarXiv:2211.02048
nunchaku, nunchaku
Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
arXiv2 reposarXiv:2211.05105
i2p, i2p
Fast DistilBERT on CPUs
arXiv2 reposarXiv:2211.07715
inteLearn_ML, intel-extension-for-transformers
Hybrid Transformers for Music Source Separation
arXiv2 reposarXiv:2211.08553
demucs, demucs
InstructPix2Pix: Learning to Follow Image Editing Instructions
arXiv2 reposarXiv:2211.09800
instruct-pix2pix, SEED
Solving math word problems with process- and outcome-based feedback
arXiv2 reposarXiv:2211.14275
Qwen2.5-Math-7B-Instruct-PRM-0.2, Qwen2.5-Math-1.5B-Instruct-PRM-0.2
Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models
arXiv2 reposarXiv:2212.03860
cool-japan-diffusion-for-learning-2-0, picasso-diffusion-1-1
Fast Number Parsing Without Fallback
arXiv2 reposarXiv:2212.06644
ffc.h, fast_float
RTMDet: An Empirical Study of Designing Real-Time Object Detectors
arXiv2 reposarXiv:2212.07784
mmyolo, mmdetection
Discovering Language Model Behaviors with Model-Written Evaluations
arXiv2 reposarXiv:2212.09251
unintentional-unalignment, model-written-evals
MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation
arXiv2 reposarXiv:2212.09478
MMDisCo, MM-Diffusion
The case for 4-bit precision: k-bit Inference Scaling Laws
arXiv2 reposarXiv:2212.09720
airllm, GPTQ-for-LLaMa
One Embedder, Any Task: Instruction-Finetuned Text Embeddings
arXiv2 reposarXiv:2212.09741
instructor-embedding, instructor-embedding
Dataless Knowledge Fusion by Merging Weights of Language Models
arXiv2 reposarXiv:2212.09849
Mario, MergeLM
DDColor: Towards Photo-Realistic Image Colorization via Dual Decoders
arXiv2 reposarXiv:2212.11613
DDColor, DDColor-models
Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
arXiv2 reposarXiv:2212.14024
dspy, dsp
Muse: Text-To-Image Generation via Masked Generative Transformers
arXiv2 reposarXiv:2301.00704
open-muse, NTU_ADL_Team11_Final
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
arXiv2 reposarXiv:2301.02111
ChatTTS, bark
ExcelFormer: A neural network surpassing GBDTs on tabular data
arXiv2 reposarXiv:2301.02819
pytorch-frame, Trompt
FullStop:Punctuation and Segmentation Prediction for Dutch with Transformers
arXiv2 reposarXiv:2301.03319
fullstop-punctuation-multilingual-base, fullstop-dutch-punctuation-prediction
Mastering Diverse Domains through World Models
arXiv2 reposarXiv:2301.04104
d-dreamerv3-world-model, dreamerv3
SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
arXiv2 reposarXiv:2301.04883
sa2va_eval, VisRAG-Ret-Train-In-domain-data
MusicLM: Generating Music From Text
arXiv2 reposarXiv:2301.11325
MusicCaps, open-musiclm
Multimodal Chain-of-Thought Reasoning in Language Models
arXiv2 reposarXiv:2302.00923
Awesome-Multimodal-Prompts, MUStReason
Black Box Adversarial Prompting for Foundation Models
arXiv2 reposarXiv:2302.04237
adversarial_prompting, JailbreakLab
Q-Diffusion: Quantizing Diffusion Models
arXiv2 reposarXiv:2302.04304
nunchaku, nunchaku
A Text-guided Protein Design Framework
arXiv2 reposarXiv:2302.04611
ChatDrug, ProteinCLAP_pretrain_EBM_NCE_downstream_property_prediction
Large Language Models for Code: Security Hardening and Adversarial Testing
arXiv2 reposarXiv:2302.05319
sven_modified, sven
Scaling Vision Transformers to 22 Billion Parameters
arXiv2 reposarXiv:2302.05442
idefics-80b-instruct, idefics-9b-instruct
Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
arXiv2 reposarXiv:2302.09664
moralchoice, Spnda
On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
arXiv2 reposarXiv:2302.12095
Julia_bench, robustlearn
ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
arXiv2 reposarXiv:2302.12288
MiDaS, prisma
WhisperX: Time-Accurate Speech Transcription of Long-Form Audio
arXiv2 reposarXiv:2303.00747
whisperX, dissertation-project
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
arXiv2 reposarXiv:2303.04137
diffusion_pusht, stable-worldmodel
Eliciting Latent Predictions from Transformers with the Tuned Lens
arXiv2 reposarXiv:2303.08112
TransformerLens, notebooks
GPT-4 Technical Report
arXiv2 reposarXiv:2303.08774
openchat_3.5, openchat-3.5-0106
EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation
arXiv2 reposarXiv:2303.11089
EmoTalk_release, 3DFaceAnimation
CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion
arXiv2 reposarXiv:2303.11916
CompoDiff, CompoDiff-Charlie
VideoXum: Cross-modal Visual and Textural Summarization of Videos
arXiv2 reposarXiv:2303.12060
videoxum, videoxum
On the De-duplication of LAION-2B
arXiv2 reposarXiv:2303.12733
idefics-80b-instruct, idefics-9b-instruct
EVA-CLIP: Improved Training Techniques for CLIP at Scale
arXiv2 reposarXiv:2303.15389
eva02_large_patch14_448.mim_m38m_ft_in1k, EVA-CLIP
The Stable Signature: Rooting Watermarks in Latent Diffusion Models
arXiv2 reposarXiv:2303.15435
impossibility-watermark, WMCopier
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
arXiv2 reposarXiv:2303.17580
HuggingGPT, JARVIS
Generative Agents: Interactive Simulacra of Human Behavior
arXiv2 reposarXiv:2304.03442
khms-memory, generative_agents
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
arXiv2 reposarXiv:2304.06767
LMFlow, FsfairX-LLaMA3-RM-v0.1
Chinese Open Instruction Generalist: A Preliminary Release
arXiv2 reposarXiv:2304.07987
COIG-CQIA, COIG
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
arXiv2 reposarXiv:2304.08177
Chinese-LLaMA-Alpaca-2, Chinese-LLaMA-Alpaca-3
UPGPT: Universal Diffusion Model for Person Image Generation, Editing and Pose Transfer
arXiv2 reposarXiv:2304.08870
upgpt, upgpt
Measuring Massive Multitask Chinese Understanding
arXiv2 reposarXiv:2304.12986
ChatGLM-6B, ChatGLM-6B
Towards Automated Circuit Discovery for Mechanistic Interpretability
arXiv2 reposarXiv:2304.14997
TransformerLens, Automatic-Circuit-Discovery
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
arXiv2 reposarXiv:2304.15010
LLaMA-Adapter, LLaMA2-Accessory
LMEye: An Interactive Perception Network for Large Language Models
arXiv2 reposarXiv:2305.03701
LingCloud, Multimodal_Instruction_data_v1
Otter: A Multi-Modal Model with In-Context Instruction Tuning
arXiv2 reposarXiv:2305.03726
Otter, Otter
VCSUM: A Versatile Chinese Meeting Summarization Dataset
arXiv2 reposarXiv:2305.05280
Qwen-7B-Chat, Qwen-14B-Chat
InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
arXiv2 reposarXiv:2305.05662
InternGPT, InternChat
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
arXiv2 reposarXiv:2305.06500
LAVIS, instructblip-vicuna-7b
AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages
arXiv2 reposarXiv:2305.06897
afriqa, afriqa
Evaluating Object Hallucination in Large Vision-Language Models
arXiv2 reposarXiv:2305.10355
POPE, POPE
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
arXiv2 reposarXiv:2305.10415
MedAI-project, ClinicalNLP_PMCVQA
arXiv:2305.10427
arXiv2 reposarXiv:2305.10427
LookaheadDecoding, LookAhead
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
arXiv2 reposarXiv:2305.10601
tree-of-thought-prompting, minihf
LDM3D: Latent Diffusion Model for 3D
arXiv2 reposarXiv:2305.10853
MiDaS, ldm3d-4c
UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
arXiv2 reposarXiv:2305.11147
MultiGen-20M_train, UniControl
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
arXiv2 reposarXiv:2305.11408
WhisperLiveKit, naist-simulst
TheoremQA: A Theorem-driven Question Answering dataset
arXiv2 reposarXiv:2305.12524
RLPR-Evaluation, KOpen-platypus
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
arXiv2 reposarXiv:2305.13971
syncode, syncode-cypher-lark
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
arXiv2 reposarXiv:2305.14292
WikiChat, wikipedia
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
arXiv2 reposarXiv:2305.14610
borderlines, borderlines
NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario
arXiv2 reposarXiv:2305.14836
nuscenes-qa-mini, NuScenes-QA
Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models
arXiv2 reposarXiv:2305.15023
LaVIN, A-Lightweight-Unified-Autoregressive-MLLM
ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers
arXiv2 reposarXiv:2305.15272
vitmatte-small-composition-1k, ViTMatte
Expand, Rerank, and Retrieve: Query Reranking for Open-Domain Question Answering
arXiv2 reposarXiv:2305.17080
EAR, query_decomposer
Trompt: Towards a Better Deep Neural Network for Tabular Data
arXiv2 reposarXiv:2305.18446
pytorch-frame, Trompt
Grammar Prompting for Domain-Specific Language Generation with Large Language Models
arXiv2 reposarXiv:2305.19234
grammar-prompting, blendsql
VideoComposer: Compositional Video Synthesis with Motion Controllability
arXiv2 reposarXiv:2306.02018
videocomposer, i2vgen-xl
Large-Scale Cell Representation Learning via Divide-and-Conquer Contrastive Learning
arXiv2 reposarXiv:2306.04371
OpenBioMed, OpenBioMed_new
DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text
arXiv2 reposarXiv:2306.05540
L2D, AdaDetectGPT
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
arXiv2 reposarXiv:2306.07691
Kokoro-82M, kokoro-82M-onnx-opt
CMMLU: Measuring massive multitask language understanding in Chinese
arXiv2 reposarXiv:2306.09212
cmmlu, PodGPT
Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
arXiv2 reposarXiv:2306.09341
HPDv2, banana100-additional-iqa-models
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
arXiv2 reposarXiv:2306.14565
LRV-Instruction, prismatic-vlms
FunQA: Towards Surprising Video Comprehension
arXiv2 reposarXiv:2306.14899
FunQA, FunQA
3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement
arXiv2 reposarXiv:2306.15354
3D-Speaker, 3d-speaker
One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization
arXiv2 reposarXiv:2306.16928
torchsparse, One-2-3-45
Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models
arXiv2 reposarXiv:2306.17203
Diff-Foley, Diff-Foley
BatGPT: A Bidirectional Autoregessive Talker from Generative Pre-trained Transformer
arXiv2 reposarXiv:2307.00360
CMMLU, cmmlu-debug
DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models
arXiv2 reposarXiv:2307.02421
DragonDiffusion, DragonDiffusion
Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers
arXiv2 reposarXiv:2307.03183
whisper-at, whisper-at
RADAR: Robust AI-Text Detection via Adversarial Learning
arXiv2 reposarXiv:2307.03838
L2D, AdaDetectGPT
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++
arXiv2 reposarXiv:2307.07686
HPC_Fortran_CPP, OpenMP-Fortran-CPP-Translation
Planting a SEED of Vision in Large Language Model
arXiv2 reposarXiv:2307.08041
SEED, SEED
MolFM: A Multimodal Molecular Foundation Model
arXiv2 reposarXiv:2307.09484
OpenBioMed, OpenBioMed_new
Evaluating the Moral Beliefs Encoded in LLMs
arXiv2 reposarXiv:2307.14324
moralchoice, contextual_moralchoice
Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers
arXiv2 reposarXiv:2307.14367
Prot2Text-Data, Prot2Text
MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation
arXiv2 reposarXiv:2307.14460
Depth-Estimation, MiDaS
Robust Distortion-free Watermarks for Language Models
arXiv2 reposarXiv:2307.15593
impossibility-watermark, watermark
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
arXiv2 reposarXiv:2307.16125
SEED-Bench, SEED-Bench
arXiv:2308.00264
arXiv2 reposarXiv:2308.00264
MMML, eval-moshi
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
arXiv2 reposarXiv:2308.01263
jailbreakbench, reward-bench
Tweet Insights: A Visualization Platform to Extract Temporal Insights from Twitter
arXiv2 reposarXiv:2308.02142
twitter-roberta-base-2022-154m, twitter-roberta-large-2022-154m
PIPPA: A Partially Synthetic Conversational Dataset
arXiv2 reposarXiv:2308.05884
PIPPA-shareGPT, PIPPA
BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents
arXiv2 reposarXiv:2308.05960
BOLAA, AgentLite
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
arXiv2 reposarXiv:2308.06721
IP-Adapter, IP-Adapter-FaceID
MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions
arXiv2 reposarXiv:2308.08544
MeViSv2, MeViS
Chinese Spelling Correction as Rephrasing Language Model
arXiv2 reposarXiv:2308.08796
lemon, ReLM
CMB: A Comprehensive Medical Benchmark in Chinese
arXiv2 reposarXiv:2308.08833
CMB, CMB
Steering Language Models With Activation Engineering
arXiv2 reposarXiv:2308.10248
OBLITERATUS, obliteratus
arXiv:2308.11276
arXiv2 reposarXiv:2308.11276
MU-LLaMA, MU-LLaMA
Large Language Models Vote: Prompting for Rare Disease Identification
arXiv2 reposarXiv:2308.12890
llms-vote, llms-vote
SAM-Med2D
arXiv2 reposarXiv:2308.16184
SA-Med2D-20M, SAM-Med2D
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
arXiv2 reposarXiv:2308.16705
CREHate, CREHate
Baseline Defenses for Adversarial Attacks Against Aligned Language Models
arXiv2 reposarXiv:2309.00614
jailbreakbench, llm-jailbreaking-defense
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
arXiv2 reposarXiv:2309.00952
Bridge_Diffusion_Model, BDM1.0
Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models
arXiv2 reposarXiv:2309.02218
DeepFakeFace, DeepFakeFace
Textbooks Are All You Need II: phi-1.5 technical report
arXiv2 reposarXiv:2309.05463
cosmopedia, phi-1_5
Natural Language Supervision for General-Purpose Audio Representations
arXiv2 reposarXiv:2309.05767
msclap, CLAP
Annotating Data for Fine-Tuning a Neural Ranker? Current Active Learning Strategies are not Better than Random Selection
arXiv2 reposarXiv:2309.06131
RAG, RAGatouille
Mitigating Hallucinations and Off-target Machine Translation with Source-Contrastive and Language-Contrastive Decoding
arXiv2 reposarXiv:2309.07098
ContraDecode, ContraDecode
Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context
arXiv2 reposarXiv:2309.08105
libriheavy, libriheavy
LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
arXiv2 reposarXiv:2309.12307
LongLoRA, LongLoRA
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
arXiv2 reposarXiv:2309.13007
opencrabs, rightmind
Aligning Large Multimodal Models with Factually Augmented RLHF
arXiv2 reposarXiv:2309.14525
vision-feedback-mix-binarized, vision-feedback-mix-binarized
ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers
arXiv2 reposarXiv:2309.16119
llmtools, llmtools
Data Filtering Networks
arXiv2 reposarXiv:2309.17425
ml-mobileclip, hashing-baseline
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ
arXiv2 reposarXiv:2310.00367
AutomaTikZ, AutomaTikZ
TimeGPT-1
arXiv2 reposarXiv:2310.03589
moment, nixtla
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
arXiv2 reposarXiv:2310.03684
jailbreakbench, llm-jailbreaking-defense
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
arXiv2 reposarXiv:2310.03714
dspy, dsp
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
arXiv2 reposarXiv:2310.03739
CogVideoX-Fun-V1.1-Reward-LoRAs, EasyAnimateV5-Reward-LoRAs
Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
arXiv2 reposarXiv:2310.04378
TCD-SDXL-LoRA, PixArt-LCM-XL-2-1024-MS
RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation
arXiv2 reposarXiv:2310.04408
recomp, Prompt-Compression
UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
arXiv2 reposarXiv:2310.05126
LLaVA-OneVision-Data, mPLUG-DocOwl
Scaling Laws of RoPE-based Extrapolation
arXiv2 reposarXiv:2310.05209
Llama-3-70B-Instruct-Gradient-262k, Llama-3-70B-Instruct-Gradient-1048k
Compressing Context to Enhance Inference Efficiency of Large Language Models
arXiv2 reposarXiv:2310.06201
Selective_Context, trimwise
DKEC: Domain Knowledge Enhanced Multi-Label Classification for Diagnosis Prediction
arXiv2 reposarXiv:2310.07059
EMS-Pipeline, DKEC
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
arXiv2 reposarXiv:2310.07240
CacheGen, CacheGen-CV
BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations
arXiv2 reposarXiv:2310.07276
OpenBioMed, OpenBioMed_new
arXiv:2310.07641
arXiv2 reposarXiv:2310.07641
LLMBar, reward-bench
Jailbreaking Black Box Large Language Models in Twenty Queries
arXiv2 reposarXiv:2310.08419
jailbreakbench, llm-jailbreaking-defense
BitNet: Scaling 1-bit Transformers for Large Language Models
arXiv2 reposarXiv:2310.11453
BitNet, zeta
GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
arXiv2 reposarXiv:2310.11513
geneval, banana100-additional-iqa-models
SALMONN: Towards Generic Hearing Abilities for Large Language Models
arXiv2 reposarXiv:2310.13289
Speech-IFEval, SALMONN
Wonder3D: Single Image to 3D using Cross-Domain Diffusion
arXiv2 reposarXiv:2310.15008
Wonder3D, wonder3d_archive
DISC-FinLLM: A Chinese Financial Large Language Model based on Multiple Experts Fine-tuning
arXiv2 reposarXiv:2310.15205
DISC-FinLLM, DISC-FinLLMa
DALE: Generative Data Augmentation for Low-Resource Legal NLP
arXiv2 reposarXiv:2310.15799
DALE, DALE
ControlLLM: Augment Language Models with Tools by Searching on Graphs
arXiv2 reposarXiv:2310.17796
ControlLLM, ControlLLM
Punica: Multi-Tenant LoRA Serving
arXiv2 reposarXiv:2310.18547
lorax, llm-lora-hotswap
Foundation Models for Generalist Geospatial Artificial Intelligence
arXiv2 reposarXiv:2310.18660
Prithvi-100M, granite-geospatial-biomass
Efficient LLM Inference on CPUs
arXiv2 reposarXiv:2311.00502
inteLearn_ML, intel-extension-for-transformers
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
arXiv2 reposarXiv:2311.03285
S-LoRA, llm-lora-hotswap
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
arXiv2 reposarXiv:2311.03348
JBB-Behaviors, jailbreakbench
OtterHD: A High-Resolution Multi-modality Model
arXiv2 reposarXiv:2311.04219
Otter, Otter
Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
arXiv2 reposarXiv:2311.04378
watermarks-remover, impossibility-watermark
LRM: Large Reconstruction Model for Single Image to 3D
arXiv2 reposarXiv:2311.04400
TripoSR, Real3D
SA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks
arXiv2 reposarXiv:2311.11969
SA-Med2D-20M, SAM-Med2D
Text-Guided Texturing by Synchronized Multi-View Diffusion
arXiv2 reposarXiv:2311.12891
FlexiSyncMVD, SyncMVD
GeoChat: Grounded Large Vision-Language Model for Remote Sensing
arXiv2 reposarXiv:2311.15826
GeoChat, GeoChat_Instruct
Self-correcting LLM-controlled Diffusion Models
arXiv2 reposarXiv:2311.16090
Omost, SLD
Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine
arXiv2 reposarXiv:2311.16452
UltraMedical, promptbase
Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding
arXiv2 reposarXiv:2311.16922
VideoLLaMA2, AVProunRLForVideoLLaMa2
Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models
arXiv2 reposarXiv:2311.17919
visual_anagrams, LatentGenerativeAnamorphoses
Sequential Modeling Enables Scalable Learning for Large Vision Models
arXiv2 reposarXiv:2312.00785
LVM, LVM_ckpts
OpenVoice: Versatile Instant Voice Cloning
arXiv2 reposarXiv:2312.01479
OpenVoiceV2_Webui_resemble_enhance, OpenVoice
With Great Humor Comes Great Developer Engagement
arXiv2 reposarXiv:2312.01680
faker, faker
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
arXiv2 reposarXiv:2312.02051
TimeChat-7b, TimeIT
PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth Estimation
arXiv2 reposarXiv:2312.02284
PatchFusion, PatchFusion
Analyzing and Improving the Training Dynamics of Diffusion Models
arXiv2 reposarXiv:2312.02696
EDM2-diffusers, edm2
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
arXiv2 reposarXiv:2312.03641
MotionCtrl, MotionCtrl
Alpha-CLIP: A CLIP Model Focusing on Wherever You Want
arXiv2 reposarXiv:2312.03818
AlphaCLIP, MaskImageNet
PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding
arXiv2 reposarXiv:2312.04461
PhotoMaker, PhotoMaker-V2
Grounded Question-Answering in Long Egocentric Videos
arXiv2 reposarXiv:2312.06505
GroundVQA, GroundVQA
Steering Llama 2 via Contrastive Activation Addition
arXiv2 reposarXiv:2312.06681
OBLITERATUS, obliteratus
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
arXiv2 reposarXiv:2312.06722
embodied-eval, behaviour_subtask
SGLang: Efficient Execution of Structured Language Model Programs
arXiv2 reposarXiv:2312.07104
sgl-learning-materials, sglang-vla
Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI
arXiv2 reposarXiv:2312.07886
mPnP-LLM, nuscenes-qa-mini
Distributed Inference and Fine-tuning of Large Language Models Over The Internet
arXiv2 reposarXiv:2312.08361
petals, bloombee_add_models
Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers
arXiv2 reposarXiv:2312.09147
TriplaneGaussian, TriplaneGaussian
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages
arXiv2 reposarXiv:2312.09508
RAG, RAGatouille
VidToMe: Video Token Merging for Zero-Shot Video Editing
arXiv2 reposarXiv:2312.10656
VidToMe, VidToMe
Silkie: Preference Distillation for Large Visual Language Models
arXiv2 reposarXiv:2312.10665
vision-feedback-mix-binarized, vision-feedback-mix-binarized
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
arXiv2 reposarXiv:2312.12456
prosparse-llama-2-13b, prosparse-llama-2-7b
Generative Multimodal Models are In-Context Learners
arXiv2 reposarXiv:2312.13286
Emu2, CoBSAT
HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models
arXiv2 reposarXiv:2312.14091
HD-Painter, HD-Painter
LingoQA: Visual Question Answering for Autonomous Driving
arXiv2 reposarXiv:2312.14115
LingoQA, Cosmos-Reason2-32B
LangSplat: 3D Language Gaussian Splatting
arXiv2 reposarXiv:2312.16084
LangSplat, LangSurf
Towards Better Monolingual Japanese Retrievers with Multi-Vector Models
arXiv2 reposarXiv:2312.16144
RAG, RAGatouille
COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
arXiv2 reposarXiv:2401.00849
cosmo, Howto-Interlink7M
A Comprehensive Study of Knowledge Editing for Large Language Models
arXiv2 reposarXiv:2401.01286
KnowLM, KnowEdit
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI
arXiv2 reposarXiv:2401.01651
AdaptiveDiffusion, Sampled_AIGCBench_text2image_ar_0.625
TinyLlama: An Open-Source Small Language Model
arXiv2 reposarXiv:2401.02385
TinyLlama_v1.1, tinyllama-embed
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
arXiv2 reposarXiv:2401.02731
speechless, speechless-sparsetral-16x7b-MoE
InstantID: Zero-shot Identity-Preserving Generation in Seconds
arXiv2 reposarXiv:2401.07519
InstantID, InstantID
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
arXiv2 reposarXiv:2401.10774
LLM-Sampling, llm_project
In-Context Learning for Extreme Multi-Label Classification
arXiv2 reposarXiv:2401.12178
dspy, dsp
A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation
arXiv2 reposarXiv:2401.12208
RadPhi-2, CheXagent-2-3b
Raidar: geneRative AI Detection viA Rewriting
arXiv2 reposarXiv:2401.12970
RAFT, L2D
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
arXiv2 reposarXiv:2401.13919
UI-TARS, WebVoyager
TURNA: A Turkish Encoder-Decoder Language Model for Enhanced Understanding and Generation
arXiv2 reposarXiv:2401.14373
bert5urk, turkish-lm-tuner
MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
arXiv2 reposarXiv:2401.15391
GraphKV, GraphKV
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
arXiv2 reposarXiv:2401.16420
internlm-xcomposer2-vl-7b, internlm-xcomposer2-4khd-7b
RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval
arXiv2 reposarXiv:2401.18059
drbrain, LARS
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
arXiv2 reposarXiv:2402.00159
dolma, llm_project
DiffEditor: Boosting Accuracy and Flexibility on Diffusion-based Image Editing
arXiv2 reposarXiv:2402.02583
DragonDiffusion, DragonDiffusion
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
arXiv2 reposarXiv:2402.02750
glq, KIVI
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models
arXiv2 reposarXiv:2402.03049
EasyInstruct, KnowLM
ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
arXiv2 reposarXiv:2402.03804
prosparse-llama-2-13b, prosparse-llama-2-7b
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
arXiv2 reposarXiv:2402.04252
EVA-CLIP-8B, EVA-CLIP-18B
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
arXiv2 reposarXiv:2402.04396
quip-sharp, llmtools
EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss
arXiv2 reposarXiv:2402.05008
efficientvit, efficientvit-sam
LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation
arXiv2 reposarXiv:2402.05054
LGM, LGM
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
arXiv2 reposarXiv:2402.05668
JailbreakRadar_Backup, JailbreakRadar
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
arXiv2 reposarXiv:2402.06044
OpenToM, OpenToM
Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations
arXiv2 reposarXiv:2402.07023
Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
arXiv2 reposarXiv:2402.08846
slam_asr_pytorch, SMIT
Generative Representational Instruction Tuning
arXiv2 reposarXiv:2402.09906
sgpt, MEDI2
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
arXiv2 reposarXiv:2402.11411
vision-feedback-mix-binarized, vision-feedback-mix-binarized
FiT: Flexible Vision Transformer for Diffusion Model
arXiv2 reposarXiv:2402.12376
Open-Sora-Plan-v1.2.0, FiT-diffusers
The Revolution of Multimodal Large Language Models: A Survey
arXiv2 reposarXiv:2402.12451
LLaVA-MORE, LLaVA_MORE-gemma_2_9b-finetuning
Visual Style Prompting with Swapping Self-Attention
arXiv2 reposarXiv:2402.12974
StyleKeeper, visual-style-prompting
Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models
arXiv2 reposarXiv:2402.13064
safe-guard-prompt-injection, slm-innovator-lab
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
arXiv2 reposarXiv:2402.14804
MathVision, MATH-V
PALO: A Polyglot Large Multimodal Model for 5B People
arXiv2 reposarXiv:2402.14818
palo_multilingual_dataset, PALO
arXiv:2402.15391
arXiv2 reposarXiv:2402.15391
Jazz, 1xgpt
Nemotron-4 15B Technical Report
arXiv2 reposarXiv:2402.16819
Curator, NeMo-Curator
Transparent Image Layer Diffusion using Latent Transparency
arXiv2 reposarXiv:2402.17113
Diffuser-layerdiffuse, layer_diffusers
Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
arXiv2 reposarXiv:2402.17245
playground-v2.5-1024px-aesthetic, MJHQ-30K
Evaluating Very Long-Term Conversational Memory of LLM Agents
arXiv2 reposarXiv:2402.17753
REALTALK, memex
ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models
arXiv2 reposarXiv:2403.02084
res-adapter, res-adapter
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
arXiv2 reposarXiv:2403.05135
ELLA, LaVi-Bridge
DeepSeek-VL: Towards Real-World Vision-Language Understanding
arXiv2 reposarXiv:2403.05525
deepseek-vl-7b-chat, deepseek-vl-1.3b-chat
SPLADE-v3: New baselines for SPLADE
arXiv2 reposarXiv:2403.06789
splade-v3-lexical-mlx, splade-v3-distilbert
ORPO: Monolithic Preference Optimization without Reference Model
arXiv2 reposarXiv:2403.07691
tunix, MedicalGPT
AutoDev: Automated AI-Driven Development
arXiv2 reposarXiv:2403.08299
pentagi, codel
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
arXiv2 reposarXiv:2403.08350
Continual-NExT, CoIN_Refined
Can We Talk Models Into Seeing the World Differently?
arXiv2 reposarXiv:2403.09193
vlm_shapebias, frequency-cue-conflict
DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
arXiv2 reposarXiv:2403.10266
MindSpeed-MM, OpenDiT
BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics
arXiv2 reposarXiv:2403.10380
BirdSet, BirdSet
arXiv:2403.13164
arXiv2 reposarXiv:2403.13164
VL-ICL, VL-ICL
Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
arXiv2 reposarXiv:2403.14403
Adaptive-RAG, raglite
MyVLM: Personalizing VLMs for User-Specific Queries
arXiv2 reposarXiv:2403.14599
MyVLM, MyVLM
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
arXiv2 reposarXiv:2403.15377
InternVideo, internvideo-d2a11ea9
Protecting Copyrighted Material with Unique Identifiers in Large Language Model Training
arXiv2 reposarXiv:2403.15740
RefAlign, Llama-2-7b-hf-conf-refalign
Explore until Confident: Efficient Exploration for Embodied Question Answering
arXiv2 reposarXiv:2403.15941
embodied-eval, behaviour_subtask
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
arXiv2 reposarXiv:2403.18058
COIG-CQIA, ruozhiba_gpt4
SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens
arXiv2 reposarXiv:2403.18647
SDSAT, CodeLlama-SDSAT_L7_13B
Are We on the Right Way for Evaluating Large Vision-Language Models?
arXiv2 reposarXiv:2403.20330
MMStar, MMStar
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
arXiv2 reposarXiv:2403.20331
UPD, MM-UPD
PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning
arXiv2 reposarXiv:2404.00776
pytorch-frame, Trompt
Release of Pre-Trained Models for the Japanese Language
arXiv2 reposarXiv:2404.01657
bilingual-gpt-neox-4b-minigpt4, youri-7b-chat-gptq
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
arXiv2 reposarXiv:2404.01833
GA, TUP-detection
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
arXiv2 reposarXiv:2404.02905
VAR, var
ReFT: Representation Finetuning for Language Models
arXiv2 reposarXiv:2404.03592
reft_ethos, reft_chat7b_1k
Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese
arXiv2 reposarXiv:2404.07824
heron-chat-git-ja-stablelm-base-7b-v1, Japanese-Heron-Bench
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
arXiv2 reposarXiv:2404.07972
UI-TARS, OSWorld
View Selection for 3D Captioning via Diffusion Ranking
arXiv2 reposarXiv:2404.07984
DiffSplat, Cap3D
Toward a Theory of Tokenization in LLMs
arXiv2 reposarXiv:2404.08335
translation-api, writing-assistance-apis
Magic Clothing: Controllable Garment-Driven Image Synthesis
arXiv2 reposarXiv:2404.09512
MagicClothing, MagicClothing
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model
arXiv2 reposarXiv:2404.09967
Ctrl-Adapter, Ctrl-Adapter
HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
arXiv2 reposarXiv:2404.09990
HQ-Edit, HQ-Edit
Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology
arXiv2 reposarXiv:2404.10242
maes_microscopy, rxrx3-core
MathWriting: A Dataset For Handwritten Mathematical Expression Recognition
arXiv2 reposarXiv:2404.10690
MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash
A Study of Undefined Behavior Across Foreign Function Boundaries in Rust Libraries
arXiv2 reposarXiv:2404.11671
miri, rust
STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases
arXiv2 reposarXiv:2404.13207
stark, stark
From Local to Global: A Graph RAG Approach to Query-Focused Summarization
arXiv2 reposarXiv:2404.16130
graphrag, agentic-graphrag
"Ask Me Anything": How Comcast Uses LLMs to Assist Agents in Real Time
arXiv2 reposarXiv:2405.00801
haystack, haystack-ai
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
arXiv2 reposarXiv:2405.01535
prometheus-7b-v2.0, JudgeBench
What matters when building vision-language models?
arXiv2 reposarXiv:2405.02246
idefics2-8b, Idefics3-8B-Llama3
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
arXiv2 reposarXiv:2405.04007
SEED-X, SEED-Data-Edit
Granite Code Models: A Family of Open Foundation Models for Code Intelligence
arXiv2 reposarXiv:2405.04324
granite-20b-code-instruct-8k, granite-20b-code-base-8k
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
arXiv2 reposarXiv:2405.04434
minimind, Tiny-R2
SketchDream: Sketch-based Text-to-3D Generation and Editing
arXiv2 reposarXiv:2405.06461
SketchDream, sketch-to-multiview
Exploring the Capabilities of Large Multimodal Models on Dense Text
arXiv2 reposarXiv:2405.06706
MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash
LangCell: Language-Cell Pre-training for Cell Identity Understanding
arXiv2 reposarXiv:2405.06708
OpenBioMed, OpenBioMed_new
A Comprehensive Analysis of Static Word Embeddings for Turkish
arXiv2 reposarXiv:2405.07778
Word-Embeddings-Repository-for-Turkish, bert-turkish-x2static
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
arXiv2 reposarXiv:2405.07940
sloptotal, raid
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection
arXiv2 reposarXiv:2405.10300
efficientvit, Rex-Omni
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
arXiv2 reposarXiv:2405.11273
UMOE-Scaling-Unified-Multimodal-LLMs, Uni-MoE
RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search
arXiv2 reposarXiv:2405.12497
turbovec, RaBitQ-Library
OLAPH: Improving Factuality in Biomedical Long-form Question Answering
arXiv2 reposarXiv:2405.12701
OLAPH, MedLFQA
Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models
arXiv2 reposarXiv:2405.16645
Diffusion4D, Diffusion4D
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
arXiv2 reposarXiv:2405.16700
ima-lmms, IMA-DePALM
GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
arXiv2 reposarXiv:2405.17251
genwarp, genwarp
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
arXiv2 reposarXiv:2405.17842
MMDisCo, MMDisCo
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
arXiv2 reposarXiv:2405.18503
soundctm, soundctm
Xwin-LM: Strong and Scalable Alignment Practice for LLMs
arXiv2 reposarXiv:2405.20335
Xwin-LM, Xwin-Math-70B-V1.0
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
arXiv2 reposarXiv:2405.21060
mamba2-minimal, nugie-jax-nemotron-3-nano
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
arXiv2 reposarXiv:2405.21075
video-tt, Video-MME
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
arXiv2 reposarXiv:2406.00356
AudioLCM, AudioLCM
YODAS: Youtube-Oriented Dataset for Audio and Speech
arXiv2 reposarXiv:2406.00899
yodas, yodas2
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
arXiv2 reposarXiv:2406.01006
semcoder_s_1030, semcoder_1030
V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation
arXiv2 reposarXiv:2406.02511
V-Express, V-Express
Parrot: Multilingual Visual Instruction Tuning
arXiv2 reposarXiv:2406.02539
Ovis, Ovis
Wings: Learning Multimodal LLMs without Text-only Forgetting
arXiv2 reposarXiv:2406.03496
Ovis, Ovis
VideoPhy: Evaluating Physical Commonsense for Video Generation
arXiv2 reposarXiv:2406.03520
videophy, videophy_autoeval_scores
SilentCipher: Deep Audio Watermarking
arXiv2 reposarXiv:2406.03822
silentcipher, SilentCipher
VideoTetris: Towards Compositional Text-to-Video Generation
arXiv2 reposarXiv:2406.04277
VideoTetris, VideoTetris-long
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval
arXiv2 reposarXiv:2406.04292
bge-visualized, VISTA_Evaluation_FineTuning
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
arXiv2 reposarXiv:2406.04334
granite-vision-4.1-4b, granite-4.0-3b-vision
MAIRA-2: Grounded Radiology Report Generation
arXiv2 reposarXiv:2406.04449
libra-maira-2, llava-rad
GenAI Arena: An Open Evaluation Platform for Generative Models
arXiv2 reposarXiv:2406.04485
ImagenHub, VideoGenHub
F-LMM: Grounding Frozen Large Multimodal Models
arXiv2 reposarXiv:2406.05821
F-LMM, F-LMM
MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
arXiv2 reposarXiv:2406.06565
MixEval-Archon, MixEval
Spectrum: Targeted Training on Signal to Noise Ratio
arXiv2 reposarXiv:2406.06623
spectrum, donutloop-genesis
PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation
arXiv2 reposarXiv:2406.06679
PatchRefiner, PatchFusion
Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
arXiv2 reposarXiv:2406.06890
mcm, mcm
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
arXiv2 reposarXiv:2406.07209
MS-Bench, MS-Diffusion
Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?
arXiv2 reposarXiv:2406.07546
Commonsense-T2I, CommonsensenT2I
An Image is Worth 32 Tokens for Reconstruction and Generation
arXiv2 reposarXiv:2406.07550
tokenizer_titok_l32_imagenet, clustermark_1d-tokenizer
What If We Recaption Billions of Web Images with LLaMA-3?
arXiv2 reposarXiv:2406.08478
Recap-DataComp-1B, ViT-L-16-HTxt-Recap-CLIP
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
arXiv2 reposarXiv:2406.08820
disfluency_speech_german, disfluency_speech_english
Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs
arXiv2 reposarXiv:2406.10216
GRM-Llama3.2-3B-rewardmodel-ft, GRM-Gemma-2B-rewardmodel-ft
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
arXiv2 reposarXiv:2406.11171
scpp, SugarCrepe_pp
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
arXiv2 reposarXiv:2406.11230
MMNeedle, multimodal-needle-in-a-haystack
Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
arXiv2 reposarXiv:2406.11695
dspy, dsp
Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels
arXiv2 reposarXiv:2406.11706
dspy, dsp
DataComp-LM: In search of the next generation of training sets for language models
arXiv2 reposarXiv:2406.11794
dclm-baseline-1.0, dclm-baseline-1.0-parquet
WebCanvas: Benchmarking Web Agents in Online Environments
arXiv2 reposarXiv:2406.12373
Mind2Web_Live_SeeAct_V, mind2web-live-seeact-v
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
arXiv2 reposarXiv:2406.12742
MIRB, MIRB
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
arXiv2 reposarXiv:2406.13352
felonybench, moltshield
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
arXiv2 reposarXiv:2406.14553
xCOMET-lite, XCOMET-lite
AEM: Attention Entropy Maximization for Multiple Instance Learning based Whole Slide Image Classification
arXiv2 reposarXiv:2406.15303
AEM, AEM-dataset
Improving Text-To-Audio Models with Synthetic Captions
arXiv2 reposarXiv:2406.15487
tango-af-ac-ft-ac, tango-music-af-ft-mc
HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis
arXiv2 reposarXiv:2406.16192
HEST, hest-tissue-seg
arXiv:2406.18521
arXiv2 reposarXiv:2406.18521
CharXiv, CharXiv
arXiv:2406.18665
arXiv2 reposarXiv:2406.18665
RouteLLM, llm-router
EmPO: Emotion Grounding for Empathetic Response Generation through Preference Optimization
arXiv2 reposarXiv:2406.19071
zephyr-7b-sft-full124, zephyr-7b-sft-full124_d270
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
arXiv2 reposarXiv:2406.19232
RuBLiMP, rublimp
ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
arXiv2 reposarXiv:2406.19392
ReXTime, ReXTime
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
arXiv2 reposarXiv:2407.04923
omchat-v2.0-13B-single-beta_hf, omchat
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
arXiv2 reposarXiv:2407.05134
Formulate_and_Solve, BeyondX
MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
arXiv2 reposarXiv:2407.06358
MiraData, MiraData
Vision language models are blind: Failing to translate detailed visual features into words
arXiv2 reposarXiv:2407.06581
vision-llms-are-blind, vlmsareblind
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
arXiv2 reposarXiv:2407.06723
ml-gbc, GBC10M-PromptGen-200M
Cue Point Estimation using Object Detection
arXiv2 reposarXiv:2407.06823
cue-detr, edm-cue
WildGaussians: 3D Gaussian Splatting in the Wild
arXiv2 reposarXiv:2407.08447
wild-gaussians, nerfonthego-undistorted
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together
arXiv2 reposarXiv:2407.10930
dspy, dsp
Goldfish: Vision-Language Understanding of Arbitrarily Long Videos
arXiv2 reposarXiv:2407.12679
MiniGPT4-Video, TVQA-Long
IMAGDressing-v1: Customizable Virtual Dressing
arXiv2 reposarXiv:2407.12705
IMAGDressing, IMAGDressing
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
arXiv2 reposarXiv:2407.12772
UniG2U, lmms-eval-mmllm
Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle
arXiv2 reposarXiv:2407.13833
Phi-3.5-MoE-instruct, Phi-4-multimodal-instruct
arXiv:2407.14435
arXiv2 reposarXiv:2407.14435
dictionary_learning, CLT-Forge
NV-Retriever: Improving text embedding models with effective hard-negative mining
arXiv2 reposarXiv:2407.15831
LateOn-Code, LateOn-Code-edge
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
arXiv2 reposarXiv:2407.18391
UOUO, UOUO-Bench
Towards A Generalizable Pathology Foundation Model via Unified Knowledge Distillation
arXiv2 reposarXiv:2407.18449
GPFM, GPFM
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
arXiv2 reposarXiv:2407.19669
gte-large-en-v1.5, gte-multilingual-base
Towards Localized Fine-Grained Control for Facial Expression Generation
arXiv2 reposarXiv:2407.20175
fineface, fineface
Learning Feature-Preserving Portrait Editing from Generated Pairs
arXiv2 reposarXiv:2407.20455
feature-preserve-portrait-editing, feature-preserve-portrait-editing
Black-Box Adversarial Attacks on LLM-Based Code Completion
arXiv2 reposarXiv:2408.02509
insec, insec-vulnerability
MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
arXiv2 reposarXiv:2408.02718
MMIU, MMIU-Benchmark
EXAONE 3.0 7.8B Instruction Tuned Language Model
arXiv2 reposarXiv:2408.03541
KoMT-Bench, KoMT-Bench
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
arXiv2 reposarXiv:2408.03695
openstorypp, OpenstoryPlusPlus
CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases
arXiv2 reposarXiv:2408.03910
ms-agent, modelscope-agent
Med42-v2: A Suite of Clinical LLMs
arXiv2 reposarXiv:2408.06142
Llama3-Med42-70B, Llama3-Med42-8B
ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area
arXiv2 reposarXiv:2408.07246
ChemVLM_test_data, ChemVLM-26B-1-2
Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
arXiv2 reposarXiv:2408.09600
Vaccine, Lisa
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
arXiv2 reposarXiv:2408.12109
Vision-LLM-Alignment, robust_visual_reward_model
Sapiens: Foundation for Human Vision Models
arXiv2 reposarXiv:2408.12569
sapiens-pose-1b-torchscript, sapiens
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
arXiv2 reposarXiv:2408.13257
Video-MME, sa2va_eval
Benchmarking foundation models as feature extractors for weakly-supervised computational pathology
arXiv2 reposarXiv:2408.15823
STAMP, STAMP_attention_ui
CSGO: Content-Style Composition in Text-to-Image Generation
arXiv2 reposarXiv:2408.16766
CSGO, CSGO
VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
arXiv2 reposarXiv:2408.17253
uni2ts, uni2ts
OnlySportsLM: Optimizing Sports-Domain Language Models with SOTA Performance under Billion Parameters
arXiv2 reposarXiv:2409.00286
OnlySportsLM, OnlySports_Dataset
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
arXiv2 reposarXiv:2409.00750
MaskGCT-Windows, maskgct
ToolACE: Winning the Points of LLM Function Calling
arXiv2 reposarXiv:2409.00920
ToolACE, ToolACE-8B
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
arXiv2 reposarXiv:2409.02813
MMMU, MMMU_Pro
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
arXiv2 reposarXiv:2409.03797
NESTFUL, nestful
LLaMA-Omni: Seamless Speech Interaction with Large Language Models
arXiv2 reposarXiv:2409.06666
LLaMA-Omni, Llama-3.1-8B-Omni
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
arXiv2 reposarXiv:2409.08425
SoloAudio, SoloAudio
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
arXiv2 reposarXiv:2409.09788
SpaceThinker-Qwen2.5VL-3B, Q-Spatial-Bench-code
Eureka: Evaluating and Understanding Large Foundation Models
arXiv2 reposarXiv:2409.10566
eureka-ml-insights, Eureka-Bench-Logs
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
arXiv2 reposarXiv:2409.10819
EzAudio, EzAudio
OmniGen: Unified Image Generation
arXiv2 reposarXiv:2409.11340
OmniGen, OmniGen-v1
Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
arXiv2 reposarXiv:2409.11406
Phidias-Diffusion, Phidias-Diffusion
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
arXiv2 reposarXiv:2409.12122
Qwen2.5-Math-7B-Instruct, Qwen2.5-Math-RM-72B
StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
arXiv2 reposarXiv:2409.12576
StoryMaker, StoryMaker
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
arXiv2 reposarXiv:2409.12951
SimSIMD, numkong
MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding
arXiv2 reposarXiv:2409.14818
mobilevlm_test, Mobile3M
DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion
arXiv2 reposarXiv:2409.17145
DreamWaltz-G, DreamWaltz-G
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
arXiv2 reposarXiv:2409.18125
LLaVA-3D-7B, LLaVA-3D
arXiv:2409.19425
arXiv2 reposarXiv:2409.19425
freeze-align, concept_coverage_laion_6m
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation
arXiv2 reposarXiv:2409.19483
MedCLIP-SAMv2, MedCLIP-SAMv2
arXiv:2409.19603
arXiv2 reposarXiv:2409.19603
VideoLISA, VideoLISA-3.8B
Preserving Generalization of Language models in Few-shot Continual Relation Extraction
arXiv2 reposarXiv:2410.00334
sirus_v2, Minion
Do Music Generation Models Encode Music Theory?
arXiv2 reposarXiv:2410.00872
syntheory, syntheory
nGPT: Normalized Transformer with Representation Learning on the Hypersphere
arXiv2 reposarXiv:2410.01131
SimSIMD, numkong
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
arXiv2 reposarXiv:2410.01744
Leopard-Idefics2, Leopard-LLaVA
SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration
arXiv2 reposarXiv:2410.02367
SageAttention, FastVideo
Recent Advances in Speech Language Models: A Survey
arXiv2 reposarXiv:2410.03751
VoxEval, VoxEval
Timer-XL: Long-Context Transformers for Unified Time Series Forecasting
arXiv2 reposarXiv:2410.04803
timer-base-84m, Large-Time-Series-Model
Accelerating Diffusion Transformers with Token-wise Feature Caching
arXiv2 reposarXiv:2410.05317
ToCa, DuCa
T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design
arXiv2 reposarXiv:2410.05677
t2v-turbo, T2V-Turbo-v2-no-MG
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
arXiv2 reposarXiv:2410.06940
budget-flow-matching, REPA
Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow
arXiv2 reposarXiv:2410.07303
Rectified-Diffusion, Rectified-Diffusion
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts
arXiv2 reposarXiv:2410.07348
zen5, Chat-UniVi
Lost in Time: A New Temporal Benchmark for VideoLLMs
arXiv2 reposarXiv:2410.07752
TVBench, tvbench
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
arXiv2 reposarXiv:2410.08208
SPA, SPA
Reconstructive Visual Instruction Tuning
arXiv2 reposarXiv:2410.09575
ross-qwen2-7b, Visual-Region
Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy
arXiv2 reposarXiv:2410.09873
AdaptiveDiffusion, Sampled_AIGCBench_text2image_ar_0.625
GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
arXiv2 reposarXiv:2410.10393
gift-eval, GiftEval
Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
arXiv2 reposarXiv:2410.10469
uni2ts, uni2ts
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI
arXiv2 reposarXiv:2410.11623
embodied-eval, behaviour_subtask
CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos
arXiv2 reposarXiv:2410.11831
co-tracker, CoTracker3_Kubric
JudgeBench: A Benchmark for Evaluating LLM-based Judges
arXiv2 reposarXiv:2410.12784
JudgeBench, JudgeBench
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
arXiv2 reposarXiv:2410.12787
VideoLLaMA2, AVProunRLForVideoLLaMa2
D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement
arXiv2 reposarXiv:2410.13842
D-FINE-seg, pydfine
BrainTransformers: SNN-LLM
arXiv2 reposarXiv:2410.14687
BrainTransformers-SNN-LLM, BrainTransformers-3B-Chat
Allegro: Open the Black Box of Commercial-Level Video Generation Model
arXiv2 reposarXiv:2410.15458
Allegro, Allegro-TI2V
Frontiers in Intelligent Colonoscopy
arXiv2 reposarXiv:2410.17241
Project-Imaging-X, VPS
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
arXiv2 reposarXiv:2410.17578
MMQA, prometheus-eval
3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation
arXiv2 reposarXiv:2410.18974
MVEdit, 3D-Adapter
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
arXiv2 reposarXiv:2410.20280
OpenVid-1M, OpenVid
ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation
arXiv2 reposarXiv:2410.20502
OpenVid-1M, OpenVid
InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models
arXiv2 reposarXiv:2410.22770
moltshield, NotInject
MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering
arXiv2 reposarXiv:2410.22949
OpenBioMed, OpenBioMed_new
In-Context LoRA for Diffusion Transformers
arXiv2 reposarXiv:2410.23775
catvton-flux, catvton-unstudio-flux
$π_0$: A Vision-Language-Action Flow Model for General Robot Control
arXiv2 reposarXiv:2410.24164
GR00T-N1.5-3B, GR00T-N1.6-3B
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
arXiv2 reposarXiv:2411.00836
DynaMath_Sample, DynaMath
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
arXiv2 reposarXiv:2411.01156
fish-speech, fish-speech-1.5
xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
arXiv2 reposarXiv:2411.01738
xDiT, mochi-xdit
GenXD: Generating Any 3D and 4D Scenes
arXiv2 reposarXiv:2411.02319
OpenVid-1M, OpenVid
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
arXiv2 reposarXiv:2411.04928
OpenVid-1M, OpenVid
LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
arXiv2 reposarXiv:2411.04997
LLM2CLIP, LLM2CLIP-Llama-3-8B-Instruct-CC-Finetuned
FlexCAD: Unified and Versatile Controllable CAD Generation with Fine-tuned Large Language Models
arXiv2 reposarXiv:2411.05823
FlexCAD, FlexCAD
Improved Video VAE for Latent Video Diffusion Model
arXiv2 reposarXiv:2411.06449
OpenVid-1M, OpenVid
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
arXiv2 reposarXiv:2411.07199
OmniEdit-Filtered-1.2M, OmniEdit
Zero-shot Voice Conversion with Diffusion Transformers
arXiv2 reposarXiv:2411.09943
maestro-seedvc, seed-vc-test
EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation
arXiv2 reposarXiv:2411.10061
echomimic_v2, EchoMimicV2
CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation
arXiv2 reposarXiv:2411.10086
CorrCLIPv2, CorrCLIP
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
arXiv2 reposarXiv:2411.10501
OnlyFlow, onlyflow
Multimodal Autoregressive Pre-training of Large Vision Encoders
arXiv2 reposarXiv:2411.14402
ml-aim, aimv2-large-patch14-native
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI
arXiv2 reposarXiv:2411.14522
GMAI-VL-5.5M, GMAI-VL
MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model
arXiv2 reposarXiv:2411.16157
MVGenMaster, MVGenMaster
BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment
arXiv2 reposarXiv:2411.16300
BayLing, BayLing
One Diffusion to Generate Them All
arXiv2 reposarXiv:2411.16318
OneDiffusion, OneDiffusion
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
arXiv2 reposarXiv:2411.16495
AtomR, BlendQA
All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages
arXiv2 reposarXiv:2411.16508
ALM-Bench, ALM-Bench
DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters
arXiv2 reposarXiv:2411.17423
DRiVE, DRiVE
Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
arXiv2 reposarXiv:2411.17525
quant.cpp, turboquant-vllm
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation
arXiv2 reposarXiv:2411.17945
MARVEL-FX3D, MARVEL-40M
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
arXiv2 reposarXiv:2411.19108
TeaCache, FastVideo
Video Depth without Video Models
arXiv2 reposarXiv:2411.19189
RollingDepth, rollingdepth-v1-0
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
arXiv2 reposarXiv:2411.19325
GEO-Bench-VLM, GEOBench-VLM
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
arXiv2 reposarXiv:2412.00733
hallo3, hallo3_training_data
CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking
arXiv2 reposarXiv:2412.01007
LateOn-Code, LateOn-Code-edge
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
arXiv2 reposarXiv:2412.02592
OHR-Bench, OHR-Bench
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
arXiv2 reposarXiv:2412.03304
Global-MMLU-Lite, Global-MMLU-Lite
MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation
arXiv2 reposarXiv:2412.03558
MIDI-3D, MIDI-3D
MV-Adapter: Multi-view Consistent Image Generation Made Easy
arXiv2 reposarXiv:2412.03632
mv-adapter, Objaverse-Rand6View
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
arXiv2 reposarXiv:2412.04447
embodied-eval, behaviour_subtask
The BrowserGym Ecosystem for Web Agent Research
arXiv2 reposarXiv:2412.05467
AL_for_hallucination, AgentLab
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
arXiv2 reposarXiv:2412.05496
SparseD, szl-khipu
LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations
arXiv2 reposarXiv:2412.06322
LLaVA-SpaceSGG, LLaVA-SpaceSGG
arXiv:2412.06410
arXiv2 reposarXiv:2412.06410
dictionary_learning, notebooks
Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation
arXiv2 reposarXiv:2412.06781
plonk, PLONK
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
arXiv2 reposarXiv:2412.07626
OmniDocBench, granite-4.0-3b-vision
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
arXiv2 reposarXiv:2412.07759
3DTrajMaster, 3DTrajMaster
From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
arXiv2 reposarXiv:2412.07772
CausVid, CausVid
TryOffAnyone: Tiled Cloth Generation from a Dressed Person
arXiv2 reposarXiv:2412.08573
try-off-anyone, tryOffAnyone
FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models
arXiv2 reposarXiv:2412.08629
FlowEdit, Wan2.1
Arbitrary-steps Image Super-resolution via Diffusion Inversion
arXiv2 reposarXiv:2412.09013
InvSR, InvSR
SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos
arXiv2 reposarXiv:2412.09401
slam3r_i2p, slam3r_l2w
Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
arXiv2 reposarXiv:2412.09585
VisPer-LM, OLA-VLM
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
arXiv2 reposarXiv:2412.09616
InternVL3-2B, InternVL3-38B
WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model
arXiv2 reposarXiv:2412.09951
WiseAD, WiseAD_training_data
AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era
arXiv2 reposarXiv:2412.10255
Index-anisora, Index-anisora
Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
arXiv2 reposarXiv:2412.10432
L2D, AdaDetectGPT
SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device
arXiv2 reposarXiv:2412.10494
OpenVid-1M, OpenVid
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
arXiv2 reposarXiv:2412.12693
SPHERE-VLM, SPHERE-VLM
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
arXiv2 reposarXiv:2412.12888
ArtAug-lora-FLUX.1dev-v1, diffSynth-studio-notes
Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis
arXiv2 reposarXiv:2412.13126
KEEP, PathPT
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
arXiv2 reposarXiv:2412.13702
llama3.1-typhoon2-audio-8b-instruct, typhoon2-audio
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues
arXiv2 reposarXiv:2412.15190
EarthDial, Land-Change-Detection
Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks
arXiv2 reposarXiv:2412.15605
CAG, Cache-Augmented-Generation-Granite
Personalized Representation from Personalized Generation
arXiv2 reposarXiv:2412.16156
personalized-rep, PODS
DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder
arXiv2 reposarXiv:2412.17644
DreamFit, Dreamfit
YuLan-Mini: An Open Data-efficient Language Model
arXiv2 reposarXiv:2412.17743
ProX, program-every-example
Towards Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST)
arXiv2 reposarXiv:2412.18367
MultilingualAITerminology, multilingual-terminology
Open-Sora: Democratizing Efficient Video Production for All
arXiv2 reposarXiv:2412.20404
Open-Sora, Open-Sora-v2
LTX-Video: Realtime Video Latent Diffusion
arXiv2 reposarXiv:2501.00103
LTX-Video, ltx2-vidgen-skill
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models
arXiv2 reposarXiv:2501.00874
lusifer, LUSIFER
EliGen: Entity-Level Controlled Image Generation with Regional Attention
arXiv2 reposarXiv:2501.01097
EliGen, diffSynth-studio-notes
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
arXiv2 reposarXiv:2501.01957
VITA, VITA-1.5
Cosmos World Foundation Model Platform for Physical AI
arXiv2 reposarXiv:2501.03575
Cosmos1GP, Cosmos-Tokenizer
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
arXiv2 reposarXiv:2501.04670
colva_internvl2_4b, CoLVA
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
arXiv2 reposarXiv:2501.04962
VoxEval, VoxEval
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
arXiv2 reposarXiv:2501.05122
Centurio, Synthdog-Multilingual-100
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
arXiv2 reposarXiv:2501.08225
FramePainter, FramePainter
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
arXiv2 reposarXiv:2501.08453
Vchitect-2.0, Vchitect_T2V_DataVerse
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
arXiv2 reposarXiv:2501.12327
VARGPT, VARGPT_datasets
Kimi k1.5: Scaling Reinforcement Learning with LLMs
arXiv2 reposarXiv:2501.12599
MathVision, MATH-V
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
arXiv2 reposarXiv:2501.13826
lmms-eval, UniG2U
Zep: A Temporal Knowledge Graph Architecture for Agent Memory
arXiv2 reposarXiv:2501.13956
graphiti, post-graph-rag
Visual Generation Without Guidance
arXiv2 reposarXiv:2501.15420
GFT, GFT
Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
arXiv2 reposarXiv:2501.15907
kani-tts-2-en, kani-tts-2-pt
Molecular-driven Foundation Model for Oncologic Pathology
arXiv2 reposarXiv:2501.16652
TridentEdited, aegis
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
arXiv2 reposarXiv:2501.16764
DiffSplat, DiffSplat
BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
arXiv2 reposarXiv:2501.17790
BreezyVoice, BreezyVoice
Diverse Preference Optimization
arXiv2 reposarXiv:2501.18101
DiversityTuning, fairseq2
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
arXiv2 reposarXiv:2501.18362
MedXpertQA, MedXpertQA
GuardReasoner: Towards Reasoning-based LLM Safeguards
arXiv2 reposarXiv:2501.18492
GuardReasoner-Omni-7B, GuardReasoner-Omni-3B
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
arXiv2 reposarXiv:2501.18954
LLMDet, LLMDet
Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
arXiv2 reposarXiv:2501.19054
CADFusion, CADFusion
Process Reinforcement through Implicit Rewards
arXiv2 reposarXiv:2502.01456
OpenRLHF, PRIME
Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation
arXiv2 reposarXiv:2502.02464
reranking-datasets, reranking-datasets-light
LIMO: Less is More for Reasoning
arXiv2 reposarXiv:2502.03387
open-korean-instructions, ko-limo
Detecting Strategic Deception Using Linear Probes
arXiv2 reposarXiv:2502.03407
FabricationGuard-linearprobe-qwen36-27b, ReasoningGuard-linearprobe-qwen36-27b
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
arXiv2 reposarXiv:2502.03930
VoxCPM, dots.tts
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
arXiv2 reposarXiv:2502.05171
ProX, program-every-example
LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models
arXiv2 reposarXiv:2502.06352
LANTERN, anole_drafter
Accelerating Data Processing and Benchmarking of AI Models for Pathology
arXiv2 reposarXiv:2502.06750
TRIDENT, TridentEdited
History-Guided Video Diffusion
arXiv2 reposarXiv:2502.06764
DFoT, diffusion-forcing-transformer
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
arXiv2 reposarXiv:2502.06782
Lumina-Video, Lumina-Video-f24R960
Self-Supervised Prompt Optimization
arXiv2 reposarXiv:2502.06855
MetaGPT, MetaGPT
WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point
arXiv2 reposarXiv:2502.08047
WorldGUI, WorldGUI-Bench
Universal Model Routing for Efficient LLM Inference
arXiv2 reposarXiv:2502.08773
Aurora-AI-local-router, router
EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
arXiv2 reposarXiv:2502.09509
EQ-SDXL-VAE, HakuLatent
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
arXiv2 reposarXiv:2502.09560
behaviour_subtask, embodied-eval
HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation
arXiv2 reposarXiv:2502.09838
VL-Health, HealthGPT
KernelBench: Can LLMs Write Efficient GPU Kernels?
arXiv2 reposarXiv:2502.10517
autokernel, KernelBench
arXiv:2502.10841
arXiv2 reposarXiv:2502.10841
SkyReels-A1, SkyReels-A1
Phantom: Subject-consistent video generation via cross-modal alignment
arXiv2 reposarXiv:2502.11079
Phantom, Phantom
How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training
arXiv2 reposarXiv:2502.11196
KnowledgeCircuits, DynamicKnowledgeCircuits
BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model
arXiv2 reposarXiv:2502.11798
BackdoorDM, BackdoorDM
Atom of Thoughts for Markov LLM Test-Time Scaling
arXiv2 reposarXiv:2502.12018
MetaGPT, MetaGPT
FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
arXiv2 reposarXiv:2502.12138
FLARE_NVS, FLARE
HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation
arXiv2 reposarXiv:2502.12148
HermesFlow, HermesFlow
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
arXiv2 reposarXiv:2502.12558
MomentSeeker, MomentSeeker
MMTEB: Massive Multilingual Text Embedding Benchmark
arXiv2 reposarXiv:2502.13595
mteb, ru_sci_bench_mteb
Language Model Re-rankers are Fooled by Lexical Similarities
arXiv2 reposarXiv:2502.17036
rerankers-and-lexical-similarities, rerankers-and-lexical-similarities
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
arXiv2 reposarXiv:2502.17420
OBLITERATUS, obliteratus
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
arXiv2 reposarXiv:2502.17422
mllms_know, TextVQA_GT_bbox
Chain of Draft: Thinking Faster by Writing Less
arXiv2 reposarXiv:2502.18600
fabric, CoDE-Stop
UniTok: A Unified Tokenizer for Visual Generation and Understanding
arXiv2 reposarXiv:2502.20321
unitok_mllm, unitok_tokenizer
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
arXiv2 reposarXiv:2503.00912
HiBench, HiBench
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
arXiv2 reposarXiv:2503.01710
Spark-TTS-0.5B, Spark-TTS
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
arXiv2 reposarXiv:2503.01743
Phi-4-multimodal-instruct, Phi-4-mini-instruct
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval
arXiv2 reposarXiv:2503.04644
IFIR, IFIR
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
arXiv2 reposarXiv:2503.04721
Full-Duplex-Bench, personaplex
Dynamic Knowledge Integration for Evidence-Driven Counter-Argument Generation with Large Language Models
arXiv2 reposarXiv:2503.05328
counter-argument, counter-argument-generation
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
arXiv2 reposarXiv:2503.06157
UrbanVideo-Bench, UrbanVideo-Bench.code
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
arXiv2 reposarXiv:2503.06800
videophy, Cosmos-Reason2-32B
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
arXiv2 reposarXiv:2503.07265
UniWorld-V1, UniWorld-V1-NF4
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
arXiv2 reposarXiv:2503.07365
MMK12, MM-EUREKA
NullFace: Training-Free Localized Face Anonymization
arXiv2 reposarXiv:2503.08478
nullface, nullface-test-set
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
arXiv2 reposarXiv:2503.08686
OmniMamba, OmniMamba
Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k
arXiv2 reposarXiv:2503.09642
Open-Sora, Open-Sora-v2
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
arXiv2 reposarXiv:2503.10291
VisualPRM-8B, VisualPRM400K
Distilling Diversity and Control in Diffusion Models
arXiv2 reposarXiv:2503.10637
stable-confusion, distillation
Taming Knowledge Conflicts in Language Models
arXiv2 reposarXiv:2503.10996
JUICE, ParaConfilct
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
arXiv2 reposarXiv:2503.11073
PURE, PURE
Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering
arXiv2 reposarXiv:2503.11117
behaviour_subtask, embodied-eval
Context-Aware Rule Mining Using a Dynamic Transformer-Based Framework
arXiv2 reposarXiv:2503.11125
awesome-datascience, awesome-datascience
FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis
arXiv2 reposarXiv:2503.13265
FlexWorld, FlexWorld
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
arXiv2 reposarXiv:2503.13444
VideoMind, VideoMind-Dataset
Where do Large Vision-Language Models Look at when Answering Questions?
arXiv2 reposarXiv:2503.13891
LVLM_Interpretation, LVLM_Interpretation
AIGVE-Tool: AI-Generated Video Evaluation Toolkit with Multifaceted Benchmark
arXiv2 reposarXiv:2503.14064
AIGVE_Tool, AIGVE-Bench
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
arXiv2 reposarXiv:2503.14935
FAVOR, FAVOR-Bench
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
arXiv2 reposarXiv:2503.15661
UI-Vision, ui-vision
InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
arXiv2 reposarXiv:2503.16418
InfiniteYou, InfiniteYou
AMD-Hummingbird: Towards an Efficient Text-to-Video Model
arXiv2 reposarXiv:2503.18559
Hummingbird, AMD-Hummingbird-T2V
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
arXiv2 reposarXiv:2503.19470
ReSearch, ReCall
Gemma 3 Technical Report
arXiv2 reposarXiv:2503.19786
SLU_pipeline, open-value
Understanding R1-Zero-Like Training: A Critical Perspective
arXiv2 reposarXiv:2503.20783
tunix, understand-r1-zero
Empowering Retrieval-based Conversational Recommendation with Contrasting User Preferences
arXiv2 reposarXiv:2503.22005
CORAL, CORAL
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
arXiv2 reposarXiv:2503.22020
VLAC, AI539_NLP
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
arXiv2 reposarXiv:2503.22152
behaviour_subtask, embodied-eval
VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
arXiv2 reposarXiv:2503.23064
VGRP-Bench, VGRP-Bench
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
arXiv2 reposarXiv:2504.00072
chapter-llama, chapter-llama
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
arXiv2 reposarXiv:2504.00294
eureka-ml-insights, Eureka-Bench-Logs
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
arXiv2 reposarXiv:2504.00470
LIMA, SMDL-Attribution
BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing
arXiv2 reposarXiv:2504.01786
BlenderGym-Open, BG_bench_data
arXiv:2504.02436
arXiv2 reposarXiv:2504.02436
SkyReels-A2, SkyReels-A2
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
arXiv2 reposarXiv:2504.02821
sae-for-vlm, sae-for-vlm
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
arXiv2 reposarXiv:2504.02949
VARGPT, VARGPT_datasets
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models
arXiv2 reposarXiv:2504.03624
NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-3-Nano-4B-GGUF
Training state-of-the-art pathology foundation models with orders of magnitude less data
arXiv2 reposarXiv:2504.05186
midnight, Midnight
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
arXiv2 reposarXiv:2504.05692
POMATO, POMATO
DDT: Decoupled Diffusion Transformer
arXiv2 reposarXiv:2504.05741
DDT-XL-22en6de-R512, DDT-XL-22en6de-R256
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography
arXiv2 reposarXiv:2504.07083
GenDoP, DataDoP
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
arXiv2 reposarXiv:2504.07981
UI-TARS, LocateAnything-3B
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
arXiv2 reposarXiv:2504.09421
MedFound-176B, ClinicalGPT-R1-Qwen-7B-EN-preview
SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
arXiv2 reposarXiv:2504.09644
EarthReason, SegEarth-R1
OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape Generation
arXiv2 reposarXiv:2504.09975
octgpt, OctGPT
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
arXiv2 reposarXiv:2504.10514
ColorBench, ColorBench
DataDecide: How to Predict Best Pretraining Data with Small Experiments
arXiv2 reposarXiv:2504.11393
ProX, program-every-example
ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
arXiv2 reposarXiv:2504.11536
Open-AgentRL, ReTool
arXiv:2504.13074
arXiv2 reposarXiv:2504.13074
SkyCaptioner-V1, SkyReels-V2-I2V-14B-720P
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
arXiv2 reposarXiv:2504.13125
Awesome-Chinese-LLM, Awesome-Chinese-LLM
$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
arXiv2 reposarXiv:2504.13143
Complex-Edit, Complex-Edit
St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
arXiv2 reposarXiv:2504.13152
St4RTrack, St4rTrack
Sleep-time Compute: Beyond Inference Scaling at Test-time
arXiv2 reposarXiv:2504.13171
letta-code, khms-memory
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
arXiv2 reposarXiv:2504.15271
Eagle, GR00T-N1.6-3B
Towards Understanding Camera Motions in Any Video
arXiv2 reposarXiv:2504.15376
t2v_metrics, CameraBench
TTRL: Test-Time Reinforcement Learning
arXiv2 reposarXiv:2504.16084
TTRL, EMPO
DreamO: A Unified Framework for Image Customization
arXiv2 reposarXiv:2504.16915
DreamO, dreamo
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
arXiv2 reposarXiv:2504.17432
ms-agent, modelscope-agent
HalluLens: LLM Hallucination Benchmark
arXiv2 reposarXiv:2504.17550
KoHalluLens, HalluLens
Kimi-Audio Technical Report
arXiv2 reposarXiv:2504.18425
Kimi-Audio, dissertation-project
PixelHacker: Image Inpainting with Structural and Semantic Consistency
arXiv2 reposarXiv:2504.20438
PixelHacker, PixelHacker
YoChameleon: Personalized Vision and Language Generation
arXiv2 reposarXiv:2504.20998
YoChameleon, Mini-YoChameleon-Data
Phi-4-reasoning Technical Report
arXiv2 reposarXiv:2504.21318
Phi-4-reasoning, Eureka-Bench-Logs
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
arXiv2 reposarXiv:2505.01456
UnLOK-VQA, UnLOK-VQA
A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP)
arXiv2 reposarXiv:2505.02279
guaca, awesome-agentic-payments
arXiv:2505.02387
arXiv2 reposarXiv:2505.02387
RM-R1, RM-R1-Qwen2.5-Instruct-7B
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
arXiv2 reposarXiv:2505.03318
VideoDPO, ShareGPTVideo-DPO
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
arXiv2 reposarXiv:2505.03730
FlexiAct, FlexiAct
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
arXiv2 reposarXiv:2505.05427
Ultra-FineWeb-L1, Ultra-FineWeb
Flow-GRPO: Training Flow Matching Models via Online RL
arXiv2 reposarXiv:2505.05470
HY-SOAR, GRPO
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
arXiv2 reposarXiv:2505.06371
zeus, benchmark
DanceGRPO: Unleashing GRPO on Visual Generation
arXiv2 reposarXiv:2505.07818
DanceGRPO, GRPO
HealthBench: Evaluating Large Language Models Towards Improved Human Health
arXiv2 reposarXiv:2505.08775
AntAngelMed, AntAngelMed
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
arXiv2 reposarXiv:2505.10292
QwenStoryteller, StoryReasoning
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
arXiv2 reposarXiv:2505.11049
GuardReasoner-Omni-7B, GuardReasoner-Omni-3B
BLEUBERI: BLEU is a surprisingly effective reward for instruction following
arXiv2 reposarXiv:2505.11080
BLEUBERI, BLEUBERI-Tulu3-50k
Video-GPT via Next Clip Diffusion
arXiv2 reposarXiv:2505.12489
Video-GPT, Video-GPT
CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models
arXiv2 reposarXiv:2505.12504
MM-EUREKA, CPGD-7B
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
arXiv2 reposarXiv:2505.12632
monday, MONDAY
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
arXiv2 reposarXiv:2505.13439
VTBench, VTBench
This Time is Different: An Observability Perspective on Time Series Foundation Models
arXiv2 reposarXiv:2505.14766
toto, toto
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
arXiv2 reposarXiv:2505.15404
LRM-Safety-Study, LRM-Safety-Study
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
arXiv2 reposarXiv:2505.15517
robo2VLM, Robo2VLM-1
Efficient PRM Training Data Synthesis via Formal Verification
arXiv2 reposarXiv:2505.15960
Qwen-2.5-7B-FoVer-PRM-2026, Llama-3.1-8B-FoVer-PRM-2026
FreshRetailNet-50K: A Stockout-Annotated Censored Demand Dataset for Latent Demand Recovery and Forecasting in Fresh Retail
arXiv2 reposarXiv:2505.16319
FreshRetailNet-50K, frn-50k-baseline
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
arXiv2 reposarXiv:2505.16933
LLaDA-V, LLaDA-V
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
arXiv2 reposarXiv:2505.17426
DistilCodec, DistilCodec-v1.0
MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation
arXiv2 reposarXiv:2505.17613
MMMG, MMMG
Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
arXiv2 reposarXiv:2505.17952
AlphaMed-7B-instruct-rl, AlphaMed-8B-instruct-rl
Enhancing Training Data Attribution with Representational Optimization
arXiv2 reposarXiv:2505.18513
AirRep, AirRep-Flan-Small
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
arXiv2 reposarXiv:2505.19223
LLaDA-1.5, LLaDA-1.5
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models
arXiv2 reposarXiv:2505.19743
MARA_AGENTS, MARA
What Can RL Bring to VLA Generalization? An Empirical Study
arXiv2 reposarXiv:2505.19789
RL4VLA, TTT_VLARL
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
arXiv2 reposarXiv:2505.20256
Omni-R1, Omni-R1
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
arXiv2 reposarXiv:2505.20732
SPA-RL-Agent, mobile-agent-rl
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
arXiv2 reposarXiv:2505.20793
detikzify-v2.5-8b, star-vector
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
arXiv2 reposarXiv:2505.22232
JQL-Annotation-Pipeline, JQL-Edu-Heads
ATI: Any Trajectory Instruction for Controllable Video Generation
arXiv2 reposarXiv:2505.22944
ATI, ATI
Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
arXiv2 reposarXiv:2505.22954
codegraff, darwinian_evolver
EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge
arXiv2 reposarXiv:2505.23009
higgs-tts-2-3b-base, higgs-audio-v2-generation-3B-base
FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
arXiv2 reposarXiv:2505.23145
FlowEdit, Wan2.1
TrackVLA: Embodied Visual Tracking in the Wild
arXiv2 reposarXiv:2505.23189
OmTrackVLA, TrackVLA
Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
arXiv2 reposarXiv:2505.24857
diffusion-gemma-lab, ParallelBench
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
arXiv2 reposarXiv:2505.24875
ReasonGen-R1, ReasonGen-R1-SFT
COSMIC: Generalized Refusal Direction Identification in LLM Activations
arXiv2 reposarXiv:2506.00085
OBLITERATUS, obliteratus
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
arXiv2 reposarXiv:2506.02448
ShotPlan, shotplan
ORV: 4D Occupancy-centric Robot Video Generation
arXiv2 reposarXiv:2506.03079
orv-gen-model, ORV
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
arXiv2 reposarXiv:2506.03135
SpaceQwen2.5-VL-3B-Instruct, SpaceThinker-Qwen2.5VL-3B
A Foundation Model for Spatial Proteomics
arXiv2 reposarXiv:2506.03373
CARTA, KRONOS
HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation
arXiv2 reposarXiv:2506.04421
HMAR, HMAR
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
arXiv2 reposarXiv:2506.04779
Qwen2.5-Omni, Qwen-2.5-7b
MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
arXiv2 reposarXiv:2506.05218
Monkey, MonkeyOCR
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
arXiv2 reposarXiv:2506.06211
PuzzleWorld, PuzzleWorld
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
arXiv2 reposarXiv:2506.06276
ml-starflow, starflow
CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
arXiv2 reposarXiv:2506.06290
CellCLIP, CellCLIP
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
arXiv2 reposarXiv:2506.06295
dLLM-cache, dLLM_Cache
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
arXiv2 reposarXiv:2506.06962
AR-RAG, arrag_faid
Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models
arXiv2 reposarXiv:2506.07334
GraphKV, GraphKV
Real-Time Execution of Action Chunking Flow Policies
arXiv2 reposarXiv:2506.07339
RLDX-1, RLDX-FineAct
Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse
arXiv2 reposarXiv:2506.07639
Adaptive-CoT-in-VLA, Fast-ECoT
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
arXiv2 reposarXiv:2506.07977
OneIG-Benchmark, OneIG-Bench
PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
arXiv2 reposarXiv:2506.10741
PosterCraft, Poster100K
Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
arXiv2 reposarXiv:2506.10912
ToxiMol, ToxiMol-benchmark
Towards Building General Purpose Embedding Models for Industry 4.0 Agents
arXiv2 reposarXiv:2506.12607
FailureSensorIQ, AssetOpsBench
LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D
arXiv2 reposarXiv:2506.13766
LHM-plusplus, LHM-plusplus
Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification
arXiv2 reposarXiv:2506.14074
cvdp_client, cvdp_benchmark
Sekai: A Video Dataset towards World Exploration
arXiv2 reposarXiv:2506.15675
Sekai, sekai-codebase
Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
arXiv2 reposarXiv:2506.16504
Hunyuan3D-2, qsdsd
Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
arXiv2 reposarXiv:2506.16962
Chiron-o1-8B, Chiron-o1-2B
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
arXiv2 reposarXiv:2506.17221
GPT4Scene, GPT4Scene-and-VLN-R1
RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
arXiv2 reposarXiv:2506.18088
robotwin2.0-fastwam, RoboTwin
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
arXiv2 reposarXiv:2506.18095
ShareGPT-4o-Image, Janus-4o-7B
LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
arXiv2 reposarXiv:2506.18841
LongWriter, LongWriter-Zero-32B
TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer
arXiv2 reposarXiv:2506.18904
TC-Light, TC-Light
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning
arXiv2 reposarXiv:2506.21355
smmile, SMMILE
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
arXiv2 reposarXiv:2506.22646
multitalker-parakeet-streaming-0.6b-v1, multitalker-parakeet-streaming-0.6b-v1-onnx-int8
Token Activation Map to Visually Explain Multimodal LLMs
arXiv2 reposarXiv:2506.23270
TAM, TAM
arXiv:2506.23869
arXiv2 reposarXiv:2506.23869
aria, aria-medium-base
Understanding and Improving Length Generalization in Recurrent Models
arXiv2 reposarXiv:2507.02782
AI21-Jamba-Reasoning-3B-GGUF, AI21-Jamba-Reasoning-3B
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
arXiv2 reposarXiv:2507.02813
LangScene-X, LangScene-X
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
arXiv2 reposarXiv:2507.02962
AWorld, AWorld-RL
EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
arXiv2 reposarXiv:2507.03905
EchoMimicV3, echomimic_v3
Neural-Driven Image Editing
arXiv2 reposarXiv:2507.05397
loongx, L-Mind
Scaling RL to Long Videos
arXiv2 reposarXiv:2507.07966
Long-RL, LongVILA-R1-7B
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
arXiv2 reposarXiv:2507.08801
Lumos, Lumos-1
Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
arXiv2 reposarXiv:2507.10524
mixture_of_recursions, R3-ViT
Text-ADBench: Text Anomaly Detection Benchmark Based on LLM Embeddings
arXiv2 reposarXiv:2507.12295
Text-Anomaly-Detection-Benchmark, Text-ADBench
Voxtral
arXiv2 reposarXiv:2507.13264
dissertation-project, Voxtral-Small-24B-2507
$π^3$: Permutation-Equivariant Visual Geometry Learning
arXiv2 reposarXiv:2507.13347
Pi3, Pi3X
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
arXiv2 reposarXiv:2507.13353
VideoITG, VideoITG-8B
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
arXiv2 reposarXiv:2507.16746
Zebra-CoT, Reasoning-Visual-World
Towards Robust Foundation Models for Digital Pathology
arXiv2 reposarXiv:2507.17845
PathoROB, croma
Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
arXiv2 reposarXiv:2507.19427
Step-3.5-Flash, Step-3.5-Flash
Agentic Reinforced Policy Optimization
arXiv2 reposarXiv:2507.19849
Tool-Star, Tool-Star
arXiv:2507.20534
arXiv2 reposarXiv:2507.20534
checkpoint-engine, Emerging-Optimizers
Latent Inter-User Difference Modeling for LLM Personalization
arXiv2 reposarXiv:2507.20849
DEP, DEP-model
SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
arXiv2 reposarXiv:2507.20984
SmallThinker-4BA0.6B-Instruct, SmallThinker-21BA3B-Instruct
TTS-1 Technical Report
arXiv2 reposarXiv:2507.21138
tts, Anime-XCodec2-44.1kHz-v2
Benchmarking LLMs for Unit Test Generation from Real-World Functions
arXiv2 reposarXiv:2508.00408
UnLeakedTestBench, unleakedtestbench
LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
arXiv2 reposarXiv:2508.01617
LLaDA-MedV, LLaDA-MedV
Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
arXiv2 reposarXiv:2508.03643
Uni3R, Uni3R
FlowState: Sampling-Rate-Equivariant Time-Series Forecasting
arXiv2 reposarXiv:2508.05287
flowstate, granite-timeseries-flowstate-r1
Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
arXiv2 reposarXiv:2508.07901
Stand-In, Stand-In
Ovis2.5 Technical Report
arXiv2 reposarXiv:2508.11737
Ovis2.5-9B, Ovis2.5-2B
Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
arXiv2 reposarXiv:2508.12945
Lumen, Lumen
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
arXiv2 reposarXiv:2508.14033
InfiniteTalk, modal-infinitetalk
Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration
arXiv2 reposarXiv:2508.14483
Vivid-VR, Vivid-VR
Dream 7B: Diffusion Large Language Models
arXiv2 reposarXiv:2508.15487
d3LLM, dllm
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
arXiv2 reposarXiv:2508.15601
SimSIMD, numkong
AWorld: Orchestrating the Training Recipe for Agentic AI
arXiv2 reposarXiv:2508.20404
Qwen3-32B-AWorld, AWorld
MSRS: Evaluating Multi-Source Retrieval-Augmented Generation
arXiv2 reposarXiv:2508.20867
MSRS, MSRS
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
arXiv2 reposarXiv:2508.20869
OLMoASR, OLMoASR-Pool
RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching
arXiv2 reposarXiv:2508.21258
open-r-lens, RelP
Is this chart lying to me? Automating the detection of misleading visualizations
arXiv2 reposarXiv:2508.21675
acl2026-misleading-visualizations, acl2026-misviz
CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders
arXiv2 reposarXiv:2509.00691
CE-Bench, contrastive-stories-v4
REFRAG: Rethinking RAG based Decoding
arXiv2 reposarXiv:2509.01092
ruvector, AXRU
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
arXiv2 reposarXiv:2509.02544
UI-TARS, VeOmni
Real-Time Detection of Hallucinated Entities in Long-Form Generation
arXiv2 reposarXiv:2509.03531
hallucination_probes, hallucination-probes
CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values
arXiv2 reposarXiv:2509.03740
BiomedCoOp, CLIP-SVD
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
arXiv2 reposarXiv:2509.03959
WenetSpeech-Yue, WSYue-TTS
Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
arXiv2 reposarXiv:2509.05952
MixGRPO, flow_grpo
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
arXiv2 reposarXiv:2509.06818
UMO, UMO
Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis
arXiv2 reposarXiv:2509.11526
MHIM-MIL, CPathPatchFeature
Embodied Navigation Foundation Model
arXiv2 reposarXiv:2509.12129
OmTrackVLA, OmTrackVLA-0.6B
Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
arXiv2 reposarXiv:2509.14295
AEGIS, AEGIS
Lynx: Towards High-Fidelity Personalized Video Generation
arXiv2 reposarXiv:2509.15496
lynx, lynx
Discovering Top-k Periodic and High-Utility Patterns
arXiv2 reposarXiv:2509.15732
awesome-datascience, awesome-datascience
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
arXiv2 reposarXiv:2509.17052
Sidon, sidon_raw_weight
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
arXiv2 reposarXiv:2509.18362
FastMTP, speculators
Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets
arXiv2 reposarXiv:2509.21245
HY3D-Bench, Hunyuan3D-Omni
LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer
arXiv2 reposarXiv:2509.22414
LucidFlux, LucidFlux
Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
arXiv2 reposarXiv:2509.25052
deepscaler, rllm
Fine-Grained GRPO for Precise Preference Alignment in Flow Models
arXiv2 reposarXiv:2510.01982
Granular-GRPO, G2RPO
xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity
arXiv2 reposarXiv:2510.02228
xlstm_scaling_laws, xlstm_scaling_laws
VideoNSA: Native Sparse Attention Scales Video Understanding
arXiv2 reposarXiv:2510.02295
VideoNSA, VideoNSA
EditLens: Quantifying the Extent of AI Editing in Text
arXiv2 reposarXiv:2510.03154
sloptotal, EditLens
ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
arXiv2 reposarXiv:2510.04706
Arc2Face, Arc2Face
Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
arXiv2 reposarXiv:2510.04786
ttc, verifiable-corpus
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
arXiv2 reposarXiv:2510.06308
Lumina-DiMOO, Lumina-DiMOO
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
arXiv2 reposarXiv:2510.09507
PhysToolBench, PhysToolBench
Don't Just Fine-tune the Agent, Tune the Environment
arXiv2 reposarXiv:2510.10197
AWorld, AWorld-RL
Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning
arXiv2 reposarXiv:2510.10973
chart-rvr-3b, chart-rvr-hard-3b
arXiv:2510.12747
arXiv2 reposarXiv:2510.12747
FlashVSR, FlashVSR-v1.1
UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
arXiv2 reposarXiv:2510.13745
UniCalli, UniCalli_Dev
Agentic Entropy-Balanced Policy Optimization
arXiv2 reposarXiv:2510.14545
Tool-Star, Tool-Star
WithAnyone: Towards Controllable and ID Consistent Image Generation
arXiv2 reposarXiv:2510.14975
WithAnyone, MultiID-Bench
UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis
arXiv2 reposarXiv:2510.15710
UniMedVL, UniMedVL
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
arXiv2 reposarXiv:2510.16888
UniWorld, UniWorld-V1
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
arXiv2 reposarXiv:2510.17532
Clinical-Reasoning-LLMs, cancer-reasoning-traces
From Charts to Code: A Hierarchical Benchmark for Multimodal Models
arXiv2 reposarXiv:2510.17932
Chart2Code, Chart2Code
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
arXiv2 reposarXiv:2510.20064
hedgespec, hedgespec_eagle_drafters
Generative Reasoning Recommendation via LLMs
arXiv2 reposarXiv:2510.20815
GRRM, GREAM_data
Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation
arXiv2 reposarXiv:2510.22115
Ling-1T, ArtifactsBenchmark
LongCat-Video Technical Report
arXiv2 reposarXiv:2510.22200
LongCat-Video, LongCat-Video
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
arXiv2 reposarXiv:2510.22319
flow_grpo, GRPO
LimRank: Less is More for Reasoning-Intensive Information Reranking
arXiv2 reposarXiv:2510.23544
limrank, limrank
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
arXiv2 reposarXiv:2510.23607
Concerto, Concerto
FullPart: Generating each 3D Part at Full Resolution
arXiv2 reposarXiv:2510.26140
fullpart, partversexl
Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification
arXiv2 reposarXiv:2510.26777
TiRex, tirex
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
arXiv2 reposarXiv:2511.00916
Fleming-VL-8B, Fleming-VL-38B
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
arXiv2 reposarXiv:2511.04727
IndicVisionBench, IndicVisionBench
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
arXiv2 reposarXiv:2511.05489
TimeSearch-R, TimeSearch-R
Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering
arXiv2 reposarXiv:2511.10900
EMS-MCQA, EMS-Knowledge
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
arXiv2 reposarXiv:2511.11139
SAP2-ASR, SAP2-ASR
Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection
arXiv2 reposarXiv:2511.13027
Skills, NeMo-Skills
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
arXiv2 reposarXiv:2511.14439
AntAngelMed, AntAngelMed
$π^{*}_{0.6}$: a VLA That Learns From Experience
arXiv2 reposarXiv:2511.14759
open-value, lerobot
arXiv:2511.15186
arXiv2 reposarXiv:2511.15186
ROSALIA, ROSALIA-7B-v1
arXiv:2511.15684
arXiv2 reposarXiv:2511.15684
walrus, walrus
SAM 3: Segment Anything with Concepts
arXiv2 reposarXiv:2511.16719
geti-instant-learn, rf-detr
ChineseErrorCorrector3-4B: State-of-the-Art Chinese Spelling and Grammar Corrector
arXiv2 reposarXiv:2511.17562
ChineseErrorCorrector3-4B, ChineseErrorCorrector
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
arXiv2 reposarXiv:2511.18050
UltraFlux-v1, UltraFlux-v1-1-Transformer
MedVision: Benchmarking Quantitative Medical Image Analysis
arXiv2 reposarXiv:2511.18676
MedVision, MedVision
Eevee: Towards Close-up High-resolution Video-based Virtual Try-on
arXiv2 reposarXiv:2511.18957
Eevee, Eevee
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
arXiv2 reposarXiv:2511.20462
ml-starflow, starflow
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
arXiv2 reposarXiv:2511.21475
MobileI2V, MobileI2V
ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
arXiv2 reposarXiv:2511.22625
Step1X-Edit, Step1X-Edit-v1p2
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
arXiv2 reposarXiv:2511.22663
AIA, AIA
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
arXiv2 reposarXiv:2511.22699
Z-Image-Turbo, Z-Image
PanFlow: Decoupled Motion Control for Panoramic Video Generation
arXiv2 reposarXiv:2512.00832
PanFlow, PanFlow
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
arXiv2 reposarXiv:2512.00891
VidCom2, STC
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
arXiv2 reposarXiv:2512.01342
InternVideo, internvideo-d2a11ea9
Improved Mean Flows: On the Challenges of Fastforward Generative Models
arXiv2 reposarXiv:2512.02012
iMF-diffusers, imeanflow
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
arXiv2 reposarXiv:2512.03514
ColNetraEmbed, NetraEmbed
Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning
arXiv2 reposarXiv:2512.03667
Project-Imaging-X, VPS
SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs
arXiv2 reposarXiv:2512.04746
auto-round, Qwen3.8-27B-bpw2.8-AutoRound
Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models
arXiv2 reposarXiv:2512.04981
fairpro, fairpro
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
arXiv2 reposarXiv:2512.05091
Sa2VA, Sa2VA
Training-Time Action Conditioning for Efficient Real-Time Chunking
arXiv2 reposarXiv:2512.05964
RLDX-1, RLDX-FineAct
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
arXiv2 reposarXiv:2512.06688
PersonaMem-v2, PersonaMem-v2
Unified Camera Positional Encoding for Controlled Video Generation
arXiv2 reposarXiv:2512.07237
PanShot, prope
LongCat-Image Technical Report
arXiv2 reposarXiv:2512.07584
LongCat-Image-Edit-Turbo, GenAI-Caption-Pipeline
InfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain Generation
arXiv2 reposarXiv:2512.08309
terrain-diffusion-30m, terrain-diffusion
Deterministic and Exact Fully-dynamic Minimum Cut of Superpolylogarithmic Size in Subpolynomial Time
arXiv2 reposarXiv:2512.13105
AXRU, ruvector
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
arXiv2 reposarXiv:2512.13604
LongVie, LongVie2
RePo: Language Models with Context Re-Positioning
arXiv2 reposarXiv:2512.14391
repo, RePo-OLMo2-1B-stage2-L5
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
arXiv2 reposarXiv:2512.16378
hearing2translate, hearing2translate-humeval
Gabliteration: Adaptive Multi-Directional Neural Weight Modification for Selective Behavioral Alteration in Large Language Models
arXiv2 reposarXiv:2512.18901
OBLITERATUS, obliteratus
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
arXiv2 reposarXiv:2512.20848
nugie-jax-nemotron-3-nano, NVIDIA-Nemotron-3-Nano-4B-GGUF
dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning
arXiv2 reposarXiv:2512.21446
dUltra-os, dUltra-math-b128
DiRL: An Efficient Post-Training Framework for Diffusion Language Models
arXiv2 reposarXiv:2512.22234
DiRL, DiRL-8B-Instruct
TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
arXiv2 reposarXiv:2512.23065
TabiBERT, Tabibert
SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation
arXiv2 reposarXiv:2512.23379
SoulX-FlashTalk, SoulX-FlashTalk-14B
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
arXiv2 reposarXiv:2512.23705
DKT, TransPhy3D
DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
arXiv2 reposarXiv:2601.01528
DrivingGen, DrivingGen
Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications
arXiv2 reposarXiv:2601.01718
Yuan3.0-Flash, Yuan3.0-Flash-4bit
K-EXAONE Technical Report
arXiv2 reposarXiv:2601.01739
K-EXAONE-236B-A23B, K-EXAONE
Pearmut: Human Evaluation of Translation Made Trivial
arXiv2 reposarXiv:2601.02933
hearing2translate-humeval, pearmut
LTX-2: Efficient Joint Audio-Visual Foundation Model
arXiv2 reposarXiv:2601.03233
LTX-2, LTX-2
MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents
arXiv2 reposarXiv:2601.03236
rlm-claude, memcp
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
arXiv2 reposarXiv:2601.05138
VerseCrafter, VerseCrafter
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
arXiv2 reposarXiv:2601.05593
Step-3.5-Flash, Step-3.5-Flash
Affostruction: 3D Affordance Grounding with Generative Reconstruction
arXiv2 reposarXiv:2601.09211
Affostruction, Affostruction
HeartMuLa: A Family of Open Sourced Music Foundation Models
arXiv2 reposarXiv:2601.10547
heartlib, HeartMuLa_ComfyUI
ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation
arXiv2 reposarXiv:2601.12983
acl2026-misleading-visualizations, acl2026-misviz
CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning
arXiv2 reposarXiv:2601.13262
cure-med, CUREMED-BENCH
Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics
arXiv2 reposarXiv:2601.14027
open-atp, lean-lsp-mcp
SAMTok: Representing Any Mask with Two Words
arXiv2 reposarXiv:2601.16093
Sa2VA, Sa2VA
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
arXiv2 reposarXiv:2601.16148
ActionMesh, actionbench
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
arXiv2 reposarXiv:2601.16208
scale-rae-data, Scale-RAE
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
arXiv2 reposarXiv:2601.19194
TS-ASR-Whisper, DiCoW_v3_3_large
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
arXiv2 reposarXiv:2601.19506
Pref_Restore, Pref-Restore-PhaseA-Fidelity
Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
arXiv2 reposarXiv:2601.19834
VisWorld-Eval, Reasoning-Visual-World
One-step Latent-free Image Generation with Pixel Mean Flows
arXiv2 reposarXiv:2601.22158
pMF-diffusers, imeanflow
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
arXiv2 reposarXiv:2601.22599
FlowSep-hive, AudioSep-hive
arXiv:2601.22710
arXiv2 reposarXiv:2601.22710
AlienLM, llama3-8b-instruct-alienlm-full
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
arXiv2 reposarXiv:2601.23161
DIFFA, DIFFA-2
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
arXiv2 reposarXiv:2602.01025
UltraBreak, UltraBreak-Repro
P-EAGLE: Parallel-Drafting EAGLE with Scalable Training
arXiv2 reposarXiv:2602.01469
SpecForge, Qwen3-8B-speculator.peagle
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
arXiv2 reposarXiv:2602.02537
WorldVQA, WorldVQA
InfMem: Learning System-2 Memory Control for Long-Context Agent
arXiv2 reposarXiv:2602.02704
InfMem, infmem_superlong
SWE-World: Building Software Engineering Agents in Docker-Free Environments
arXiv2 reposarXiv:2602.03419
SWE-Master, SWE-World
See-through: Single-image Layer Decomposition for Anime Characters
arXiv2 reposarXiv:2602.03749
see-through-demo, ComfyUI-See-through
Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval
arXiv2 reposarXiv:2602.03992
vllm-factory, nemotron-colembed-vl-4b-v2
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
arXiv2 reposarXiv:2602.05176
model_collaboration, AmongUs
Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models
arXiv2 reposarXiv:2602.06909
patchtst-fm-r1, granite-timeseries-patchtst-fm-r1
VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning
arXiv2 reposarXiv:2602.08828
Veritas, VideoVeritas
Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters
arXiv2 reposarXiv:2602.10604
Step-3.5-Flash, Step-3.5-Flash
Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
arXiv2 reposarXiv:2602.11149
data-repetition, olmo3-7b_data-repetition
arXiv:2602.11910
arXiv2 reposarXiv:2602.11910
steer-audio, patching-music-musiccaps-prompts
GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics
arXiv2 reposarXiv:2602.12617
GeoAgent, GeoAgent
Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges
arXiv2 reposarXiv:2602.13576
Rubrics-as-an-Attack-Surface, ripd-dataset
Experiential Reinforcement Learning
arXiv2 reposarXiv:2602.13949
deepscaler, rllm
VLANeXt: Recipes for Building Strong VLA Models
arXiv2 reposarXiv:2602.18532
VLANeXt, VLANeXt
WildOS: Open-Vocabulary Object Search in the Wild
arXiv2 reposarXiv:2602.19308
nebula2-wildos, wildos
arXiv:2602.20113
arXiv2 reposarXiv:2602.20113
StyleStream, StyleStream
PreScience: A Dataset and Benchmark for Scientific Forecasting
arXiv2 reposarXiv:2602.20459
prescience, prescience
D-FINE-seg: Object Detection and Instance Segmentation Framework with multi-backend deployment
arXiv2 reposarXiv:2602.23043
D-FINE-seg, D-FINE-seg
SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching
arXiv2 reposarXiv:2602.24208
maxdiffusion, ltx2-vidgen-skill
LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model
arXiv2 reposarXiv:2603.01068
LLaDA-o, LLaDA-V
According to Me: Long-Term Personalized Referential Memory QA
arXiv2 reposarXiv:2603.01990
ATM-Bench, ATM-Bench
Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training
arXiv2 reposarXiv:2603.02208
reasoning-core, reasoning_core
LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
arXiv2 reposarXiv:2603.03269
LoGeR, LoGeR
Utonia: Toward One Encoder for All Point Clouds
arXiv2 reposarXiv:2603.03283
Utonia, Utonia
$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
arXiv2 reposarXiv:2603.04304
deepscaler, rllm
MSpoofTTS: Multi-Resolution Spoof-Guided Inference for Discrete Speech Synthesis
arXiv2 reposarXiv:2603.05373
MSpoofTTS, MSpoofTTS
Scalable Training of Mixture-of-Experts Models with Megatron Core
arXiv2 reposarXiv:2603.07685
Megatron-LM, geodesic-megatron
DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
arXiv2 reposarXiv:2603.08216
dualturn, dualturn-qwen2.5-mimi-0.5B
LLM2Vec-Gen: Generative Embeddings from Large Language Models
arXiv2 reposarXiv:2603.10913
llm2vec, LLM2Vec-Gen-Qwen3-8B
arXiv:2603.11661
arXiv2 reposarXiv:2603.11661
Resonate, Resonate
Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
arXiv2 reposarXiv:2603.12217
track_on, track_on_r
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization
arXiv2 reposarXiv:2603.13375
InfiniteDance, InfiniteDance
GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis
arXiv2 reposarXiv:2603.14965
GeoNVS, GeoNVS
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
arXiv2 reposarXiv:2603.15118
varex-bench, VAREX
Gym-V: A Unified Vision Environment System for Agentic Vision Research
arXiv2 reposarXiv:2603.15432
Game-RL, GameQA-140K
SegviGen: Repurposing 3D Generative Model for Part Segmentation
arXiv2 reposarXiv:2603.16869
SegviGen, SegviGen
Beyond String Matching: Semantic Evaluation of PDF Table Extraction
arXiv2 reposarXiv:2603.18652
pdf-parse-bench, pdf-parse-bench
Breeze Taigi: Benchmarks and Models for Taiwanese Hokkien Speech Recognition and Synthesis
arXiv2 reposarXiv:2603.19259
Breeze-ASR-26, faster-whisper-Breeze-ASR-26
LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
arXiv2 reposarXiv:2603.19312
mlx-tune, lewm-pusht
The Universal Normal Embedding
arXiv2 reposarXiv:2603.21786
UNE, NoiseZoo
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
arXiv2 reposarXiv:2603.22016
ROM, ROM
MemDLM: Memory-Enhanced DLM Training
arXiv2 reposarXiv:2603.22241
LLaDA-MoE-7B-A1B-Base-MemDLM, LLaDA2.1-mini-MemDLM
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
arXiv2 reposarXiv:2603.23516
MSA, MSA-4B
SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks
arXiv2 reposarXiv:2603.24755
unlazy, scb-check
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
arXiv2 reposarXiv:2603.26008
FairLLaVA, FairLLaVA
SonoWorld: From One Image to a 3D Audio-Visual Scene
arXiv2 reposarXiv:2603.28757
sonoworld, SonoScene360
Generative World Renderer
arXiv2 reposarXiv:2604.02329
AlayaRenderer, AlayaRenderer
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
arXiv2 reposarXiv:2604.03016
Agentic-MME, Agentic-MME
TORA: Topological Representation Alignment for 3D Shape Assembly
arXiv2 reposarXiv:2604.04050
tora, tora
Synthetic Sandbox for Training Machine Learning Engineering Agents
arXiv2 reposarXiv:2604.04872
deepscaler, rllm
Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents
arXiv2 reposarXiv:2604.04979
tool-output-extraction-swebench, squeez-2b
Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space
arXiv2 reposarXiv:2604.05030
npcpy, qllm2
The Art of Building Verifiers for Computer Use Agents
arXiv2 reposarXiv:2604.06240
CUAVerifierBench, WebTailBench
Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM
arXiv2 reposarXiv:2604.06832
Fast-dLLM, Fast_dVLM_3B
PhysInOne: Visual Physics Learning and Reasoning in One Suite
arXiv2 reposarXiv:2604.09415
PhysInOne, PhysInOne
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
arXiv2 reposarXiv:2604.10708
Audio-Omni, Audio-Omni
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
arXiv2 reposarXiv:2604.11804
OmniShow, HOIVG-Bench
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
arXiv2 reposarXiv:2604.13016
MiniCPM5-1B-GGUF, MiniCPM5-1B
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
arXiv2 reposarXiv:2604.18486
OneVL_training, onevl
SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing
arXiv2 reposarXiv:2604.19587
SmartPhotoCrafter, SmartPhotoCrafter
Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers
arXiv2 reposarXiv:2604.21592
Sculpt4D, Sculpt4D
Building a Precise Video Language with Human-AI Oversight
arXiv2 reposarXiv:2604.21718
t2v_metrics, CHAI_testset
CADFit: Precise Mesh-to-CAD Program Generation with Hybrid Optimization
arXiv2 reposarXiv:2605.01171
CADFit, CADFit
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
arXiv2 reposarXiv:2605.04128
JoyAI-Image, JoyAI-Image-Edit
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
arXiv2 reposarXiv:2605.06068
vibesys, vibe-train-sample
Relit-LiVE: Relight Video by Jointly Learning Environment Video
arXiv2 reposarXiv:2605.06658
Relit-LiVE, Relit-LiVE
Pixal3D: Pixel-Aligned 3D Generation from Images
arXiv2 reposarXiv:2605.10922
ComfyUI-Pixal3D, Pixal3D
Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning
arXiv2 reposarXiv:2605.14386
Darwin-TTS-1.7B-Cross-Qwen3Tokenizer, Darwin-TTS-1.7B-Cross
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
arXiv2 reposarXiv:2605.17757
OSCAR, OSCAR-RotationZoo
PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis
arXiv2 reposarXiv:2605.17916
PanoWorld, PanoWorld
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
arXiv2 reposarXiv:2605.20179
TIDE, TIDE_DATA_COLLECTION
AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
arXiv2 reposarXiv:2605.22715
AnyMo, AnyMo-Bench
Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving
arXiv2 reposarXiv:2605.23163
Fast-dDrive, Fast-dLLM
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
arXiv2 reposarXiv:2605.23904
darwin-skill, SkillOpt
Raon-Speech Technical Report
arXiv2 reposarXiv:2605.23912
Raon-Speech-9B, Raon-SpeechChat-9B
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
arXiv2 reposarXiv:2605.27365
LocateAnything-3B, LocateAnything
Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation
arXiv2 reposarXiv:2605.28642
ESRT-4B, esrt
DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions
arXiv2 reposarXiv:2605.31271
DriveMA-2B, DriveMA_Datasets
MOSS-Audio Technical Report
arXiv2 reposarXiv:2606.01802
MOSS-Audio, MOSS-audio-swift
Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
arXiv2 reposarXiv:2606.02373
harness-1, harness-1
TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering
arXiv2 reposarXiv:2606.02624
TadA-Bench, TadA-Bench
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
arXiv2 reposarXiv:2606.09079
FlashMemory-Deepseek-V4, FlashMemory-Deepseek-V4
DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction
arXiv2 reposarXiv:2606.09186
DuplexOmni, DuplexOmni-Data
Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
arXiv2 reposarXiv:2606.12886
Game-RL, GameQA-140K
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
arXiv2 reposarXiv:2606.14516
every_eval_ever, EEE_datastore
ReportQA: QA-Based Radiology Report Evaluation
arXiv2 reposarXiv:2606.15037
ReportQA, ReportQA
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
arXiv2 reposarXiv:2606.19047
AWorld-RL, Qwen3-4B-RODS
PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction
arXiv2 reposarXiv:2606.19096
amalia-vl-eval, PorTEXTO
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
arXiv2 reposarXiv:2606.19195
Moebius, Moebius
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
arXiv2 reposarXiv:2606.19348
DeepSeek-V4-Flash-0731, DeepSeek-V4-Pro
Improved Large Language Diffusion Models
arXiv2 reposarXiv:2606.25331
iLLaDA-8B-Base, iLLaDA-8B-Instruct
Frequency-Aware Self-Supervised Music Representation Learning
arXiv2 reposarXiv:2606.25713
MERT-v2-30s, MERT-v2-FullSong
Orca: The World is in Your Mind
arXiv2 reposarXiv:2606.30534
Orca, Orca-4B
CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes
arXiv2 reposarXiv:2606.31435
data-juicer-hub, CDR-Bench
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
arXiv2 reposarXiv:2606.31986
CoLT, CoLT-8B
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
arXiv2 reposarXiv:2607.04064
speaker_disentangled_hubert, SylReg-LM-7B
HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
arXiv2 reposarXiv:2607.04884
HunyuanOCR, HunyuanOCR
ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching
arXiv2 reposarXiv:2607.07119
ColorFM, ColorFM
Persona Cartography: Charting Language Model Personality Traits in Weight Space
arXiv2 reposarXiv:2607.07916
persona-cartography, monorepo
MuScriptor: An Open Model for Multi-Instrument Music Transcription
arXiv2 reposarXiv:2607.08168
qinglong-captions, HOT-Step-CPP
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
arXiv2 reposarXiv:2607.08716
deepscaler, rllm
MonkeyOCRv2: A Visual-Text Foundation Model for Document AI
arXiv2 reposarXiv:2607.11562
MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash
Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate
arXiv2 reposarXiv:2607.12266
Qwen3.8-27B-K4, Qwen3.8-27B-EXL3-K5K6
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
arXiv2 reposarXiv:2607.14935
VideoChat-Flash, Ask-Anything
ShotPlan: Cinematic Video Generation with Learnable Planning Token
arXiv2 reposarXiv:2607.17675
ShotPlan-Wan2.2-T2V-A14B-HighNoise, ShotPlan-Wan2.1-T2V-14B
A Controlled Study of Attention-Only Transformers
arXiv2 reposarXiv:2607.18363
mimimodel, needle
Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
arXiv2 reposarXiv:2607.18934
CrisperWhisper, CrisperWhisper2.0_large
AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
arXiv2 reposarXiv:2607.19223
AdaFlash, Qwen3-8B-AdaFlash
VibeVoice-ASR-BitNet Technical Report
arXiv2 reposarXiv:2607.21075
VibeVoice, VibeVoice-ASR-BitNet
ID-V2V: Identity-Preserving Video Restylization
arXiv2 reposarXiv:2607.22830
ID-V2V, ID-V2V
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
arXiv2 reposarXiv:2607.23855
OmniVAE, OmniVAE
AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
arXiv2 reposarXiv:2607.25852
TorchSpec, TorchSpec
Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection
arXiv2 reposarXiv:2607.27113
Veritas, VideoVeritas
Beacon: Knowing When and How to Perform Agentic Visual Reasoning
arXiv2 reposarXiv:2607.28595
Game-RL, GameQA-140K
PhiZero: A World Model Built Around Physical Language
arXiv2 reposarXiv:2607.28624
PhiZero, PhiZero
Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection
arXiv2 reposarXiv:2608.00207
Adaptation, TLoRA-Adaptation
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
arXiv2 reposarXiv:2608.02673
dots.tts, dots.tts.edit
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
arXiv2 reposarXiv:2608.04205
MatrAIx-Persona-8B, MatrAIx_Persona_1M
DarwinX: Evolving Agent Harnesses Through Natural Selection
arXiv2 reposarXiv:2608.07545
Beagle, darwinx
Instruction-Based Video Editing by Repurposing an Image Editing Model
arXiv2 reposarXiv:2608.14790
Qwen-Video-Edit, Qwen-Video-Edit
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
arXiv2 reposarXiv:2608.16157
FreeToken, sparklab
FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors
arXiv2 reposarXiv:2608.23549
fix-anything, fix-anything
SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking
arXiv2 reposarXiv:2609.03047
shelf-benchmark, SHELF
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
arXiv2 reposarXiv:2609.05405
WearableQA, WearableQA
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
arXiv2 reposarXiv:2609.08936
AuK, AuK-Flash
Essential Incompleteness of Arithmetic Verified by Coq
arXiv2 reposarXiv:cs/0505034
hydra-battles, hydra-battles
MAPS: pathologist-level cell type annotation from tissue images through machine learning
Nature2 reposNature:s41467-023-44188-w
CORAL, KRONOS2
GrandQC: A comprehensive solution to quality control problem in digital pathology
Nature2 reposNature:s41467-024-54769-y
TRIDENT, TridentEdited
Explainable machine-learning predictions for the prevention of hypoxaemia during surgery
Nature2 reposNature:s41551-018-0304-0
shap, shap
Data-efficient and weakly supervised computational pathology on whole-slide images
Nature2 reposNature:s41551-020-00682-w
CLAM, Mussel
Large language models encode clinical knowledge
Nature2 reposNature:s41586-023-06291-2
Awesome-Medical-Dataset, healthsearchqa
World and Human Action Models towards gameplay ideation
Nature2 reposNature:s41586-025-08600-3
Awesome-From-Video-Generation-to-World-Model, wham
The Virtual Tissues foundation model resolves spatial proteomics across scales
Nature2 reposNature:s41586-026-10884-y
virtues, virtues
A multimodal whole-slide foundation model for pathology
Nature2 reposNature:s41591-025-03982-3
AtlasPatch, TITAN
NeuroMechFly v2: simulating embodied sensorimotor control in adult Drosophila
Nature2 reposNature:s41592-024-02497-y
flygym, fly
From whole-slide image to biomarker prediction: end-to-end weakly supervised deep learning in computational pathology
Nature2 reposNature:s41596-024-01047-2
STAMP, STAMP_attention_ui
HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy
Nature2 reposNature:s41597-020-00622-y
Endo-FM, hyper-kvasir
A high-resolution temporal transcriptomic and imaging dataset of porcine wound healing
Nature2 reposNature:s41597-025-05921-w
wound-forecasting, porcine-wound-forecasting-processed
From local explanations to global understanding with explainable AI for trees
Nature2 reposNature:s42256-019-0138-9
shap, shap
A Conversation with Mary Lou Jepsen
ACM1 repoACM:1331287.1331291
Low-power-E-Paper-OS
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
ACM1 repoACM:1571941.1572114
ogham-mcp
OXenstored
ACM1 repoACM:1631687.1596581
irmin
Printing floating-point numbers quickly and accurately with integers
ACM1 repoACM:1806596.1806623
godot
Manifold bootstrapping for SVBRDF capture
ACM1 repoACM:1833349.1778835
Awesome-InverseRendering
Learning to efficiently rank
ACM1 repoACM:1835449.1835475
RLT4Reranking
A standardised polarisation visualisation for images
ACM1 repoACM:1925059.1925070
Awesome-Polarization
A mixed reality system for virtual glasses try-on
ACM1 repoACM:2087756.2087816
awesome-virtual-try-on
Virtual try-on of eyeglasses using 3D model of the head
ACM1 repoACM:2087756.2087838
awesome-virtual-try-on
Minimizing row displacement dispatch tables
ACM1 repoACM:217839.217851
sdk
Printing floating-point numbers quickly and accurately
ACM1 repoACM:231379.231397
godot
Beyond random walk and metropolis-hastings samplers
ACM1 repoACM:2318857.2254795
littleballoffur
A volumetric method for building complex models from range images
ACM1 repoACM:237170.237269
awesome-mvs
Practical SVBRDF capture in the frequency domain
ACM1 repoACM:2461912.2461978
Awesome-InverseRendering
Metric convergence in social network sampling
ACM1 repoACM:2491159.2491168
littleballoffur
Network Sampling
ACM1 repoACM:2601438
littleballoffur
Garment Replacement in Monocular Video Sequences
ACM1 repoACM:2634212
awesome-virtual-try-on
ESC
ACM1 repoACM:2733373.2806390
BEANS-Zero
Two-shot SVBRDF capture for stationary materials
ACM1 repoACM:2766967
Awesome-InverseRendering
Learning to Represent Knowledge Graphs with Gaussian Embedding
ACM1 repoACM:2806416.2806502
pykeen
The use of MMR, diversity-based reranking for reordering documents and producing summaries
ACM1 repoACM:290941.291025
ogham-mcp
Asymmetric Transitivity Preserving Graph Embedding
ACM1 repoACM:2939672.2939751
karateclub
ACM:296806.296824
ACM1 repoACM:296806.296824
szl-mesh
KickStarter
ACM1 repoACM:3037697.3037748
awesome-dynamic-graphs
GPU Virtualization and Scheduling Methods
ACM1 repoACM:3068281
awesome-gpu-engineering
Anserini
ACM1 repoACM:3077136.3080721
anserini
Selfie and the basics
ACM1 repoACM:3133850.3133857
selfie
Verifying strong eventual consistency in distributed systems
ACM1 repoACM:3133933
awesome-local-first
FreeGuard
ACM1 repoACM:3133956.3133957
mimalloc-bench
Random sampling with a reservoir
ACM1 repoACM:3147.3165
qsv
TinyLFU
ACM1 repoACM:3149371
caffeine
ACM:3164135.3164139
ACM1 repoACM:3164135.3164139
awesome-dynamic-graphs
On optimistic methods for concurrency control
ACM1 repoACM:319566.319567
miniflare
Graphtides: a framework for evaluating stream-based graph processing platforms
ACM1 repoACM:3210259.3210262
awesome-dynamic-graphs
Partisan
ACM1 repoACM:3231104.3231106
partisan
Plan3D
ACM1 repoACM:3233794
awesome-mvs
Anserini
ACM1 repoACM:3239571
anserini
Adaptive Software Cache Management
ACM1 repoACM:3274808.3274816
caffeine
Software multiplexing: share your libraries and statically link them too
ACM1 repoACM:3276524
linux_distro_tests
GraphBolt
ACM1 repoACM:3302424.3303974
awesome-dynamic-graphs
Learning to Represent the Evolution of Dynamic Graphs with Recurrent Models
ACM1 repoACM:3308560.3316581
pytorch_geometric_temporal
Gen: a general-purpose probabilistic programming system with programmable inference
ACM1 repoACM:3314221.3314642
Gen.jl
PASE: PostgreSQL Ultra-High-Dimensional Approximate Nearest Neighbor Search Extension
ACM1 repoACM:3318464.3386131
pgvector
Deeper Text Understanding for IR with Contextual Neural Language Modeling
ACM1 repoACM:3331184.3331303
RLT4Reranking
Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations
ACM1 repoACM:3336191.3371856
pypantera
Massively Parallel ANS Decoding on GPUs
ACM1 repoACM:3337821.3337888
dietgpu
Enhancing Graph Neural Network-based Fraud Detectors against Camouflaged Fraudsters
ACM1 repoACM:3340531.3411903
antifraud
ElderReact: A Multimodal Dataset for Recognizing Emotional Response in Aging Adults
ACM1 repoACM:3340555.3353747
ElderReact
CATSLU: The 1st Chinese Audio-Textual Spoken Language Understanding Challenge
ACM1 repoACM:3340555.3356098
ChineseNLPCorpus
Cubical agda: a dependently typed programming language with univalence and higher inductive types
ACM1 repoACM:3341691
cubical
An Assumption-Free Approach to the Dynamic Truncation of Ranked Lists
ACM1 repoACM:3341981.3344234
RLT4Reranking
Virtually Trying on New Clothing with Arbitrary Poses
ACM1 repoACM:3343031.3350946
awesome-virtual-try-on
A Framework for Effective Known-item Search in Video
ACM1 repoACM:3343031.3351046
TransNetV2
FashionOn
ACM1 repoACM:3343031.3351075
awesome-virtual-try-on
G raph O ne
ACM1 repoACM:3364180
awesome-dynamic-graphs
Extracting Knowledge from Web Text with Monte Carlo Tree Search
ACM1 repoACM:3366423.3380010
awesome-monte-carlo-tree-search-papers
Overview of the HASOC track at FIRE 2019
ACM1 repoACM:3368567.3368584
Tutorial-Resources
An equational theory for weak bisimulation via generalized parameterized coinduction
ACM1 repoACM:3372885.3373813
paco
The Deceptive Potential of Common Design Tactics Used in Data Visualizations
ACM1 repoACM:3380851.3416762
acl2026-misleading-visualizations
A robust and flexible operating system compatibility architecture
ACM1 repoACM:3381052.3381327
kerla
IMACS - an <u>i</u>nteractive cognitive assistant <u>m</u>odule for <u>c</u>ardiac <u>a</u>rrest cases in emergency medical <u>s</u>ervice
ACM1 repoACM:3384419.3430451
EMS-Pipeline
Down to the Last Detail
ACM1 repoACM:3394171.3413514
awesome-virtual-try-on
VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments Generation
ACM1 repoACM:3394171.3413890
VideoIC
Geodesic Forests
ACM1 repoACM:3394486.3403094
awesome-decision-tree-papers
Predicting Temporal Sets with Deep Neural Networks
ACM1 repoACM:3394486.3403152
pytorch_geometric_temporal
Choppy: Cut Transformer for Ranked List Truncation
ACM1 repoACM:3397271.3401188
RLT4Reranking
Evidence Weighted Tree Ensembles for Text Classification
ACM1 repoACM:3397271.3401229
awesome-decision-tree-papers
Huffman Coding with Gap Arrays for GPU Acceleration
ACM1 repoACM:3404397.3404429
dietgpu
Propensity-scored Probabilistic Label Trees
ACM1 repoACM:3404835.3463084
napkinXC
Tools, Tricks, and Hacks: Exploring Novel Digital Fabrication Workflows on #PlotterTwitter
ACM1 repoACM:3411764.3445653
awesome-plotters
Offsite aerial path planning for efficient urban scene reconstruction
ACM1 repoACM:3414685.3417791
awesome-mvs
Single image portrait relighting via explicit multiple reflectance channel modeling
ACM1 repoACM:3414685.3417824
Awesome-InverseRendering
Studying Politeness across Cultures using English Twitter and Mandarin Weibo
ACM1 repoACM:3415190
funNLP
Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic Regression
ACM1 repoACM:3447548.3467278
awesome-decision-tree-papers
BLOCKSET (Block-Aligned Serialized Trees)
ACM1 repoACM:3447548.3467368
awesome-decision-tree-papers
ControlBurn
ACM1 repoACM:3447548.3467387
awesome-decision-tree-papers
Unikraft
ACM1 repoACM:3447786.3456248
awesome-os
Retrofitting effect handlers onto OCaml
ACM1 repoACM:3453483.3454039
effects-examples
Efficient large-scale language model training on GPU clusters using megatron-LM
ACM1 repoACM:3458817.3476209
awesome-gpu-engineering
Learning to Pack
ACM1 repoACM:3459637.3481933
awesome-monte-carlo-tree-search-papers
Geometric Heuristics for Transfer Learning in Decision Trees
ACM1 repoACM:3459637.3482259
awesome-decision-tree-papers
Fairness-Aware Training of Decision Trees by Abstract Interpretation
ACM1 repoACM:3459637.3482342
awesome-decision-tree-papers
Unsupervised Domain Adaptation for Static Malware Detection based on Gradient Boosting Trees
ACM1 repoACM:3459637.3482400
awesome-gradient-boosting-papers
BNN
ACM1 repoACM:3459637.3482414
awesome-gradient-boosting-papers
CrossVul: a cross-language vulnerability dataset with commit data
ACM1 repoACM:3468264.3473122
FinalYearProject
Marcelle: Composing Interactive Machine Learning Workflows and Interfaces
ACM1 repoACM:3472749.3474734
IR-Lens
Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning
ACM1 repoACM:3472749.3474765
RefinedVision
Relevance under the Iceberg
ACM1 repoACM:3477495.3531767
pecos
From Distillation to Hard Negative Sampling
ACM1 repoACM:3477495.3531857
Rankify
InPars: Unsupervised Dataset Generation for Information Retrieval
ACM1 repoACM:3477495.3531863
InPars
Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive Training
ACM1 repoACM:3477495.3531867
GEMEL
Incorporating Retrieval Information into the Truncation of Ranking Lists for Better Legal Search
ACM1 repoACM:3477495.3531998
RLT4Reranking
Trimming Data Sets: a Verified Algorithm for Robust Mean Estimation
ACM1 repoACM:3479394.3479412
infotheo
Enterprise-Scale Search: Accelerating Inference for Sparse Extreme Multi-Label Ranking Trees
ACM1 repoACM:3485447.3511973
pecos
GPU Accelerated Boosted Trees and Deep Neural Networks for Better Recommender Systems
ACM1 repoACM:3487572.3487605
xgboost
MtCut
ACM1 repoACM:3488560.3498466
RLT4Reranking
Transform, Warp, and Dress: A New Transformation-guided Model for Virtual Try-on
ACM1 repoACM:3491226
awesome-virtual-try-on
D3
ACM1 repoACM:3492321.3519576
erdos
Lightweight Robust Size Aware Cache Management
ACM1 repoACM:3507920
caffeine
Cross-category Virtual Try-on Technology Research Based on PF-AFN
ACM1 repoACM:3511176.3511201
awesome-virtual-try-on
Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?
ACM1 repoACM:3531146.3533192
Q16
Integrity Authentication in Tree Models
ACM1 repoACM:3534678.3539428
awesome-decision-tree-papers
Online Clustering
ACM1 repoACM:3534678.3542600
river
SciTS: A Benchmark for Time-Series Databases in Scientific Experiments and Industrial Internet of Things
ACM1 repoACM:3538712.3538723
ClickBench
Surprise: Result List Truncation via Extreme Value Theory
ACM1 repoACM:3539618.3592066
RLT4Reranking
FINGER: Fast Inference for Graph-based Approximate Nearest Neighbor Search
ACM1 repoACM:3543507.3583318
pecos
Filtered-DiskANN: Graph Algorithms for Approximate Nearest Neighbor Search with Filters
ACM1 repoACM:3543507.3583552
pgvectorscale
EasySpider: A No-Code Visual System for Crawling the Web
ACM1 repoACM:3543873.3587345
EasySpider
CALVI: Critical Thinking Assessment for Literacy in Visualizations
ACM1 repoACM:3544548.3581406
acl2026-misleading-visualizations
Aeneas: Rust verification by functional translation
ACM1 repoACM:3547647
aeneas
Make Your Own Sprites
ACM1 repoACM:3550454.3555482
Pixelization
Optimization Techniques for GPU Programming
ACM1 repoACM:3570638
awesome-gpu-engineering
Omnisemantics: Smooth Handling of Nondeterminism
ACM1 repoACM:3579834
bedrock2
Taming the Domain Shift in Multi-source Learning for Energy Disaggregation
ACM1 repoACM:3580305.3599910
awesome-nilm
Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data Poisoning
ACM1 repoACM:3581783.3612108
BackdoorDM
MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning
ACM1 repoACM:3581783.3612836
MERTools
FORGE: Pre-Training Open Foundation Models for Science
ACM1 repoACM:3581784.3613215
gpt-neox
EMSAssist: An End-to-End Mobile Voice Assistant at the Edge for Emergency Medical Services
ACM1 repoACM:3581791.3596853
EMS-Pipeline
Interactive Latent Diffusion Model
ACM1 repoACM:3583131.3590471
sdstudio
Practice on Effectively Extracting NLP Features for Click-Through Rate Prediction
ACM1 repoACM:3583780.3614707
BAIU
PyABSA: A Modularized Framework for Reproducible Aspect-based Sentiment Analysis
ACM1 repoACM:3583780.3614752
PyABSA
SE-PQA: Personalized Community Question Answering
ACM1 repoACM:3589335.3651445
SE-PQA
PG-Schema: Schemas for Property Graphs
ACM1 repoACM:3589778
rudof
CryptOpt: Verified Compilation with Randomized Program Search for Cryptographic Primitives
ACM1 repoACM:3591272
CryptOpt
Delilah: eBPF-offload on Computational Storage
ACM1 repoACM:3592980.3595319
awesome-ebpf
SSProve: A Foundational Framework for Modular Cryptographic Proofs in Coq
ACM1 repoACM:3594735
ssprove
Detail-Preserving Video-based Virtual Try-On (DPV-VTON)
ACM1 repoACM:3599589.3599599
awesome-virtual-try-on
PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation
ACM1 repoACM:3600006.3613139
MInference
A Grounded Conceptual Model for Ownership Types in Rust
ACM1 repoACM:3622841
aquascope
Understanding Any Time Series Classifier with a Subsequence-based Explainer
ACM1 repoACM:3624480
Awesome-Time-Series-Explainability
CIVSCOPE: Analyzing Potential Memory Corruption Bugs in Compartment Interfaces
ACM1 repoACM:3625275.3625399
publications
ACM TechBrief: Generative Artificial Intelligence
ACM1 repoACM:3626110
awesome-ai4lam
OpenIVM: a SQL-to-SQL Compiler for Incremental Computations
ACM1 repoACM:3626246.3654743
openivm
ALP: Adaptive Lossless floating-Point Compression
ACM1 repoACM:3626717
ALP
MACRec: A Multi-Agent Collaboration Framework for Recommendation
ACM1 repoACM:3626772.3657669
MACRec
Resources for Brewing BEIR: Reproducible Reference Models and Statistical Analyses
ACM1 repoACM:3626772.3657862
beir
Ranked List Truncation for Large Language Model-based Re-Ranking
ACM1 repoACM:3626772.3657864
RLT4Reranking
Fine-Tuning LLaMA for Multi-Stage Text Retrieval
ACM1 repoACM:3626772.3657951
reranker-as-judge
pyPANTERA: A Python PAckage for Natural language obfuscaTion Enforcing pRivacy & Anonymization
ACM1 repoACM:3627673.3679173
pypantera
Distributed Boosting: An Enhancing Method on Dataset Distillation
ACM1 repoACM:3627673.3679897
awesome-gradient-boosting-papers
Exploring Performance and Cost Optimization with ASIC-Based CXL Memory
ACM1 repoACM:3627703.3650061
lightllm
Guided Equality Saturation
ACM1 repoACM:3632900
equational_theories
Endoprocess: Programmable and Extensible Subprocess Isolation
ACM1 repoACM:3633500.3633507
publications
PEMBOT: Pareto-Ensembled Multi-task Boosted Trees
ACM1 repoACM:3637528.3671619
awesome-gradient-boosting-papers
ImputeFormer: Low Rankness-Induced Transformers for Generalizable Spatiotemporal Imputation
ACM1 repoACM:3637528.3671751
frn-50k-baseline
Iterative Weak Learnability and Multiclass AdaBoost
ACM1 repoACM:3637528.3671842
awesome-gradient-boosting-papers
Uplift Modelling via Gradient Boosting
ACM1 repoACM:3637528.3672019
awesome-gradient-boosting-papers
FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems
ACM1 repoACM:3639477.3639754
awesome-LLM-AIOps
Hypermedia Controls: Feral to Formal
ACM1 repoACM:3648188.3675127
fixi
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
ACM1 repoACM:3658644.3670388
jailbreak_llms
ClassID: Enabling Student Behavior Attribution from Ambient Classroom Sensing Systems
ACM1 repoACM:3659586
ClassID
Leveraging Large Language Models for the Auto-remediation of Microservice Applications: An Experimental Study
ACM1 repoACM:3663529.3663855
awesome-LLM-AIOps
Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders
ACM1 repoACM:3664190.3672528
IR-Lens
BrainRAM: Cross-Modality Retrieval-Augmented Image Reconstruction from Human Brain Activity
ACM1 repoACM:3664647.3681296
BrainRAM
Sound Borrow-Checking for Rust via Symbolic Semantics
ACM1 repoACM:3674640
aeneas
Design and Implementation of a Coverage-Guided Ruby Fuzzer
ACM1 repoACM:3675741.3675749
ruzzy
Past-Future Scheduler for LLM Serving under SLA Guarantees
ACM1 repoACM:3676641.3716011
lightllm
Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Trustworthy Response Generation in Chinese
ACM1 repoACM:3686807
Huatuo-Llama-Med-Chinese
Reproducibility Report for ACM SIGMOD 2024 Paper: 'ALP: Adaptive Lossless Floating-Point Compression'
ACM1 repoACM:3687998.3717057
ALP
MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
ACM1 repoACM:3689092.3689959
MERTools
StarMalloc: Verifying a Modern, Hardened Memory Allocator
ACM1 repoACM:3689773
mimalloc-bench
CoqPilot, a plugin for LLM-based generation of proofs
ACM1 repoACM:3691620.3695357
coqpilot
The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small Classifier
ACM1 repoACM:3691620.3695475
awesome-LLM-AIOps
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
ACM1 repoACM:3694715.3695948
lightllm
Reducing Energy Bloat in Large Model Training
ACM1 repoACM:3694715.3695970
zeus
Unearthing Semantic Checks for Cloud Infrastructure-as-Code Programs
ACM1 repoACM:3694715.3695974
awesome-LLM-AIOps
Grad: Guided Relation Diffusion Generation for Graph Augmentation in Graph Fraud Detection
ACM1 repoACM:3696410.3714520
antifraud
Hybrid, Unified and Iterative: A Novel Framework for Text-based Person Anomaly Retrieval
ACM1 repoACM:3701716.3717653
Hybrid-Unified-and-Iterative-A-Novel-Framework-for-Text-based-Person-Anomaly-Retrieval
A Survey of Geometric Optimization for Deep Learning: From Euclidean Space to Riemannian Manifold
ACM1 repoACM:3708498
ai-agent-book
SymphonyQG: Towards Symphonious Integration of Quantization and Graph for Approximate Nearest Neighbor Search
ACM1 repoACM:3709730
RaBitQ-Library
Helping the Helper : Supporting Peer Counselors via AI-Empowered Practice and Feedback
ACM1 repoACM:3710993
CARE
Multi-modal Time Series Analysis: A Tutorial and Survey
ACM1 repoACM:3711896.3736567
TS-RAG
Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models
ACM1 repoACM:3726302.3730059
PAD
TourMLLM: A Retrieval-Augmented Multimodal Large Language Model for Multitask Learning in the Tourism Domain
ACM1 repoACM:3731715.3733450
TourMLLM
Semantics of Probabilistic Programs Using s-Finite Kernels in Dependent Type Theory
ACM1 repoACM:3732291
analysis
G-ALP: Rethinking Light-weight Encodings for GPUs
ACM1 repoACM:3736227.3736242
fastlanes
MER 2025: When Affective Computing Meets Large Language Models
ACM1 repoACM:3746027.3762007
MERTools
Federated Gradient Boosting for Financial Fraud Detection: An Empirical Study in the Banking Sector
ACM1 repoACM:3746252.3760891
awesome-gradient-boosting-papers
FairRegBoost: An End-to-End Data Processing Framework for Fair and Scalable Regression
ACM1 repoACM:3746252.3761277
awesome-gradient-boosting-papers
Cloud Infrastructure Management in the Age of AI Agents
ACM1 repoACM:3759441.3759443
awesome-LLM-AIOps
A Joint Classification Method for Traditional Chinese Medicine Diseases and Syndromes Based on BertChinese-RCNNATTN
ACM1 repoACM:3759972.3759979
SwanLab
Multi-Agent LLM Reasoning for Clinical Procedure Sequencing from High-Granularity EHR Data
ACM1 repoACM:3765612.3767238
Multiagent_Procedure_MIMIC-III
AgentSight: System-Level Observability for AI Agents Using eBPF
ACM1 repoACM:3766882.3767169
agentsight
Representation-Aware Root Cause Analysis with Large Language Models (Position Paper)
ACM1 repoACM:3777911.3801108
awesome-LLM-AIOps
Cylindrical Algebraic Decomposition in Coq/Rocq
ACM1 repoACM:3779031.3779100
analysis
Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization
ACM1 repoACM:3786161.3788460
boatse-extractor
A Replicability Study of Joint Product Quantisation for Effective Space-Efficient Dense Retrieval
ACM1 repoACM:3805712.3808565
pyterrier_dr
Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior in Generative Retrieval
ACM1 repoACM:3805712.3808567
lost-in-decoding
Programming languages for distributed computing systems
ACM1 repoACM:72551.72552
nextflow
Format-directed list processing in LISP
ACM1 repoACM:800005.807968
regal
Random sampling from hash files
ACM1 repoACM:93605.98746
mongo
Light Logics and Optimal Reduction: Completeness and Complexity
arXiv1 repoarXiv:0704.2448
ESCoC
Power-law distributions in empirical data
arXiv1 repoarXiv:0706.1062
tokenizer-flores-validation
Steganography of VoIP Streams
arXiv1 repoarXiv:0805.2938
awesome-rtc-hacking
Estimating and Sampling Graphs with Multidimensional Random Walks
arXiv1 repoarXiv:1002.1751
littleballoffur
Type Classes for Mathematics in Type Theory
arXiv1 repoarXiv:1102.1323
math-classes
Deciding Kleene Algebras in Coq
arXiv1 repoarXiv:1105.4537
atbr
RTED: A Robust Algorithm for the Tree Edit Distance
arXiv1 repoarXiv:1201.0230
opendataloader-bench
A Synthesis of the Procedural and Declarative Styles of Interactive Theorem Proving
arXiv1 repoarXiv:1201.3601
hol-light
Harmony Explained: Progress Towards A Scientific Theory of Music
arXiv1 repoarXiv:1202.4212
awesome-music-production
Programming with Algebraic Effects and Handlers
arXiv1 repoarXiv:1203.1539
eff
Automatic facial feature extraction and expression recognition based on neural network
arXiv1 repoarXiv:1204.2073
awesome-affective-computing
BPR: Bayesian Personalized Ranking from Implicit Feedback
arXiv1 repoarXiv:1205.2618
implicit
Public Key Cryptography Standards: PKCS
arXiv1 repoarXiv:1207.5446
awesome-standards
PaxosLease: Diskless Paxos for Leases
arXiv1 repoarXiv:1209.4187
translations
Fast Packed String Matching for Short Patterns
arXiv1 repoarXiv:1209.6449
gecko-dev
Sequence Transduction with Recurrent Neural Networks
arXiv1 repoarXiv:1211.3711
GigaAM
A Multilingual Semantic Wiki Based on Attempto Controlled English and Grammatical Framework
arXiv1 repoarXiv:1303.4293
logicmoo_workspace
Functional Package Management with Guix
arXiv1 repoarXiv:1305.4584
Functional-Programming
Bounding the Estimation Error of Sampling-based Shapley Value Approximation
arXiv1 repoarXiv:1306.4265
shapley
Generating Sequences With Recurrent Neural Networks
arXiv1 repoarXiv:1308.0850
char-rnn
Distributed Representations of Words and Phrases and their Compositionality
arXiv1 repoarXiv:1310.4546
Word-Embeddings-Repository-for-Turkish
Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding
arXiv1 repoarXiv:1311.2540
FiniteStateEntropy
Computing in Operations Research using Julia
arXiv1 repoarXiv:1312.1431
JuMP.jl
Experience Implementing a Performant Category-Theory Library in Coq
arXiv1 repoarXiv:1401.7694
agda-categories
Better bitmap performance with Roaring bitmaps
arXiv1 repoarXiv:1402.6407
roaring-rs
Principles of Antifragile Software
arXiv1 repoarXiv:1404.3056
awesome-chaos-engineering
NILMTK: An Open Source Toolkit for Non-intrusive Load Monitoring
arXiv1 repoarXiv:1404.3878
awesome-nilm
Static Analysis for Regular Expression Exponential Runtime via Substructural Logics (Extended)
arXiv1 repoarXiv:1405.7058
fancy-regex
Koka: Programming with Row Polymorphic Effect Types
arXiv1 repoarXiv:1406.2061
koka
Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
arXiv1 repoarXiv:1406.2227
Meta-SelfLearning
Convolutional Neural Networks for Sentence Classification
arXiv1 repoarXiv:1408.5882
simple-ntc
Very Deep Convolutional Networks for Large-Scale Image Recognition
arXiv1 repoarXiv:1409.1556
materialgan
Going Deeper with Convolutions
arXiv1 repoarXiv:1409.4842
caffe
Memory Networks
arXiv1 repoarXiv:1410.3916
nmt
Conditional Generative Adversarial Nets
arXiv1 repoarXiv:1411.1784
ocaml-torch
Show and Tell: A Neural Image Caption Generator
arXiv1 repoarXiv:1411.4555
mscoco-it
Learn Physics by Programming in Haskell
arXiv1 repoarXiv:1412.4880
Functional-Programming
ORB-SLAM: a Versatile and Accurate Monocular SLAM System
arXiv1 repoarXiv:1502.00956
stella_vslam
Gated Feedback Recurrent Neural Networks
arXiv1 repoarXiv:1502.02367
parallel-ss-dep
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
arXiv1 repoarXiv:1502.03044
LaTeX_OCR
arXiv:1503.03465
arXiv1 repoarXiv:1503.03465
clhash
U-Net: Convolutional Networks for Biomedical Image Segmentation
arXiv1 repoarXiv:1505.04597
Landslide4Sense-2022
Spatial Transformer Networks
arXiv1 repoarXiv:1506.02025
stn3d
Teaching Machines to Read and Comprehend
arXiv1 repoarXiv:1506.03340
rc-data
Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
arXiv1 repoarXiv:1506.06724
DiagnosisCoding
Skip-Thought Vectors
arXiv1 repoarXiv:1506.06726
SentEval
Causal Decision Trees
arXiv1 repoarXiv:1508.03812
Data-Science
Effective Approaches to Attention-based Neural Machine Translation
arXiv1 repoarXiv:1508.04025
nmt
A large annotated corpus for learning natural language inference
arXiv1 repoarXiv:1508.05326
roberta-large-mnli
A Neural Algorithm of Artistic Style
arXiv1 repoarXiv:1508.06576
ocaml-torch
Continuous control with deep reinforcement learning
arXiv1 repoarXiv:1509.02971
cleanrl
Spatially Encoding Temporal Correlations to Classify Temporal Data Using Convolutional Neural Networks
arXiv1 repoarXiv:1509.07481
AI-assisted-chemical-sensing
Evasion and Hardening of Tree Ensemble Classifiers
arXiv1 repoarXiv:1509.07892
RobustTrees
Fast Algorithms for Convolutional Neural Networks
arXiv1 repoarXiv:1509.09308
mace
Semi-supervised Sequence Learning
arXiv1 repoarXiv:1511.01432
bert
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
arXiv1 repoarXiv:1511.06434
ocaml-torch
Rethinking the Inception Architecture for Computer Vision
arXiv1 repoarXiv:1512.00567
materialgan
ShapeNet: An Information-Rich 3D Model Repository
arXiv1 repoarXiv:1512.03012
Cap3D
Programming in logic without logic programming
arXiv1 repoarXiv:1601.00529
logicmoo_workspace
Convolutional Pose Machines
arXiv1 repoarXiv:1602.00134
openpose
Are Elephants Bigger than Butterflies? Reasoning about Sizes of Objects
arXiv1 repoarXiv:1602.00753
Kosmos-X
EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos
arXiv1 repoarXiv:1602.03012
TF-Cholec80
Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
arXiv1 repoarXiv:1602.07332
LRV-Instruction
XGBoost: A Scalable Tree Boosting System
arXiv1 repoarXiv:1603.02754
xgboost
The sphere packing problem in dimension 8
arXiv1 repoarXiv:1603.04246
glq
Incorporating Copying Mechanism in Sequence-to-Sequence Learning
arXiv1 repoarXiv:1603.06393
nl2bash
Consistently faster and smaller compressed bitmaps with Roaring
arXiv1 repoarXiv:1603.06549
RoaringFormatSpec
A Diagram Is Worth A Dozen Images
arXiv1 repoarXiv:1603.07396
ai2d
S-hull: a fast radial sweep-hull routine for Delaunay triangulation
arXiv1 repoarXiv:1604.01428
torch_delaunay
GLEU Without Tuning
arXiv1 repoarXiv:1605.02592
gec-metrics
Neural Network Translation Models for Grammatical Error Correction
arXiv1 repoarXiv:1606.00189
CTCResources
OpenAI Gym
arXiv1 repoarXiv:1606.01540
gym
Natural Language Generation enhances human decision-making with uncertain information
arXiv1 repoarXiv:1606.03254
awesome-nlg
Neural Generation of Regular Expressions from Natural Language with Minimal Domain Knowledge
arXiv1 repoarXiv:1608.03000
llm-jepa
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
arXiv1 repoarXiv:1608.04207
SentEval
SandBlaster: Reversing the Apple Sandbox
arXiv1 repoarXiv:1608.04303
sandblaster
Using the Output Embedding to Improve Language Models
arXiv1 repoarXiv:1608.05859
GPT-2
Polysemous codes
arXiv1 repoarXiv:1609.01882
faiss
Predicting the future relevance of research institutions - The winning solution of the KDD Cup 2016
arXiv1 repoarXiv:1609.02728
xgboost
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
arXiv1 repoarXiv:1609.04802
MAX-Image-Resolution-Enhancer
Image-to-Markup Generation with Coarse-to-Fine Attention
arXiv1 repoarXiv:1609.04938
LaTeX-OCR
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
arXiv1 repoarXiv:1609.08144
nmt
Non-Intrusive Load Monitoring: A Review and Outlook
arXiv1 repoarXiv:1610.01191
awesome-nilm
Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models
arXiv1 repoarXiv:1610.02424
fairseq
ORB-SLAM2: an Open-Source SLAM System for Monocular, Stereo and RGB-D Cameras
arXiv1 repoarXiv:1610.06475
stella_vslam
Towards Automatic Resource Bound Analysis for OCaml
arXiv1 repoarXiv:1611.00692
plutus
Cubical Type Theory: a constructive interpretation of the univalence axiom
arXiv1 repoarXiv:1611.02108
cubical
Image-to-Image Translation with Conditional Adversarial Networks
arXiv1 repoarXiv:1611.07004
pachyderm
Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
arXiv1 repoarXiv:1611.08050
openpose
NewsQA: A Machine Comprehension Dataset
arXiv1 repoarXiv:1611.09830
CoConflictQA
Self-critical Sequence Training for Image Captioning
arXiv1 repoarXiv:1612.00563
vilmedic
Overcoming catastrophic forgetting in neural networks
arXiv1 repoarXiv:1612.00796
Continual-NExT
FMA: A Dataset For Music Analysis
arXiv1 repoarXiv:1612.01840
fma
FastText.zip: Compressing text classification models
arXiv1 repoarXiv:1612.03651
fasttext-language-identification
Fast keyed hash/pseudo-random function using SIMD multiply and permute
arXiv1 repoarXiv:1612.06257
highwayhash
OpenNMT: Open-Source Toolkit for Neural Machine Translation
arXiv1 repoarXiv:1701.02810
nmt
Fast Exact k-Means, k-Medians and Bregman Divergence Clustering in 1D
arXiv1 repoarXiv:1701.07204
ml-stable-diffusion
Emotion Recognition From Speech With Recurrent Neural Networks
arXiv1 repoarXiv:1701.08071
awesome-affective-computing
arXiv:1701.08398
arXiv1 repoarXiv:1701.08398
TIL-2023
New cardinality estimation algorithms for HyperLogLog sketches
arXiv1 repoarXiv:1702.01284
hash4j
Software Engineering at Google
arXiv1 repoarXiv:1702.01715
awesome-cto
Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning
arXiv1 repoarXiv:1702.03118
tiny-vllm
Chaos Engineering
arXiv1 repoarXiv:1702.05843
awesome-chaos-engineering
A Platform for Automating Chaos Experiments
arXiv1 repoarXiv:1702.05849
awesome-chaos-engineering
ERA: A Framework for Economic Resource Allocation for the Cloud
arXiv1 repoarXiv:1702.07311
papers-notebook
Billion-scale similarity search with GPUs
arXiv1 repoarXiv:1702.08734
faiss
Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
arXiv1 repoarXiv:1703.01619
nmt
Massive Exploration of Neural Machine Translation Architectures
arXiv1 repoarXiv:1703.03906
nmt
Automated Hate Speech Detection and the Problem of Offensive Language
arXiv1 repoarXiv:1703.04009
toxic-bert
Understanding Black-box Predictions via Influence Functions
arXiv1 repoarXiv:1703.04730
influence_boosting
Prototypical Networks for Few-shot Learning
arXiv1 repoarXiv:1703.05175
dialog-mteb
Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks
arXiv1 repoarXiv:1703.07015
Timeseries-PILE
Deep Photo Style Transfer
arXiv1 repoarXiv:1703.07511
Github-Ranking
Transfer learning for music classification and regression tasks
arXiv1 repoarXiv:1703.09179
fma
Towards Automatic Learning of Procedures from Web Instructional Videos
arXiv1 repoarXiv:1703.09788
TimeChat-Online-139K
Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation
arXiv1 repoarXiv:1703.09902
awesome-nlg
Reading Wikipedia to Answer Open-Domain Questions
arXiv1 repoarXiv:1704.00051
odqa_baseline_code
A Survey of Distributed Message Broker Queues
arXiv1 repoarXiv:1704.00411
awesome-scalability
Faster Base64 Encoding and Decoding Using AVX2 Instructions
arXiv1 repoarXiv:1704.00605
simdutf
Get To The Point: Summarization with Pointer-Generator Networks
arXiv1 repoarXiv:1704.04368
cnn-dailymail
RACE: Large-scale ReAding Comprehension Dataset From Examinations
arXiv1 repoarXiv:1704.04683
lares
Learning to Reason: End-to-End Module Networks for Visual Question Answering
arXiv1 repoarXiv:1704.05526
n2nmn
Accelerated Nearest Neighbor Search with Quick ADC
arXiv1 repoarXiv:1704.07355
RaBitQ-Library
Automatic Anomaly Detection in the Cloud Via Statistical Learning
arXiv1 repoarXiv:1704.07706
AnomalyDetection.rb
Hand Keypoint Detection in Single Images using Multiview Bootstrapping
arXiv1 repoarXiv:1704.07809
openpose
A Novel Hybrid Quicksort Algorithm Vectorized using AVX-512 on Intel Skylake
arXiv1 repoarXiv:1704.08579
x86-simd-sort
Dense-Captioning Events in Videos
arXiv1 repoarXiv:1705.00754
ActivityNet_Captions
TALL: Temporal Activity Localization via Language Query
arXiv1 repoarXiv:1705.02101
TALL
Supervised Learning of Universal Sentence Representations from Natural Language Inference Data
arXiv1 repoarXiv:1705.02364
SentEval
Large-scale, Fast and Accurate Shot Boundary Detection through Spatio-temporal Convolutional Neural Networks
arXiv1 repoarXiv:1705.03281
TransNetV2
Learning how to explain neural networks: PatternNet and PatternAttribution
arXiv1 repoarXiv:1705.05598
dianna
Engineering Record And Replay For Deployability: Extended Technical Report
arXiv1 repoarXiv:1705.05937
rr
Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
arXiv1 repoarXiv:1705.07115
Fine-Grained_Features_Alignment_via_Constrastive_Learning
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
arXiv1 repoarXiv:1705.07750
JavisBench
pix2code: Generating Code from a Graphical User Interface Screenshot
arXiv1 repoarXiv:1705.07962
Screenshot-to-code
arXiv:1705.08926
arXiv1 repoarXiv:1705.08926
smac
Marmara Turkish Coreference Corpus and Coreference Resolution Baseline
arXiv1 repoarXiv:1706.01863
g4t0r2-nlp
Deep reinforcement learning from human preferences
arXiv1 repoarXiv:1706.03741
prompt-engineering
On Calibration of Modern Neural Networks
arXiv1 repoarXiv:1706.04599
ogham-mcp
SuperMinHash - A New Minwise Hashing Algorithm for Jaccard Similarity Estimation
arXiv1 repoarXiv:1706.05698
hash4j
Developing Bug-Free Machine Learning Systems With Formal Mathematics
arXiv1 repoarXiv:1706.08605
certigrad
Causal Structure Learning
arXiv1 repoarXiv:1706.09141
Data-Science
CatBoost: unbiased boosting with categorical features
arXiv1 repoarXiv:1706.09516
catboost
SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
arXiv1 repoarXiv:1708.00055
t5-large-encoder-only-bf16
Localizing Moments in Video with Natural Language
arXiv1 repoarXiv:1708.01641
LocalizingMoments
StarCraft II: A New Challenge for Reinforcement Learning
arXiv1 repoarXiv:1708.04782
pysc2
Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
arXiv1 repoarXiv:1709.00103
MirageTVQA
ProSLAM: Graph SLAM from a Programmer's Perspective
arXiv1 repoarXiv:1709.04377
stella_vslam
The First Evaluation of Chinese Human-Computer Dialogue Technology
arXiv1 repoarXiv:1709.10217
ChineseNLPCorpus
DisSent: Sentence Representation Learning from Explicit Discourse Relations
arXiv1 repoarXiv:1710.04334
SentEval
Representation Learning of Music Using Artist Labels
arXiv1 repoarXiv:1710.06648
fma
FigureQA: An Annotated Figure Dataset for Visual Reasoning
arXiv1 repoarXiv:1710.07300
PlotQA
Souper: A Synthesizing Superoptimizer
arXiv1 repoarXiv:1711.04422
souper
Emotional End-to-End Neural Speech Synthesizer
arXiv1 repoarXiv:1711.05447
FastSpeech2-Plus
Total Haskell is Reasonable Coq
arXiv1 repoarXiv:1711.09286
hs-to-coq
Are GANs Created Equal? A Large-Scale Study
arXiv1 repoarXiv:1711.10337
pytorch-fid
Occam's razor is insufficient to infer the preferences of irrational agents
arXiv1 repoarXiv:1712.05812
minihf
SuperPoint: Self-Supervised Interest Point Detection and Description
arXiv1 repoarXiv:1712.07629
gtsfm
Demystifying MMD GANs
arXiv1 repoarXiv:1801.01401
stylegan2-ada-pytorch
MobileNetV2: Inverted Residuals and Linear Bottlenecks
arXiv1 repoarXiv:1801.04381
mobilenet_v2_1.4_224
Universal Language Model Fine-tuning for Text Classification
arXiv1 repoarXiv:1801.06146
indonesian-language-models
Generating Wikipedia by Summarizing Long Sequences
arXiv1 repoarXiv:1801.10198
transformer-tricks
DensePose: Dense Human Pose Estimation In The Wild
arXiv1 repoarXiv:1802.00434
DensePose
On Higher Inductive Types in Cubical Type Theory
arXiv1 repoarXiv:1802.01170
cubical
Classification and Disease Localization in Histopathology Using Only Global Labels: A Weakly-Supervised Approach
arXiv1 repoarXiv:1802.02212
HistoSSLscaling
Revisiting the Inverted Indices for Billion-Scale Approximate Nearest Neighbors
arXiv1 repoarXiv:1802.02422
hnswlib
Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
arXiv1 repoarXiv:1802.02611
spectra
Partisan: Enabling Cloud-Scale Erlang Applications
arXiv1 repoarXiv:1802.02652
partisan
Attention-based Deep Multiple Instance Learning
arXiv1 repoarXiv:1802.04712
HistoSSLscaling
Deep contextualized word representations
arXiv1 repoarXiv:1802.05365
Word-Embeddings-Repository-for-Turkish
Finding Influential Training Samples for Gradient Boosted Decision Trees
arXiv1 repoarXiv:1802.06640
influence_boosting
Learning Word Vectors for 157 Languages
arXiv1 repoarXiv:1802.06893
fasttext-language-identification
Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration
arXiv1 repoarXiv:1802.08802
android_world_seeact_v
Back to Basics: Benchmarking Canonical Evolution Strategies for Playing Atari
arXiv1 repoarXiv:1802.08842
Canonical_ES_Atari
Addressing Function Approximation Error in Actor-Critic Methods
arXiv1 repoarXiv:1802.09477
cleanrl
Chest X-Ray Analysis of Tuberculosis by Deep Learning with Segmentation and Augmentation
arXiv1 repoarXiv:1803.01199
Project-Imaging-X
GONet: A Semi-Supervised Deep Learning Approach For Traversability Estimation
arXiv1 repoarXiv:1803.03254
UniWM_Dataset
Deep-FSMN for Large Vocabulary Continuous Speech Recognition
arXiv1 repoarXiv:1803.05030
fsmn-vad-onnx
OSINT Analysis of the TOR Foundation
arXiv1 repoarXiv:1803.05201
non-typical-OSINT-guide
Learning to Recognize Musical Genre from Audio
arXiv1 repoarXiv:1803.05337
fma
SentEval: An Evaluation Toolkit for Universal Sentence Representations
arXiv1 repoarXiv:1803.05449
SentEval
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
arXiv1 repoarXiv:1803.05457
ai2_arc
Complex-YOLO: Real-time 3D Object Detection on Point Clouds
arXiv1 repoarXiv:1803.06199
Complex-YOLOv4-Pytorch
Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
arXiv1 repoarXiv:1803.09017
MockingBird
Universal Sentence Encoder
arXiv1 repoarXiv:1803.11175
SentEval
arXiv:1803.11485
arXiv1 repoarXiv:1803.11485
smac
ESPnet: End-to-End Speech Processing Toolkit
arXiv1 repoarXiv:1804.00015
kan-bayashi_ljspeech_vits
Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
arXiv1 repoarXiv:1804.00079
SentEval
SBFT: a Scalable and Decentralized Trust Infrastructure
arXiv1 repoarXiv:1804.01626
concord-bft
Flexible and Scalable Deep Learning with MMLSpark
arXiv1 repoarXiv:1804.04031
SynapseML
arXiv:1804.05839
arXiv1 repoarXiv:1804.05839
llm_test
Phrase-Based & Neural Unsupervised Machine Translation
arXiv1 repoarXiv:1804.07755
XLM
An Aggregated Multicolumn Dilated Convolution Network for Perspective-Free Counting
arXiv1 repoarXiv:1804.07821
marker
Gender Bias in Coreference Resolution
arXiv1 repoarXiv:1804.09301
model-written-evals
Link and code: Fast indexing with graphs and compact regression codes
arXiv1 repoarXiv:1804.09996
faiss
Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates
arXiv1 repoarXiv:1804.10959
sentencepiece
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
arXiv1 repoarXiv:1805.01070
SentEval
Online normalizer calculation for softmax
arXiv1 repoarXiv:1805.02867
GPT-2
TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation
arXiv1 repoarXiv:1805.04699
asr_consilium
A Chaos Engineering System for Live Analysis and Falsification of Exception-handling in the JVM
arXiv1 repoarXiv:1805.05246
awesome-chaos-engineering
arXiv:1805.08318
arXiv1 repoarXiv:1805.08318
SkinDeep
COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval
arXiv1 repoarXiv:1805.08661
coco-cn
AutoAugment: Learning Augmentation Policies from Data
arXiv1 repoarXiv:1805.09501
Chinese-CLIP
Neural Network Acceptability Judgments
arXiv1 repoarXiv:1805.12471
t5-large-encoder-only-bf16
Digging Into Self-Supervised Monocular Depth Estimation
arXiv1 repoarXiv:1806.01260
Depth-Estimation
Generative Adversarial Networks for Realistic Synthesis of Hyperspectral Samples
arXiv1 repoarXiv:1806.02583
Data-Science
Cell Detection with Star-convex Polygons
arXiv1 repoarXiv:1806.03535
stardist
Know What You Don't Know: Unanswerable Questions for SQuAD
arXiv1 repoarXiv:1806.03822
squad_v2
The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems
arXiv1 repoarXiv:1806.09514
EmoV-DB
The relativistic discriminator: a key element missing from standard GAN
arXiv1 repoarXiv:1807.00734
ocaml-torch
CAIL2018: A Large-Scale Legal Dataset for Judgment Prediction
arXiv1 repoarXiv:1807.02478
CAIL
The Double Sphere Camera Model
arXiv1 repoarXiv:1807.08957
kalibr
HiDDeN: Hiding Data With Deep Networks
arXiv1 repoarXiv:1807.09937
WMCopier
Unified Perceptual Parsing for Scene Understanding
arXiv1 repoarXiv:1807.10221
FoodSeg103-Benchmark-v1
Speaker Recognition from Raw Waveform with SincNet
arXiv1 repoarXiv:1808.00158
NIPS4Bplus
One Billion Apples' Secret Sauce: Recipe for the Apple Wireless Direct Link Ad hoc Protocol
arXiv1 repoarXiv:1808.03156
stop-stutter
Fast Video Shot Transition Localization with Deep Structured Models
arXiv1 repoarXiv:1808.04234
TransNetV2
WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations
arXiv1 repoarXiv:1808.09121
t5-large-encoder-only-bf16
Evaluating Theory of Mind in Question Answering
arXiv1 repoarXiv:1808.09352
ToMi
Story Ending Generation with Incremental Encoding and Commonsense Knowledge
arXiv1 repoarXiv:1808.10113
AwesomeSEG
Toward a Standardized and More Accurate Indonesian Part-of-Speech Tagging
arXiv1 repoarXiv:1809.03391
indonlu
Addressing the Fundamental Tension of PCGML with Discriminative Learning
arXiv1 repoarXiv:1809.04432
WaveFunctionCollapse
BRAVO -- Biased Locking for Reader-Writer Locks
arXiv1 repoarXiv:1810.01553
VictoriaMetrics
The UCR Time Series Archive
arXiv1 repoarXiv:1810.07758
UTSD
Don't Unroll Adjoint: Differentiating SSA-Form Programs
arXiv1 repoarXiv:1810.07951
Zygote.jl
From Louvain to Leiden: guaranteeing well-connected communities
arXiv1 repoarXiv:1810.08473
sweet-search
MMLSpark: Unifying Machine Learning Ecosystems at Massive Scales
arXiv1 repoarXiv:1810.08744
SynapseML
pair2vec: Compositional Word-Pair Embeddings for Cross-Sentence Inference
arXiv1 repoarXiv:1810.08854
relbert
3D MRI brain tumor segmentation using autoencoder regularization
arXiv1 repoarXiv:1810.11654
TriALS
Audio inpainting of music by means of neural networks
arXiv1 repoarXiv:1810.12138
fma
ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
arXiv1 repoarXiv:1810.12885
t5-large-encoder-only-bf16
Exploration by Random Network Distillation
arXiv1 repoarXiv:1810.12894
cleanrl
WaveGlow: A Flow-based Generative Network for Speech Synthesis
arXiv1 repoarXiv:1811.00002
waveglow
Image Chat: Engaging Grounded Conversations
arXiv1 repoarXiv:1811.00945
sirius-spring2021-image2chat
Extended Isolation Forest
arXiv1 repoarXiv:1811.02141
isolation-forest
Subtask Gated Networks for Non-Intrusive Load Monitoring
arXiv1 repoarXiv:1811.06692
awesome-nilm
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
arXiv1 repoarXiv:1811.06965
awsome-llm-papers
A New Cervical Cytology Dataset for Nucleus Detection and Image Classification (Cervix93) and Methods for Cervical Nucleus Detection
arXiv1 repoarXiv:1811.09651
cytology_dataset
Visual Entailment Task for Visually-Grounded Language Learning
arXiv1 repoarXiv:1811.10582
SNLI-VE
CCNet: Criss-Cross Attention for Semantic Segmentation
arXiv1 repoarXiv:1811.11721
FoodSeg103-Benchmark-v1
Learning from a tiny dataset of manual annotations: a teacher/student approach for surgical phase recognition
arXiv1 repoarXiv:1812.00033
TF-Cholec80
Scalable Graph Learning for Anti-Money Laundering: A First Look
arXiv1 repoarXiv:1812.00076
AMLSim
Transferring Knowledge across Learning Processes
arXiv1 repoarXiv:1812.01054
xfer
A micro Lie theory for state estimation in robotics
arXiv1 repoarXiv:1812.01537
optik
Towards Accurate Generative Models of Video: A New Metric & Challenges
arXiv1 repoarXiv:1812.01717
VideoGPT
Soft Actor-Critic Algorithms and Applications
arXiv1 repoarXiv:1812.05905
cleanrl
Inverse Cooking: Recipe Generation from Food Images
arXiv1 repoarXiv:1812.06164
inversecooking
The Adverse Effects of Code Duplication in Machine Learning Models of Code
arXiv1 repoarXiv:1812.06469
Project_CodeNet
OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
arXiv1 repoarXiv:1812.08008
openpose
Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
arXiv1 repoarXiv:1812.08466
fadtk
Pansori: ASR Corpus Generation from Open Online Video Contents
arXiv1 repoarXiv:1812.09798
kabooks
A Poisson-Gaussian Denoising Dataset with Real Fluorescence Microscopy Images
arXiv1 repoarXiv:1812.10366
denoising-fluorescence
TripleAgent: Monitoring, Perturbation and Failure-obliviousness for Automated Resilience Improvement in Java Applications
arXiv1 repoarXiv:1812.10706
awesome-chaos-engineering
A Comprehensive Survey on Graph Neural Networks
arXiv1 repoarXiv:1901.00596
Data-Science
Learning From Less Data: A Unified Data Subset Selection and Active Learning Framework for Computer Vision
arXiv1 repoarXiv:1901.01151
cords
Panoptic Feature Pyramid Networks
arXiv1 repoarXiv:1901.02446
FoodSeg103-Benchmark-v1
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
arXiv1 repoarXiv:1901.02860
transfo-xl-wt103
From Plots to Endings: A Reinforced Pointer Generator for Story Ending Generation
arXiv1 repoarXiv:1901.03459
AwesomeSEG
Passage Re-ranking with BERT
arXiv1 repoarXiv:1901.04085
pygaggle
TensorFlow.js: Machine Learning for the Web and Beyond
arXiv1 repoarXiv:1901.05350
awesome-tensorflow-js
Visual Entailment: A Novel Task for Fine-Grained Image Understanding
arXiv1 repoarXiv:1901.06706
SNLI-VE
MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs
arXiv1 repoarXiv:1901.07042
baple
Cross-lingual Language Model Pretraining
arXiv1 repoarXiv:1901.07291
XLM
Personalized Dialogue Generation with Diversified Traits
arXiv1 repoarXiv:1901.09672
BoB
Information Operations Recognition: from Nonlinear Analysis to Decision-making
arXiv1 repoarXiv:1901.10876
non-typical-OSINT-guide
UcoSLAM: Simultaneous Localization and Mapping by Fusion of KeyPoints and Squared Planar Markers
arXiv1 repoarXiv:1902.03729
stella_vslam
BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
arXiv1 repoarXiv:1902.04094
temporal-robustness
Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology
arXiv1 repoarXiv:1902.06543
Midnight
Parsing Gigabytes of JSON per Second
arXiv1 repoarXiv:1902.08318
simdjson
Wavenilm: A causal neural network for power disaggregation from the complex power signal
arXiv1 repoarXiv:1902.08736
awesome-nilm
A large annotated medical image dataset for the development and evaluation of segmentation algorithms
arXiv1 repoarXiv:1902.09063
TotalSegmentator
Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression
arXiv1 repoarXiv:1902.09630
Complex-YOLOv4-Pytorch
EvolveGCN: Evolving Graph Convolutional Networks for Dynamic Graphs
arXiv1 repoarXiv:1902.10191
AMLSim
Accelerating Self-Play Learning in Go
arXiv1 repoarXiv:1902.10565
KataGo
Robust Decision Trees Against Adversarial Examples
arXiv1 repoarXiv:1902.10660
RobustTrees
COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
arXiv1 repoarXiv:1903.02874
TimeChat-Online-139K
Fine-tune BERT for Extractive Summarization
arXiv1 repoarXiv:1903.10318
BertSum
arXiv:1903.11269
arXiv1 repoarXiv:1903.11269
css10
Learning Discrete Structures for Graph Neural Networks
arXiv1 repoarXiv:1903.11960
fma
Habitat: A Platform for Embodied AI Research
arXiv1 repoarXiv:1904.01201
habitat-lab
Character Region Awareness for Text Detection
arXiv1 repoarXiv:1904.01941
EasyOCR
arXiv:1904.02285
arXiv1 repoarXiv:1904.02285
Jellyfish-13B
Speech Model Pre-training for End-to-End Spoken Language Understanding
arXiv1 repoarXiv:1904.03670
openWakeWord
StegaStamp: Invisible Hyperlinks in Physical Photographs
arXiv1 repoarXiv:1904.05343
WMCopier
The Android Platform Security Model (2023)
arXiv1 repoarXiv:1904.05572
apparmor.d
wav2vec: Unsupervised Pre-training for Speech Recognition
arXiv1 repoarXiv:1904.05862
dissertation-project
DocBERT: BERT for Document Classification
arXiv1 repoarXiv:1904.08398
NLP-DocBERT-financial-news-trading
SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
arXiv1 repoarXiv:1904.08779
s2t-small-librispeech-asr
Cantor-Bernstein implies Excluded Middle
arXiv1 repoarXiv:1904.09193
topology
ERNIE: Enhanced Representation through Knowledge Integration
arXiv1 repoarXiv:1904.09223
ernie-1.0-base-zh
SocialIQA: Commonsense Reasoning about Social Interactions
arXiv1 repoarXiv:1904.09728
DeepEnlighten
arXiv:1904.09751
arXiv1 repoarXiv:1904.09751
llama2.zig
Generating Long Sequences with Sparse Transformers
arXiv1 repoarXiv:1904.10509
VideoGPT
Genet: A Quickly Scalable Fat-Tree Overlay for Personal Volunteer Computing using WebRTC
arXiv1 repoarXiv:1904.11402
simple-peer
Local Relation Networks for Image Recognition
arXiv1 repoarXiv:1904.11491
Swin-Transformer
Style Transfer by Relaxed Optimal Transport and Self-Similarity
arXiv1 repoarXiv:1904.12785
wise
Drug-Drug Adverse Effect Prediction with Graph Co-Attention
arXiv1 repoarXiv:1905.00534
chemicalx
RetinaFace: Single-stage Dense Face Localisation in the Wild
arXiv1 repoarXiv:1905.00641
face-alignment
Searching for MobileNetV3
arXiv1 repoarXiv:1905.02244
yas
Taming Pretrained Transformers for Extreme Multi-label Text Classification
arXiv1 repoarXiv:1905.02331
pecos
Does Environmental Economics lead to patentable research?
arXiv1 repoarXiv:1905.02875
Kosmos-X
Harvey: A Greybox Fuzzer for Smart Contracts
arXiv1 repoarXiv:1905.06944
echidna
MR-GNN: Multi-Resolution and Dual Graph Neural Network for Predicting Structured Entity Interactions
arXiv1 repoarXiv:1905.09558
chemicalx
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
arXiv1 repoarXiv:1905.10044
t5-large-encoder-only-bf16
CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition
arXiv1 repoarXiv:1905.11235
ASR-Knowledge-Transferring
Mixed Precision DNNs: All you need is a good parametrization
arXiv1 repoarXiv:1905.11452
ai-research-code
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
arXiv1 repoarXiv:1905.11946
yolov5
Racial Bias in Hate Speech and Abusive Language Detection Datasets
arXiv1 repoarXiv:1905.12516
toxic-bert
Latent Retrieval for Weakly Supervised Open Domain Question Answering
arXiv1 repoarXiv:1906.00300
llm-jepa
Coresets for Data-efficient Training of Machine Learning Models
arXiv1 repoarXiv:1906.01827
cords
Gradio: Hassle-Free Sharing and Testing of ML Models in the Wild
arXiv1 repoarXiv:1906.02569
gradio
Visualizing and Measuring the Geometry of BERT
arXiv1 repoarXiv:1906.02715
adv-glue-plus-plus
Mesh R-CNN
arXiv1 repoarXiv:1906.02739
pytorch3d
Multi-hop Reading Comprehension through Question Decomposition and Rescoring
arXiv1 repoarXiv:1906.02916
query_decomposer
TransNet: A deep network for fast detection of common shot transitions
arXiv1 repoarXiv:1906.03363
TransNetV2
Robustness Verification of Tree-based Models
arXiv1 repoarXiv:1906.03849
RobustTrees
GluonTS: Probabilistic Time Series Models in Python
arXiv1 repoarXiv:1906.05264
gluonts
Image Captioning: Transforming Objects into Words
arXiv1 repoarXiv:1906.05963
object-relation-transformer
Scheduled Sampling for Transformers
arXiv1 repoarXiv:1906.07651
final-project-level3-nlp-02
The Second DIHARD Diarization Challenge: Dataset, task, and baselines
arXiv1 repoarXiv:1906.07839
speaker-diarization-benchmark
Multi-Span Acoustic Modelling using Raw Waveform Signals
arXiv1 repoarXiv:1906.11047
ami
The Indirect Convolution Algorithm
arXiv1 repoarXiv:1907.02129
XNNPACK
Zero-shot Learning for Audio-based Music Classification and Tagging
arXiv1 repoarXiv:1907.02670
fma
XGBoostLSS -- An extension of XGBoost to probabilistic forecasting
arXiv1 repoarXiv:1907.03178
LightGBMLSS
BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs
arXiv1 repoarXiv:1907.05047
face-alignment
Large Memory Layers with Product Keys
arXiv1 repoarXiv:1907.05242
XLM
Hello, It's GPT-2 -- How Can I Help You? Towards the Use of Pretrained Language Models for Task-Oriented Dialogue Systems
arXiv1 repoarXiv:1907.05774
KoGPT2-chatbot
Facebook FAIR's WMT19 News Translation Task Submission
arXiv1 repoarXiv:1907.06616
wmt19-en-ru
Efficient Pipeline for Camera Trap Image Review
arXiv1 repoarXiv:1907.06772
Depth-Estimation
SpanBERT: Improving Pre-training by Representing and Predicting Spans
arXiv1 repoarXiv:1907.10529
odqa_baseline_code
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
arXiv1 repoarXiv:1907.10641
Kosmos-X
ConCert: A Smart Contract Certification Framework in Coq
arXiv1 repoarXiv:1907.10674
ConCert
MaskGAN: Towards Diverse and Interactive Facial Image Manipulation
arXiv1 repoarXiv:1907.11922
CelebAMask-HQ
Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
arXiv1 repoarXiv:1907.11932
DeBERTa_TxtClassifier
Observability and Chaos Engineering on System Calls for Containerized Applications in Docker
arXiv1 repoarXiv:1907.13039
awesome-chaos-engineering
arXiv:1907.13440
arXiv1 repoarXiv:1907.13440
minerl
VisualBERT: A Simple and Performant Baseline for Vision and Language
arXiv1 repoarXiv:1908.03557
visual-spatial-reasoning
Neural Text Generation with Unlikelihood Training
arXiv1 repoarXiv:1908.04319
calibrating-summaries
Aspect and Opinion Terms Extraction Using Double Embeddings and Attention Mechanism for Indonesian Hotel Reviews
arXiv1 repoarXiv:1908.04899
indonlu
LXMERT: Learning Cross-Modality Encoder Representations from Transformers
arXiv1 repoarXiv:1908.07490
visual-spatial-reasoning
An Empirical Study of Incorporating Pseudo Data into Grammatical Error Correction
arXiv1 repoarXiv:1909.00502
CTCResources
Robust Invisible Video Watermarking with Attention
arXiv1 repoarXiv:1909.01285
WMCopier
Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop
arXiv1 repoarXiv:1909.01362
ConvoKit
An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction
arXiv1 repoarXiv:1909.02027
task-aware-embedding-refinement
Neural Machine Translation with Byte-Level Subwords
arXiv1 repoarXiv:1909.03341
vocab-coverage
KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension
arXiv1 repoarXiv:1909.07005
squad_kor_v1
Ludwig: a type-based declarative deep learning toolbox
arXiv1 repoarXiv:1909.07930
ludwig
Aspect and Opinion Term Extraction for Hotel Reviews using Transfer Learning and Auxiliary Labels
arXiv1 repoarXiv:1909.11879
indonlu
A Pilot Study for Chinese SQL Semantic Parsing
arXiv1 repoarXiv:1909.13293
ChineseNLPCorpus
Interpretations are useful: penalizing explanations to align neural networks with prior knowledge
arXiv1 repoarXiv:1909.13584
imodelsX
Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
arXiv1 repoarXiv:1910.00177
trax
OpenVSLAM: A Versatile Visual SLAM Framework
arXiv1 repoarXiv:1910.01122
stella_vslam
MLPerf Training Benchmark
arXiv1 repoarXiv:1910.01500
training
NGBoost: Natural Gradient Boosting for Probabilistic Prediction
arXiv1 repoarXiv:1910.03225
ngboost
Base64 encoding and decoding at almost the speed of a memory copy
arXiv1 repoarXiv:1910.05109
simdutf
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
arXiv1 repoarXiv:1910.05453
dissertation-project
JSDoop and TensorFlow.js: Volunteer Distributed Web Browser-Based Neural Network Training
arXiv1 repoarXiv:1910.07402
awesome-tensorflow-js
MLQA: Evaluating Cross-lingual Extractive Question Answering
arXiv1 repoarXiv:1910.07475
Multilingual-MiniLM-L12-H384
Can I teach a robot to replicate a line art
arXiv1 repoarXiv:1910.07860
awesome-plotters
Using Speech Synthesis to Train End-to-End Spoken Language Understanding Models
arXiv1 repoarXiv:1910.09463
openWakeWord
Hierarchical Transformers for Long Document Classification
arXiv1 repoarXiv:1910.10781
NLP-DocBERT-financial-news-trading
A Unified MRC Framework for Named Entity Recognition
arXiv1 repoarXiv:1910.11476
nlp-bazel-tutorial
Confident Learning: Estimating Uncertainty in Dataset Labels
arXiv1 repoarXiv:1911.00068
cleanlab
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data
arXiv1 repoarXiv:1911.00359
falcon-refinedweb
On the Measure of Intelligence
arXiv1 repoarXiv:1911.01547
logicmoo_workspace
arXiv:1911.01601
arXiv1 repoarXiv:1911.01601
MTP-Codebase
Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports
arXiv1 repoarXiv:1911.02541
vilmedic
MLPerf Inference Benchmark
arXiv1 repoarXiv:1911.02549
inference
Scalable Zero-shot Entity Linking with Dense Entity Retrieval
arXiv1 repoarXiv:1911.03814
trusted_ke
Knowledge Guided Text Retrieval and Reading for Open Domain Question Answering
arXiv1 repoarXiv:1911.03868
DPR
Smart Contract Interactions in Coq
arXiv1 repoarXiv:1911.04732
ConCert
Identification of Rhetorical Roles of Sentences in Indian Legal Judgments
arXiv1 repoarXiv:1911.05405
InLegalBERT
CASTER: Predicting Drug Interactions with Chemical Substructure Representation
arXiv1 repoarXiv:1911.06446
chemicalx
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
arXiv1 repoarXiv:1911.08265
xwm
Semantic Hierarchy Emerges in Deep Generative Representations for Scene Synthesis
arXiv1 repoarXiv:1911.09267
materialgan
Fast Sparse ConvNets
arXiv1 repoarXiv:1911.09723
XNNPACK
Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack Integration
arXiv1 repoarXiv:1911.09925
chipyard
Causality for Machine Learning
arXiv1 repoarXiv:1911.10500
Data-Science
Image-based table recognition: data, model, and evaluation
arXiv1 repoarXiv:1911.10683
opendataloader-bench
PIQA: Reasoning about Physical Commonsense in Natural Language
arXiv1 repoarXiv:1911.11641
Kosmos-X
SuperGlue: Learning Feature Matching with Graph Neural Networks
arXiv1 repoarXiv:1911.11763
gtsfm
CSPNet: A New Backbone that can Enhance Learning Capability of CNN
arXiv1 repoarXiv:1911.11929
darknet
Pythia: AI-assisted Code Completion System
arXiv1 repoarXiv:1912.00742
CodeXGLUE
Analyzing and Improving the Image Quality of StyleGAN
arXiv1 repoarXiv:1912.04958
materialgan
Common Voice: A Massively-Multilingual Speech Corpus
arXiv1 repoarXiv:1912.06670
covost
SynSin: End-to-end View Synthesis from a Single Image
arXiv1 repoarXiv:1912.08804
pytorch3d
Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting
arXiv1 repoarXiv:1912.09363
frn-50k-baseline
PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
arXiv1 repoarXiv:1912.10211
audio-embeddings
Big Transfer (BiT): General Visual Representation Learning
arXiv1 repoarXiv:1912.11370
bit-50
Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro
arXiv1 repoarXiv:1912.11554
numpyro
A Gentle Introduction to Deep Learning for Graphs
arXiv1 repoarXiv:1912.12693
Data-Science
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
arXiv1 repoarXiv:1912.13318
layoutlm-base-cased
The LDBC Social Network Benchmark
arXiv1 repoarXiv:2001.02299
duckdb
Efficient Memory Management for Deep Neural Net Inference
arXiv1 repoarXiv:2001.03288
XNNPACK
The Two-Pass Softmax Algorithm
arXiv1 repoarXiv:2001.04438
XNNPACK
Reformer: The Efficient Transformer
arXiv1 repoarXiv:2001.04451
trax
SQLFlow: A Bridge between SQL and Machine Learning
arXiv1 repoarXiv:2001.06846
sqlflow
Fast Sequence-Based Embedding with Diffusion Graphs
arXiv1 repoarXiv:2001.07463
littleballoffur
Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference
arXiv1 repoarXiv:2001.07676
pet
Scaling Laws for Neural Language Models
arXiv1 repoarXiv:2001.08361
parameter-golf
Towards Measuring Supply Chain Attacks on Package Managers for Interpreted Languages
arXiv1 repoarXiv:2002.01139
guarddog
CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus
arXiv1 repoarXiv:2002.01320
covost
Training Keyword Spotters with Limited and Synthesized Speech Data
arXiv1 repoarXiv:2002.01322
openWakeWord
Using Fractal Neural Networks to Play SimCity 1 and Conway's Game of Life at Variable Scales
arXiv1 repoarXiv:2002.03896
MicropolisCore
A Simple General Approach to Balance Task Difficulty in Multi-Task Learning
arXiv1 repoarXiv:2002.04792
Fine-Grained_Features_Alignment_via_Constrastive_Learning
Causality in cognitive neuroscience: concepts, challenges, and distributional robustness
arXiv1 repoarXiv:2002.06060
Data-Science
CodeBERT: A Pre-Trained Model for Programming and Natural Languages
arXiv1 repoarXiv:2002.08155
CodeXGLUE
Scalable Second Order Optimization for Deep Learning
arXiv1 repoarXiv:2002.09018
modded-nanogpt
Unsupervised Question Decomposition for Question Answering
arXiv1 repoarXiv:2002.09758
query_decomposer
Resources for Turkish Dependency Parsing: Introducing the BOUN Treebank and the BoAT Annotation Tool
arXiv1 repoarXiv:2002.10416
Word-Embeddings-Repository-for-Turkish
CausalML: Python Package for Causal Machine Learning
arXiv1 repoarXiv:2002.11631
causalml
Gradient Boosted Normalizing Flows
arXiv1 repoarXiv:2002.11896
gradient-boosted-normalizing-flows
OpEn: Code Generation for Embedded Nonconvex Optimization
arXiv1 repoarXiv:2003.00292
optimization-engine
CheXclusion: Fairness gaps in deep chest X-ray classifiers
arXiv1 repoarXiv:2003.00827
Fine-Grained_Features_Alignment_via_Constrastive_Learning
Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap
arXiv1 repoarXiv:2003.02405
wordcab-transcribe
Combining GHOST and Casper
arXiv1 repoarXiv:2003.03052
consensus-specs
Document Ranking with a Pretrained Sequence-to-Sequence Model
arXiv1 repoarXiv:2003.06713
pygaggle
Stanza: A Python Natural Language Processing Toolkit for Many Human Languages
arXiv1 repoarXiv:2003.07082
stanza
Dash: Scalable Hashing on Persistent Memory
arXiv1 repoarXiv:2003.07302
dragonfly
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
arXiv1 repoarXiv:2003.08934
nerf
A Trustful Monad for Axiomatic Reasoning with Probability and Nondeterminism
arXiv1 repoarXiv:2003.09993
monae
Two-stage Discriminative Re-ranking for Large-scale Landmark Retrieval
arXiv1 repoarXiv:2003.11211
google-landmark
Similarity of Neural Networks with Gradients
arXiv1 repoarXiv:2003.11498
xfer
COVID-19 Image Data Collection
arXiv1 repoarXiv:2003.11597
covid-chestxray-dataset
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
arXiv1 repoarXiv:2003.12039
prisma
SaccadeNet: A Fast and Accurate Object Detector
arXiv1 repoarXiv:2003.12125
foveate
Computer Aided Detection for Pulmonary Embolism Challenge (CAD-PE)
arXiv1 repoarXiv:2003.13440
Project-Imaging-X
Information Leakage in Embedding Models
arXiv1 repoarXiv:2004.00053
langtest
Merkle-CRDTs: Merkle-DAGs meet CRDTs
arXiv1 repoarXiv:2004.00107
defradb
CurricularFace: Adaptive Curriculum Learning Loss for Deep Face Recognition
arXiv1 repoarXiv:2004.00288
FLUXSynID
arXiv:2004.00584
arXiv1 repoarXiv:2004.00584
Jellyfish-13B
RisGraph: A Real-Time Streaming System for Evolving Graphs to Support Sub-millisecond Per-update Analysis at Millions Ops/s
arXiv1 repoarXiv:2004.00803
awesome-dynamic-graphs
Google Landmarks Dataset v2 -- A Large-Scale Benchmark for Instance-Level Recognition and Retrieval
arXiv1 repoarXiv:2004.01804
google-landmark
TAPAS: Weakly Supervised Table Parsing via Pre-training
arXiv1 repoarXiv:2004.02349
tapas-base-finetuned-wtq
Deep Learning Based Text Classification: A Comprehensive Review
arXiv1 repoarXiv:2004.03705
OpenTextClassification
On the Effect of Dropping Layers of Pre-trained Transformer Models
arXiv1 repoarXiv:2004.03844
final-project-level3-nlp-02
Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence
arXiv1 repoarXiv:2004.03974
turftopic
Unveiling COVID-19 from Chest X-ray with deep learning: a hurdles race with small data
arXiv1 repoarXiv:2004.05405
covid-chestxray-dataset
Minimizing FLOPs to Learn Efficient Sparse Representations
arXiv1 repoarXiv:2004.05665
splade-ecommerce-esci
Benchmarking Unsupervised Outlier Detection with Realistic Synthetic Data
arXiv1 repoarXiv:2004.06947
Data-Science
Image Quality Assessment: Unifying Structure and Texture Similarity
arXiv1 repoarXiv:2004.07728
DISTS
Classification Benchmarks for Under-resourced Bengali Language based on Multichannel Convolutional-LSTM Network
arXiv1 repoarXiv:2004.07807
bangla-electra
SongNet: Rigid Formats Controlled Text Generation
arXiv1 repoarXiv:2004.08022
textgen
arXiv:2004.08483
arXiv1 repoarXiv:2004.08483
Block-Sparse-Attention
Data Efficient and Weakly Supervised Computational Pathology on Whole Slide Images
arXiv1 repoarXiv:2004.09666
CLAM
DIET: Lightweight Language Understanding for Dialogue Systems
arXiv1 repoarXiv:2004.09936
datacopilot
Yoga-82: A New Dataset for Fine-grained Classification of Human Poses
arXiv1 repoarXiv:2004.10362
ai-yoga-trainer
ivis Dimensionality Reduction Framework for Biomacromolecular Simulations
arXiv1 repoarXiv:2004.10718
SciencePlots
DuReader_robust: A Chinese Dataset Towards Evaluating Robustness and Generalization of Machine Reading Comprehension in Real-World Applications
arXiv1 repoarXiv:2004.11142
DuReader
Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
arXiv1 repoarXiv:2004.11867
opus-100
Formal Adventures in Convex and Conical Spaces
arXiv1 repoarXiv:2004.12713
infotheo
"Call me sexist, but...": Revisiting Sexism Detection Using Psychological Scales and Adversarial Samples
arXiv1 repoarXiv:2004.12764
DeepLearningProject
A Critic Evaluation of Methods for COVID-19 Automatic Detection from X-Ray Images
arXiv1 repoarXiv:2004.12823
covid-chestxray-dataset
Reevaluating Adversarial Examples in Natural Language
arXiv1 repoarXiv:2004.14174
TextAttack
VGGSound: A Large-scale Audio-Visual Dataset
arXiv1 repoarXiv:2004.14368
VGGSound
Fact or Fiction: Verifying Scientific Claims
arXiv1 repoarXiv:2004.14974
scifact
UnifiedQA: Crossing Format Boundaries With a Single QA System
arXiv1 repoarXiv:2005.00700
mmlu
Feature Selection Methods for Uplift Modeling and Heterogeneous Treatment Effect
arXiv1 repoarXiv:2005.03447
causalml
Beyond Accuracy: Behavioral Testing of NLP models with CheckList
arXiv1 repoarXiv:2005.04118
langtest
arXiv:2005.05110
arXiv1 repoarXiv:2005.05110
misp-galaxy
TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
arXiv1 repoarXiv:2005.05909
TextAttack
Arabic Dialect Identification in the Wild
arXiv1 repoarXiv:2005.06557
Glot500
ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
arXiv1 repoarXiv:2005.07143
Displace2024_baseline_updated
IntelliCode Compose: Code Generation Using Transformer
arXiv1 repoarXiv:2005.08025
CodeXGLUE
Efficient Wait-k Models for Simultaneous Machine Translation
arXiv1 repoarXiv:2005.08595
translation-api
U$^2$-Net: Going Deeper with Nested U-Structure for Salient Object Detection
arXiv1 repoarXiv:2005.09007
U-2-Net
Backstabber's Knife Collection: A Review of Open Source Software Supply Chain Attacks
arXiv1 repoarXiv:2005.09535
guarddog
Lung Segmentation from Chest X-rays using Variational Data Imputation
arXiv1 repoarXiv:2005.10052
covid-chestxray-dataset
BERTweet: A pre-trained language model for English Tweets
arXiv1 repoarXiv:2005.10200
tweeteval
arXiv:2005.10356
arXiv1 repoarXiv:2005.10356
ReVOS-api
LibriMix: An Open-Source Dataset for Generalizable Speech Separation
arXiv1 repoarXiv:2005.11262
LibriMix
Predicting COVID-19 Pneumonia Severity on Chest X-ray with Deep Learning
arXiv1 repoarXiv:2005.11856
covid-chestxray-dataset
NDD20: A large-scale few-shot dolphin dataset for coarse and fine-grained categorisation
arXiv1 repoarXiv:2005.13359
RobustSAM
NuClick: A Deep Learning Framework for Interactive Segmentation of Microscopy Images
arXiv1 repoarXiv:2005.14511
DinoBloom
Massive Choice, Ample Tasks (MaChAmp): A Toolkit for Multi-task Learning in NLP
arXiv1 repoarXiv:2005.14672
machamp
A Scalable and Cloud-Native Hyperparameter Tuning System
arXiv1 repoarXiv:2006.02085
awesome-kubeflow
Unsupervised Translation of Programming Languages
arXiv1 repoarXiv:2006.03511
CodeXGLUE
Little Ball of Fur: A Python Library for Graph Sampling
arXiv1 repoarXiv:2006.04311
littleballoffur
Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection
arXiv1 repoarXiv:2006.04388
nanodet
Maximizing Error Injection Realism for Chaos Engineering with System Calls
arXiv1 repoarXiv:2006.04444
awesome-chaos-engineering
BS-Net: learning COVID-19 pneumonia severity on a large Chest X-Ray dataset
arXiv1 repoarXiv:2006.04603
covid-chestxray-dataset
Active Invariant Causal Prediction: Experiment Selection through Stability
arXiv1 repoarXiv:2006.05690
Data-Science
TableQA: a Large-Scale Chinese Text-to-SQL Dataset for Table-Aware SQL Generation
arXiv1 repoarXiv:2006.06434
ChineseNLPCorpus
Training Generative Adversarial Networks with Limited Data
arXiv1 repoarXiv:2006.06676
stylegan2-ada-pytorch
FastPitch: Parallel Text-to-speech with Pitch Prediction
arXiv1 repoarXiv:2006.06873
tts_en_fastpitch
Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
arXiv1 repoarXiv:2006.09092
openchat
New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles
arXiv1 repoarXiv:2006.11138
ViQG
COVID-19 Image Data Collection: Prospective Predictions Are the Future
arXiv1 repoarXiv:2006.11988
covid-chestxray-dataset
Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules
arXiv1 repoarXiv:2006.16981
parallel-ss-dep
Playing with Words at the National Library of Sweden -- Making a Swedish BERT
arXiv1 repoarXiv:2007.01658
bert-base-swedish-cased
Detailed spectroscopy of doubly magic $^{132}$Sn
arXiv1 repoarXiv:2007.03029
awsome-llm-papers
MosAIc: Finding Artistic Connections across Culture with Conditional Image Retrieval
arXiv1 repoarXiv:2007.07177
SynapseML
LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning
arXiv1 repoarXiv:2007.08124
logiqa
Accelerating 3D Deep Learning with PyTorch3D
arXiv1 repoarXiv:2007.08501
pytorch3d
CoVoST 2 and Massively Multilingual Speech-to-Text Translation
arXiv1 repoarXiv:2007.10310
covost
Biomedical and Clinical English Model Packages in the Stanza Python NLP Library
arXiv1 repoarXiv:2007.14640
stanza
MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering
arXiv1 repoarXiv:2007.15207
gte-multilingual-base
Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
arXiv1 repoarXiv:2007.15779
BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext
A Survey on Text Classification: From Shallow to Deep Learning
arXiv1 repoarXiv:2008.00364
OpenTextClassification
Aligning AI With Shared Human Values
arXiv1 repoarXiv:2008.02275
mmlu
Shonan Rotation Averaging: Global Optimality by Surfing $SO(p)^n$
arXiv1 repoarXiv:2008.02737
gtsfm
A Parallel Evaluation Data Set of Software Documentation with Document Structure Annotation
arXiv1 repoarXiv:2008.04550
Glot500
Object Detection for Graphical User Interface: Old Fashioned or Deep Learning or a Combination?
arXiv1 repoarXiv:2008.05132
transformer-final-proj
An Experimental Study of Deep Neural Network Models for Vietnamese Multiple-Choice Reading Comprehension
arXiv1 repoarXiv:2008.08810
ViQG
Top2Vec: Distributed Representations of Topics
arXiv1 repoarXiv:2008.09470
turftopic
A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
arXiv1 repoarXiv:2008.10010
Wav2Lip
An Ensemble of Simple Convolutional Neural Network Models for MNIST Digit Recognition
arXiv1 repoarXiv:2008.10400
MaaAI
Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization
arXiv1 repoarXiv:2008.11293
mslr-shared-task
Automatic Yara Rule Generation Using Biclustering
arXiv1 repoarXiv:2009.03779
awesome-ai-security-tools
Phasic Policy Gradient
arXiv1 repoarXiv:2009.04416
cleanrl
An Open-Source Platform for High-Performance Non-Coherent On-Chip Communication
arXiv1 repoarXiv:2009.05334
axi
Searching for a Search Method: Benchmarking Search Algorithms for Generating NLP Adversarial Examples
arXiv1 repoarXiv:2009.06368
TextAttack
It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
arXiv1 repoarXiv:2009.07118
pet
DLBCL-Morph: Morphological features computed using deep learning for an annotated digital DLBCL image set
arXiv1 repoarXiv:2009.08123
DLBCL-Morph
FewJoint: A Few-shot Learning Benchmark for Joint Language Understanding
arXiv1 repoarXiv:2009.08138
ChineseNLPCorpus
Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering
arXiv1 repoarXiv:2009.10791
odqa_baseline_code
Qlib: An AI-oriented Quantitative Investment Platform
arXiv1 repoarXiv:2009.11189
qlib
Probabilistic Label Trees for Extreme Multi-label Classification
arXiv1 repoarXiv:2009.11218
napkinXC
Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress
arXiv1 repoarXiv:2009.13807
Large-Time-Series-Model
PettingZoo: Gym for Multi-Agent Reinforcement Learning
arXiv1 repoarXiv:2009.14471
PettingZoo
A Vietnamese Dataset for Evaluating Machine Reading Comprehension
arXiv1 repoarXiv:2009.14725
ViQG
Understanding tables with intermediate pre-training
arXiv1 repoarXiv:2010.00571
tapas-base-finetuned-wtq
Contrastive Learning of Medical Visual Representations from Paired Images and Text
arXiv1 repoarXiv:2010.00747
clip-image-search
Sharpness-Aware Minimization for Efficiently Improving Generalization
arXiv1 repoarXiv:2010.01412
vision_transformer
Learning from Context or Names? An Empirical Study on Neural Relation Extraction
arXiv1 repoarXiv:2010.01923
RE-Context-or-Names
Constraining Logits by Bounded Function for Adversarial Robustness
arXiv1 repoarXiv:2010.02558
frn-50k-baseline
Inductive Entity Representations from Text via Link Prediction
arXiv1 repoarXiv:2010.03496
mkb
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
arXiv1 repoarXiv:2010.03768
alfworld
All for One and One for All: Improving Music Separation by Bridging Networks
arXiv1 repoarXiv:2010.04228
ai-research-code
Load What You Need: Smaller Versions of Multilingual BERT
arXiv1 repoarXiv:2010.05609
smaller-transformers
PECOS: Prediction for Enormous and Correlated Output Spaces
arXiv1 repoarXiv:2010.05878
pecos
arXiv:2010.07115
arXiv1 repoarXiv:2010.07115
WasmEdge
Dimsum @LaySumm 20: BART-based Approach for Scientific Document Summarization
arXiv1 repoarXiv:2010.09252
Laysumm
DiDiSpeech: A Large Scale Mandarin Speech Corpus
arXiv1 repoarXiv:2010.09275
seed-tts-eval
PySBD: Pragmatic Sentence Boundary Disambiguation
arXiv1 repoarXiv:2010.09657
pySBD
Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation
arXiv1 repoarXiv:2010.10042
vilmedic
German's Next Language Model
arXiv1 repoarXiv:2010.10906
berts
Uncovering the Hidden Dangers: Finding Unsafe Go Code in the Wild
arXiv1 repoarXiv:2010.11242
go-recipes
Self-training and Pre-training are Complementary for Speech Recognition
arXiv1 repoarXiv:2010.11430
fairseq
Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text Classification
arXiv1 repoarXiv:2010.12512
PIXIU
Pre-trained Summarization Distillation
arXiv1 repoarXiv:2010.13002
kotoba-whisper
Out-of-core Training for Extremely Large-Scale Neural Networks With Adaptive Window-Based Scheduling
arXiv1 repoarXiv:2010.14109
ai-research-code
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
arXiv1 repoarXiv:2010.14406
saycanpay
Generating Radiology Reports via Memory-driven Transformer
arXiv1 repoarXiv:2010.16056
vilmedic
Joint Masked CPC and CTC Training for ASR
arXiv1 repoarXiv:2011.00093
voxpopuli
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP
arXiv1 repoarXiv:2011.00677
indobert-base-uncased
Tabular Transformers for Modeling Multivariate Time Series
arXiv1 repoarXiv:2011.01843
TabFormer
Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
arXiv1 repoarXiv:2011.02523
ml-hypersim
A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text Systems
arXiv1 repoarXiv:2011.03992
awesome-nlg
Dual-stream Multiple Instance Learning Network for Whole Slide Image Classification with Self-supervised Contrastive Learning
arXiv1 repoarXiv:2011.08939
HistoSSLscaling
QuerYD: A video dataset with high-quality text and audio narrations
arXiv1 repoarXiv:2011.11071
TimeChat-Online-139K
Densely connected multidilated convolutional networks for dense prediction tasks
arXiv1 repoarXiv:2011.11844
ai-research-code
CPM: A Large-scale Generative Chinese Pre-trained Language Model
arXiv1 repoarXiv:2012.00413
CPM-Generate
Exploring the Effect of Image Enhancement Techniques on COVID-19 Detection using Chest X-rays Images
arXiv1 repoarXiv:2012.02238
baple
FloodNet: A High Resolution Aerial Imagery Dataset for Post Flood Scene Understanding
arXiv1 repoarXiv:2012.02951
FloodNet-Challenge-EARTHVISION2021
MLS: A Large-Scale Multilingual Dataset for Speech Research
arXiv1 repoarXiv:2012.03411
multilingual_librispeech
Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
arXiv1 repoarXiv:2012.07436
informer-tourism-monthly
Session-Aware Query Auto-completion using Extreme Multi-label Ranking
arXiv1 repoarXiv:2012.07654
pecos
Taming Transformers for High-Resolution Image Synthesis
arXiv1 repoarXiv:2012.09841
dalle-mini
SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
arXiv1 repoarXiv:2012.09852
deepcompressor
Few-Shot Text Generation with Pattern-Exploiting Training
arXiv1 repoarXiv:2012.11926
pet
Learning Dense Representations of Phrases at Scale
arXiv1 repoarXiv:2012.12624
SimCSE
Training data-efficient image transformers & distillation through attention
arXiv1 repoarXiv:2012.12877
Vim
Towards Fully Automated Manga Translation
arXiv1 repoarXiv:2012.14271
manga-image-translator
arXiv:2012.14913
arXiv1 repoarXiv:2012.14913
KEditVis-LLM-Editing
How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
arXiv1 repoarXiv:2012.15613
hgiyt
Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection
arXiv1 repoarXiv:2012.15761
roberta-hate-speech-dynabench-r4-target
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
arXiv1 repoarXiv:2012.15840
FoodSeg103-Benchmark-v1
I-BERT: Integer-only BERT Quantization
arXiv1 repoarXiv:2101.01321
gemmini
Dynamic Hybrid Relation Network for Cross-Domain Context-Dependent Semantic Parsing
arXiv1 repoarXiv:2101.01686
DAMO-ConvAI
The Shapley Value of Classifiers in Ensemble Games
arXiv1 repoarXiv:2101.02153
shapley
Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
arXiv1 repoarXiv:2101.02235
query_decomposer
Towards Real-World Blind Face Restoration with Generative Facial Prior
arXiv1 repoarXiv:2101.04061
GFPGAN
MLGO: a Machine Learning Guided Compiler Optimizations Framework
arXiv1 repoarXiv:2101.04808
ml-compiler-opt
The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
arXiv1 repoarXiv:2101.05667
MonoQwen2-VL-v0.1
Persistent Anti-Muslim Bias in Large Language Models
arXiv1 repoarXiv:2101.05783
clip-italian
UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
arXiv1 repoarXiv:2101.07597
unispeech-large-1500h-cv
WangchanBERTa: Pretraining transformer-based Thai Language Models
arXiv1 repoarXiv:2101.09635
SEA-PILE-v1
VisualMRC: Machine Reading Comprehension on Document Images
arXiv1 repoarXiv:2101.11272
VisualMRC
MultiRocket: Multiple pooling operators and transformations for fast and effective time series classification
arXiv1 repoarXiv:2102.00457
FM4Motor
Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details
arXiv1 repoarXiv:2102.01066
GLIP
PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation
arXiv1 repoarXiv:2102.01243
ast
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
arXiv1 repoarXiv:2102.04664
CodeXGLUE
Proof Artifact Co-training for Theorem Proving with Language Models
arXiv1 repoarXiv:2102.06203
llmstep-mathlib4-pythia2.8b
Neural Network Libraries: A Deep Learning Framework Designed from Engineers' Perspectives
arXiv1 repoarXiv:2102.06725
nnabla
Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm
arXiv1 repoarXiv:2102.07350
awesome-prompt-engineering
Top-$k$ eXtreme Contextual Bandits with Arm Hierarchy
arXiv1 repoarXiv:2102.07800
pecos
TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models
arXiv1 repoarXiv:2102.07988
xDiT
SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
arXiv1 repoarXiv:2102.09542
LLaDA-MedV
Pyserini: An Easy-to-Use Python Toolkit to Support Replicable IR Research with Sparse and Dense Representations
arXiv1 repoarXiv:2102.10073
odqa_baseline_code
LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
arXiv1 repoarXiv:2102.10815
univnet
LogME: Practical Assessment of Pre-trained Models for Transfer Learning
arXiv1 repoarXiv:2102.11005
FM4Motor
Design and Analysis of a Logless Dynamic Reconfiguration Protocol
arXiv1 repoarXiv:2102.11960
mongo
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions
arXiv1 repoarXiv:2102.12122
FoodSeg103-Benchmark-v1
PyCG: Practical Call Graph Generation in Python
arXiv1 repoarXiv:2103.00587
PyCG
Fast Adaptation with Linearized Neural Networks
arXiv1 repoarXiv:2103.01439
xfer
Predicting Video with VQVAE
arXiv1 repoarXiv:2103.01950
Kosmos-X
Who Can Find My Devices? Security and Privacy of Apple's Crowd-Sourced Bluetooth Location Tracking System
arXiv1 repoarXiv:2103.02282
openhaystack
Ribbon filter: practically smaller than Bloom and Xor
arXiv1 repoarXiv:2103.02515
gecko-dev
Catala: A Programming Language for the Law
arXiv1 repoarXiv:2103.03198
catala
Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food
arXiv1 repoarXiv:2103.03375
MacroScope-Portion-Scout
Measuring Mathematical Problem Solving With the MATH Dataset
arXiv1 repoarXiv:2103.03874
FTTT
ModelingToolkit: A Composable Graph Transformation System For Equation-Based Modeling
arXiv1 repoarXiv:2103.05244
ModelingToolkit.jl
Unknown Object Segmentation from Stereo Images
arXiv1 repoarXiv:2103.06796
humanoid_grasping
EXSCLAIM! -- An automated pipeline for the construction of labeled materials imaging datasets from literature
arXiv1 repoarXiv:2103.10631
exsclaim2.0
MuRIL: Multilingual Representations for Indian Languages
arXiv1 repoarXiv:2103.10730
IndicAbusive
Data Cleansing for Deep Neural Networks with Storage-efficient Approximation of Influence Functions
arXiv1 repoarXiv:2103.11807
ai-research-code
Open Domain Question Answering over Tables via Dense Retrieval
arXiv1 repoarXiv:2103.12011
tapas
Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes
arXiv1 repoarXiv:2103.14127
humanoid_grasping
Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
arXiv1 repoarXiv:2103.14749
cleanlab
VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization
arXiv1 repoarXiv:2103.16874
VITON-HD
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
arXiv1 repoarXiv:2104.00650
webvid
LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
arXiv1 repoarXiv:2104.01136
MiDaS
A flexible and fast PyTorch toolkit for simulating training and inference on analog crossbar arrays
arXiv1 repoarXiv:2104.02184
aihwkit
Image Composition Assessment with Saliency-augmented Multi-pattern Pooling
arXiv1 repoarXiv:2104.03133
compositio_nn
Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings
arXiv1 repoarXiv:2104.03502
tmh
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
arXiv1 repoarXiv:2104.04473
MindSpeed-MM
arXiv:2104.04691
arXiv1 repoarXiv:2104.04691
ReVOS-api
An Efficient 2D Method for Training Super-Large Deep Learning Models
arXiv1 repoarXiv:2104.05343
ColossalAI
Getting to the Point. Index Sets and Parallelism-Preserving Autodiff for Pointful Array Programming
arXiv1 repoarXiv:2104.05372
dex-lang
Rapid Exploration for Open-World Navigation with Latent Goal Models
arXiv1 repoarXiv:2104.05859
UniWM_Dataset
Learning and Planning in Complex Action Spaces
arXiv1 repoarXiv:2104.06303
xwm
MS2: Multi-Document Summarization of Medical Studies
arXiv1 repoarXiv:2104.06486
mslr-shared-task
TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning
arXiv1 repoarXiv:2104.06979
Vorbereitung
EAT: Enhanced ASR-TTS for Self-supervised Speech Recognition
arXiv1 repoarXiv:2104.07474
parrots
The Power of Scale for Parameter-Efficient Prompt Tuning
arXiv1 repoarXiv:2104.08691
Continual-NExT
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
arXiv1 repoarXiv:2104.08758
falcon-refinedweb
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
arXiv1 repoarXiv:2104.08860
CLIP4Clip
Evaluating Deep Neural Networks Trained on Clinical Images in Dermatology with the Fitzpatrick 17k Dataset
arXiv1 repoarXiv:2104.09957
fitzpatrick17k
VideoGPT: Video Generation using VQ-VAE and Transformers
arXiv1 repoarXiv:2104.10157
VideoGPT
ImageNet-21K Pretraining for the Masses
arXiv1 repoarXiv:2104.10972
convnext
Multiscale Vision Transformers
arXiv1 repoarXiv:2104.11227
pytorchvideo
Morph Call: Probing Morphosyntactic Content of Multilingual Transformers
arXiv1 repoarXiv:2104.12847
morph-call
TRECVID 2020: A comprehensive campaign for evaluating video retrieval tasks across multiple application domains
arXiv1 repoarXiv:2104.13473
ladi-overview
arXiv:2104.13921
arXiv1 repoarXiv:2104.13921
Echo-ViLD
Emerging Properties in Self-Supervised Vision Transformers
arXiv1 repoarXiv:2104.14294
dino
Conversational Machine Reading Comprehension for Vietnamese Healthcare Texts
arXiv1 repoarXiv:2105.01542
ViQG
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism
arXiv1 repoarXiv:2105.02446
fish-diffusion
What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
arXiv1 repoarXiv:2105.02732
whatsinthebox
A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers
arXiv1 repoarXiv:2105.03011
qasper
ResMLP: Feedforward networks for image classification with data-efficient training
arXiv1 repoarXiv:2105.03404
Swin-Transformer
High-performance symbolic-numerics via multiple dispatch
arXiv1 repoarXiv:2105.03949
Symbolics.jl
Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning
arXiv1 repoarXiv:2105.04165
InterGPS
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
arXiv1 repoarXiv:2105.04906
xwm
Diffusion Models Beat GANs on Image Synthesis
arXiv1 repoarXiv:2105.05233
guided-diffusion
Snipuzz: Black-box Fuzzing of IoT Firmware via Message Snippet Inference
arXiv1 repoarXiv:2105.05445
awesome-connected-things-sec
Monash Time Series Forecasting Archive
arXiv1 repoarXiv:2105.06643
UTSD
A cost-benefit analysis of cross-lingual transfer methods
arXiv1 repoarXiv:2105.06813
mmarco
Few-NERD: A Few-Shot Named Entity Recognition Dataset
arXiv1 repoarXiv:2105.07464
Few-NERD
NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions
arXiv1 repoarXiv:2105.08276
NExT-QA
Progressively Normalized Self-Attention Network for Video Polyp Segmentation
arXiv1 repoarXiv:2105.08468
VPS
Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech
arXiv1 repoarXiv:2105.09040
UniSpeech
Methods for Detoxification of Texts for the Russian Language
arXiv1 repoarXiv:2105.09052
rudetoxifier
DeepCAD: A Deep Generative Network for Computer-Aided Design Models
arXiv1 repoarXiv:2105.09492
DeepCAD
FreshDiskANN: A Fast and Accurate Graph-Based ANN Index for Streaming Similarity Search
arXiv1 repoarXiv:2105.09613
slater
Unsupervised Speech Recognition
arXiv1 repoarXiv:2105.11084
fairseq
Read, Listen, and See: Leveraging Multimodal Information Helps Chinese Spell Checking
arXiv1 repoarXiv:2105.12306
CTCResources
Sequence Parallelism: Long Sequence Training from System Perspective
arXiv1 repoarXiv:2105.13120
ColossalAI
ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and Explanation
arXiv1 repoarXiv:2105.13562
InLegalBERT
An Attention Free Transformer
arXiv1 repoarXiv:2105.14103
rwkv
Maximizing Parallelism in Distributed Training for Huge Neural Networks
arXiv1 repoarXiv:2105.14450
ColossalAI
Tesseract: Parallelize the Tensor Parallelism Efficiently
arXiv1 repoarXiv:2105.14500
ColossalAI
Exploration and Exploitation: Two Ways to Improve Chinese Spelling Correction Models
arXiv1 repoarXiv:2105.14813
CTCResources
DoT: An efficient Double Transformer for NLP tasks with tables
arXiv1 repoarXiv:2106.00479
tapas
SpanNER: Named Entity Re-/Recognition as Span Prediction
arXiv1 repoarXiv:2106.00641
nlp-bazel-tutorial
Enabling Efficiency-Precision Trade-offs for Label Trees in Extreme Classification
arXiv1 repoarXiv:2106.00730
pecos
TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification
arXiv1 repoarXiv:2106.00908
HistoSSLscaling
NVC-Net: End-to-End Adversarial Voice Conversion
arXiv1 repoarXiv:2106.00992
ai-research-code
When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
arXiv1 repoarXiv:2106.01548
vision_transformer
Fre-GAN: Adversarial Frequency-consistent Audio Synthesis
arXiv1 repoarXiv:2106.02297
MockingBird
Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation
arXiv1 repoarXiv:2106.03153
vad_score_prediction
arXiv:2106.03609
arXiv1 repoarXiv:2106.03609
HEBO
End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering
arXiv1 repoarXiv:2106.05346
art
Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
arXiv1 repoarXiv:2106.06103
vitsgpt-vits
TrafficStream: A Streaming Traffic Flow Forecasting Framework Based on Graph Neural Networks and Continual Learning
arXiv1 repoarXiv:2106.06273
A2TTA
RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised Learning
arXiv1 repoarXiv:2106.07760
cords
UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation
arXiv1 repoarXiv:2106.07889
univnet
CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
arXiv1 repoarXiv:2106.08087
SDAK
BEiT: BERT Pre-Training of Image Transformers
arXiv1 repoarXiv:2106.08254
MiDaS
End-to-End Semi-Supervised Object Detection with Soft Teacher
arXiv1 repoarXiv:2106.09018
Swin-Transformer
Large-Scale Chemical Language Representations Capture Molecular Structure and Properties
arXiv1 repoarXiv:2106.09553
molformer
XCiT: Cross-Covariance Image Transformers
arXiv1 repoarXiv:2106.09681
dino
Bad Characters: Imperceptible NLP Attacks
arXiv1 repoarXiv:2106.09898
garak
Distributed Deep Learning in Open Collaborations
arXiv1 repoarXiv:2106.10207
DeDLOC
How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers
arXiv1 repoarXiv:2106.10270
vision_transformer
Surgical data science for safe cholecystectomy: a protocol for segmentation of hepatocystic anatomy and assessment of the critical view of safety
arXiv1 repoarXiv:2106.10916
Endoscapes
BARTScore: Evaluating Generated Text as Text Generation
arXiv1 repoarXiv:2106.11520
BARTScore
It's All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning
arXiv1 repoarXiv:2106.12066
xwinograd
Alias-Free Generative Adversarial Networks
arXiv1 repoarXiv:2106.12423
ffhq-dataset
Extreme Multi-label Learning for Semantic Matching in Product Search
arXiv1 repoarXiv:2106.12657
pecos
Label Disentanglement in Partition-based Extreme Multilabel Classification
arXiv1 repoarXiv:2106.12751
pecos
Unsupervised Topic Segmentation of Meetings with BERT Embeddings
arXiv1 repoarXiv:2106.12978
E2E-Video-Processing-system
Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting
arXiv1 repoarXiv:2106.13008
autoformer-tourism-monthly
AudioCLIP: Extending CLIP to Image, Text and Audio
arXiv1 repoarXiv:2106.13043
JavisBench
Video Swin Transformer
arXiv1 repoarXiv:2106.13230
Swin-Transformer
panda-gym: Open-source goal-conditioned environments for robotic learning
arXiv1 repoarXiv:2106.13687
panda-gym
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages
arXiv1 repoarXiv:2106.13822
Glot500
Habitat 2.0: Training Home Assistants to Rearrange their Habitat
arXiv1 repoarXiv:2106.14405
habitat-lab
R-Drop: Regularized Dropout for Neural Networks
arXiv1 repoarXiv:2106.14448
final-project-level3-nlp-02
A Few Brief Notes on DeepImpact, COIL, and a Conceptual Framework for Information Retrieval Techniques
arXiv1 repoarXiv:2106.14807
bge-m3
AutoFormer: Searching Transformers for Visual Recognition
arXiv1 repoarXiv:2107.00651
Cream
Data Centric Domain Adaptation for Historical Text with OCR Errors
arXiv1 repoarXiv:2107.00927
historic-domain-adaptation-icdar
Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors
arXiv1 repoarXiv:2107.01545
UniSpeech
DeepDDS: deep graph neural network with attention mechanism to predict synergistic drug combinations
arXiv1 repoarXiv:2107.02467
chemicalx
Depth-supervised NeRF: Fewer Views and Faster Training for Free
arXiv1 repoarXiv:2107.02791
svraster
SoundStream: An End-to-End Neural Audio Codec
arXiv1 repoarXiv:2107.03312
moshi
Trusting RoBERTa over BERT: Insights from CheckListing the Natural Language Inference Task
arXiv1 repoarXiv:2107.07229
awesome-cybersecurity-agentic-ai
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
arXiv1 repoarXiv:2107.07651
LAVIS
Declarative Machine Learning Systems
arXiv1 repoarXiv:2107.08148
ludwig
YOLOX: Exceeding YOLO Series in 2021
arXiv1 repoarXiv:2107.08430
YOLOX
Frequency-Domain Data-Driven Controller Synthesis for Unstable LPV Systems
arXiv1 repoarXiv:2107.09712
speech-emotion-recognition
Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data
arXiv1 repoarXiv:2107.10833
Real-ESRGAN
SurfaceNet: Adversarial SVBRDF Estimation from a Single Image
arXiv1 repoarXiv:2107.11298
surfacenet
Open-Ended Learning Leads to Generally Capable Agents
arXiv1 repoarXiv:2107.12808
WaveFunctionCollapse
Domain-matched Pre-training Tasks for Dense Retrieval
arXiv1 repoarXiv:2107.13602
dpr-scale
Perceiver IO: A General Architecture for Structured Inputs & Outputs
arXiv1 repoarXiv:2107.14795
FM4Motor
PyEuroVoc: A Tool for Multilingual Legal Document Classification with EuroVoc Descriptors
arXiv1 repoarXiv:2108.01139
pyeurovoc
Object Wake-up: 3D Object Rigging from a Single Image
arXiv1 repoarXiv:2108.02708
object-wakeup
BERT-based distractor generation for Swedish reading comprehension questions using a small-scale dataset
arXiv1 repoarXiv:2108.03973
rc-answer-generation
Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
arXiv1 repoarXiv:2108.04024
CIRR
Util::Lookup: Exploiting key decoding in cryptographic libraries
arXiv1 repoarXiv:2108.04600
bc-rust
PatrickStar: Parallel Training of Pre-trained Models via Chunk-based Memory Management
arXiv1 repoarXiv:2108.05818
ColossalAI
Semantic Answer Similarity for Evaluating Question Answering Models
arXiv1 repoarXiv:2108.06130
Medical-Assistant
Conditional DETR for Fast Training Convergence
arXiv1 repoarXiv:2108.06152
DINO
Pixel Difference Networks for Efficient Edge Detection
arXiv1 repoarXiv:2108.07009
pidinet
On the Opportunities and Risks of Foundation Models
arXiv1 repoarXiv:2108.07258
gpttools
EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning
arXiv1 repoarXiv:2108.08842
llama.cpp-eden
A Survey on Common Threats in npm and PyPi Registries
arXiv1 repoarXiv:2108.09576
guarddog
An Empirical Assessment of Endpoint Security Systems Against Advanced Persistent Threats Attack Vectors
arXiv1 repoarXiv:2108.10422
awesome-edr-bypass
Rewrite Rule Inference Using Equality Saturation
arXiv1 repoarXiv:2108.10436
equational_theories
One TTS Alignment To Rule Them All
arXiv1 repoarXiv:2108.10447
tts_en_fastpitch
LLVIP: A Visible-infrared Paired Dataset for Low-light Vision
arXiv1 repoarXiv:2108.10831
LLVIP
Meta Self-Learning for Multi-Source Domain Adaptation: A Benchmark
arXiv1 repoarXiv:2108.10840
Meta-SelfLearning
mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset
arXiv1 repoarXiv:2108.13897
mmarco
FinQA: A Dataset of Numerical Reasoning over Financial Data
arXiv1 repoarXiv:2109.00122
MirageTVQA
WebQA: Multihop and Multimodal QA
arXiv1 repoarXiv:2109.00590
ReMuQ
Learning to Prompt for Vision-Language Models
arXiv1 repoarXiv:2109.01134
BiomedCoOp
Similarity of Sentence Representations in Multilingual LMs: Resolving Conflicting Literature and Case Study of Baltic Languages
arXiv1 repoarXiv:2109.01207
xsim
Fast Succinct Retrieval and Approximate Membership using Ribbon
arXiv1 repoarXiv:2109.01892
gecko-dev
MATE: Multi-view Attention for Table Transformer Efficiency
arXiv1 repoarXiv:2109.04312
tapas
Smoothed Contrastive Learning for Unsupervised Sentence Embedding
arXiv1 repoarXiv:2109.04321
sentemb
ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding
arXiv1 repoarXiv:2109.04380
sentemb
Packed Levitated Marker for Entity and Relation Extraction
arXiv1 repoarXiv:2109.06067
trusted_ke
LM-Critic: Language Models for Unsupervised Grammatical Error Correction
arXiv1 repoarXiv:2109.06822
CTCResources
Resolution-robust Large Mask Inpainting with Fourier Convolutions
arXiv1 repoarXiv:2109.07161
lama
Towards Zero-shot Cross-lingual Image Retrieval and Tagging
arXiv1 repoarXiv:2109.07622
Multilingual-CLIP
ROS-X-Habitat: Bridging the ROS Ecosystem with Embodied AI
arXiv1 repoarXiv:2109.07703
habitat-lab
Towards Zero and Few-shot Knowledge-seeking Turn Detection in Task-orientated Dialogue Systems
arXiv1 repoarXiv:2109.08820
KoPrivateGPT
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models
arXiv1 repoarXiv:2109.10282
trocr-base-handwritten
Transcoding Billions of Unicode Characters per Second with SIMD Instructions
arXiv1 repoarXiv:2109.10433
simdutf
Pix2seq: A Language Modeling Framework for Object Detection
arXiv1 repoarXiv:2109.10852
LocateAnything-3B
Transferring Knowledge from Vision to Language: How to Achieve it and how to Measure it?
arXiv1 repoarXiv:2109.11321
Kosmos-X
Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations
arXiv1 repoarXiv:2109.13059
trans-encoder
Robust SLAM Systems: Are We There Yet?
arXiv1 repoarXiv:2109.13160
slambench
VoiceFixer: Toward General Speech Restoration with Neural Vocoder
arXiv1 repoarXiv:2109.13731
voicefixer
VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
arXiv1 repoarXiv:2109.14084
fairseq
Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text Classification
arXiv1 repoarXiv:2110.00685
pecos
arXiv:2110.01200
arXiv1 repoarXiv:2110.01200
MTP-Codebase
WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
arXiv1 repoarXiv:2110.03370
WenetSpeech
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
arXiv1 repoarXiv:2110.04544
BiomedCoOp
Vector-quantized Image Modeling with Improved VQGAN
arXiv1 repoarXiv:2110.04627
echo-vqgan
Large-scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification
arXiv1 repoarXiv:2110.05777
UniSpeech
ABO: Dataset and Benchmarks for Real-World 3D Object Understanding
arXiv1 repoarXiv:2110.06199
Cap3D
S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations
arXiv1 repoarXiv:2110.06280
s3prl-vc
Salient Phrase Aware Dense Retrieval: Can a Dense Retriever Imitate a Sparse One?
arXiv1 repoarXiv:2110.06918
dpr-scale
Toward Degradation-Robust Voice Conversion
arXiv1 repoarXiv:2110.07537
RobustVC
P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
arXiv1 repoarXiv:2110.07602
P-tuning-v2
HumBugDB: A Large-scale Acoustic Mosquito Dataset
arXiv1 repoarXiv:2110.07607
BEANS-Zero
Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models
arXiv1 repoarXiv:2110.08173
medlama
BBQ: A Hand-Built Bias Benchmark for Question Answering
arXiv1 repoarXiv:2110.08193
langtest
LSA: Modeling Aspect Sentiment Coherency via Local Sentiment Aggregation
arXiv1 repoarXiv:2110.08604
News-Sentiment-Analysis
NormFormer: Improved Transformer Pretraining with Extra Normalization
arXiv1 repoarXiv:2110.09456
CLIP-ViT-L-14-laion2B-s32B-b82K
SSAST: Self-Supervised Audio Spectrogram Transformer
arXiv1 repoarXiv:2110.09784
ast
Wacky Weights in Learned Sparse Representations and the Revenge of Score-at-a-Time Query Evaluation
arXiv1 repoarXiv:2110.11540
SPLADERunner
Lhotse: a speech data representation library for the modern deep learning ecosystem
arXiv1 repoarXiv:2110.12561
lhotse
Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations
arXiv1 repoarXiv:2110.14513
NANSY
Distilling Relation Embeddings from Pre-trained Language Models
arXiv1 repoarXiv:2110.15705
relbert
Chaos Engineering of Ethereum Blockchain Clients
arXiv1 repoarXiv:2111.00221
awesome-chaos-engineering
RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and Intensity Responses
arXiv1 repoarXiv:2111.00962
vocoder
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
arXiv1 repoarXiv:2111.02114
mlcd-vit-large-patch14-336
An Empirical Study of Training End-to-End Vision-and-Language Transformers
arXiv1 repoarXiv:2111.02387
VLE
WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image
arXiv1 repoarXiv:2111.02403
WORD
A Unified View of Relational Deep Learning for Drug Pair Scoring
arXiv1 repoarXiv:2111.02916
chemicalx
Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
arXiv1 repoarXiv:2111.03930
BiomedCoOp
arXiv:2111.06178
arXiv1 repoarXiv:2111.06178
HEBO
Adding more data does not always help: A study in medical conversation summarization with PEGASUS
arXiv1 repoarXiv:2111.07564
curai-research
XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale
arXiv1 repoarXiv:2111.09296
wav2vec2-xls-r-300m
MEDCOD: A Medically-Accurate, Emotive, Diverse, and Controllable Dialog System
arXiv1 repoarXiv:2111.09381
curai-research
ClipCap: CLIP Prefix for Image Captioning
arXiv1 repoarXiv:2111.09734
webapp
Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions
arXiv1 repoarXiv:2111.10337
TimeChat-Online-139K
arXiv:2111.12704
arXiv1 repoarXiv:2111.12704
RealBasicVSR
Replicating Monotonic Payoffs Without Oracles
arXiv1 repoarXiv:2111.13740
osmosis
What Do You See in this Patient? Behavioral Testing of Clinical NLP Models
arXiv1 repoarXiv:2111.15512
langtest
DKPLM: Decomposable Knowledge-enhanced Pre-trained Language Model for Natural Language Understanding
arXiv1 repoarXiv:2112.01047
EasyNLP
BERTMap: A BERT-based Ontology Alignment System
arXiv1 repoarXiv:2112.02682
BERTMap
Grounded Language-Image Pre-training
arXiv1 repoarXiv:2112.03857
GLIP
arXiv:2112.05251
arXiv1 repoarXiv:2112.05251
Kosmos-X
DistilCSE: Effective Knowledge Distillation For Contrastive Sentence Embeddings
arXiv1 repoarXiv:2112.05638
sentemb
Margin Calibration for Long-Tailed Visual Recognition
arXiv1 repoarXiv:2112.07225
robustlearn
Large Dual Encoders Are Generalizable Retrievers
arXiv1 repoarXiv:2112.07899
gtr-t5-base
Learning Cross-Lingual IR from an English Retriever
arXiv1 repoarXiv:2112.08185
DrDecr_XOR-TyDi_whitebox
StyleMC: Multi-Channel Based Fast Text-Guided Image Generation and Manipulation
arXiv1 repoarXiv:2112.08493
StyleMC-SentenceTransformers
DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching Models
arXiv1 repoarXiv:2112.08609
DuReader
Self-Supervised Learning for speech recognition with Intermediate layer supervision
arXiv1 repoarXiv:2112.08778
UniSpeech
PeopleSansPeople: A Synthetic Data Generator for Human-Centric Computer Vision
arXiv1 repoarXiv:2112.09290
humans
WebGPT: Browser-assisted question-answering with human feedback
arXiv1 repoarXiv:2112.09332
webgpt_comparisons
Align and Prompt: Video-and-Language Pre-training with Entity Prompts
arXiv1 repoarXiv:2112.09583
LAVIS
What are Weak Links in the npm Supply Chain?
arXiv1 repoarXiv:2112.10165
guarddog
Mask2Former for Video Instance Segmentation
arXiv1 repoarXiv:2112.10764
Mask2Former
Does MAML Only Work via Feature Re-use? A Data Centric Perspective
arXiv1 repoarXiv:2112.13137
ultimate-utils
Temporally Constrained Neural Networks (TCNN): A framework for semi-supervised video semantic segmentation
arXiv1 repoarXiv:2112.13815
Endoscapes
LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents
arXiv1 repoarXiv:2112.14731
InLegalBERT
Towards a secure API client generator for IoT devices
arXiv1 repoarXiv:2201.00270
openapi-generator
arXiv:2201.00487
arXiv1 repoarXiv:2201.00487
ReferFormer
Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images
arXiv1 repoarXiv:2201.01266
TriALS
A Survey of JSON-compatible Binary Serialization Specifications
arXiv1 repoarXiv:2201.02089
smile-format-specification
Detecting Twenty-thousand Classes using Image-level Supervision
arXiv1 repoarXiv:2201.02605
pixmo-count
A Benchmark of JSON-compatible Binary Serialization Specifications
arXiv1 repoarXiv:2201.03051
smile-format-specification
DeepKE: A Deep Learning Based Knowledge Extraction Toolkit for Knowledge Base Population
arXiv1 repoarXiv:2201.03335
DeepKE
CVSS Corpus and Massively Multilingual Speech-to-Speech Translation
arXiv1 repoarXiv:2201.03713
UniSS
UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models
arXiv1 repoarXiv:2201.05966
UnifiedSKG
Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth
arXiv1 repoarXiv:2201.07436
GLPDepth
Can Model Compression Improve NLP Fairness
arXiv1 repoarXiv:2201.08542
distilgpt2
STRIDE-based Cyber Security Threat Modeling for IoT-enabled Precision Agriculture Systems
arXiv1 repoarXiv:2201.09493
awesome-connected-things-sec
Describing Differences between Text Distributions with Natural Language
arXiv1 repoarXiv:2201.12323
imodelsX
DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
arXiv1 repoarXiv:2201.12329
DINO
Incorporating Commonsense Knowledge into Story Ending Generation via Heterogeneous Graph Networks
arXiv1 repoarXiv:2201.12538
AwesomeSEG
A Dataset for Medical Instructional Video Classification and Question Answering
arXiv1 repoarXiv:2201.12888
VPTSL
Negativity Spreads Faster: A Large-Scale Multilingual Twitter Analysis on the Role of Sentiment in Political Communication
arXiv1 repoarXiv:2202.00396
xlm-twitter-politics-sentiment
Locally Typical Sampling
arXiv1 repoarXiv:2202.00666
LLM-Sampling
data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
arXiv1 repoarXiv:2202.03555
data2vec-audio-large
MaskGIT: Masked Generative Image Transformer
arXiv1 repoarXiv:2202.04200
nanoMFM
InPars: Data Augmentation for Information Retrieval using Large Language Models
arXiv1 repoarXiv:2202.05144
InPars
TwHIN: Embedding the Twitter Heterogeneous Information Network for Personalized Recommendation
arXiv1 repoarXiv:2202.05387
the-algorithm-ml
A Contrastive Framework for Neural Text Generation
arXiv1 repoarXiv:2202.06417
4th-Bookathon-The-Unbearable-Heaviness-of-GPT
arXiv:2202.06558
arXiv1 repoarXiv:2202.06558
HEBO
Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?
arXiv1 repoarXiv:2202.06675
clip-retrieval
Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text
arXiv1 repoarXiv:2202.06935
awesome-nlg
Transformers in Time Series: A Survey
arXiv1 repoarXiv:2202.07125
time-moe
arXiv:2202.07359
arXiv1 repoarXiv:2202.07359
textlesslib
Information Extraction in Low-Resource Scenarios: Survey and Perspective
arXiv1 repoarXiv:2202.08063
DeepKE
Probing Pretrained Models of Source Code
arXiv1 repoarXiv:2202.08975
probings4code
Pseudo Numerical Methods for Diffusion Models on Manifolds
arXiv1 repoarXiv:2202.09778
stable-diffusion
Phrase-Based Affordance Detection via Cyclic Bilateral Interaction
arXiv1 repoarXiv:2202.12076
Cross-View-AG
LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding
arXiv1 repoarXiv:2202.13669
LiLT
Combining Modular Skills in Multitask Learning
arXiv1 repoarXiv:2202.13914
mttl
Mukayese: Turkish NLP Strikes Back
arXiv1 repoarXiv:2203.01215
mukayese
DN-DETR: Accelerate DETR Training by Introducing Query DeNoising
arXiv1 repoarXiv:2203.01305
DINO
Nuclei instance segmentation and classification in histopathology images with StarDist
arXiv1 repoarXiv:2203.02284
stardist
arXiv:2203.02923
arXiv1 repoarXiv:2203.02923
smalldiffusion
HEAR: Holistic Evaluation of Audio Representations
arXiv1 repoarXiv:2203.03022
usad
Improving CTC-based speech recognition via knowledge transferring from pre-trained language models
arXiv1 repoarXiv:2203.03582
ASR-Knowledge-Transferring
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
arXiv1 repoarXiv:2203.03605
DINO
A Simple Multi-Modality Transfer Learning Baseline for Sign Language Translation
arXiv1 repoarXiv:2203.04287
SLRT
Temporal Difference Learning for Model Predictive Control
arXiv1 repoarXiv:2203.04955
xwm
Conditional Prompt Learning for Vision-Language Models
arXiv1 repoarXiv:2203.05557
BiomedCoOp
CMKD: CNN/Transformer-Based Cross-Model Knowledge Distillation for Audio Classification
arXiv1 repoarXiv:2203.06760
ast
Block-STM: Scaling Blockchain Execution by Turning Ordering Curse to a Performance Blessing
arXiv1 repoarXiv:2203.06871
cosmos-sdk
Formalising Decentralised Exchanges in Coq
arXiv1 repoarXiv:2203.08016
ConCert
Surrogate Gap Minimization Improves Sharpness-Aware Training
arXiv1 repoarXiv:2203.08065
vision_transformer
AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter Tuning
arXiv1 repoarXiv:2203.08212
cords
Multilingual Pre-training with Language and Task Adaptation for Multilingual Text Style Transfer
arXiv1 repoarXiv:2203.08552
multilingual_tst_copy
PosePipe: Open-Source Human Pose Estimation Pipeline for Clinical Research
arXiv1 repoarXiv:2203.08792
PosePipeline
A Survey of Multi-Tenant Deep Learning Inference on GPU
arXiv1 repoarXiv:2203.09040
awesome-gpu-engineering
CaRTS: Causality-driven Robot Tool Segmentation from Vision and Kinematics Data
arXiv1 repoarXiv:2203.09475
CaRTS
Learning Affordance Grounding from Exocentric Images
arXiv1 repoarXiv:2203.09905
Cross-View-AG
DuReader_retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine
arXiv1 repoarXiv:2203.10232
DuReader
Mitigating Gender Bias in Distilled Language Models via Counterfactual Role Reversal
arXiv1 repoarXiv:2203.12574
distilgpt2
Video Polyp Segmentation: A Deep Learning Perspective
arXiv1 repoarXiv:2203.14291
VPS
Certified Mergeable Replicated Data Types
arXiv1 repoarXiv:2203.14518
irmin
Socially Compliant Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation
arXiv1 repoarXiv:2203.15041
UniWM_Dataset
X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval
arXiv1 repoarXiv:2203.15086
xpool
Killing Two Birds with One Stone:Efficient and Robust Training of Face Recognition CNNs by Partial FC
arXiv1 repoarXiv:2203.15565
insightface
Earnings-22: A Practical Benchmark for Accents in the Wild
arXiv1 repoarXiv:2203.15591
earnings22
Parameter-efficient Model Adaptation for Vision Transformers
arXiv1 repoarXiv:2203.16329
MoA
MMER: Multimodal Multi-task Learning for Speech Emotion Recognition
arXiv1 repoarXiv:2203.16794
MMER
BRIO: Bringing Order to Abstractive Summarization
arXiv1 repoarXiv:2203.16804
BRIO
Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data
arXiv1 repoarXiv:2203.17113
SpeechT5
Making Pre-trained Language Models End-to-end Few-shot Learners with Contrastive Prompt Tuning
arXiv1 repoarXiv:2204.00166
EasyNLP
PriMock57: A Dataset Of Primary Care Mock Consultations
arXiv1 repoarXiv:2204.00333
primock57
Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation
arXiv1 repoarXiv:2204.00447
primock57
Learning Audio-Video Modalities from Image Captions
arXiv1 repoarXiv:2204.00679
videoCC-data
Distributional Gradient Boosting Machines
arXiv1 repoarXiv:2204.00778
LightGBMLSS
HLDC: Hindi Legal Documents Corpus
arXiv1 repoarXiv:2204.00806
HLDC
arXiv:2204.01715
arXiv1 repoarXiv:2204.01715
llm_test
Temporal Alignment Networks for Long-term Video
arXiv1 repoarXiv:2204.02968
TemporalAlignNet
BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model
arXiv1 repoarXiv:2204.03905
Fengshenbang-LM
ASQA: Factoid Questions Meet Long-Form Answers
arXiv1 repoarXiv:2204.06092
KoPrivateGPT
WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity Types
arXiv1 repoarXiv:2204.06347
GEMEL
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
arXiv1 repoarXiv:2204.06745
level3_nlp_finalproject-nlp-12
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
arXiv1 repoarXiv:2204.08387
LayoutLMv3-DocVQA
Dress Code: High-Resolution Multi-Category Virtual Try-On
arXiv1 repoarXiv:2204.08532
dress-code
ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models
arXiv1 repoarXiv:2204.08790
GLIP
ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers
arXiv1 repoarXiv:2204.09224
contentvec
SciTS: A Benchmark for Time-Series Databases in Scientific Experiments and Industrial Internet of Things
arXiv1 repoarXiv:2204.09795
ClickBench
Learning to Revise References for Faithful Summarization
arXiv1 repoarXiv:2204.10290
summary-reference-revision
TorchSparse: Efficient Point Cloud Inference Engine
arXiv1 repoarXiv:2204.10319
torchsparse
arXiv:2204.12260
arXiv1 repoarXiv:2204.12260
TIL-2023
DoPose-6D dataset for object segmentation and 6D pose estimation
arXiv1 repoarXiv:2204.13613
image_agnostic_segmentation
SVTR: Scene Text Recognition with a Single Visual Model
arXiv1 repoarXiv:2205.00159
yas
MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
arXiv1 repoarXiv:2205.00445
math
Sequencer: Deep LSTM for Image Classification
arXiv1 repoarXiv:2205.01972
FM4Motor
Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition
arXiv1 repoarXiv:2205.03433
vocalsound
Reducing Activation Recomputation in Large Transformer Models
arXiv1 repoarXiv:2205.05198
MindSpeed-MM
Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning
arXiv1 repoarXiv:2205.05638
level3_nlp_finalproject-nlp-12
ViT5: Pretrained Text-to-Text Transformer for Vietnamese Language Generation
arXiv1 repoarXiv:2205.06457
ViT5
Consistent Human Evaluation of Machine Translation across Language Pairs
arXiv1 repoarXiv:2205.08533
pearmut
Summarization as Indirect Supervision for Relation Extraction
arXiv1 repoarXiv:2205.09837
SuRE
DeepStruct: Pretraining of Language Models for Structure Prediction
arXiv1 repoarXiv:2205.10475
trusted_ke
Deep Learning Workload Scheduling in GPU Datacenters: Taxonomy, Challenges and Vision
arXiv1 repoarXiv:2205.11913
awesome-gpu-engineering
RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder
arXiv1 repoarXiv:2205.12035
ember-v1
FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
arXiv1 repoarXiv:2205.12446
fleurs
End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models
arXiv1 repoarXiv:2205.12487
Mocheg
arXiv:2205.13603
arXiv1 repoarXiv:2205.13603
web-stable-diffusion
Multimodal Masked Autoencoders Learn Transferable Representations
arXiv1 repoarXiv:2205.14204
m3ae_public
Controllable Text Generation with Neurally-Decomposed Oracle
arXiv1 repoarXiv:2205.14219
constrDecoding
Multimodal Fake News Detection via CLIP-Guided Learning
arXiv1 repoarXiv:2205.14304
Multi-fake-detective
EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction
arXiv1 repoarXiv:2205.14756
efficientvit
Prompt-aligned Gradient for Prompt Tuning
arXiv1 repoarXiv:2205.14865
BiomedCoOp
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
arXiv1 repoarXiv:2205.15868
CogVideo
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages
arXiv1 repoarXiv:2205.15960
NusaX-senti
Text2Human: Text-Driven Controllable Human Image Generation
arXiv1 repoarXiv:2205.15996
DeepFashion-MultiModal
Squeezeformer: An Efficient Transformer for Automatic Speech Recognition
arXiv1 repoarXiv:2206.00888
kaggle-asl-fingerspelling-1st-place-solution
Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate Progress
arXiv1 repoarXiv:2206.01626
cleanrl
A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
arXiv1 repoarXiv:2206.01718
sa2va_eval
Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning
arXiv1 repoarXiv:2206.02647
HistoSSLscaling
arXiv:2206.02675
arXiv1 repoarXiv:2206.02675
HEBO
Tutel: Adaptive Mixture-of-Experts at Scale
arXiv1 repoarXiv:2206.03382
Swin-Transformer
JuMP 1.0: Recent improvements to a modeling language for mathematical optimization
arXiv1 repoarXiv:2206.03866
JuMP.jl
Sparse Mixture-of-Experts are Domain Generalizable Learners
arXiv1 repoarXiv:2206.04046
MoA
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
arXiv1 repoarXiv:2206.04615
BIG-bench
Factuality Enhanced Language Models for Open-Ended Text Generation
arXiv1 repoarXiv:2206.04624
FasterTransformer
On Data Scaling in Masked Image Modeling
arXiv1 repoarXiv:2206.04664
Swin-Transformer
SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views
arXiv1 repoarXiv:2206.05737
torchsparse
The YiTrans End-to-End Speech Translation System for IWSLT 2022 Offline Shared Task
arXiv1 repoarXiv:2206.05777
SpeechT5
GLIPv2: Unifying Localization and Vision-Language Understanding
arXiv1 repoarXiv:2206.05836
GLIP
Semantic-Discriminative Mixup for Generalizable Sensor-based Cross-domain Activity Recognition
arXiv1 repoarXiv:2206.06629
robustlearn
Aeneas: Rust Verification by Functional Translation
arXiv1 repoarXiv:2206.07185
aeneas
arXiv:2206.07293
arXiv1 repoarXiv:2206.07293
TIL-2023
PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
arXiv1 repoarXiv:2206.10498
LLMs-Planning
Questions Are All You Need to Train a Dense Passage Retriever
arXiv1 repoarXiv:2206.10658
art
ProGen2: Exploring the Boundaries of Protein Language Models
arXiv1 repoarXiv:2206.13517
jaxformer
Feature Refinement to Improve High Resolution Image Inpainting
arXiv1 repoarXiv:2206.13644
lama
TweetNLP: Cutting-Edge Natural Language Processing for Social Media
arXiv1 repoarXiv:2206.14774
tweetnlp
Solving Quantitative Reasoning Problems with Language Models
arXiv1 repoarXiv:2206.14858
RLPR-Evaluation
Dissecting Self-Supervised Learning Methods for Surgical Computer Vision
arXiv1 repoarXiv:2207.00449
Endoscapes
An Efficiency Study for SPLADE Models
arXiv1 repoarXiv:2207.03834
LLM_Web_search
Improving Entity Disambiguation by Reasoning over a Knowledge Base
arXiv1 repoarXiv:2207.04106
trusted_ke
ReFinED: An Efficient Zero-shot-capable Approach to End-to-End Entity Linking
arXiv1 repoarXiv:2207.04108
trusted_ke
arXiv:2207.04296
arXiv1 repoarXiv:2207.04296
web-stable-diffusion
A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
arXiv1 repoarXiv:2207.04356
s3prl-vc
Embedding Recycling for Language Models
arXiv1 repoarXiv:2207.04993
EmbeddingRecycling
Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios
arXiv1 repoarXiv:2207.05501
MiDaS
OSLAT: Open Set Label Attention Transformer for Medical Entity Retrieval and Span Extraction
arXiv1 repoarXiv:2207.05817
curai-research
ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
arXiv1 repoarXiv:2207.06389
Make-An-Audio
Confident Adaptive Language Modeling
arXiv1 repoarXiv:2207.07061
mixture_of_recursions
Parameter-Efficient Prompt Tuning Makes Generalized and Calibrated Neural Text Retrievers
arXiv1 repoarXiv:2207.07087
P-tuning-v2
CoqQ: Foundational Verification of Quantum Programs
arXiv1 repoarXiv:2207.11350
analysis
Towards Complex Document Understanding By Discrete Reasoning
arXiv1 repoarXiv:2207.11871
TAT-DQA
Finding smart contract vulnerabilities with ConCert's property-based testing framework
arXiv1 repoarXiv:2208.00758
ConCert
Prompt Tuning for Generative Multimodal Pretrained Models
arXiv1 repoarXiv:2208.02532
OFA
Investigating Efficiently Extending Transformers for Long Input Summarization
arXiv1 repoarXiv:2208.04347
pegasus-x-base-synthsumm_open-16k
Exploring Hate Speech Detection with HateXplain and BERT
arXiv1 repoarXiv:2208.04489
DeepLearningProject
TotalSegmentator: robust segmentation of 104 anatomical structures in CT images
arXiv1 repoarXiv:2208.05868
TotalSegmentator
Perspective Reconstruction of Human Faces by Joint Mesh and Landmark Regression
arXiv1 repoarXiv:2208.07142
insightface
Domain-Specific Risk Minimization for Out-of-Distribution Generalization
arXiv1 repoarXiv:2208.08661
robustlearn
Flat Multi-modal Interaction Transformer for Named Entity Recognition
arXiv1 repoarXiv:2208.11039
Fengshenbang-LM
Automatic music mixing with deep learning and out-of-domain data
arXiv1 repoarXiv:2208.11428
FxNorm-automix
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
arXiv1 repoarXiv:2208.12242
dreambooth
Measure Construction by Extension in Dependent Type Theory with Application to Integration
arXiv1 repoarXiv:2209.02345
analysis
What does a platypus look like? Generating customized prompts for zero-shot image classification
arXiv1 repoarXiv:2209.03320
CLIP_benchmark
FP8 Formats for Deep Learning
arXiv1 repoarXiv:2209.05433
ao
Pre-trained Language Models for the Legal Domain: A Case Study on Indian Law
arXiv1 repoarXiv:2209.06049
InLegalBERT
arXiv:2209.06995
arXiv1 repoarXiv:2209.06995
Patron
Out-of-Distribution Representation Learning for Time Series Classification
arXiv1 repoarXiv:2209.07027
robustlearn
Monolith: Real Time Recommendation System With Collisionless Embedding Table
arXiv1 repoarXiv:2209.07663
SIMURG
LAVIS: A Library for Language-Vision Intelligence
arXiv1 repoarXiv:2209.09019
LAVIS
Operationalizing Machine Learning: An Interview Study
arXiv1 repoarXiv:2209.09125
dtu_mlops
Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
arXiv1 repoarXiv:2209.09513
ScienceQA
Generate rather than Retrieve: Large Language Models are Strong Context Generators
arXiv1 repoarXiv:2209.10063
Ensemble-of-Retrievers
Efficient Few-Shot Learning Without Prompts
arXiv1 repoarXiv:2209.11055
ember-v1
OLIVES Dataset: Ophthalmic Labels for Investigating Visual Eye Semantics
arXiv1 repoarXiv:2209.11195
OLIVES_Dataset
All are Worth Words: A ViT Backbone for Diffusion Models
arXiv1 repoarXiv:2209.12152
U-ViT
T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition
arXiv1 repoarXiv:2209.12616
tner
WikiDes: A Wikipedia-Based Dataset for Generating Short Descriptions from Paragraphs
arXiv1 repoarXiv:2209.13101
WikiDes
A Benchmark Comparison of Python Malware Detection Approaches
arXiv1 repoarXiv:2209.13288
guarddog
COLO: A Contrastive Learning based Re-ranking Framework for One-Stage Summarization
arXiv1 repoarXiv:2209.14569
TestCoLo
SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data
arXiv1 repoarXiv:2209.15329
SpeechT5
Music-to-Text Synaesthesia: Generating Descriptive Text from Music Recordings
arXiv1 repoarXiv:2210.00434
emotion-english-distilroberta-base
Improving Sample Quality of Diffusion Models Using Self-Attention Guidance
arXiv1 repoarXiv:2210.00939
Fooocus
Omnigrok: Grokking Beyond Algorithmic Data
arXiv1 repoarXiv:2210.01117
grokfast
Explaining Patterns in Data with Language Models via Interpretable Autoprompting
arXiv1 repoarXiv:2210.01848
imodelsX
TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis
arXiv1 repoarXiv:2210.02186
frn-50k-baseline
Decomposed Prompting: A Modular Approach for Solving Complex Tasks
arXiv1 repoarXiv:2210.02406
DecomP-ODQA
arXiv:2210.02410
arXiv1 repoarXiv:2210.02410
Vendi-Score
arXiv:2210.02437
arXiv1 repoarXiv:2210.02437
MTP-Codebase
Binding Language Models in Symbolic Languages
arXiv1 repoarXiv:2210.02875
blendsql
On Distillation of Guided Diffusion Models
arXiv1 repoarXiv:2210.03142
i2p
Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding
arXiv1 repoarXiv:2210.03347
pix2struct-large
Measuring and Narrowing the Compositionality Gap in Language Models
arXiv1 repoarXiv:2210.03350
self-ask
SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training
arXiv1 repoarXiv:2210.03730
SpeechT5
ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering
arXiv1 repoarXiv:2210.03849
ConvFinQA
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
arXiv1 repoarXiv:2210.05148
DiffRoll
Foundation Transformers
arXiv1 repoarXiv:2210.06423
Kosmos-X
InfoCSE: Information-aggregated Contrastive Learning of Sentence Embeddings
arXiv1 repoarXiv:2210.06432
sentemb
arXiv:2210.06886
arXiv1 repoarXiv:2210.06886
ImaginaryNet
arXiv:2210.07229
arXiv1 repoarXiv:2210.07229
KEditVis-LLM-Editing
Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training
arXiv1 repoarXiv:2210.08773
LAVIS
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
arXiv1 repoarXiv:2210.09261
bbh
Revision Transformers: Instructing Language Models to Change their Values
arXiv1 repoarXiv:2210.10332
Revision-Transformer
TabLLM: Few-shot Classification of Tabular Data with Large Language Models
arXiv1 repoarXiv:2210.10723
TabLLM
Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and Granularity
arXiv1 repoarXiv:2210.10996
ChineseBERT-for-csc
CEFR-Based Sentence Difficulty Annotation and Assessment
arXiv1 repoarXiv:2210.11766
CEFR-SP
An Analysis of Fusion Functions for Hybrid Retrieval
arXiv1 repoarXiv:2210.11934
KoPrivateGPT
arXiv:2210.12213
arXiv1 repoarXiv:2210.12213
spabert
Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs
arXiv1 repoarXiv:2210.12283
LeanTool
BEANS: The Benchmark of Animal Sounds
arXiv1 repoarXiv:2210.12300
BEANS-Zero
Bootstrapping meaning through listening: Unsupervised learning of spoken sentence embeddings
arXiv1 repoarXiv:2210.12857
spoken_sent_embedding
NVIDIA FLARE: Federated Learning from Simulation to Real-World
arXiv1 repoarXiv:2210.13291
NVFlare
EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
arXiv1 repoarXiv:2210.14090
moshi
arXiv:2210.14648
arXiv1 repoarXiv:2210.14648
TIL-2023
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks
arXiv1 repoarXiv:2210.14712
Glot500
DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models
arXiv1 repoarXiv:2210.14896
diffusiondb
Truncation Sampling as Language Model Desmoothing
arXiv1 repoarXiv:2210.15191
LLM-Sampling
Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit
arXiv1 repoarXiv:2210.17016
wespeaker
Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation
arXiv1 repoarXiv:2210.17027
SpeechT5
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
arXiv1 repoarXiv:2211.00593
ARENA_2.0
Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022
arXiv1 repoarXiv:2211.00815
wespeaker
DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models
arXiv1 repoarXiv:2211.01095
TCD-SDXL-LoRA
Two-Stream Network for Sign Language Recognition and Translation
arXiv1 repoarXiv:2211.01367
SLRT
MPCFormer: fast, performant and private Transformer inference with MPC
arXiv1 repoarXiv:2211.01452
MPCFormer
Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
arXiv1 repoarXiv:2211.02001
bloom
Multi-Head Adapter Routing for Cross-Task Generalization
arXiv1 repoarXiv:2211.03831
mttl
Will we run out of data? Limits of LLM scaling based on human-generated data
arXiv1 repoarXiv:2211.04325
gigatoken
InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions
arXiv1 repoarXiv:2211.05778
InternImage
A Novel Sampling Scheme for Text- and Image-Conditional Image Synthesis in Quantized Latent Spaces
arXiv1 repoarXiv:2211.07292
i2p
Towards a Mathematics Formalisation Assistant using Large Language Models
arXiv1 repoarXiv:2211.07524
LeanAide
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale
arXiv1 repoarXiv:2211.07636
EVA
UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction
arXiv1 repoarXiv:2211.09039
trusted_ke
Holistic Evaluation of Language Models
arXiv1 repoarXiv:2211.09110
lmms-eval
Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
arXiv1 repoarXiv:2211.09808
InternImage
CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval
arXiv1 repoarXiv:2211.10411
dpr-scale
PAL: Program-aided Language Models
arXiv1 repoarXiv:2211.10435
gsm-hard
BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision
arXiv1 repoarXiv:2211.10439
InternImage
CryptOpt: Verified Compilation with Randomized Program Search for Cryptographic Primitives (full version)
arXiv1 repoarXiv:2211.10665
CryptOpt
An Empirical Study On Contrastive Search And Contrastive Decoding For Open-ended Text Generation
arXiv1 repoarXiv:2211.10797
Adaptive-Contrastive-Search
You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model
arXiv1 repoarXiv:2211.11152
OFA
L3Cube-MahaSBERT and HindSBERT: Sentence BERT Models and Benchmarking BERT Sentence Representations for Hindi and Marathi
arXiv1 repoarXiv:2211.11187
bengali-sentence-similarity-sbert
UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition
arXiv1 repoarXiv:2211.11256
UniMSE
VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning
arXiv1 repoarXiv:2211.11275
SpeechT5
Visual Programming: Compositional visual reasoning without training
arXiv1 repoarXiv:2211.11559
visprog
Coreference Resolution through a seq2seq Transition-Based System
arXiv1 repoarXiv:2211.12142
trusted_ke
Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition
arXiv1 repoarXiv:2211.12368
RAD-NeRF
EDICT: Exact Diffusion Inversion via Coupled Transformations
arXiv1 repoarXiv:2211.12446
DOODL
G^3: Geolocation via Guidebook Grounding
arXiv1 repoarXiv:2211.15521
diff-mining
Connecting the Dots: Floorplan Reconstruction Using Two-Level Queries
arXiv1 repoarXiv:2211.15658
RoomFormer
On Word Error Rate Definitions and their Efficient Computation for Multi-Speaker Speech Recognition Systems
arXiv1 repoarXiv:2211.16112
meeteval
NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° Views
arXiv1 repoarXiv:2211.16431
NeuralLift-360
Rationale-Guided Few-Shot Classification to Detect Abusive Language
arXiv1 repoarXiv:2211.17046
Rationale_predictor
Rethinking Causality-driven Robot Tool Segmentation with Temporal Constraints
arXiv1 repoarXiv:2212.00072
CaRTS
MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition
arXiv1 repoarXiv:2212.00500
OFA
Scaling Language-Image Pre-training via Masking
arXiv1 repoarXiv:2212.00794
Chinese-CLIP
Moving Beyond Downstream Task Accuracy for Information Retrieval Benchmarking
arXiv1 repoarXiv:2212.01340
ColBERT
Melody transcription via generative pre-training
arXiv1 repoarXiv:2212.01884
SheetSage2
One-shot Implicit Animatable Avatars with Model-based Priors
arXiv1 repoarXiv:2212.02469
ELICIT
3DGazeNet: Generalizing Gaze Estimation with Weak-Supervision from Synthetic Views
arXiv1 repoarXiv:2212.02997
insightface
NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors
arXiv1 repoarXiv:2212.03267
zero123
Discovering Latent Knowledge in Language Models Without Supervision
arXiv1 repoarXiv:2212.03827
language_exploration
Latent Graph Representations for Critical View of Safety Assessment
arXiv1 repoarXiv:2212.04155
Endoscapes
Diffusion Guided Domain Adaptation of Image Generators
arXiv1 repoarXiv:2212.04473
styleganfusion
Multi-Concept Customization of Text-to-Image Diffusion
arXiv1 repoarXiv:2212.04488
custom-diffusion
Learning Video Representations from Large Language Models
arXiv1 repoarXiv:2212.04501
VideoTree
Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
arXiv1 repoarXiv:2212.05032
Structured-Diffusion-Guidance
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
arXiv1 repoarXiv:2212.05055
smolMoELM-custom
Transcoding Unicode Characters with AVX-512 Instructions
arXiv1 repoarXiv:2212.05098
simdutf
MAGVIT: Masked Generative Video Transformer
arXiv1 repoarXiv:2212.05199
magvit
How to Backdoor Diffusion Models?
arXiv1 repoarXiv:2212.05400
BadDiffusion
Unfolding Local Growth Rate Estimates for (Almost) Perfect Adversarial Detection
arXiv1 repoarXiv:2212.06776
deepfake_multiLID
SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning
arXiv1 repoarXiv:2212.07489
smacv2
Attention as a Guide for Simultaneous Speech Translation
arXiv1 repoarXiv:2212.07850
naist-simulst
Constitutional AI: Harmlessness from AI Feedback
arXiv1 repoarXiv:2212.08073
minihf
Transferring General Multimodal Pretrained Models to Text Recognition
arXiv1 repoarXiv:2212.09297
OFA
Rethinking Label Smoothing on Multi-hop Question Answering
arXiv1 repoarXiv:2212.09512
Smoothing-R3
Visconde: Multi-document QA with GPT-3 and Neural Reranking
arXiv1 repoarXiv:2212.09656
KoPrivateGPT
Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model
arXiv1 repoarXiv:2212.09811
nllb-pruning
Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
arXiv1 repoarXiv:2212.10509
ircot
From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models
arXiv1 repoarXiv:2212.10846
LAVIS
Training language models to summarize narratives improves brain alignment
arXiv1 repoarXiv:2212.10898
brain_language_summarization
OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
arXiv1 repoarXiv:2212.12017
opt-iml-max-1.3b
MAUVE Scores for Generative Models: Theory and Practice
arXiv1 repoarXiv:2212.14578
ssharoff.github.io
Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting
arXiv1 repoarXiv:2301.00493
torchsparse
ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders
arXiv1 repoarXiv:2301.00808
ConvNeXt-V2
Iterated Decomposition: Improving Science Q&A by Supervising Reasoning Processes
arXiv1 repoarXiv:2301.01751
ice
InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval
arXiv1 repoarXiv:2301.01820
InPars
Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
arXiv1 repoarXiv:2301.03598
flashinfer
Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing
arXiv1 repoarXiv:2301.04558
BiomedVLP-BioViL-T
Progress measures for grokking via mechanistic interpretability
arXiv1 repoarXiv:2301.05217
TransformerLens
Domain Expansion of Image Generators
arXiv1 repoarXiv:2301.05225
ControlNet
GLIGEN: Open-Set Grounded Text-to-Image Generation
arXiv1 repoarXiv:2301.07093
GLIGEN
How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
arXiv1 repoarXiv:2301.07597
humanizer-skill
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
arXiv1 repoarXiv:2301.08243
xwm
Blind Spots: Automatically detecting ignored program inputs
arXiv1 repoarXiv:2301.08700
publications
Set-Theoretic and Type-Theoretic Ordinals Coincide
arXiv1 repoarXiv:2301.10696
TypeTopology
DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature
arXiv1 repoarXiv:2301.11305
AdaDetectGPT
3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion Models
arXiv1 repoarXiv:2301.11445
3DShape2VecSet
SEGA: Instructing Text-to-Image Models using Semantic Guidance
arXiv1 repoarXiv:2301.12247
ControlNet
Domain Theory in Constructive and Predicative Univalent Foundations
arXiv1 repoarXiv:2301.12405
TypeTopology
Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
arXiv1 repoarXiv:2301.12661
Make-An-Audio
Edge-guided Multi-domain RGB-to-TIR image Translation for Training Vision Tasks with Challenging Labels
arXiv1 repoarXiv:2301.12689
sRGB-TIR
SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer
arXiv1 repoarXiv:2301.12811
san
arXiv:2301.12844
arXiv1 repoarXiv:2301.12844
HEBO
Icicle: A Re-Designed Emulator for Grey-Box Firmware Fuzzing
arXiv1 repoarXiv:2301.13346
awesome-connected-things-sec
GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis
arXiv1 repoarXiv:2301.13430
GeneFace
In-Context Retrieval-Augmented Language Models
arXiv1 repoarXiv:2302.00083
TinyRAG
Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization
arXiv1 repoarXiv:2302.00275
StreetCLIP
Mixture of Diffusers for scene composition and high resolution image generation
arXiv1 repoarXiv:2302.02412
mixture-of-diffusers
Colossal-Auto: Unified Automation of Parallelization and Activation Checkpoint for Large-scale Models
arXiv1 repoarXiv:2302.02599
ColossalAI
Structure and Content-Guided Video Synthesis with Diffusion Models
arXiv1 repoarXiv:2302.03011
videophy
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations
arXiv1 repoarXiv:2302.03025
TransformerLens
Zero-shot Image-to-Image Translation
arXiv1 repoarXiv:2302.03027
pix2pix-zero
Exploring the Benefits of Training Expert Language Models over Instruction Tuning
arXiv1 repoarXiv:2302.03202
ELM
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
arXiv1 repoarXiv:2302.03540
whisperspeech
A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
arXiv1 repoarXiv:2302.04023
Julia_bench
GPTScore: Evaluate as You Desire
arXiv1 repoarXiv:2302.04166
GPTScore
Will ChatGPT get you caught? Rethinking of Plagiarism Detection
arXiv1 repoarXiv:2302.04335
verify-ai
MaskSketch: Unpaired Structure-guided Masked Image Generation
arXiv1 repoarXiv:2302.05496
ControlNet
Level Generation Through Large Language Models
arXiv1 repoarXiv:2302.05817
lm-pcg
Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking
arXiv1 repoarXiv:2302.07189
LM-ontology-concept-placement
Cliff-Learning
arXiv1 repoarXiv:2302.07348
scaling
How to Train Your DRAGON: Diverse Augmentation Towards Generalizable Dense Retrieval
arXiv1 repoarXiv:2302.07452
dpr-scale
Slapo: A Schedule Language for Progressive Optimization of Large Deep Learning Model Training
arXiv1 repoarXiv:2302.08005
slapo
Composer: Creative and Controllable Image Synthesis with Composable Conditions
arXiv1 repoarXiv:2302.09778
videocomposer
Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels
arXiv1 repoarXiv:2302.10586
U-ViT
FiNER-ORD: Financial Named Entity Recognition Open Research Dataset
arXiv1 repoarXiv:2302.11157
flare-finer-ord
Guiding Large Language Models via Directional Stimulus Prompting
arXiv1 repoarXiv:2302.11520
Directional-Stimulus-Prompting
Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models
arXiv1 repoarXiv:2302.12228
e4t-diffusion
Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple Views
arXiv1 repoarXiv:2302.12297
temporal-robustness
VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining
arXiv1 repoarXiv:2302.12584
VivesDebate-Speech
FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning
arXiv1 repoarXiv:2302.13485
robustlearn
Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation
arXiv1 repoarXiv:2303.00440
EMA-VFI
Do Machine Learning Models Learn Statistical Rules Inferred from Data?
arXiv1 repoarXiv:2303.01433
sqrl
Chasing Low-Carbon Electricity for Practical and Sustainable DNN Training
arXiv1 repoarXiv:2303.02508
zeus
Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes
arXiv1 repoarXiv:2303.02760
HumanArt
Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding
arXiv1 repoarXiv:2303.03267
speech-adapters
CoTEVer: Chain of Thought Prompting Annotation Toolkit for Explanation Verification
arXiv1 repoarXiv:2303.03628
CoTEVer
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
arXiv1 repoarXiv:2303.03926
SpeechT5
Scaling up GANs for Text-to-Image Synthesis
arXiv1 repoarXiv:2303.05511
GigaGAN
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
arXiv1 repoarXiv:2303.06458
CLFM
Query2doc: Query Expansion with Large Language Models
arXiv1 repoarXiv:2303.07678
SPLADERunner
A Simple Framework for Open-Vocabulary Segmentation and Detection
arXiv1 repoarXiv:2303.08131
DINO
SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference
arXiv1 repoarXiv:2303.08308
Moonlit
VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation
arXiv1 repoarXiv:2303.08320
MuseV
PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs
arXiv1 repoarXiv:2303.08954
presto
TriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice Conversion
arXiv1 repoarXiv:2303.09057
TriAAN-VC
BanglaCoNER: Towards Robust Bangla Complex Named Entity Recognition
arXiv1 repoarXiv:2303.09306
Bangla-Complex-Named-Entity-Recognition-Challenge
ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices
arXiv1 repoarXiv:2303.09730
Moonlit
On the rise of fear speech in online social media
arXiv1 repoarXiv:2303.10311
Fearspeech-project
AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
arXiv1 repoarXiv:2303.10512
Continual-NExT
EVA-02: A Visual Representation for Neon Genesis
arXiv1 repoarXiv:2303.11331
eva02_large_patch14_448.mim_m38m_ft_in1k
Reflexion: Language Agents with Verbal Reinforcement Learning
arXiv1 repoarXiv:2303.11366
ReWOO
Stable Bias: Analyzing Societal Representations in Diffusion Models
arXiv1 repoarXiv:2303.11408
BDM1.0
arXiv:2303.11866
arXiv1 repoarXiv:2303.11866
LilT
Natural Language-Assisted Sign Language Recognition
arXiv1 repoarXiv:2303.12080
SLRT
Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions
arXiv1 repoarXiv:2303.12789
instruct-nerf2nerf
CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive Learning
arXiv1 repoarXiv:2303.12793
SLRT
Visual-Language Prompt Tuning with Knowledge-guided Context Optimization
arXiv1 repoarXiv:2303.13283
BiomedCoOp
Xplainer: From X-Ray Observations to Explainable Zero-Shot Diagnosis
arXiv1 repoarXiv:2303.13391
Fine-Grained_Features_Alignment_via_Constrastive_Learning
End-to-End Diffusion Latent Optimization Improves Classifier Guidance
arXiv1 repoarXiv:2303.13703
DOODL
Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation
arXiv1 repoarXiv:2303.13873
Fantasia3D
Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior
arXiv1 repoarXiv:2303.14184
Make-It-3D
Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels
arXiv1 repoarXiv:2303.14307
Llama-AVSR
Human Preference Score: Better Aligning Text-to-Image Models with Human Preference
arXiv1 repoarXiv:2303.14420
align_sd
ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks
arXiv1 repoarXiv:2303.15056
autolabel
Fine-grained Audible Video Description
arXiv1 repoarXiv:2303.15616
FAVDBench
A Pilot Study of Query-Free Adversarial Attack against Stable Diffusion
arXiv1 repoarXiv:2303.16378
Fine-Grained_Features_Alignment_via_Constrastive_Learning
Hierarchical Video-Moment Retrieval and Step-Captioning
arXiv1 repoarXiv:2303.16406
TimeChat-Online-139K
Fairlearn: Assessing and Improving Fairness of AI Systems
arXiv1 repoarXiv:2303.16626
lares
AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators
arXiv1 repoarXiv:2303.16854
autolabel
DERA: Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents
arXiv1 repoarXiv:2303.17071
curai-research
arXiv:2303.17602
arXiv1 repoarXiv:2303.17602
TIL-2023
Token Merging for Fast Stable Diffusion
arXiv1 repoarXiv:2303.17604
Text2Video-Zero
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
arXiv1 repoarXiv:2303.17760
camel
Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations
arXiv1 repoarXiv:2303.18027
JMedBench
ViMMRC 2.0 -- Enhancing Machine Reading Comprehension on Vietnamese Literature Text
arXiv1 repoarXiv:2303.18162
ViQG
Assessing Language Model Deployment with Risk Cards
arXiv1 repoarXiv:2303.18190
garak
arXiv:2304.01715
arXiv1 repoarXiv:2304.01715
LVVIS
AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content Creation
arXiv1 repoarXiv:2304.01961
wit
Dialogue-Contextualized Re-ranking for Medical History-Taking
arXiv1 repoarXiv:2304.01974
curai-research
DoUnseen: Tuning-Free Class-Adaptive Object Detection of Unseen Objects for Robotic Grasping
arXiv1 repoarXiv:2304.02833
image_agnostic_segmentation
arXiv:2304.02970
arXiv1 repoarXiv:2304.02970
Echo-ViLD
Instruction Tuning with GPT-4
arXiv1 repoarXiv:2304.03277
Instruct-SkillMix-SDD
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
arXiv1 repoarXiv:2304.03279
machiavelli
Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder
arXiv1 repoarXiv:2304.04052
Partial-Attention-Language-Model
OpenAGI: When LLM Meets Domain Experts
arXiv1 repoarXiv:2304.04370
OpenAGI
VARS: Video Assistant Referee System for Automated Soccer Decision Making from Multiple Views
arXiv1 repoarXiv:2304.04617
thesis_automatic_faul_recognition
On the Possibilities of AI-Generated Text Detection
arXiv1 repoarXiv:2304.04736
verify-ai
Multi-step Jailbreaking Privacy Attacks on ChatGPT
arXiv1 repoarXiv:2304.05197
LLM-Multistep-Jailbreak
Improving Diffusion Models for Scene Text Editing with Dual Encoders
arXiv1 repoarXiv:2304.05568
DiffSTE
Unicom: Universal and Compact Representation Learning for Image Retrieval
arXiv1 repoarXiv:2304.05884
unicom
DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion
arXiv1 repoarXiv:2304.06025
DreamPose
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
arXiv1 repoarXiv:2304.06364
XVERSE-7B
HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge
arXiv1 repoarXiv:2304.06975
Huatuo-Llama-Med-Chinese
OpenAssistant Conversations -- Democratizing Large Language Model Alignment
arXiv1 repoarXiv:2304.07327
oasst1
ArguGPT: evaluating, understanding and identifying argumentative essays generated by GPT models
arXiv1 repoarXiv:2304.07666
verify-ai
Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation
arXiv1 repoarXiv:2304.07854
BiLLa
DETRs Beat YOLOs on Real-time Object Detection
arXiv1 repoarXiv:2304.08069
rt-detr-huggingface
MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
arXiv1 repoarXiv:2304.08247
DiagnosisCoding
The MiniPile Challenge for Data-Efficient Language Models
arXiv1 repoarXiv:2304.08442
minipile
MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing
arXiv1 repoarXiv:2304.08465
MasaCtrl
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
arXiv1 repoarXiv:2304.08818
NeuroClips
Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents
arXiv1 repoarXiv:2304.09542
RankGPT
Anything-3D: Towards Single-view Anything Reconstruction in the Wild
arXiv1 repoarXiv:2304.10261
Anything-3D
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
arXiv1 repoarXiv:2304.10592
bilingual-gpt-neox-4b-minigpt4
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
arXiv1 repoarXiv:2304.11277
MindSpeed-MM
L3Cube-IndicSBERT: A simple approach for learning cross-lingual sentence representations using multilingual BERT
arXiv1 repoarXiv:2304.11434
bengali-sentence-similarity-sbert
Unlocking Context Constraints of LLMs: Enhancing Context Efficiency of LLMs with Self-Information-Based Content Filtering
arXiv1 repoarXiv:2304.12102
Selective_Context
Segment Anything in Medical Images
arXiv1 repoarXiv:2304.12306
SAMReg
A Static Pruning Study on Sparse Neural Retrievers
arXiv1 repoarXiv:2304.12702
splade
Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
arXiv1 repoarXiv:2304.13731
tango
Customized Segment Anything Model for Medical Image Segmentation
arXiv1 repoarXiv:2304.13785
TriALS
DataComp: In search of the next generation of multimodal datasets
arXiv1 repoarXiv:2304.14108
datacomp
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
arXiv1 repoarXiv:2304.14178
A-Lightweight-Unified-Autoregressive-MLLM
PMC-LLaMA: Towards Building Open-source Language Models for Medicine
arXiv1 repoarXiv:2304.14454
PMC-LLaMA
MinMaxLTTB: Leveraging MinMax-Preselection to Scale LTTB
arXiv1 repoarXiv:2305.00332
flot-downsample
The Art of the Fugue: Minimizing Interleaving in Collaborative Text Editing
arXiv1 repoarXiv:2305.00583
loro
Self-Evaluation Guided Beam Search for Reasoning
arXiv1 repoarXiv:2305.00633
minihf
TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion Synthesis
arXiv1 repoarXiv:2305.00976
TMR-SOMA-RP-v1
VPGTrans: Transfer Visual Prompt Generator across LLMs
arXiv1 repoarXiv:2305.01278
A-Lightweight-Unified-Autoregressive-MLLM
Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation
arXiv1 repoarXiv:2305.01569
banana100-additional-iqa-models
Finding Neurons in a Haystack: Case Studies with Sparse Probing
arXiv1 repoarXiv:2305.01610
TransformerLens
CodeGen2: Lessons for Training LLMs on Programming and Natural Languages
arXiv1 repoarXiv:2305.02309
CodeGen
Caption Anything: Interactive Image Description with Diverse Multimodal Controls
arXiv1 repoarXiv:2305.02677
Caption-Anything
HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
arXiv1 repoarXiv:2305.02765
AcademiCodec
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
arXiv1 repoarXiv:2305.03047
minihf
Personalize Segment Anything Model with One Shot
arXiv1 repoarXiv:2305.03048
geti-instant-learn
Automatic Prompt Optimization with "Gradient Descent" and Beam Search
arXiv1 repoarXiv:2305.03495
PRL-Prompts-from-Reinforcement-Learning
Exploring One-shot Semi-supervised Federated Learning with A Pre-trained Diffusion Model
arXiv1 repoarXiv:2305.04063
FedDISC
Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
arXiv1 repoarXiv:2305.04091
gpt-researcher
Flex-SFU: Accelerating DNN Activation Functions by Non-Uniform Piecewise Approximation
arXiv1 repoarXiv:2305.04546
flex-sfu
WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset
arXiv1 repoarXiv:2305.05432
wit
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
arXiv1 repoarXiv:2305.05726
VRSBench
VideoChat: Chat-Centric Video Understanding
arXiv1 repoarXiv:2305.06355
Ask-Anything
Active Retrieval Augmented Generation
arXiv1 repoarXiv:2305.06983
FLARE
Exploiting Diffusion Prior for Real-World Image Super-Resolution
arXiv1 repoarXiv:2305.07015
StableSR-TestSets
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
arXiv1 repoarXiv:2305.07895
OCRBench
Multilingual Previously Fact-Checked Claim Retrieval
arXiv1 repoarXiv:2305.07991
claim-retrival
Instance-Aware Repeat Factor Sampling for Long-Tailed Object Detection
arXiv1 repoarXiv:2305.08069
camie-tagger-v2
arXiv:2305.08227
arXiv1 repoarXiv:2305.08227
TIL-2023
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
arXiv1 repoarXiv:2305.08283
modular_pluralism
Large Language Model Guided Tree-of-Thought
arXiv1 repoarXiv:2305.08291
tree-of-thought-prompting
Document Understanding Dataset and Evaluation (DUDE)
arXiv1 repoarXiv:2305.08455
sa2va_eval
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
arXiv1 repoarXiv:2305.08850
Make-A-Protagonist
SatLM: Satisfiability-Aided Language Models Using Declarative Prompting
arXiv1 repoarXiv:2305.09656
aurora-m2
Explaining black box text modules in natural language with language models
arXiv1 repoarXiv:2305.09863
imodelsX
arXiv:2305.10037
arXiv1 repoarXiv:2305.10037
model_swarm
UniEX: An Effective and Efficient Framework for Unified Information Extraction via a Span-extractive Perspective
arXiv1 repoarXiv:2305.10306
Fengshenbang-LM
SLiC-HF: Sequence Likelihood Calibration with Human Feedback
arXiv1 repoarXiv:2305.10425
pair-preference-model-LLaMA3-8B
Statistical Knowledge Assessment for Large Language Models
arXiv1 repoarXiv:2305.10519
DistFactAssessLM
Structural Pruning for Diffusion Models
arXiv1 repoarXiv:2305.10924
Diff-Pruning
SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
arXiv1 repoarXiv:2305.11000
SpeechGPT
Going Denser with Open-Vocabulary Part Segmentation
arXiv1 repoarXiv:2305.11173
grounded-segment-any-parts
VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks
arXiv1 repoarXiv:2305.11175
VisionLLM
Self-QA: Unsupervised Knowledge Guided Language Model Alignment
arXiv1 repoarXiv:2305.11952
XuanYuan
XuanYuan 2.0: A Large Chinese Financial Chat Model with Hundreds of Billions Parameters
arXiv1 repoarXiv:2305.12002
XuanYuan
Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
arXiv1 repoarXiv:2305.12031
med42
Movie101: A New Movie Understanding Benchmark
arXiv1 repoarXiv:2305.12140
TimeChat-Online-139K
Are Your Explanations Reliable? Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial Attack
arXiv1 repoarXiv:2305.12351
XAIFooler
Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design
arXiv1 repoarXiv:2305.13035
siglip-so400m-patch14-384
RWKV: Reinventing RNNs for the Transformer Era
arXiv1 repoarXiv:2305.13048
rwkv
Making Language Models Better Tool Learners with Execution Feedback
arXiv1 repoarXiv:2305.13068
KnowLM
Editing Large Language Models: Problems, Methods, and Opportunities
arXiv1 repoarXiv:2305.13172
KnowEdit
RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text
arXiv1 repoarXiv:2305.13304
Recurrent-LLM
Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching
arXiv1 repoarXiv:2305.13310
geti-instant-learn
Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach
arXiv1 repoarXiv:2305.13579
ProFusion
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
arXiv1 repoarXiv:2305.13655
Omost
Aligning Large Language Models through Synthetic Feedback
arXiv1 repoarXiv:2305.13735
minihf
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
arXiv1 repoarXiv:2305.14004
Saamayik
Parts of Speech-Grounded Subspaces in Vision-Language Models
arXiv1 repoarXiv:2305.14053
PoS-subspaces
Language Models with Rationality
arXiv1 repoarXiv:2305.14250
llm-mysteries
Linear Cross-Lingual Mapping of Sentence Embeddings
arXiv1 repoarXiv:2305.14256
Wikinews-multilingual
Query Rewriting for Retrieval-Augmented Large Language Models
arXiv1 repoarXiv:2305.14283
RAG-query-rewriting
Connecting Multi-modal Contrastive Representations
arXiv1 repoarXiv:2305.14381
20251R0136COSE40500
CGCE: A Chinese Generative Chat Evaluation Benchmark for General and Financial Domains
arXiv1 repoarXiv:2305.14471
XuanYuan
Optimal Linear Subspace Search: Learning to Construct Fast and High-Quality Schedulers for Diffusion Models
arXiv1 repoarXiv:2305.14677
diffSynth-studio-notes
Investigating Table-to-Text Generation Capabilities of LLMs in Real-World Information Seeking Scenarios
arXiv1 repoarXiv:2305.14987
LLM-T2T
Gorilla: Large Language Model Connected with Massive APIs
arXiv1 repoarXiv:2305.15334
gorilla
arXiv:2305.15717
arXiv1 repoarXiv:2305.15717
alpaca_eval
On the Planning Abilities of Large Language Models : A Critical Investigation
arXiv1 repoarXiv:2305.15771
LLMs-Planning
arXiv:2305.15778
arXiv1 repoarXiv:2305.15778
awesome-ai-sre
arXiv:2305.15930
arXiv1 repoarXiv:2305.15930
HEBO
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
arXiv1 repoarXiv:2305.16107
SpeechT5
ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
arXiv1 repoarXiv:2305.16213
prolificdreamer
Landmark Attention: Random-Access Infinite Context Length for Transformers
arXiv1 repoarXiv:2305.16300
landmark-attention
IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages
arXiv1 repoarXiv:2305.16307
IndicTrans2
Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models
arXiv1 repoarXiv:2305.16322
Uni-ControlNet
An Empirical Comparison of LM-based Question and Answer Generation Methods
arXiv1 repoarXiv:2305.17002
lm-question-generation
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks
arXiv1 repoarXiv:2305.17100
DermFM-Zero
PromptNER: Prompt Locating and Typing for Named Entity Recognition
arXiv1 repoarXiv:2305.17104
PromptNER
arXiv:2305.17216
arXiv1 repoarXiv:2305.17216
CoBSAT
Fine-Tuning Language Models with Just Forward Passes
arXiv1 repoarXiv:2305.17333
MeZO
DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
arXiv1 repoarXiv:2305.17359
AdaDetectGPT
A Practical Toolkit for Multilingual Question and Answer Generation
arXiv1 repoarXiv:2305.17416
lm-question-generation
SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created Through Human-Machine Collaboration
arXiv1 repoarXiv:2305.17696
korean-safety-benchmarks
KoSBi: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Application
arXiv1 repoarXiv:2305.17701
korean-safety-benchmarks
arXiv:2305.17926
arXiv1 repoarXiv:2305.17926
alpaca_eval
ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback
arXiv1 repoarXiv:2305.18090
ChatDrug
Do Large Language Models Know What They Don't Know?
arXiv1 repoarXiv:2305.18153
SelfAware
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
arXiv1 repoarXiv:2305.18752
GPT4Tools
Voice Conversion With Just Nearest Neighbors
arXiv1 repoarXiv:2305.18975
knn-vc
Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models
arXiv1 repoarXiv:2305.19187
UQ-NLG
infoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-information
arXiv1 repoarXiv:2305.19344
infoVerse
CryptOpt: Automatic Optimization of Straightline Code
arXiv1 repoarXiv:2305.19586
CryptOpt
Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust
arXiv1 repoarXiv:2305.20030
tree-ring-watermark
Let's Verify Step by Step
arXiv1 repoarXiv:2305.20050
KOpen-platypus
Humans in 4D: Reconstructing and Tracking Humans with Transformers
arXiv1 repoarXiv:2305.20091
4D-Humans
Preference-grounded Token-level Guidance for Language Model Fine-tuning
arXiv1 repoarXiv:2306.00398
DenseRewardRLHF-PPO
End-to-end Knowledge Retrieval with Multi-modal Queries
arXiv1 repoarXiv:2306.00424
ReMuQ
Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners
arXiv1 repoarXiv:2306.00561
j2a
arXiv:2306.00745
arXiv1 repoarXiv:2306.00745
Jellyfish-13B
Continual Learning for Abdominal Multi-Organ and Tumor Segmentation
arXiv1 repoarXiv:2306.00988
awesome-cybersecurity-agentic-ai
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
arXiv1 repoarXiv:2306.01442
ai-research-code
SACSoN: Scalable Autonomous Control for Social Navigation
arXiv1 repoarXiv:2306.01874
UniWM_Dataset
Can Contextual Biasing Remain Effective with Whisper and GPT-2?
arXiv1 repoarXiv:2306.01942
WhisperBiasing
Exploring the Optimal Choice for Generative Processes in Diffusion Models: Ordinary vs Stochastic Differential Equations
arXiv1 repoarXiv:2306.02063
LakonLab
Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
arXiv1 repoarXiv:2306.02416
adapt_med_seg
Benchmarking Large Language Models on CMExam -- A Comprehensive Chinese Medical Exam Dataset
arXiv1 repoarXiv:2306.03030
CMExam
Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs
arXiv1 repoarXiv:2306.03081
llamppl
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias
arXiv1 repoarXiv:2306.03509
MagVITS
GEO-Bench: Toward Foundation Models for Earth Monitoring
arXiv1 repoarXiv:2306.03831
geo-bench
CL-UZH at SemEval-2023 Task 10: Sexism Detection through Incremental Fine-Tuning and Multi-Task Learning with Label Descriptions
arXiv1 repoarXiv:2306.03907
CL-UZH-EDOS-2023
M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning
arXiv1 repoarXiv:2306.04387
M3IT
arXiv:2306.04618
arXiv1 repoarXiv:2306.04618
OOD_NLP
On the Reliability of Watermarks for Large Language Models
arXiv1 repoarXiv:2306.04634
lm-watermarking
Generalizable Low-Resource Activity Recognition with Diverse and Discriminative Representation Learning
arXiv1 repoarXiv:2306.04641
robustlearn
Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
arXiv1 repoarXiv:2306.04723
L2D
arXiv:2306.04848
arXiv1 repoarXiv:2306.04848
smalldiffusion
K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization
arXiv1 repoarXiv:2306.05064
k2
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
arXiv1 repoarXiv:2306.05350
peft-ser
PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance
arXiv1 repoarXiv:2306.05443
PIXIU
Prompt Injection attack against LLM-integrated Applications
arXiv1 repoarXiv:2306.05499
Awesome-OpenClaw
There's Plenty of Room in the Middle: The Unsung Revolution of the Renormalization Group
arXiv1 repoarXiv:2306.06020
Low-power-E-Paper-OS
Image Vectorization: a Review
arXiv1 repoarXiv:2306.06441
vtracer
detrex: Benchmarking Detection Transformers
arXiv1 repoarXiv:2306.07265
detrex
Scalable 3D Captioning with Pretrained Models
arXiv1 repoarXiv:2306.07279
Cap3D
Resources for Brewing BEIR: Reproducible Reference Models and an Official Leaderboard
arXiv1 repoarXiv:2306.07471
beir
TART: A plug-and-play Transformer module for task-agnostic reasoning
arXiv1 repoarXiv:2306.07536
TART
SqueezeLLM: Dense-and-Sparse Quantization
arXiv1 repoarXiv:2306.07629
turboquant-vllm
WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences
arXiv1 repoarXiv:2306.07906
WebGLM
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
arXiv1 repoarXiv:2306.07954
Rerender_A_Video
LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting
arXiv1 repoarXiv:2306.08259
Traffic_Forecast_Benchmark
TryOnDiffusion: A Tale of Two UNets
arXiv1 repoarXiv:2306.08276
opentryon
arXiv:2306.08645
arXiv1 repoarXiv:2306.08645
Video-Infinity
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations
arXiv1 repoarXiv:2306.08658
Babel-ImageNet
LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
arXiv1 repoarXiv:2306.09265
Multi-Modality-Arena
DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data
arXiv1 repoarXiv:2306.09344
dreamsim
Scaling Open-Vocabulary Object Detection
arXiv1 repoarXiv:2306.09683
owlv2-base-patch16-ensemble
arXiv:2306.09803
arXiv1 repoarXiv:2306.09803
HEBO
ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation
arXiv1 repoarXiv:2306.09968
ClinicalGPT-base-zh
Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects
arXiv1 repoarXiv:2306.10125
time-moe
MARBLE: Music Audio Representation Benchmark for Universal Evaluation
arXiv1 repoarXiv:2306.10548
usad
Guiding Language Models of Code with Global Context using Monitors
arXiv1 repoarXiv:2306.10763
monitors4codegen
SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking Faces
arXiv1 repoarXiv:2306.10799
SelfTalk_release
Textbooks Are All You Need
arXiv1 repoarXiv:2306.11644
cosmopedia
A Simple and Effective Pruning Approach for Large Language Models
arXiv1 repoarXiv:2306.11695
cosyvoice3-inference-acceleration
Training Transformers with 4-bit Integers
arXiv1 repoarXiv:2306.11987
JobList
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
arXiv1 repoarXiv:2306.12420
LMFlow
Generative Multimodal Entity Linking
arXiv1 repoarXiv:2306.12725
GEMEL
An overview on the evaluated video retrieval tasks at TRECVID 2022
arXiv1 repoarXiv:2306.13118
ladi-overview
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
arXiv1 repoarXiv:2306.13394
Video-MME
Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data
arXiv1 repoarXiv:2306.13840
ultimate-utils
H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
arXiv1 repoarXiv:2306.14048
mlx-flash
Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference
arXiv1 repoarXiv:2306.14393
Moonlit
SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality
arXiv1 repoarXiv:2306.14610
SugarCrepe_pp
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
arXiv1 repoarXiv:2306.15195
shikra
arXiv:2306.15895
arXiv1 repoarXiv:2306.15895
ReGen
UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data
arXiv1 repoarXiv:2306.16083
UnitSpeech
SE-PQA: Personalized Community Question Answering
arXiv1 repoarXiv:2306.16261
SE-PQA
Foundation Model for Endoscopy Video Analysis via Large-scale Self-supervised Pre-train
arXiv1 repoarXiv:2306.16741
Endo-FM
Tokenization and the Noiseless Channel
arXiv1 repoarXiv:2306.16842
tokenizer-flores-validation
Hierarchical Neural Coding for Controllable CAD Model Generation
arXiv1 repoarXiv:2307.00149
FlexCAD
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models
arXiv1 repoarXiv:2307.01379
SAR
Detecting Images Generated by Deep Diffusion Models using their Local Intrinsic Dimensionality
arXiv1 repoarXiv:2307.02347
deepfake_multiLID
Jailbroken: How Does LLM Safety Training Fail?
arXiv1 repoarXiv:2307.02483
JailbreakLab
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
arXiv1 repoarXiv:2307.03601
gradio-box
A Survey on Graph Neural Networks for Time Series: Forecasting, Classification, Imputation, and Anomaly Detection
arXiv1 repoarXiv:2307.03759
time-moe
InPars Toolkit: A Unified and Reproducible Synthetic Data Generation Pipeline for Neural Information Retrieval
arXiv1 repoarXiv:2307.04601
InPars
arXiv:2307.05222
arXiv1 repoarXiv:2307.05222
CoBSAT
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
arXiv1 repoarXiv:2307.06942
InternVid-Full
MGit: A Model Versioning and Management System
arXiv1 repoarXiv:2307.07507
mgit
CoTracker: It is Better to Track Together
arXiv1 repoarXiv:2307.07635
co-tracker
CA-LoRA: Adapting Existing LoRA for Compressed LLMs to Enable Efficient Multi-Tasking on Personal Devices
arXiv1 repoarXiv:2307.07705
CA-LoRA
ChatDev: Communicative Agents for Software Development
arXiv1 repoarXiv:2307.07924
ChatDev
Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
arXiv1 repoarXiv:2307.08487
latent-jailbreak
Generative Type Inference for Python
arXiv1 repoarXiv:2307.09163
TypeGen
CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
arXiv1 repoarXiv:2307.09705
COIG-CQIA
arXiv:2307.10373
arXiv1 repoarXiv:2307.10373
TokenFlow
L-Eval: Instituting Standardized Evaluation for Long Context Language Models
arXiv1 repoarXiv:2307.11088
Qwen
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
arXiv1 repoarXiv:2307.11394
meeteval
A Change of Heart: Improving Speech Emotion Recognition through Speech-to-Text Modality Conversion
arXiv1 repoarXiv:2307.11584
dissertation-project
Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry
arXiv1 repoarXiv:2307.12868
Diffusion-Pullback
FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
arXiv1 repoarXiv:2307.13528
SDAK
Measuring Faithfulness in Chain-of-Thought Reasoning
arXiv1 repoarXiv:2307.13702
CoTFaithChecker
Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image Classification
arXiv1 repoarXiv:2307.15254
MHIM-MIL
ChatHome: Development and Evaluation of a Domain-Specific Language Model for Home Renovation
arXiv1 repoarXiv:2307.15290
BELLE
A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats
arXiv1 repoarXiv:2307.15517
mase
The Hydra Effect: Emergent Self-repair in Language Model Computations
arXiv1 repoarXiv:2307.15771
tfmlens
Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection
arXiv1 repoarXiv:2307.16888
virtual-prompt-injection
Three Bricks to Consolidate Watermarks for Large Language Models
arXiv1 repoarXiv:2308.00113
lm-watermarking
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
arXiv1 repoarXiv:2308.00352
codel
From Sparse to Soft Mixtures of Experts
arXiv1 repoarXiv:2308.00951
PETL_AST
Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data
arXiv1 repoarXiv:2308.02463
RadFM
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
arXiv1 repoarXiv:2308.02490
MM-Vet
Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning
arXiv1 repoarXiv:2308.02533
robustlearn
Studying Large Language Model Generalization with Influence Functions
arXiv1 repoarXiv:2308.03296
bergson
DiffSynth: Latent In-Iteration Deflickering for Realistic Video Synthesis
arXiv1 repoarXiv:2308.03463
diffSynth-studio-notes
AgentBench: Evaluating LLMs as Agents
arXiv1 repoarXiv:2308.03688
Awesome-OpenClaw
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
arXiv1 repoarXiv:2308.03729
Multi-Modality-Arena
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
arXiv1 repoarXiv:2308.04430
silo-lm
Classification of Human- and AI-Generated Texts: Investigating Features for ChatGPT
arXiv1 repoarXiv:2308.05341
verify-ai
Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow
arXiv1 repoarXiv:2308.06101
DCI-VTON-Virtual-Try-On
DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion Models
arXiv1 repoarXiv:2308.06160
DatasetDM
Self-Alignment with Instruction Backtranslation
arXiv1 repoarXiv:2308.06259
EasyInstruct
Foundation Model is Efficient Multimodal Multitask Model Selector
arXiv1 repoarXiv:2308.06262
Multitask-Model-Selector
Tiny and Efficient Model for the Edge Detection Generalization
arXiv1 repoarXiv:2308.06468
ComfyUI-Anyline
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models
arXiv1 repoarXiv:2308.07269
KnowEdit
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
arXiv1 repoarXiv:2308.07308
llm-self-defense
DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory
arXiv1 repoarXiv:2308.08089
dragnuwa-pruned-safetensors
Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes
arXiv1 repoarXiv:2308.08769
Chat-Scene
Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?
arXiv1 repoarXiv:2308.10168
DistFactAssessLM
ROSGPT_Vision: Commanding Robots Using Only Language Models' Prompts
arXiv1 repoarXiv:2308.11236
ROSGPT_Vision
Efficient Benchmarking of Language Models
arXiv1 repoarXiv:2308.11696
helm
arXiv:2308.11957
arXiv1 repoarXiv:2308.11957
hashing-baseline
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
arXiv1 repoarXiv:2308.12038
OmniLMM-12B
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
arXiv1 repoarXiv:2308.12067
InstructionGPT-4
MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning
arXiv1 repoarXiv:2308.13218
MultiCapCLIP
arXiv:2308.13418
arXiv1 repoarXiv:2308.13418
nougat
Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models
arXiv1 repoarXiv:2308.13437
pvit
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
arXiv1 repoarXiv:2308.14508
Marathon
CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation
arXiv1 repoarXiv:2308.15226
CLIPTrans
When Do Program-of-Thoughts Work for Reasoning?
arXiv1 repoarXiv:2308.15452
EasyInstruct
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
arXiv1 repoarXiv:2308.15459
ParaGuide
arXiv:2308.16361
arXiv1 repoarXiv:2308.16361
Jellyfish-13B
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
arXiv1 repoarXiv:2308.16884
M-AbstainQA
TouchStone: Evaluating Vision-Language Models by Language Models
arXiv1 repoarXiv:2308.16890
TouchStone
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
arXiv1 repoarXiv:2309.00267
YiVal
NLLB-CLIP -- train performant multilingual image retrieval model on a budget
arXiv1 repoarXiv:2309.01859
CLIP_benchmark
arXiv:2309.02243
arXiv1 repoarXiv:2309.02243
improvnet
Certifying LLM Safety against Adversarial Prompting
arXiv1 repoarXiv:2309.02705
certified-llm-safety
Norm Tweaking: High-performance Low-bit Quantization of Large Language Models
arXiv1 repoarXiv:2309.02784
norm-tweaking
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
arXiv1 repoarXiv:2309.02836
bigvsan
Matcha-TTS: A fast TTS architecture with conditional flow matching
arXiv1 repoarXiv:2309.03199
Matcha-TTS
Large Language Models as Optimizers
arXiv1 repoarXiv:2309.03409
YiVal
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
arXiv1 repoarXiv:2309.03883
SLED
Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis
arXiv1 repoarXiv:2309.03904
Aurora
Large-Scale Automatic Audiobook Creation
arXiv1 repoarXiv:2309.03926
SynapseML
Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese
arXiv1 repoarXiv:2309.04175
Huatuo-Llama-Med-Chinese
Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain
arXiv1 repoarXiv:2309.04198
Huatuo-Llama-Med-Chinese
From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting
arXiv1 repoarXiv:2309.04269
YiVal
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
arXiv1 repoarXiv:2309.04662
madlad400-10b-mt
arXiv:2309.05519
arXiv1 repoarXiv:2309.05519
NExT-GPT
Kani: A Lightweight and Highly Hackable Framework for Building Language Model Applications
arXiv1 repoarXiv:2309.05542
kani
An Empirical Study of NetOps Capability of Pre-Trained Large Language Models
arXiv1 repoarXiv:2309.05557
neteval-exam
InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation
arXiv1 repoarXiv:2309.06380
InstaFlow
An Image Dataset for Benchmarking Recommender Systems with Raw Pixels
arXiv1 repoarXiv:2309.06789
PixelRec
RAIN: Your Language Models Can Align Themselves without Finetuning
arXiv1 repoarXiv:2309.07124
RAIN
PromptASR for contextualized ASR with controllable style
arXiv1 repoarXiv:2309.07414
libriheavy
VerilogEval: Evaluating Large Language Models for Verilog Code Generation
arXiv1 repoarXiv:2309.07544
verilog-eval
Generative AI Text Classification using Ensemble LLM Approaches
arXiv1 repoarXiv:2309.07755
verify-ai
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
arXiv1 repoarXiv:2309.07875
med-safety-bench
Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
arXiv1 repoarXiv:2309.08125
Cornstarch
Advancing the Evaluation of Traditional Chinese Language Models: Towards a Comprehensive Benchmark Suite
arXiv1 repoarXiv:2309.08448
TCEval-v2
Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens
arXiv1 repoarXiv:2309.08531
Image-to-Speech
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
arXiv1 repoarXiv:2309.08532
PRL-Prompts-from-Reinforcement-Learning
HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform
arXiv1 repoarXiv:2309.09493
kanade-tokenizer
Adapting Large Language Models to Domains via Reading Comprehension
arXiv1 repoarXiv:2309.09530
KoCommercial-Dataset
Baichuan 2: Open Large-scale Language Models
arXiv1 repoarXiv:2309.10305
Baichuan2
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
arXiv1 repoarXiv:2309.10400
EasyContext
A Configurable Library for Generating and Manipulating Maze Datasets
arXiv1 repoarXiv:2309.10498
maze-dataset
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
arXiv1 repoarXiv:2309.10691
mint-bench
DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Services
arXiv1 repoarXiv:2309.11325
DISC-FinLLM
You Only Look at Screens: Multimodal Chain-of-Action Agents
arXiv1 repoarXiv:2309.11436
carl_distrl
SignBank+: Preparing a Multilingual Sign Language Dataset for Machine Translation Using Large Language Models
arXiv1 repoarXiv:2309.11566
signbank-plus
VoiceLDM: Text-to-Speech with Environmental Context
arXiv1 repoarXiv:2309.13664
VoiceLDM
VidChapters-7M: Video Chapters at Scale
arXiv1 repoarXiv:2309.13952
VidChapters
UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human Generation
arXiv1 repoarXiv:2309.14335
CosmicMan
Joint Audio and Speech Understanding
arXiv1 repoarXiv:2309.14405
Speech-IFEval
Motions in Microseconds via Vectorized Sampling-Based Planning
arXiv1 repoarXiv:2309.14545
mdtensor
Updated Corpora and Benchmarks for Long-Form Speech Recognition
arXiv1 repoarXiv:2309.15013
speech-datasets
RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models
arXiv1 repoarXiv:2309.15088
pyterrier_genrank
A Content-Driven Micro-Video Recommendation Dataset at Scale
arXiv1 repoarXiv:2309.15379
MicroLens
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
arXiv1 repoarXiv:2309.16058
papagei-foundation-model
Masked Autoencoders are Scalable Learners of Cellular Morphology
arXiv1 repoarXiv:2309.16064
maes_microscopy
Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
arXiv1 repoarXiv:2309.16240
f-divergence-dpo
MHG-GNN: Combination of Molecular Hypergraph Grammar with Graph Neural Network
arXiv1 repoarXiv:2309.16374
materials
Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation
arXiv1 repoarXiv:2309.16429
JavisBench
Vision Transformers Need Registers
arXiv1 repoarXiv:2309.16588
geti-instant-learn
DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation
arXiv1 repoarXiv:2309.16653
controlled-dreamgaussian
Demystifying CLIP Data
arXiv1 repoarXiv:2309.16671
MetaCLIP
Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities
arXiv1 repoarXiv:2309.16739
SplitFM
Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks
arXiv1 repoarXiv:2309.17002
robustlearn
Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive Feedback
arXiv1 repoarXiv:2309.17078
RLCF
Guiding Instruction-based Image Editing via Multimodal Large Language Models
arXiv1 repoarXiv:2309.17102
ml-mgie
arXiv:2309.17444
arXiv1 repoarXiv:2309.17444
LLM-groundedVideoDiffusion
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
arXiv1 repoarXiv:2309.17452
NuminaMath-7B-TIR
FELM: Benchmarking Factuality Evaluation of Large Language Models
arXiv1 repoarXiv:2310.00741
felm
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
arXiv1 repoarXiv:2310.01107
Ground-A-Video
Quantifying the Plausibility of Context Reliance in Neural Machine Translation
arXiv1 repoarXiv:2310.01188
LRP-eXplains-Transformers
Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy
arXiv1 repoarXiv:2310.01334
MC-SMoE
Compressing LLMs: The Truth is Rarely Pure and Never Simple
arXiv1 repoarXiv:2310.01382
llm-kick
arXiv:2310.01405
arXiv1 repoarXiv:2310.01405
drowse
ImagenHub: Standardizing the evaluation of conditional image generation models
arXiv1 repoarXiv:2310.01596
ImagenHub
CAT-LM: Training Language Models on Aligned Code And Tests
arXiv1 repoarXiv:2310.01602
CAT-LM
Fool Your (Vision and) Language Model With Embarrassingly Simple Permutations
arXiv1 repoarXiv:2310.01651
FoolyourVLLMs
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
arXiv1 repoarXiv:2310.02255
MathVista
Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
arXiv1 repoarXiv:2310.02279
ctm
Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
arXiv1 repoarXiv:2310.02949
shadow-alignment
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
arXiv1 repoarXiv:2310.03294
EasyContext
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
arXiv1 repoarXiv:2310.03708
modpo
Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts Conversion
arXiv1 repoarXiv:2310.04361
MoFEbaseD2D
Amortizing intractable inference in large language models
arXiv1 repoarXiv:2310.04363
gfn-lm-tuning
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
arXiv1 repoarXiv:2310.04406
PromptingTools.jl
What's the Magic Word? A Control Theory of LLM Prompting
arXiv1 repoarXiv:2310.04444
Magic_Words
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
arXiv1 repoarXiv:2310.04451
GA
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
arXiv1 repoarXiv:2310.04673
FunCodec
DORIS-MAE: Scientific Document Retrieval using Multi-level Aspect-based Queries
arXiv1 repoarXiv:2310.04678
Doris-Mae-Dataset
Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages
arXiv1 repoarXiv:2310.04799
deccp
SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF
arXiv1 repoarXiv:2310.05344
Llama-3_3-Nemotron-Super-49B-GenRM
Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel Decoding
arXiv1 repoarXiv:2310.05424
mixture_of_recursions
Generative Judge for Evaluating Alignment
arXiv1 repoarXiv:2310.05470
JudgeBench
UAVs and Neural Networks for search and rescue missions
arXiv1 repoarXiv:2310.05512
argus
Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution
arXiv1 repoarXiv:2310.05634
SAFE
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
arXiv1 repoarXiv:2310.05736
Prompt-Compression
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
arXiv1 repoarXiv:2310.05737
Pyramid-Flow
Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models
arXiv1 repoarXiv:2310.06313
PCDMs
Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
arXiv1 repoarXiv:2310.06387
llm-jailbreaking-defense
iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
arXiv1 repoarXiv:2310.06625
frn-50k-baseline
Text Embeddings Reveal (Almost) As Much As Text
arXiv1 repoarXiv:2310.06816
align-text-encoders
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
arXiv1 repoarXiv:2310.06839
Prompt-Compression
Violation of Expectation via Metacognitive Prompting Reduces Theory of Mind Prediction Error in Large Language Models
arXiv1 repoarXiv:2310.06983
tutor-gpt
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
arXiv1 repoarXiv:2310.06987
Jailbreak_LLM
Composite Backdoor Attacks Against Large Language Models
arXiv1 repoarXiv:2310.07676
CBA
MatFormer: Nested Transformer for Elastic Inference
arXiv1 repoarXiv:2310.07707
mlx-flash
DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model
arXiv1 repoarXiv:2310.07771
DrivingDiffusion
Large Language Models Are Zero-Shot Time Series Forecasters
arXiv1 repoarXiv:2310.07820
llmtime
MemGPT: Towards LLMs as Operating Systems
arXiv1 repoarXiv:2310.08560
letta-code
arXiv:2310.08588
arXiv1 repoarXiv:2310.08588
Octopus
The Data Lakehouse: Data Warehousing and More
arXiv1 repoarXiv:2310.08697
data-engineer-handbook
Extending Multi-modal Contrastive Representations
arXiv1 repoarXiv:2310.08884
20251R0136COSE40500
SeqXGPT: Sentence-Level AI-Generated Text Detection
arXiv1 repoarXiv:2310.08903
SeqXGPT
A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models
arXiv1 repoarXiv:2310.09497
llm-rankers
TEQ: Trainable Equivalent Transformation for Quantization of LLMs
arXiv1 repoarXiv:2310.10944
auto-round
Domain Generalization Using Large Pretrained Models with Mixture-of-Adapters
arXiv1 repoarXiv:2310.11031
MoA
Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis
arXiv1 repoarXiv:2310.11722
SDAK
A General Theoretical Paradigm to Understand Learning from Human Preferences
arXiv1 repoarXiv:2310.12036
Reinforcement-Learning-Full-Pipeline
Enhancing High-Resolution 3D Generation through Pixel-wise Gradient Clipping
arXiv1 repoarXiv:2310.12474
PGC-3D
arXiv:2310.12537
arXiv1 repoarXiv:2310.12537
Jellyfish-13B
arXiv:2310.12952
arXiv1 repoarXiv:2310.12952
Vendi-Score
GraphGPT: Graph Instruction Tuning for Large Language Models
arXiv1 repoarXiv:2310.13023
GraphGPT
Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
arXiv1 repoarXiv:2310.13724
habitat-lab
Contrast Everything: A Hierarchical Contrastive Framework for Medical Time-Series
arXiv1 repoarXiv:2310.14017
UTSD
Tree Prompting: Efficient Task Adaptation without Fine-Tuning
arXiv1 repoarXiv:2310.14034
imodelsX
Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases
arXiv1 repoarXiv:2310.14303
red-instruct
arXiv:2310.14478
arXiv1 repoarXiv:2310.14478
geolm
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
arXiv1 repoarXiv:2310.14566
LRV-Instruction
Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model
arXiv1 repoarXiv:2310.15110
zero123plus
Open-Set Image Tagging with Multi-Grained Text Supervision
arXiv1 repoarXiv:2310.15200
recognize-anything
Efficient Online String Matching through Linked Weak Factors
arXiv1 repoarXiv:2310.15711
HashChain
In-Context Learning Creates Task Vectors
arXiv1 repoarXiv:2310.15916
icl_task_vectors
Woodpecker: Hallucination Correction for Multimodal Large Language Models
arXiv1 repoarXiv:2310.16045
Woodpecker
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
arXiv1 repoarXiv:2310.16049
llm-mysteries
Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge
arXiv1 repoarXiv:2310.16112
Fine-Grained_Features_Alignment_via_Constrastive_Learning
Attention Lens: A Tool for Mechanistically Interpreting the Attention Head Information Retrieval Mechanism
arXiv1 repoarXiv:2310.16270
AttentionLens
A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation
arXiv1 repoarXiv:2310.16656
tuxemon
DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior
arXiv1 repoarXiv:2310.16818
DreamCraft3D
TD-MPC2: Scalable, Robust World Models for Continuous Control
arXiv1 repoarXiv:2310.16828
xwm
How do Language Models Bind Entities in Context?
arXiv1 repoarXiv:2310.17191
binding-iclr
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
arXiv1 repoarXiv:2310.17631
JudgeBench
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
arXiv1 repoarXiv:2310.17884
confaide
FormalGeo: An Extensible Formalized Framework for Olympiad Geometric Problem Solving
arXiv1 repoarXiv:2310.18021
FormalGeo
DUMA: a Dual-Mind Conversational Agent with Fast and Slow Thinking
arXiv1 repoarXiv:2310.18075
BELLE
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
arXiv1 repoarXiv:2310.18235
DSG
Image Clustering Conditioned on Text Criteria
arXiv1 repoarXiv:2310.18297
ICTC
LLMSTEP: LLM proofstep suggestions in Lean
arXiv1 repoarXiv:2310.18457
llmstep
EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images
arXiv1 repoarXiv:2310.18652
blendsql
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
arXiv1 repoarXiv:2310.19102
deepcompressor
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
arXiv1 repoarXiv:2310.19512
VideoCrafter
arXiv:2310.20145
arXiv1 repoarXiv:2310.20145
HEBO
Integrating curation into scientific publishing to train AI models
arXiv1 repoarXiv:2310.20440
SourceData
JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models
arXiv1 repoarXiv:2311.00286
jade-db
TopicGPT: A Prompt-based Topic Modeling Framework
arXiv1 repoarXiv:2311.01449
turkish-complaint-topic-clustering
Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization
arXiv1 repoarXiv:2311.01544
who_what_benchmark
Adapting Frechet Audio Distance for Generative Music Evaluation
arXiv1 repoarXiv:2311.01616
fadtk
SAC3: Reliable Hallucination Detection in Black-Box Language Models via Semantic-aware Cross-check Consistency
arXiv1 repoarXiv:2311.01740
sac3
AnyText: Multilingual Visual Text Generation And Editing
arXiv1 repoarXiv:2311.03054
AnyText
GLaMM: Pixel Grounding Large Multimodal Model
arXiv1 repoarXiv:2311.03356
groundingLMM
I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models
arXiv1 repoarXiv:2311.04145
i2vgen-xl
Holistic Evaluation of Text-To-Image Models
arXiv1 repoarXiv:2311.04287
helm
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
arXiv1 repoarXiv:2311.04934
prompt-cache
LooGLE: Can Long-Context Language Models Understand Long Contexts?
arXiv1 repoarXiv:2311.04939
LooGLE
BeLLM: Backward Dependency Enhanced Large Language Model for Sentence Embeddings
arXiv1 repoarXiv:2311.05296
AnglE
Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model
arXiv1 repoarXiv:2311.06214
DiffSplat
Knowledge-Augmented Large Language Models for Personalized Contextual Query Suggestion
arXiv1 repoarXiv:2311.06318
azure-openai-llm-cookbook
PerceptionGPT: Effectively Fusing Visual Perception into LLM
arXiv1 repoarXiv:2311.06612
medllm
Embarassingly Simple Dataset Distillation
arXiv1 repoarXiv:2311.07025
PoDD
Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification
arXiv1 repoarXiv:2311.07125
AEM-dataset
arXiv:2311.07574
arXiv1 repoarXiv:2311.07574
prismatic-vlms
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
arXiv1 repoarXiv:2311.07575
idefics2-8b
Large Language Models can Strategically Deceive their Users when Put Under Pressure
arXiv1 repoarXiv:2311.07590
how-to-catch-an-ai-liar
Computing Implicitizations of Multi-Graded Polynomial Maps
arXiv1 repoarXiv:2311.07678
smash
Fair Abstractive Summarization of Diverse Perspectives
arXiv1 repoarXiv:2311.07884
FairSumm
One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion
arXiv1 repoarXiv:2311.07885
torchsparse
REST: Retrieval-Based Speculative Decoding
arXiv1 repoarXiv:2311.08252
REST
Learning to Filter Context for Retrieval-Augmented Generation
arXiv1 repoarXiv:2311.08377
TinyRAG
Towards Open-Ended Visual Recognition with Large Language Model
arXiv1 repoarXiv:2311.08400
OmniScient-Model
Fine-tuning Language Models for Factuality
arXiv1 repoarXiv:2311.08401
llm_factuality_tuning
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
arXiv1 repoarXiv:2311.08590
PEMA
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
arXiv1 repoarXiv:2311.08640
MCKD-semantic-segmentation
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
arXiv1 repoarXiv:2311.08849
ofa
FastBlend: a Powerful Model-Free Toolkit Making Video Stylization Easier
arXiv1 repoarXiv:2311.09265
diffSynth-studio-notes
HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM
arXiv1 repoarXiv:2311.09528
Llama-3_3-Nemotron-Super-49B-GenRM
MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification
arXiv1 repoarXiv:2311.09761
MAFALDA
TransFusion -- A Transparency-Based Diffusion Model for Anomaly Detection
arXiv1 repoarXiv:2311.09999
awesome-cybersecurity-agentic-ai
The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
arXiv1 repoarXiv:2311.10057
song-describer-dataset
Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
arXiv1 repoarXiv:2311.10702
proxy-tuning
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
arXiv1 repoarXiv:2311.10774
LRV-Instruction
LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching
arXiv1 repoarXiv:2311.11284
LucidDreamer
Large Pre-trained time series models for cross-domain Time series analysis tasks
arXiv1 repoarXiv:2311.11413
Samay
FinanceBench: A New Benchmark for Financial Question Answering
arXiv1 repoarXiv:2311.11944
PageIndex
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
arXiv1 repoarXiv:2311.12022
RLPR-Evaluation
T-Rex: Counting by Visual Prompting
arXiv1 repoarXiv:2311.13596
Rex-Omni
ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs
arXiv1 repoarXiv:2311.13600
ComfyUI-LoRA-Optimizer
MAIRA-1: A specialised large multimodal model for radiology report generation
arXiv1 repoarXiv:2311.13668
rad-dino
SinSR: Diffusion-Based Image Super-Resolution in a Single Step
arXiv1 repoarXiv:2311.14760
SinSR
Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity Vocoder
arXiv1 repoarXiv:2311.14957
BigVGAN
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
arXiv1 repoarXiv:2311.15127
videophy
GART: Gaussian Articulated Template Models
arXiv1 repoarXiv:2311.16099
GART
Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
arXiv1 repoarXiv:2311.16103
LanguageBind
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
arXiv1 repoarXiv:2311.16484
SeeingEyeToAI
DemoFusion: Democratising High-Resolution Image Generation With No $$$
arXiv1 repoarXiv:2311.16973
DemoFusion
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
arXiv1 repoarXiv:2311.17002
Omost
arXiv:2311.17005
arXiv1 repoarXiv:2311.17005
MVBench
Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features
arXiv1 repoarXiv:2311.17024
Diffusion-3D-Features
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
arXiv1 repoarXiv:2311.17049
ml-mobileclip
DreamPropeller: Supercharge Text-to-3D Generation with Parallel Sampling
arXiv1 repoarXiv:2311.17082
DreamPropeller
SEED-Bench-2: Benchmarking Multimodal Large Language Models
arXiv1 repoarXiv:2311.17092
SEED-Bench
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
arXiv1 repoarXiv:2311.17117
MuseV
TaskWeaver: A Code-First Agent Framework
arXiv1 repoarXiv:2311.17541
TaskWeaver
SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis
arXiv1 repoarXiv:2311.17590
SyncTalk
Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving
arXiv1 repoarXiv:2311.17918
Drive-WM
VBench: Comprehensive Benchmark Suite for Video Generative Models
arXiv1 repoarXiv:2311.17982
VBench
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model
arXiv1 repoarXiv:2311.18248
mPLUG-DocOwl
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
arXiv1 repoarXiv:2311.18259
Ego4d
Splitwise: Efficient generative LLM inference using phase splitting
arXiv1 repoarXiv:2311.18677
Nanoflow
TaskBench: Benchmarking Large Language Models for Task Automation
arXiv1 repoarXiv:2311.18760
JARVIS
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
arXiv1 repoarXiv:2311.18799
LAVIS
Distributed Global Structure-from-Motion with a Deep Front-End
arXiv1 repoarXiv:2311.18801
gtsfm
Exploiting Diffusion Prior for Generalizable Dense Prediction
arXiv1 repoarXiv:2311.18832
dmp
SparseGS: Sparse View Synthesis using 3D Gaussian Splatting
arXiv1 repoarXiv:2312.00206
SparseGS
arXiv:2312.00438
arXiv1 repoarXiv:2312.00438
Dolphins
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
arXiv1 repoarXiv:2312.00849
OmniLMM-12B
DeepCache: Accelerating Diffusion Models for Free
arXiv1 repoarXiv:2312.00858
DeepCache
Segment and Caption Anything
arXiv1 repoarXiv:2312.00869
segment-caption-anything
arXiv:2312.01305
arXiv1 repoarXiv:2312.01305
vivid123
arXiv:2312.01678
arXiv1 repoarXiv:2312.01678
Jellyfish-13B
StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On
arXiv1 repoarXiv:2312.01725
StableVITON
Language-only Efficient Training of Zero-shot Composed Image Retrieval
arXiv1 repoarXiv:2312.01998
lincir
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
arXiv1 repoarXiv:2312.02119
GA
QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers
arXiv1 repoarXiv:2312.02220
QuantAttack
Recursive Visual Programming
arXiv1 repoarXiv:2312.02249
RVP
RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!
arXiv1 repoarXiv:2312.02724
pyterrier_genrank
Customization Assistant for Text-to-image Generation
arXiv1 repoarXiv:2312.03045
ProFusion
LooseControl: Lifting ControlNet for Generalized Depth Conditioning
arXiv1 repoarXiv:2312.03079
LooseControl
DiffusionSat: A Generative Foundation Model for Satellite Imagery
arXiv1 repoarXiv:2312.03606
DiffusionSat
Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers
arXiv1 repoarXiv:2312.03694
PETL_AST
Graph Convolutions Enrich the Self-Attention in Transformers!
arXiv1 repoarXiv:2312.04234
GFSA
Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models
arXiv1 repoarXiv:2312.04410
Smooth-Diffusion
DreamVideo: Composing Your Dream Videos with Customized Subject and Motion
arXiv1 repoarXiv:2312.04433
i2vgen-xl
Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation
arXiv1 repoarXiv:2312.04483
i2vgen-xl
An LLM Compiler for Parallel Function Calling
arXiv1 repoarXiv:2312.04511
generative-ai
Generating Illustrated Instructions
arXiv1 repoarXiv:2312.04552
generating-illustrated-instructions-reproduction
HuRef: HUman-REadable Fingerprint for Large Language Models
arXiv1 repoarXiv:2312.04828
HuRef
Text-to-3D Generation with Bidirectional Diffusion using both 2D and 3D priors
arXiv1 repoarXiv:2312.04963
bidiff
SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation
arXiv1 repoarXiv:2312.05239
SwiftBrush
SlimSAM: 0.1% Data Makes Segment Anything Slim
arXiv1 repoarXiv:2312.05284
SAMReg
Jumpstarting Surgical Computer Vision
arXiv1 repoarXiv:2312.05968
Endoscapes
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
arXiv1 repoarXiv:2312.06109
GOT-OCR2_0
Cataract-1K: Cataract Surgery Dataset for Scene Segmentation, Phase Recognition, and Irregularity Detection
arXiv1 repoarXiv:2312.06295
Cataract-1K
EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion
arXiv1 repoarXiv:2312.06725
EpiDiff
Honeybee: Locality-enhanced Projector for Multimodal LLM
arXiv1 repoarXiv:2312.06742
honeybee
Encoding Surgical Videos as Latent Spatiotemporal Graphs for Object and Anatomy-Driven Reasoning
arXiv1 repoarXiv:2312.06829
Endoscapes
Reducing Energy Bloat in Large Model Training
arXiv1 repoarXiv:2312.06902
zeus
BIRB: A Generalization Benchmark for Information Retrieval in Bioacoustics
arXiv1 repoarXiv:2312.07439
perch
arXiv:2312.07509
arXiv1 repoarXiv:2312.07509
Peekaboo
FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition
arXiv1 repoarXiv:2312.07536
freecontrol
PaperQA: Retrieval-Augmented Generative Agent for Scientific Research
arXiv1 repoarXiv:2312.07559
paper-qa
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
arXiv1 repoarXiv:2312.08168
Chat-Scene
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions
arXiv1 repoarXiv:2312.08578
DCI
RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution
arXiv1 repoarXiv:2312.08617
magicoder
MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning
arXiv1 repoarXiv:2312.08636
nugie-jax-nemotron-3-nano
VideoLCM: Video Latent Consistency Model
arXiv1 repoarXiv:2312.09109
i2vgen-xl
Marathon: A Race Through the Realm of Long Context with Large Language Models
arXiv1 repoarXiv:2312.09542
Marathon
MobileSAMv2: Faster Segment Anything to Everything
arXiv1 repoarXiv:2312.09579
MobileSAM
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
arXiv1 repoarXiv:2312.09608
Faster-Diffusion
Retrieval-Augmented Generation for Large Language Models: A Survey
arXiv1 repoarXiv:2312.10997
TinyRAG
G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model
arXiv1 repoarXiv:2312.11370
R1-V
arXiv:2312.11805
arXiv1 repoarXiv:2312.11805
CoBSAT
The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark
arXiv1 repoarXiv:2312.12429
Endoscapes
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
arXiv1 repoarXiv:2312.13108
WorldGUI
Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting
arXiv1 repoarXiv:2312.13271
repaint123
The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
arXiv1 repoarXiv:2312.13558
AIDocks
TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance
arXiv1 repoarXiv:2312.13632
TraceFL
Typhoon: Thai Large Language Models
arXiv1 repoarXiv:2312.13951
typhoon-7b
DUSt3R: Geometric 3D Vision Made Easy
arXiv1 repoarXiv:2312.14132
svraster
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
arXiv1 repoarXiv:2312.14197
BIPIA
TACO: Topics in Algorithmic COde generation dataset
arXiv1 repoarXiv:2312.14852
TACO
UniHuman: A Unified Model for Editing Human Images in the Wild
arXiv1 repoarXiv:2312.14985
UniHuman
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
arXiv1 repoarXiv:2312.15685
deita-10k-v0-sft
Align on the Fly: Adapting Chatbot Behavior to Established Norms
arXiv1 repoarXiv:2312.15907
OPO
One-Dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications
arXiv1 repoarXiv:2312.16145
SPM
Experiential Co-Learning of Software-Developing Agents
arXiv1 repoarXiv:2312.17025
ChatDev
DreamGaussian4D: Generative 4D Gaussian Splatting
arXiv1 repoarXiv:2312.17142
dreamgaussian4d
On-Demand JSON: A Better Way to Parse Documents?
arXiv1 repoarXiv:2312.17149
simdjson
4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency
arXiv1 repoarXiv:2312.17225
4DGen
Fast Inference of Mixture-of-Experts Language Models with Offloading
arXiv1 repoarXiv:2312.17238
mixtral-offloading
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
arXiv1 repoarXiv:2401.00396
MiniCheck
Fairness in Serving Large Language Models
arXiv1 repoarXiv:2401.00588
S-LoRA
Benchmarking Large Language Models on Controllable Generation under Diversified Instructions
arXiv1 repoarXiv:2401.00690
CoDI-Eval
ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios
arXiv1 repoarXiv:2401.00741
ToolEyes
Improving the Stability and Efficiency of Diffusion Models for Content Consistent Super-Resolution
arXiv1 repoarXiv:2401.00877
OSEDiff
TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
arXiv1 repoarXiv:2401.00896
TrailBlazer
DocLLM: A layout-aware generative language model for multimodal document understanding
arXiv1 repoarXiv:2401.00908
DocLLM
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
arXiv1 repoarXiv:2401.01325
LongLM
Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting
arXiv1 repoarXiv:2401.01339
mystreetgscar
DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection
arXiv1 repoarXiv:2401.02032
DiffusionEdge
An Open and Comprehensive Pipeline for Unified Object Grounding and Detection
arXiv1 repoarXiv:2401.02361
mmdetection
Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss
arXiv1 repoarXiv:2401.02677
Segmind-Vega
German Text Embedding Clustering Benchmark
arXiv1 repoarXiv:2401.02709
mteb-1.34.14
From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models
arXiv1 repoarXiv:2401.02777
BELLE
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
arXiv1 repoarXiv:2401.03945
SpeechGPT
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild
arXiv1 repoarXiv:2401.04210
FunnyNet-W
U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation
arXiv1 repoarXiv:2401.04722
TriALS
A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars
arXiv1 repoarXiv:2401.04730
SLRT
Towards Online Continuous Sign Language Recognition and Translation
arXiv1 repoarXiv:2401.05336
SLRT
Combating Adversarial Attacks with Multi-Agent Debate
arXiv1 repoarXiv:2401.05998
rightmind
Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
arXiv1 repoarXiv:2401.06102
AudioLens
EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
arXiv1 repoarXiv:2401.06201
JARVIS
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
arXiv1 repoarXiv:2401.06373
llm-jailbreaking-defense
TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models
arXiv1 repoarXiv:2401.06620
TransliCo
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
arXiv1 repoarXiv:2401.06643
LLM-div-incts
Don't Rank, Combine! Combining Machine Translation Hypotheses Using Quality Estimation
arXiv1 repoarXiv:2401.06688
qe-fusion
arXiv:2401.08417
arXiv1 repoarXiv:2401.08417
CPO_SIMPO
Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
arXiv1 repoarXiv:2401.08491
acl2025-contrastive-perplexity
Scalable Pre-training of Large Autoregressive Image Models
arXiv1 repoarXiv:2401.08541
ml-aim
Tuning Language Models by Proxy
arXiv1 repoarXiv:2401.08565
proxy-tuning
HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
arXiv1 repoarXiv:2401.08772
HuixiangDou
XTable in Action: Seamless Interoperability in Data Lakes
arXiv1 repoarXiv:2401.09621
data-engineer-handbook
Self-Rewarding Language Models
arXiv1 repoarXiv:2401.10020
fineweb-edu
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
arXiv1 repoarXiv:2401.10208
MM-Interleaved
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
arXiv1 repoarXiv:2401.11240
lightllm
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
arXiv1 repoarXiv:2401.11324
slater
Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
arXiv1 repoarXiv:2401.11708
RPG_models
Benchmarking Large Multimodal Models against Common Corruptions
arXiv1 repoarXiv:2401.11943
MMCBench
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
arXiv1 repoarXiv:2401.11944
Yi-34B-Chat
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
arXiv1 repoarXiv:2401.12179
TuneJury
Universal Neurons in GPT2 Language Models
arXiv1 repoarXiv:2401.12181
gemma-scope-2b-pt-transcoders
Text Embedding Inversion Security for Multilingual Language Models
arXiv1 repoarXiv:2401.12192
MultiVec2Text
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
arXiv1 repoarXiv:2401.12850
SHARC
Lumiere: A Space-Time Diffusion Model for Video Generation
arXiv1 repoarXiv:2401.12945
videophy
GALA: Generating Animatable Layered Assets from a Single Scan
arXiv1 repoarXiv:2401.12979
gala
PatternPortrait: Draw Me Like One of Your Scribbles
arXiv1 repoarXiv:2401.13001
awesome-plotters
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
arXiv1 repoarXiv:2401.13178
agentboard
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
arXiv1 repoarXiv:2401.13527
SpeechGPT
Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild
arXiv1 repoarXiv:2401.13627
SUPIR
Conformal Prediction Sets Improve Human Decision Making
arXiv1 repoarXiv:2401.13744
hitl-conformal-prediction
Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
arXiv1 repoarXiv:2401.14159
GroundingDINO
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
arXiv1 repoarXiv:2401.14196
Reinforcement-Learning-Full-Pipeline
Demystifying Chains, Trees, and Graphs of Thoughts
arXiv1 repoarXiv:2401.14295
rightmind
MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
arXiv1 repoarXiv:2401.14361
MoE-Infinity
Unearthing Large Scale Domain-Specific Knowledge from Public Corpora
arXiv1 repoarXiv:2401.14624
Knowledge_Pile
Spatial Transcriptomics Analysis of Zero-shot Gene Expression Prediction
arXiv1 repoarXiv:2401.14772
SGN
FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models
arXiv1 repoarXiv:2401.15636
FreeStyle
StableIdentity: Inserting Anybody into Anywhere at First Sight
arXiv1 repoarXiv:2401.15975
StableIdentity
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
arXiv1 repoarXiv:2401.16158
MobileAgent
Diffutoon: High-Resolution Editable Toon Shading via Diffusion Models
arXiv1 repoarXiv:2401.16224
diffSynth-studio-notes
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
arXiv1 repoarXiv:2401.17186
CLFM
Proactive Detection of Voice Cloning with Localized Watermarking
arXiv1 repoarXiv:2401.17264
audioseal
YOLO-World: Real-Time Open-Vocabulary Object Detection
arXiv1 repoarXiv:2401.17270
YOLO-World
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
arXiv1 repoarXiv:2401.17377
infini-gram
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
arXiv1 repoarXiv:2401.18079
KVQuant
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
arXiv1 repoarXiv:2402.00794
ReAGent
Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
arXiv1 repoarXiv:2402.00828
PETL_AST
EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks
arXiv1 repoarXiv:2402.00892
vocoder
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
arXiv1 repoarXiv:2402.01109
Vaccine
arXiv:2402.01293
arXiv1 repoarXiv:2402.01293
CoBSAT
KTO: Model Alignment as Prospect Theoretic Optimization
arXiv1 repoarXiv:2402.01306
Reinforcement-Learning-Full-Pipeline
Rethinking Interpretability in the Era of Large Language Models
arXiv1 repoarXiv:2402.01761
imodelsX
When Large Language Models Meet Vector Databases: A Survey
arXiv1 repoarXiv:2402.01763
TinyRAG
S2malloc: Statistically Secure Allocator for Use-After-Free Protection And More
arXiv1 repoarXiv:2402.01894
mimalloc-bench
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
arXiv1 repoarXiv:2402.02037
EffiBench
arXiv:2402.02057
arXiv1 repoarXiv:2402.02057
LookaheadDecoding
BECLR: Batch Enhanced Contrastive Few-Shot Learning
arXiv1 repoarXiv:2402.02444
awesome-cybersecurity-agentic-ai
Verifiable evaluations of machine learning models using zkSNARKs
arXiv1 repoarXiv:2402.02675
zkbc
Position: What Can Large Language Models Tell Us about Time Series Analysis
arXiv1 repoarXiv:2402.02713
time-moe
arXiv:2402.02952
arXiv1 repoarXiv:2402.02952
awesome-ai-sre
Is Mamba Capable of In-Context Learning?
arXiv1 repoarXiv:2402.03170
is_mamba_capable_of_icl
Training-Free Consistent Text-to-Image Generation
arXiv1 repoarXiv:2402.03286
experimental-consistory
arXiv:2402.03310
arXiv1 repoarXiv:2402.03310
VIRL
SpecFormer: Guarding Vision Transformer Robustness via Maximum Singular Value Penalization
arXiv1 repoarXiv:2402.03317
robustlearn
Self-Discover: Large Language Models Self-Compose Reasoning Structures
arXiv1 repoarXiv:2402.03620
awesome-dspy
MolTC: Towards Molecular Relational Modeling In Language Models
arXiv1 repoarXiv:2402.03781
MolTC
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls
arXiv1 repoarXiv:2402.04253
gbrain
LESS: Selecting Influential Data for Targeted Instruction Tuning
arXiv1 repoarXiv:2402.04333
bergson
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry
arXiv1 repoarXiv:2402.04347
LLaDA-Hybrid
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
arXiv1 repoarXiv:2402.04615
transformer-final-proj
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning
arXiv1 repoarXiv:2402.04833
OpenHermes-2.5-1k-longest
Data-efficient Large Vision Models through Sequential Autoregression
arXiv1 repoarXiv:2402.04841
DeLVM
Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding
arXiv1 repoarXiv:2402.05109
Hydra
More Agents Is All You Need
arXiv1 repoarXiv:2402.05120
rightmind
Zero-Shot Clinical Trial Patient Matching with LLMs
arXiv1 repoarXiv:2402.05125
clinical_trial_patient_matching
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
arXiv1 repoarXiv:2402.05602
mirage
Learning to Route Among Specialized Experts for Zero-Shot Generalization
arXiv1 repoarXiv:2402.05859
mttl
Let Your Graph Do the Talking: Encoding Structured Data for LLMs
arXiv1 repoarXiv:2402.05862
GNN4TaskPlan
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
arXiv1 repoarXiv:2402.06599
OOD-Generalization-of-LMMs
Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
arXiv1 repoarXiv:2402.06619
aya_dataset
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning
arXiv1 repoarXiv:2402.06954
FedLLM-Bench
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
arXiv1 repoarXiv:2402.07033
fiddler
Dólares or Dollars? Unraveling the Bilingual Prowess of Financial LLMs Between Spanish and English
arXiv1 repoarXiv:2402.07405
PIXIU
Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT
arXiv1 repoarXiv:2402.07440
gte-multilingual-base
T-RAG: Lessons from the LLM Trenches
arXiv1 repoarXiv:2402.07483
LARS
A Multinomial Canonical Decomposition Model, with emphasis on the analysis of Multivariate Binary data
arXiv1 repoarXiv:2402.07634
smash
Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model
arXiv1 repoarXiv:2402.07827
aya-101
arXiv:2402.07865
arXiv1 repoarXiv:2402.07865
prismatic-vlms
eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data
arXiv1 repoarXiv:2402.08831
eCeLLM
Attacking Large Language Models with Projected Gradient Descent
arXiv1 repoarXiv:2402.09154
reinforce-attacks-llms
Less is More: Fewer Interpretable Region via Submodular Subset Selection
arXiv1 repoarXiv:2402.09164
SMDL-Attribution
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
arXiv1 repoarXiv:2402.09181
OmniMedVQA
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
arXiv1 repoarXiv:2402.09299
TraWiC
LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset
arXiv1 repoarXiv:2402.09391
BioMedGPT-Mol
CodeMind: Evaluating Large Language Models for Code Reasoning
arXiv1 repoarXiv:2402.09664
CodeMind
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
arXiv1 repoarXiv:2402.10076
QUICK
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
arXiv1 repoarXiv:2402.10176
OpenMathInstruct-1
BitDelta: Your Fine-Tune May Only Be Worth One Bit
arXiv1 repoarXiv:2402.10193
BitDelta
Chain-of-Thought Reasoning Without Prompting
arXiv1 repoarXiv:2402.10200
LLM-Sampling
A StrongREJECT for Empty Jailbreaks
arXiv1 repoarXiv:2402.10260
karma-electric-llama31-8b
BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains
arXiv1 repoarXiv:2402.10373
BioMistral-7B
Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
arXiv1 repoarXiv:2402.10517
any-precision-llm
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models
arXiv1 repoarXiv:2402.10744
GenRES
ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages
arXiv1 repoarXiv:2402.10753
Awesome-OpenClaw
Contrastive Instruction Tuning
arXiv1 repoarXiv:2402.11138
CoIN
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
arXiv1 repoarXiv:2402.11208
agent-backdoor-attacks
ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs
arXiv1 repoarXiv:2402.11235
ZeroG
KMMLU: Measuring Massive Multitask Language Understanding in Korean
arXiv1 repoarXiv:2402.11548
KMMLU
Machine-Generated Text Localization
arXiv1 repoarXiv:2402.11744
MGT_Localization
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
arXiv1 repoarXiv:2402.11753
JailbreakLab
An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
arXiv1 repoarXiv:2402.11814
LLM_CTF
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
arXiv1 repoarXiv:2402.12201
circuit_backup
A Critical Evaluation of AI Feedback for Aligning Large Language Models
arXiv1 repoarXiv:2402.12366
dpo-rlaif
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
arXiv1 repoarXiv:2402.12374
mlx-flash
Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems
arXiv1 repoarXiv:2402.12784
vec2text-dense_retriever-threat
Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data
arXiv1 repoarXiv:2402.12869
opendataloader-bench
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
arXiv1 repoarXiv:2402.12908
RealCompo
Benchmarking Retrieval-Augmented Generation for Medicine
arXiv1 repoarXiv:2402.13178
textbooks
Transformer tricks: Precomputing the first layer
arXiv1 repoarXiv:2402.13388
transformer-tricks
SDXL-Lightning: Progressive Adversarial Diffusion Distillation
arXiv1 repoarXiv:2402.13929
SDXL-Lightning
arXiv:2402.14017
arXiv1 repoarXiv:2402.14017
RectifID
BIRCO: A Benchmark of Information Retrieval Tasks with Complex Objectives
arXiv1 repoarXiv:2402.14151
BIRCO
T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching
arXiv1 repoarXiv:2402.14167
T-Stitch
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
arXiv1 repoarXiv:2402.14334
InstructIR
Orca-Math: Unlocking the potential of SLMs in Grade School Math
arXiv1 repoarXiv:2402.14830
orca-math-word-problems-193k-korean
arXiv:2402.14891
arXiv1 repoarXiv:2402.14891
LLMBind
Watermarking Makes Language Models Radioactive
arXiv1 repoarXiv:2402.14904
watermarks-remover
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
arXiv1 repoarXiv:2402.14905
minimind
tinyBenchmarks: evaluating LLMs with fewer examples
arXiv1 repoarXiv:2402.14992
SWEBench-verified-mini
The AffectToolbox: Affect Analysis for Everyone
arXiv1 repoarXiv:2402.15195
AffectToolbox
Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency
arXiv1 repoarXiv:2402.15481
Bias-Volatility-Framework
AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System
arXiv1 repoarXiv:2402.15538
AgentLite
Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
arXiv1 repoarXiv:2402.16192
llm-jailbreaking-defense
IR2: Information Regularization for Information Retrieval
arXiv1 repoarXiv:2402.16200
Information-Regularization
CodeS: Towards Building Open-source Language Models for Text-to-SQL
arXiv1 repoarXiv:2402.16347
text2sql-demo
TOTEM: TOkenized Time Series EMbeddings for General Time Series Analysis
arXiv1 repoarXiv:2402.16412
moment
Defending LLMs against Jailbreaking Attacks via Backtranslation
arXiv1 repoarXiv:2402.16459
llm-jailbreaking-defense
A Survey on Data Selection for Language Models
arXiv1 repoarXiv:2402.16827
llm_project
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
arXiv1 repoarXiv:2402.16906
LLMDebugger
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
arXiv1 repoarXiv:2402.16929
LangGPT
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
arXiv1 repoarXiv:2402.17231
MathSensei
VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis
arXiv1 repoarXiv:2402.17300
VoCo_Downstream
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
arXiv1 repoarXiv:2402.17553
omniact
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions
arXiv1 repoarXiv:2402.17633
ytseg
arXiv:2402.17700
arXiv1 repoarXiv:2402.17700
SAEBench
Case-Based or Rule-Based: How Do Transformers Do the Math?
arXiv1 repoarXiv:2402.17709
Case_or_Rule
Tower: An Open Multilingual Large Language Model for Translation-Related Tasks
arXiv1 repoarXiv:2402.17733
TowerInstruct-7B-v0.1
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
arXiv1 repoarXiv:2402.17764
BitNet
A Language Model based Framework for New Concept Placement in Ontologies
arXiv1 repoarXiv:2402.17897
LM-ontology-concept-placement
arXiv:2402.18158
arXiv1 repoarXiv:2402.18158
qllm-eval
Bluebell: An Alliance of Relational Lifting and Independence For Probabilistic Reasoning
arXiv1 repoarXiv:2402.18708
ArkLib
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers
arXiv1 repoarXiv:2402.19479
Panda-70M
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
arXiv1 repoarXiv:2403.00231
VisRAG-Ret-Train-In-domain-data
AtP*: An efficient and scalable method for localizing LLM behaviour to components
arXiv1 repoarXiv:2403.00745
notebooks
IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact
arXiv1 repoarXiv:2403.01241
KVQuant
Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal
arXiv1 repoarXiv:2403.01244
SSR
KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations
arXiv1 repoarXiv:2403.01469
KorMedMCQA
Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral
arXiv1 repoarXiv:2403.01851
Chinese-LLaMA-Alpaca-3
An Improved Traditional Chinese Evaluation Suite for Foundation Model
arXiv1 repoarXiv:2403.01858
tmmluplus
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
arXiv1 repoarXiv:2403.02691
moltshield
KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents
arXiv1 repoarXiv:2403.03101
KnowLM
CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation
arXiv1 repoarXiv:2403.03411
CrossNet
NoiseCollage: A Layout-Aware Text-to-Image Diffusion Model Based on Noise Cropping and Merging
arXiv1 repoarXiv:2403.03485
noisecollage
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
arXiv1 repoarXiv:2403.03507
Finetune_with_GaLore
MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models
arXiv1 repoarXiv:2403.03744
med-safety-bench
Learning to Decode Collaboratively with Multiple Language Models
arXiv1 repoarXiv:2403.03870
co-llm
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
arXiv1 repoarXiv:2403.03894
acl2024-ircoder
Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
arXiv1 repoarXiv:2403.03952
AMAZON-Products-2023
Symmetry Considerations for Learning Task Symmetric Robot Policies
arXiv1 repoarXiv:2403.04359
gauss_gym
Learning to Remove Wrinkled Transparent Film with Polarized Prior
arXiv1 repoarXiv:2403.04368
FilmRemoval
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
arXiv1 repoarXiv:2403.04692
PixArt-Sigma-XL-2-512-MS
Face2Diffusion for Fast and Editable Face Personalization
arXiv1 repoarXiv:2403.05094
Face2Diffusion
LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation
arXiv1 repoarXiv:2403.05246
TriALS
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
arXiv1 repoarXiv:2403.06098
VidProM
MACE: Mass Concept Erasure in Diffusion Models
arXiv1 repoarXiv:2403.06135
MACE
No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and Benchmarks
arXiv1 repoarXiv:2403.06249
PIXIU
Editing Conceptual Knowledge for Large Language Models
arXiv1 repoarXiv:2403.06259
ConceptEdit
CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean
arXiv1 repoarXiv:2403.06412
Phi-3.5-MoE-instruct
SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
arXiv1 repoarXiv:2403.06534
sardet_100k
Leveraging Foundation Models for Content-Based Image Retrieval in Radiology
arXiv1 repoarXiv:2403.06567
foundation-models-for-cbmir
Real-Time Multimodal Cognitive Assistant for Emergency Medical Services
arXiv1 repoarXiv:2403.06734
EMS-Pipeline
Accurate Spatial Gene Expression Prediction by integrating Multi-resolution features
arXiv1 repoarXiv:2403.07592
TRIPLEX
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
arXiv1 repoarXiv:2403.07974
code_generation_lite
A bargain for mergesorts -- How to prove your mergesort correct and stable, almost for free
arXiv1 repoarXiv:2403.08173
stablesort
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
arXiv1 repoarXiv:2403.08495
MING
Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation
arXiv1 repoarXiv:2403.08605
SmallPlan
Faster Projected GAN: Towards Faster Few-Shot Image Generation
arXiv1 repoarXiv:2403.08778
Faster-Projected-GAN
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
arXiv1 repoarXiv:2403.09029
WebSight
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences
arXiv1 repoarXiv:2403.09032
magicoder
RAGGED: Towards Informed Design of Scalable and Stable RAG Systems
arXiv1 repoarXiv:2403.09040
ragged
SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models
arXiv1 repoarXiv:2403.09055
semantic-draw
Dial-insight: Fine-tuning Large Language Models with High-Quality Domain-Specific Data Preventing Capability Collapse
arXiv1 repoarXiv:2403.09167
BELLE
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
arXiv1 repoarXiv:2403.09472
easy-to-hard
Hyper-CL: Conditioning Sentence Representations with Hypernetworks
arXiv1 repoarXiv:2403.09490
Hyper-CL
Repoformer: Selective Retrieval for Repository-Level Code Completion
arXiv1 repoarXiv:2403.10059
Repoformer
Improving Medical Multi-modal Contrastive Learning with Expert Annotations
arXiv1 repoarXiv:2403.10153
eCLIP
Block Verification Accelerates Speculative Decoding
arXiv1 repoarXiv:2403.10444
Hierarchical-Speculative-Decoding
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
arXiv1 repoarXiv:2403.10517
VideoTree
S3LLM: Large-Scale Scientific Software Understanding with LLMs using Source, Metadata, and Document
arXiv1 repoarXiv:2403.10588
S3LLM
LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
arXiv1 repoarXiv:2403.11703
OmniLMM-12B
SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion
arXiv1 repoarXiv:2403.12008
DiffSplat
One-Step Image Translation with Text-to-Image Models
arXiv1 repoarXiv:2403.12036
img2img-turbo
Distilling Datasets Into Less Than One Image
arXiv1 repoarXiv:2403.12040
PoDD
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
arXiv1 repoarXiv:2403.12171
EasyJailbreak
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
arXiv1 repoarXiv:2403.12895
mPLUG-DocOwl
LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression
arXiv1 repoarXiv:2403.12968
llmlingua-2-bert-base-multilingual-cased-meetingbank
When Do We Not Need Larger Vision Models?
arXiv1 repoarXiv:2403.13043
scaling_on_scales
AdaptSFL: Adaptive Split Federated Learning in Resource-constrained Edge Networks
arXiv1 repoarXiv:2403.13101
SplitFM
Mora: Enabling Generalist Video Generation via A Multi-Agent Framework
arXiv1 repoarXiv:2403.13248
Mora
Arcee's MergeKit: A Toolkit for Merging Large Language Models
arXiv1 repoarXiv:2403.13257
donutloop-genesis
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding
arXiv1 repoarXiv:2403.14174
UniSDNet
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
arXiv1 repoarXiv:2403.14370
FlexiSyncMVD
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
arXiv1 repoarXiv:2403.14468
AnyV2V
Implicit Style-Content Separation using B-LoRA
arXiv1 repoarXiv:2403.14572
B-LoRA
Foundation Models for Time Series Analysis: A Tutorial and Survey
arXiv1 repoarXiv:2403.14735
time-moe
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
arXiv1 repoarXiv:2403.14743
VURF
KeyPoint Relative Position Encoding for Face Recognition
arXiv1 repoarXiv:2403.14852
AdaFace
Long-CLIP: Unlocking the Long-Text Capability of CLIP
arXiv1 repoarXiv:2403.15378
Long-CLIP
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
arXiv1 repoarXiv:2403.15388
LLaVA-PruMerge
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
arXiv1 repoarXiv:2403.16167
ESREAL
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
arXiv1 repoarXiv:2403.17806
KnowledgeCircuits
Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
arXiv1 repoarXiv:2403.17834
CT-CLIP
2D Gaussian Splatting for Geometrically Accurate Radiance Fields
arXiv1 repoarXiv:2403.17888
svraster
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
arXiv1 repoarXiv:2403.17919
LMFlow
arXiv:2403.18118
arXiv1 repoarXiv:2403.18118
egolifter
Annolid: Annotate, Segment, and Track Anything You Need
arXiv1 repoarXiv:2403.18690
annolid
TextCraftor: Your Text Encoder Can be Image Quality Controller
arXiv1 repoarXiv:2403.18978
textcraftor
LITA: Language Instructed Temporal-Localization Assistant
arXiv1 repoarXiv:2403.19046
LITA
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
arXiv1 repoarXiv:2403.19114
evoeval
sDPO: Don't Use Your Data All at Once
arXiv1 repoarXiv:2403.19270
SOLAR-10.7B-Instruct-v1.0
OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion
arXiv1 repoarXiv:2403.19417
OakInk-v2
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
arXiv1 repoarXiv:2403.19521
Factual-Recall-Mechanism
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
arXiv1 repoarXiv:2403.19647
notebooks
Jamba: A Hybrid Transformer-Mamba Language Model
arXiv1 repoarXiv:2403.19887
Jamba-v0.1
TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods
arXiv1 repoarXiv:2403.20150
TAB
Monocular Identity-Conditioned Facial Reflectance Reconstruction
arXiv1 repoarXiv:2404.00301
insightface
Towards Variable and Coordinated Holistic Co-Speech Motion Generation
arXiv1 repoarXiv:2404.00368
probtalk
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
arXiv1 repoarXiv:2404.00456
deepcompressor
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories
arXiv1 repoarXiv:2404.00599
EvoCodeBench
WavLLM: Towards Robust and Adaptive Speech Large Language Model
arXiv1 repoarXiv:2404.00656
SpeechT5
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
arXiv1 repoarXiv:2404.01318
JBB-Behaviors
arXiv:2404.01363
arXiv1 repoarXiv:2404.01363
awesome-ai-sre
Are large language models superhuman chemists?
arXiv1 repoarXiv:2404.01475
chembench
M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets
arXiv1 repoarXiv:2404.01753
M2SA-multimodal-multilingual-sentiment-analysis
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
arXiv1 repoarXiv:2404.02101
CameraCtrl
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
arXiv1 repoarXiv:2404.02151
jailbreakbench
LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIP
arXiv1 repoarXiv:2404.02285
BiomedCoOp
On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons
arXiv1 repoarXiv:2404.02431
lang_neuron
Sound Borrow-Checking for Rust via Symbolic Semantics (Long Version)
arXiv1 repoarXiv:2404.02680
aeneas
InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation
arXiv1 repoarXiv:2404.02733
InstantStyle
Faster Diffusion via Temporal Attention Decomposition
arXiv1 repoarXiv:2404.02747
T-GATE
arXiv:2404.02949
arXiv1 repoarXiv:2404.02949
feud
Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures
arXiv1 repoarXiv:2404.03010
TotalSegmentator
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
arXiv1 repoarXiv:2404.03413
MiniGPT4-Video
DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling
arXiv1 repoarXiv:2404.03575
DreamScene
arXiv:2404.04057
arXiv1 repoarXiv:2404.04057
ml-sid-dit
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
arXiv1 repoarXiv:2404.04125
frequency_determines_performance
Robust Depth Enhancement via Polarization Prompt Fusion Tuning
arXiv1 repoarXiv:2404.04318
Polarization-Prompt-Fusion-Tuning
Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs
arXiv1 repoarXiv:2404.04363
Idea23D
arXiv:2404.04475
arXiv1 repoarXiv:2404.04475
alpaca_eval
MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems
arXiv1 repoarXiv:2404.04735
math-solving
AI2Apps: A Visual IDE for Building LLM-based AI Agent Applications
arXiv1 repoarXiv:2404.04902
ai2apps
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts
arXiv1 repoarXiv:2404.05019
colibri
DinoBloom: A Foundation Model for Generalizable Cell Embeddings in Hematology
arXiv1 repoarXiv:2404.05022
DinoBloom
MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering
arXiv1 repoarXiv:2404.05590
PodGPT
VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain
arXiv1 repoarXiv:2404.05659
MultiMed
SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
arXiv1 repoarXiv:2404.05717
swap-anything
Learning Embeddings with Centroid Triplet Loss for Object Identification in Robotic Grasping
arXiv1 repoarXiv:2404.06277
image_agnostic_segmentation
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models
arXiv1 repoarXiv:2404.06309
ClipClap-GZSL
MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
arXiv1 repoarXiv:2404.06395
MiniCPM3-4B
Automated Federated Pipeline for Parameter-Efficient Fine-Tuning of Large Language Models
arXiv1 repoarXiv:2404.06448
SplitFM
Autonomous Evaluation and Refinement of Digital Agents
arXiv1 repoarXiv:2404.06474
Agent-Eval-Refine
Visually Descriptive Language Model for Vector Graphics Reasoning
arXiv1 repoarXiv:2404.06479
vtracer
Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?
arXiv1 repoarXiv:2404.06644
khayyam-challenge
arXiv:2404.06798
arXiv1 repoarXiv:2404.06798
uMedGround
GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM Applications
arXiv1 repoarXiv:2404.06921
gorilla
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
arXiv1 repoarXiv:2404.07143
InfiniTransformer
InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models
arXiv1 repoarXiv:2404.07191
InstantMesh
RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion
arXiv1 repoarXiv:2404.07199
realmdreamer
Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain
arXiv1 repoarXiv:2404.07613
Multilingual-Medical-Corpus
rollama: An R package for using generative large language models through Ollama
arXiv1 repoarXiv:2404.07654
rollama
High-Dimension Human Value Representation in Large Language Models
arXiv1 repoarXiv:2404.07900
UniVaR
Taming Stable Diffusion for Text to 360° Panorama Image Generation
arXiv1 repoarXiv:2404.07949
PanFusion
Manipulating Large Language Models to Increase Product Visibility
arXiv1 repoarXiv:2404.07981
llm-rank-optimizer
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models
arXiv1 repoarXiv:2404.07990
OpenBias
Revisiting Feature Prediction for Learning Visual Representations from Video
arXiv1 repoarXiv:2404.08471
xwm
Probing the 3D Awareness of Visual Foundation Models
arXiv1 repoarXiv:2404.08636
probe3d
COCONut: Modernizing COCO Segmentation
arXiv1 repoarXiv:2404.08639
coconut_cvpr2024
MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts
arXiv1 repoarXiv:2404.09027
MING
arXiv:2404.09837
arXiv1 repoarXiv:2404.09837
awesome-ai-sre
Demonstration of DB-GPT: Next Generation Data Interaction System Empowered by Large Language Models
arXiv1 repoarXiv:2404.10209
DB-GPT
MobileNetV4 -- Universal Models for the Mobile Ecosystem
arXiv1 repoarXiv:2404.10518
MaaAI
ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images
arXiv1 repoarXiv:2404.10652
ViTextVQA
Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
arXiv1 repoarXiv:2404.10719
ReaLHF
Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes
arXiv1 repoarXiv:2404.10772
mip-splatting
TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image Editing
arXiv1 repoarXiv:2404.11120
TiNO-Edit
LongEmbed: Extending Embedding Models for Long Context Retrieval
arXiv1 repoarXiv:2404.12096
mteb-1.34.14
Advancing the Robustness of Large Language Models through Self-Denoised Smoothing
arXiv1 repoarXiv:2404.12274
SelfDenoise
KV-weights are all you need for skipless transformers
arXiv1 repoarXiv:2404.12362
transformer-tricks
Lean Copilot: Large Language Models as Copilots for Theorem Proving in Lean
arXiv1 repoarXiv:2404.12534
LeanCopilot
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
arXiv1 repoarXiv:2404.13033
LLM-Tuning
SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs
arXiv1 repoarXiv:2404.13081
ICLR24_SuRe
Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis
arXiv1 repoarXiv:2404.13686
Hyper-SD
Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D Glimpses
arXiv1 repoarXiv:2404.14410
gtu
SnapKV: LLM Knows What You are Looking for Before Generation
arXiv1 repoarXiv:2404.14469
GUI-KV
Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous Decoding
arXiv1 repoarXiv:2404.14600
lost-in-decoding
FlashSpeech: Efficient Zero-Shot Speech Synthesis
arXiv1 repoarXiv:2404.14700
FlashSpeech
Setting up the Data Printer with Improved English to Ukrainian Machine Translation
arXiv1 repoarXiv:2404.15196
dragoman
CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies
arXiv1 repoarXiv:2404.15238
modular_pluralism
TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting
arXiv1 repoarXiv:2404.15264
TalkingGaussian
From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation
arXiv1 repoarXiv:2404.15267
DeepFashion-MultiModal-Parts2Whole
Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation
arXiv1 repoarXiv:2404.15506
Depth-Estimation
Semantic Routing for Enhanced Performance of LLM-Assisted Intent-Based 5G Core Network Management and Orchestration
arXiv1 repoarXiv:2404.15869
semantic-router
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
arXiv1 repoarXiv:2404.16006
MMT-Bench
GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting
arXiv1 repoarXiv:2404.16012
GaussianTalker
arXiv:2404.16014
arXiv1 repoarXiv:2404.16014
dictionary_learning
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
arXiv1 repoarXiv:2404.16019
VoteSim
MoDE: CLIP Data Experts via Clustering
arXiv1 repoarXiv:2404.16030
MetaCLIP
Validating Traces of Distributed Programs Against TLA+ Specifications
arXiv1 repoarXiv:2404.16075
formal-web
Leveraging tropical reef, bird and unrelated sounds for superior transfer learning in marine bioacoustics
arXiv1 repoarXiv:2404.16436
perch
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
arXiv1 repoarXiv:2404.16635
mPLUG-DocOwl
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
arXiv1 repoarXiv:2404.16710
mlx-flash
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
arXiv1 repoarXiv:2404.16790
SEED-Bench
When to Trust LLMs: Aligning Confidence with Response Quality
arXiv1 repoarXiv:2404.17287
CONQORD
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
arXiv1 repoarXiv:2404.17672
BlenderAlchemyOfficial
Diffusion-Aided Joint Source Channel Coding For High Realism Wireless Image Transmission
arXiv1 repoarXiv:2404.17736
DiffJSCC
KAN: Kolmogorov-Arnold Networks
arXiv1 repoarXiv:2404.19756
imodelsX
Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge
arXiv1 repoarXiv:2405.00263
clover
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
arXiv1 repoarXiv:2405.00823
WorkBench
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
arXiv1 repoarXiv:2405.01008
LocoGen
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
arXiv1 repoarXiv:2405.01413
MiniGPT-3D
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
arXiv1 repoarXiv:2405.01434
StoryDiffusion
Automatically Extracting Numerical Results from Randomized Controlled Trials with Large Language Models
arXiv1 repoarXiv:2405.01686
llm-meta-analysis
Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection
arXiv1 repoarXiv:2405.02318
NL2FOL
Labeling supervised fine-tuning data with the scaling law
arXiv1 repoarXiv:2405.02817
HuixiangDou
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
arXiv1 repoarXiv:2405.03152
WenetSpeech-Chuan
Bridging discrete and continuous state spaces: Exploring the Ehrenfest process in time-continuous diffusion models
arXiv1 repoarXiv:2405.03549
EhrenfestDiffusion
A Construct-Optimize Approach to Sparse View Synthesis without Camera Pose
arXiv1 repoarXiv:2405.03659
COGS
sqlelf: a SQL-centric Approach to ELF Analysis
arXiv1 repoarXiv:2405.03883
selfdb
Iterative Experience Refinement of Software-Developing Agents
arXiv1 repoarXiv:2405.04219
ChatDev
xLSTM: Extended Long Short-Term Memory
arXiv1 repoarXiv:2405.04517
xlstm
Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking
arXiv1 repoarXiv:2405.04685
turkish-llm
You Only Cache Once: Decoder-Decoder Architectures for Language Models
arXiv1 repoarXiv:2405.05254
LCKV
Mirage: A Multi-Level Superoptimizer for Tensor Programs
arXiv1 repoarXiv:2405.05751
mirage
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
arXiv1 repoarXiv:2405.05904
ChroKnowledge
An Investigation of Incorporating Mamba for Speech Enhancement
arXiv1 repoarXiv:2405.06573
SEMamba
LLM-Generated Black-box Explanations Can Be Adversarially Helpful
arXiv1 repoarXiv:2405.06800
adversarial_helpfulness
PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator
arXiv1 repoarXiv:2405.07510
Rectified-Diffusion
MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
arXiv1 repoarXiv:2405.07551
aimo-progress-prize
Characterizing virulence differences in a parasitoid wasp through comparative transcriptomic and proteomic
arXiv1 repoarXiv:2405.07772
PusaV1
Forecasting with Hyper-Trees
arXiv1 repoarXiv:2405.07836
Hyper-Trees
UnMarker: A Universal Attack on Defensive Image Watermarking
arXiv1 repoarXiv:2405.08363
watermarks-remover
Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
arXiv1 repoarXiv:2405.08890
PDL
MTP: A Meaning-Typed Language Abstraction for AI-Integrated Programming
arXiv1 repoarXiv:2405.08965
jac
A safety realignment framework via subspace-oriented model fusion for large language models
arXiv1 repoarXiv:2405.09055
safety_realignment
Chameleon: Mixed-Modal Early-Fusion Foundation Models
arXiv1 repoarXiv:2405.09818
YoChameleon
DocuMint: Docstring Generation for Python using Small Language Models
arXiv1 repoarXiv:2405.10243
DocuMint
PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology
arXiv1 repoarXiv:2405.10254
TITAN
One registration is worth two segmentations
arXiv1 repoarXiv:2405.10879
SAMReg
Adhesion of a nematic elastomer cylinder
arXiv1 repoarXiv:2405.11116
compagent
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
arXiv1 repoarXiv:2405.11143
selfplay-redteaming
Towards Modular LLMs by Building and Reusing a Library of LoRAs
arXiv1 repoarXiv:2405.11157
mttl
DocReLM: Mastering Document Retrieval with Language Model
arXiv1 repoarXiv:2405.11461
sci-bert-finetune
Training Data Attribution via Approximate Unrolled Differentiation
arXiv1 repoarXiv:2405.12186
bergson
Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices
arXiv1 repoarXiv:2405.12211
Slicedit
Images that Sound: Composing Images and Sounds on a Single Canvas
arXiv1 repoarXiv:2405.12221
images-that-sound
Tagengo: A Multilingual Chat Dataset
arXiv1 repoarXiv:2405.12612
suzume-llama-3-8B-japanese
Pytorch-Wildlife: A Collaborative Deep Learning Framework for Conservation
arXiv1 repoarXiv:2405.12930
Depth-Estimation
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
arXiv1 repoarXiv:2405.12981
LCKV
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
arXiv1 repoarXiv:2405.13576
FlashRAG_datasets
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
arXiv1 repoarXiv:2405.13622
auto-rag-eval
Irreducibility in generalized power series
arXiv1 repoarXiv:2405.13815
conway-refinement
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
arXiv1 repoarXiv:2405.13868
circuit_backup
Focus Anywhere for Fine-grained Multi-page Document Understanding
arXiv1 repoarXiv:2405.14295
GOT-OCR2_0
PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference
arXiv1 repoarXiv:2405.14430
xDiT
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
arXiv1 repoarXiv:2405.14573
UI-TARS
exLong: Generating Exceptional Behavior Tests with Large Language Models
arXiv1 repoarXiv:2405.14619
exLong
arXiv:2405.14677
arXiv1 repoarXiv:2405.14677
RectifID
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
arXiv1 repoarXiv:2405.14831
hippo-memory
Extracting Prompts by Inverting LLM Outputs
arXiv1 repoarXiv:2405.15012
output2prompt
Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization
arXiv1 repoarXiv:2405.15071
GrokkedTransformer
DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
arXiv1 repoarXiv:2405.15232
DEEM
arXiv:2405.15593
arXiv1 repoarXiv:2405.15593
MicroAdam
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
arXiv1 repoarXiv:2405.15793
mini-swe-agent
Intruding with Words: Towards Understanding Graph Injection Attacks at the Text Level
arXiv1 repoarXiv:2405.16405
Text-level-Graph-Attack
SpinQuant: LLM quantization with learned rotations
arXiv1 repoarXiv:2405.16406
turboquant-vllm
arXiv:2405.16444
arXiv1 repoarXiv:2405.16444
CacheBlend
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
arXiv1 repoarXiv:2405.16473
M3CoT
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
arXiv1 repoarXiv:2405.16504
UnifiedImplicitAttnRepr
Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents
arXiv1 repoarXiv:2405.16533
AutoTools
Crafting Interpretable Embeddings by Asking LLMs Questions
arXiv1 repoarXiv:2405.16714
imodelsX
Saturn: Sample-efficient Generative Molecular Design using Memory Manipulation
arXiv1 repoarXiv:2405.17066
sego
ReMoDetect: Reward Models Recognize Aligned LLM's Generations
arXiv1 repoarXiv:2405.17382
ReMoDetect-deberta
Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment
arXiv1 repoarXiv:2405.17888
Reward_learning_SFT
Knowledge Circuits in Pretrained Transformers
arXiv1 repoarXiv:2405.17969
KnowledgeCircuits
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
arXiv1 repoarXiv:2405.18392
modded-nanogpt
Learning diverse attacks on large language models for robust red-teaming and safety tuning
arXiv1 repoarXiv:2405.18540
red-teaming
Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
arXiv1 repoarXiv:2405.18641
Lisa
NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild
arXiv1 repoarXiv:2405.18715
nerfonthego-undistorted
SketchDeco: Training-Free Latent Composition for Precise Sketch Colourisation
arXiv1 repoarXiv:2405.18716
sketchdeco-code
T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback
arXiv1 repoarXiv:2405.18750
t2v-turbo
Can Graph Learning Improve Planning in LLM-based Agents?
arXiv1 repoarXiv:2405.19119
GNN4TaskPlan
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
arXiv1 repoarXiv:2405.19209
VideoTree
ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning
arXiv1 repoarXiv:2405.19237
BackdoorDM
TotalSegmentator MRI: Robust Sequence-independent Segmentation of Multiple Anatomic Structures in MRI
arXiv1 repoarXiv:2405.19492
TotalSegmentator
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
arXiv1 repoarXiv:2405.19544
CAN
EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos
arXiv1 repoarXiv:2405.19644
EgoSurgery
Grokfast: Accelerated Grokking by Amplifying Slow Gradients
arXiv1 repoarXiv:2405.20233
grokfast
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
arXiv1 repoarXiv:2405.20279
CV-VAE
Improving the Training of Rectified Flows
arXiv1 repoarXiv:2405.20320
Rectified-Diffusion
From Zero to Hero: Cold-Start Anomaly Detection
arXiv1 repoarXiv:2405.20341
ColdFusion
Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image
arXiv1 repoarXiv:2405.20343
Unique3D
Unraveling and Mitigating Retriever Inconsistencies in Retrieval-Augmented Large Language Models
arXiv1 repoarXiv:2405.20680
Ensemble-of-Retrievers
Grammar-Aligned Decoding
arXiv1 repoarXiv:2405.21047
transformers-GAD
Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
arXiv1 repoarXiv:2405.21069
LPCNet
Learning Manipulation by Predicting Interaction
arXiv1 repoarXiv:2406.00439
MPI
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources with Retrieval and Semantic Parsing
arXiv1 repoarXiv:2406.00562
wikipedia
Invisible Backdoor Attacks on Diffusion Models
arXiv1 repoarXiv:2406.00816
BackdoorDM
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation
arXiv1 repoarXiv:2406.01188
UniAnimate-DiT
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
arXiv1 repoarXiv:2406.01326
TabPedia_v1.0
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards
arXiv1 repoarXiv:2406.01364
BELLS
arXiv:2406.01561
arXiv1 repoarXiv:2406.01561
ml-sid-dit
Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA
arXiv1 repoarXiv:2406.02044
QROA
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
arXiv1 repoarXiv:2406.02069
GUI-KV
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
arXiv1 repoarXiv:2406.02265
RobustCap
GrootVL: Tree Topology is All You Need in State Space Model
arXiv1 repoarXiv:2406.02395
MindOmni
The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding
arXiv1 repoarXiv:2406.02396
mteb-1.34.14
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
arXiv1 repoarXiv:2406.02500
Unified-MoE-Compression
Guiding a Diffusion Model with a Bad Version of Itself
arXiv1 repoarXiv:2406.02507
edm2
Loki: Low-rank Keys for Efficient Sparse Attention
arXiv1 repoarXiv:2406.02542
loki
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition
arXiv1 repoarXiv:2406.02554
glimmer
Collision-Affording Point Trees: SIMD-Amenable Nearest Neighbors for Fast Collision Checking
arXiv1 repoarXiv:2406.02807
capt
Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms
arXiv1 repoarXiv:2406.02832
mbrs
DenoDet: Attention as Deformable Multi-Subspace Feature Denoising for Target Detection in SAR Images
arXiv1 repoarXiv:2406.02833
sardet_100k
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
arXiv1 repoarXiv:2406.03095
EgoSurgery
Text-to-Image Rectified Flow as Plug-and-Play Priors
arXiv1 repoarXiv:2406.03293
InstaFlow
LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection
arXiv1 repoarXiv:2406.03459
rf-detr
BLSP-Emo: Towards Empathetic Large Speech-Language Models
arXiv1 repoarXiv:2406.03872
Speech-IFEval
CDMamba: Incorporating Local Clues into Mamba for Remote Sensing Image Binary Change Detection
arXiv1 repoarXiv:2406.04207
Land-Change-Detection
ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models
arXiv1 repoarXiv:2406.04214
ValueBench
ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization
arXiv1 repoarXiv:2406.04312
ReNO
LawGPT: A Chinese Legal Knowledge-Enhanced Large Language Model
arXiv1 repoarXiv:2406.04614
LaWGPT
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
arXiv1 repoarXiv:2406.04770
WildBench
MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks
arXiv1 repoarXiv:2406.04801
Monkey
FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models
arXiv1 repoarXiv:2406.04845
FedLLM-Bench
Hibou: A Family of Foundational Vision Transformers for Pathology
arXiv1 repoarXiv:2406.05074
dpfm_factory
DALD: Improving Logits-based Detector without Logits from Black-box LLMs
arXiv1 repoarXiv:2406.05232
DALD
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
arXiv1 repoarXiv:2406.05478
ImprovedNAT
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
arXiv1 repoarXiv:2406.05551
dots.tts
PaRa: Personalizing Text-to-Image Diffusion via Parameter Rank Reduction
arXiv1 repoarXiv:2406.05641
DEFT
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
arXiv1 repoarXiv:2406.05763
F5-TTS
Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents
arXiv1 repoarXiv:2406.05870
jamming_attack
Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
arXiv1 repoarXiv:2406.05955
PowerInfer
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
arXiv1 repoarXiv:2406.05967
cvqa
EpiLearn: A Python Library for Machine Learning in Epidemic Modeling
arXiv1 repoarXiv:2406.06016
EpiLearn
PowerInfer-2: Fast Large Language Model Inference on a Smartphone
arXiv1 repoarXiv:2406.06282
PowerInfer
UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor
arXiv1 repoarXiv:2406.06519
reranker-as-judge
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
arXiv1 repoarXiv:2406.06525
LlamaGen
Achieving Sparse Activation in Small Language Models
arXiv1 repoarXiv:2406.06562
Sparse-Activation
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
arXiv1 repoarXiv:2406.06571
subllm
AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
arXiv1 repoarXiv:2406.06911
DeepCache
MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
arXiv1 repoarXiv:2406.07057
MMTrustEval
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
arXiv1 repoarXiv:2406.07096
SLU_pipeline
NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images
arXiv1 repoarXiv:2406.07111
NeRSP
Scaling Large Language Model-based Multi-Agent Collaboration
arXiv1 repoarXiv:2406.07155
ChatDev
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
arXiv1 repoarXiv:2406.07162
EmoBox
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
arXiv1 repoarXiv:2406.07368
Linearized-LLM
MINERS: Multilingual Language Models as Semantic Retrievers
arXiv1 repoarXiv:2406.07424
miners
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
arXiv1 repoarXiv:2406.07522
Samba
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
arXiv1 repoarXiv:2406.07803
FastSpeech2-Plus
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
arXiv1 repoarXiv:2406.07969
libritts-r-filtered-speaker-descriptions
One-Step Effective Diffusion Network for Real-World Image Super-Resolution
arXiv1 repoarXiv:2406.08177
OSEDiff
VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
arXiv1 repoarXiv:2406.08394
VisionLLM
Real3D: Scaling Up Large Reconstruction Models with Real-World Images
arXiv1 repoarXiv:2406.08479
Real3D
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
arXiv1 repoarXiv:2406.08568
TTDS
HelpSteer2: Open-source dataset for training top-performing reward models
arXiv1 repoarXiv:2406.08673
Llama-3_3-Nemotron-Super-49B-GenRM
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
arXiv1 repoarXiv:2406.08801
hallo
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
arXiv1 repoarXiv:2406.09162
ELLA
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval
arXiv1 repoarXiv:2406.09188
lincir
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding
arXiv1 repoarXiv:2406.09297
LCKV
WonderWorld: Interactive 3D Scene Generation from a Single Image
arXiv1 repoarXiv:2406.09394
WonderWorld
The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments
arXiv1 repoarXiv:2406.09494
Displace2024_baseline_updated
$S^3$ -- Semantic Signal Separation
arXiv1 repoarXiv:2406.09556
turftopic
Large language model validity via enhanced conformal prediction methods
arXiv1 repoarXiv:2406.09714
OLAPH
Grounding Image Matching in 3D with MASt3R
arXiv1 repoarXiv:2406.09756
svraster
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
arXiv1 repoarXiv:2406.09827
hip-ainl
Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild
arXiv1 repoarXiv:2406.09905
nymeria_dataset
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages
arXiv1 repoarXiv:2406.09948
BLEnD
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
arXiv1 repoarXiv:2406.10052
WhisperLiveKit
GenQA: Generating Millions of Instructions from a Handful of Prompts
arXiv1 repoarXiv:2406.10323
GenQA
STAR: Scale-wise Text-conditioned AutoRegressive image generation
arXiv1 repoarXiv:2406.10797
VAR
garak: A Framework for Security Probing Large Language Models
arXiv1 repoarXiv:2406.11036
garak
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
arXiv1 repoarXiv:2406.11069
vision-arena
arXiv:2406.11149
arXiv1 repoarXiv:2406.11149
GoldCoin
Liberal Entity Matching as a Compound AI Toolchain
arXiv1 repoarXiv:2406.11255
libem
VideoVista: A Versatile Benchmark for Video Understanding and Reasoning
arXiv1 repoarXiv:2406.11303
VideoVista_Train
GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement
arXiv1 repoarXiv:2406.11546
AudioBench-N
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
arXiv1 repoarXiv:2406.11617
ComfyUI-LoRA-Optimizer
Task Me Anything
arXiv1 repoarXiv:2406.11775
TaskMeAnything-v1-imageqa-random
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
arXiv1 repoarXiv:2406.11801
safety-arithmetic
VideoLLM-online: Online Video Large Language Model for Streaming Video
arXiv1 repoarXiv:2406.11816
videollm-online
MegaScenes: Scene-Level View Synthesis at Scale
arXiv1 repoarXiv:2406.11819
dataset
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
arXiv1 repoarXiv:2406.11839
mDPO
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
arXiv1 repoarXiv:2406.11931
evalplus
Transcoders Find Interpretable LLM Feature Circuits
arXiv1 repoarXiv:2406.11944
gemma-scope-2b-pt-transcoders
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
arXiv1 repoarXiv:2406.12233
gujarati-vsr
Navigating Knowledge Management Implementation Success in Government Organizations: A type-2 fuzzy approach
arXiv1 repoarXiv:2406.12345
Recap-COCO-30K
HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors
arXiv1 repoarXiv:2406.12459
humansplat
Unified Active Retrieval for Retrieval Augmented Generation
arXiv1 repoarXiv:2406.12534
UAR_qwen
DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?
arXiv1 repoarXiv:2406.12641
llm-mysteries
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
arXiv1 repoarXiv:2406.12832
ComfyUI-LoRA-Optimizer
Medical Spoken Named Entity Recognition
arXiv1 repoarXiv:2406.13337
MultiMed
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
arXiv1 repoarXiv:2406.13362
VisualRWKV
VDebugger: Harnessing Execution Feedback for Debugging Visual Programs
arXiv1 repoarXiv:2406.13444
vdebugger
SpatialBot: Precise Spatial Understanding with Vision Language Models
arXiv1 repoarXiv:2406.13642
SpatialBot
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
arXiv1 repoarXiv:2406.13663
mirage
Rethinking Abdominal Organ Segmentation (RAOS) in the clinical scenario: A robustness evaluation benchmark with challenging cases
arXiv1 repoarXiv:2406.13674
RAOS
Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation
arXiv1 repoarXiv:2406.13692
sync-ralm-faithfulness
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
arXiv1 repoarXiv:2406.13743
t2v_metrics
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms
arXiv1 repoarXiv:2406.14228
evoagent
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
arXiv1 repoarXiv:2406.14598
unintentional-unalignment
Direct Multi-Turn Preference Optimization for Language Agents
arXiv1 repoarXiv:2406.14868
DMPO
Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
arXiv1 repoarXiv:2406.15479
Twin-Merging
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
arXiv1 repoarXiv:2406.15513
PKU-SafeRLHF
Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph
arXiv1 repoarXiv:2406.15627
SAR
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
arXiv1 repoarXiv:2406.15704
SALMONN
What Matters in Transformers? Not All Attention is Needed
arXiv1 repoarXiv:2406.15786
LLM-Drop
Real-time Speech Summarization for Medical Conversations
arXiv1 repoarXiv:2406.15888
MultiMed
Teaching LLMs to Abstain across Languages via Multilingual Feedback
arXiv1 repoarXiv:2406.15948
M-AbstainQA
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
arXiv1 repoarXiv:2406.15951
modular_pluralism
AudioBench: A Universal Benchmark for Audio Large Language Models
arXiv1 repoarXiv:2406.16020
MERaLiON-AudioLLM-Whisper-SEA-LION
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
arXiv1 repoarXiv:2406.16797
lottery-ticket-adaptation
RaTEScore: A Metric for Radiology Report Generation
arXiv1 repoarXiv:2406.16845
PMC-LLaMA
StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal
arXiv1 repoarXiv:2406.16864
StableNormal
CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation
arXiv1 repoarXiv:2406.17186
CLERC
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
arXiv1 repoarXiv:2406.17378
Text_aligns_tokens
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation
arXiv1 repoarXiv:2406.17484
MING
LumberChunker: Long-Form Narrative Document Segmentation
arXiv1 repoarXiv:2406.17526
gacha
Training-Free Exponential Context Extension via Cascading KV Cache
arXiv1 repoarXiv:2406.17808
cascading_kv_cache
E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
arXiv1 repoarXiv:2406.18009
F5-TTS
RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network
arXiv1 repoarXiv:2406.18284
Sonic
GaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enhanced Quality
arXiv1 repoarXiv:2406.18462
GaussianDreamerPro
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
arXiv1 repoarXiv:2406.18495
selfplay-redteaming
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
arXiv1 repoarXiv:2406.18510
wildteaming
PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation
arXiv1 repoarXiv:2406.18528
PrExMe
Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
arXiv1 repoarXiv:2406.18583
Lumina-T2X
arXiv:2406.18925
arXiv1 repoarXiv:2406.18925
VisArgs
AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation
arXiv1 repoarXiv:2406.18958
AnyControl
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
arXiv1 repoarXiv:2406.19135
DEX-TTS
PathAlign: A vision-language model for whole slide images in histopathology
arXiv1 repoarXiv:2406.19578
VLSA
Fine-tuning of Geospatial Foundation Models for Aboveground Biomass Estimation
arXiv1 repoarXiv:2406.19888
granite-geospatial-biomass
arXiv:2407.00023
arXiv1 repoarXiv:2407.00023
preble
PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration
arXiv1 repoarXiv:2407.00203
AEM-dataset
Tarsier: Recipes for Training and Evaluating Large Video Description Models
arXiv1 repoarXiv:2407.00634
tarsier
Preserving Multilingual Quality While Tuning Query Encoder on English Only
arXiv1 repoarXiv:2407.00923
arxiv-negatives
Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs
arXiv1 repoarXiv:2407.00945
EEP
SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
arXiv1 repoarXiv:2407.00952
SplitFM
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs
arXiv1 repoarXiv:2407.01082
LLM-Sampling
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
arXiv1 repoarXiv:2407.01102
bergen
arXiv:2407.01392
arXiv1 repoarXiv:2407.01392
SkyReels-V2-I2V-14B-720P
Retrieval-augmented generation in multilingual settings
arXiv1 repoarXiv:2407.01463
bergen
FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
arXiv1 repoarXiv:2407.01494
FoleyCrafter
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
arXiv1 repoarXiv:2407.01523
MMLongBench-Doc
fVDB: A Deep-Learning Framework for Sparse, Large-Scale, and High-Performance Spatial Intelligence
arXiv1 repoarXiv:2407.01781
SCube
VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
arXiv1 repoarXiv:2407.01863
Mirage
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
arXiv1 repoarXiv:2407.02883
coir
AgentInstruct: Toward Generative Teaching with Agentic Flows
arXiv1 repoarXiv:2407.03502
aurora-m2
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
arXiv1 repoarXiv:2407.03791
m5b_v2
ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild
arXiv1 repoarXiv:2407.04172
chartgemma
arXiv:2407.04292
arXiv1 repoarXiv:2407.04292
Corki
LaRa: Efficient Large-Baseline Radiance Fields
arXiv1 repoarXiv:2407.04699
LaRa
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
arXiv1 repoarXiv:2407.05282
detikzify-v2-8b
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
arXiv1 repoarXiv:2407.05361
F5-TTS
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
arXiv1 repoarXiv:2407.05407
BreezyVoice
See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition
arXiv1 repoarXiv:2407.05417
Subspace-Tuning
SmurfCat at PAN 2024 TextDetox: Alignment of Multilingual Transformers for Text Detoxification
arXiv1 repoarXiv:2407.05449
mt0-xl-detox-orpo
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models
arXiv1 repoarXiv:2407.05965
JailbreakDiffusionBench
DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metrics
arXiv1 repoarXiv:2407.06426
debunc
OffsetBias: Leveraging Debiased Data for Tuning Evaluators
arXiv1 repoarXiv:2407.06551
offsetbias
PaliGemma: A versatile 3B VLM for transfer
arXiv1 repoarXiv:2407.07726
detikzify-v2-8b
Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
arXiv1 repoarXiv:2407.07791
KnowledgeSpread
OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training
arXiv1 repoarXiv:2407.07852
OpenDiloco
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
arXiv1 repoarXiv:2407.08296
Finetune_with_GaLore
Self-training Language Models for Arithmetic Reasoning
arXiv1 repoarXiv:2407.08400
calc-x
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
arXiv1 repoarXiv:2407.08739
Image-Generation-CoT
TAPFixer: Automatic Detection and Repair of Home Automation Vulnerabilities based on Negated-property Reasoning
arXiv1 repoarXiv:2407.09095
ComfyUI-LoRA-Optimizer
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
arXiv1 repoarXiv:2407.09447
ASTPrompter
arXiv:2407.09450
arXiv1 repoarXiv:2407.09450
HEBO
Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models
arXiv1 repoarXiv:2407.10299
AnomalyRuler
LAB-Bench: Measuring Capabilities of Language Models for Biology Research
arXiv1 repoarXiv:2407.10362
LAB-Bench
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
arXiv1 repoarXiv:2407.10627
aurora-m2
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
arXiv1 repoarXiv:2407.11214
open-atp
Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness
arXiv1 repoarXiv:2407.11229
Robust-CQA
arXiv:2407.11325
arXiv1 repoarXiv:2407.11325
VISA
Continuity Preserving Online CenterLine Graph Learning
arXiv1 repoarXiv:2407.11337
CGNet
Revisiting the Impact of Pursuing Modularity for Code Generation
arXiv1 repoarXiv:2407.11406
Revisiting-Modularity
Scaling Diffusion Transformers to 16 Billion Parameters
arXiv1 repoarXiv:2407.11633
DiT-MoE-diffusers
Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation
arXiv1 repoarXiv:2407.11820
Stepping-Stones
SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge
arXiv1 repoarXiv:2407.11906
CaRTS
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
arXiv1 repoarXiv:2407.12735
EchoSight
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
arXiv1 repoarXiv:2407.12883
BRIGHT
Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild
arXiv1 repoarXiv:2407.12927
feature-vs-text-compound-emotion
Deep Time Series Models: A Comprehensive Survey and Benchmark
arXiv1 repoarXiv:2407.13278
Time-Series-Library
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
arXiv1 repoarXiv:2407.14505
TTOM
Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models
arXiv1 repoarXiv:2407.15408
ChronAccRet
Invariance Times Transfer Properties
arXiv1 repoarXiv:2407.15460
anchormind
LLMmap: Fingerprinting For Large Language Models
arXiv1 repoarXiv:2407.15847
LLMmap
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
arXiv1 repoarXiv:2407.17274
AVG
AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents
arXiv1 repoarXiv:2407.17490
AMEX
Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models
arXiv1 repoarXiv:2407.18015
GiftEvalPretrain
Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation
arXiv1 repoarXiv:2407.18698
Adaptive-Contrastive-Search
MangaUB: A Manga Understanding Benchmark for Large Multimodal Models
arXiv1 repoarXiv:2407.19034
manga-ub
ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development
arXiv1 repoarXiv:2407.20143
VeOmni
arXiv:2407.21004
arXiv1 repoarXiv:2407.21004
Evolver
CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning
arXiv1 repoarXiv:2407.21011
CLEFT
arXiv:2407.21315
arXiv1 repoarXiv:2407.21315
SpeechCueLLM
MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
arXiv1 repoarXiv:2407.21439
RagVL
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
arXiv1 repoarXiv:2407.21705
Tora
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
arXiv1 repoarXiv:2408.00264
clover
EXAONEPath 1.0 Patch-level Foundation Model for Pathology
arXiv1 repoarXiv:2408.00380
dpfm_factory
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
arXiv1 repoarXiv:2408.01337
AudioBench-N
VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance
arXiv1 repoarXiv:2408.01432
Concept-Bottleneck-LLM
arXiv:2408.02514
arXiv1 repoarXiv:2408.02514
Stem-JEPA
Multistain Pretraining for Slide Representation Learning in Pathology
arXiv1 repoarXiv:2408.02859
MADELEINE
VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge
arXiv1 repoarXiv:2408.02865
MedUMM
UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
arXiv1 repoarXiv:2408.05939
UniPortrait
DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
arXiv1 repoarXiv:2408.06010
DEEPTalk
BMX: Entropy-weighted Similarity and Semantic-enhanced Lexical Search
arXiv1 repoarXiv:2408.06643
baguetter
Imagen 3
arXiv1 repoarXiv:2408.07009
t2v_metrics
Post-Training Sparse Attention with Double Sparsity
arXiv1 repoarXiv:2408.07092
DoubleSparse
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
arXiv1 repoarXiv:2408.07199
agent-q
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
arXiv1 repoarXiv:2408.07440
baple
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
arXiv1 repoarXiv:2408.08067
RAGChecker
ECG-Chat: A Large ECG-Language Model for Cardiac Disease Diagnosis
arXiv1 repoarXiv:2408.08849
ECG-Chat
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
arXiv1 repoarXiv:2408.08872
xgen-mm-phi3-mini-instruct-interleave-r-v1.5
Selective Prompt Anchoring for Code Generation
arXiv1 repoarXiv:2408.09121
Selective-Prompt-Anchoring
BLADE: Benchmarking Language Model Agents for Data-Driven Science
arXiv1 repoarXiv:2408.09667
BLADE
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
arXiv1 repoarXiv:2408.09744
RealCustom
LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
arXiv1 repoarXiv:2408.10343
legalbenchrag
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
arXiv1 repoarXiv:2408.11039
zen5
Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation
arXiv1 repoarXiv:2408.11053
verilog-eval
Ophthalmic Biomarker Detection: Highlights from the IEEE Video and Image Processing Cup 2023 Student Competition
arXiv1 repoarXiv:2408.11170
OLIVES_Dataset
Real-Time Video Generation with Pyramid Attention Broadcast
arXiv1 repoarXiv:2408.12588
OpenDiT
Building and better understanding vision-language models: insights and future directions
arXiv1 repoarXiv:2408.12637
Idefics3-8B-Llama3
NanoFlow: Towards Optimal Large Language Model Serving Throughput
arXiv1 repoarXiv:2408.12757
Nanoflow
SONICS: Synthetic Or Not -- Identifying Counterfeit Songs
arXiv1 repoarXiv:2408.14080
bach-or-bot
A model of generation of a jet in stratified nonequilibrium plasma
arXiv1 repoarXiv:2408.14210
SphereForge
arXiv:2408.14262
arXiv1 repoarXiv:2408.14262
s3m-aave
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
arXiv1 repoarXiv:2408.14419
acl2026-misleading-visualizations
Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations
arXiv1 repoarXiv:2408.15232
storm
LRP4RAG: Detecting Hallucinations in Retrieval-Augmented Generation via Layer-wise Relevance Propagation
arXiv1 repoarXiv:2408.15533
LRP-eXplains-Transformers
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
arXiv1 repoarXiv:2408.15585
wespeaker
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
arXiv1 repoarXiv:2408.15966
MiniGPT-3D
CogVLM2: Visual Language Models for Image and Video Understanding
arXiv1 repoarXiv:2408.16500
cogvlm2-llama3-chat-19B
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
arXiv1 repoarXiv:2408.16737
aurora-m2
VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning
arXiv1 repoarXiv:2409.01667
VProChart
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
arXiv1 repoarXiv:2409.01704
GOT-OCR2_0
Boosting Vision-Language Models for Histopathology Classification: Predict all at once
arXiv1 repoarXiv:2409.01883
Histo-TransCLIP
BEAVER: An Enterprise Benchmark for Text-to-SQL
arXiv1 repoarXiv:2409.02038
beaver
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)
arXiv1 repoarXiv:2409.02920
RoboTwin
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
arXiv1 repoarXiv:2409.03420
mPLUG-DocOwl
Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding
arXiv1 repoarXiv:2409.03757
Lexicon3D
BreachSeek: A Multi-Agent Automated Penetration Tester
arXiv1 repoarXiv:2409.03789
awesome-cybersecurity-agentic-ai
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
arXiv1 repoarXiv:2409.04701
late-chunking
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
arXiv1 repoarXiv:2409.05395
vl_mamba
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
arXiv1 repoarXiv:2409.05591
MemoRAG
FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations
arXiv1 repoarXiv:2409.05976
FederatedLLM
World-Grounded Human Motion Recovery via Gravity-View Coordinates
arXiv1 repoarXiv:2409.06662
GVHMR
gsplat: An Open-Source Library for Gaussian Splatting
arXiv1 repoarXiv:2409.06765
gsplat
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
arXiv1 repoarXiv:2409.07123
Cross-Refine
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
arXiv1 repoarXiv:2409.08091
EZIGen
arXiv:2409.08248
arXiv1 repoarXiv:2409.08248
textboost
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
arXiv1 repoarXiv:2409.08264
UI-TARS
Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation
arXiv1 repoarXiv:2409.09016
CLOVER
TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings
arXiv1 repoarXiv:2409.09564
TG-LLaVA
Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search
arXiv1 repoarXiv:2409.09913
RaBitQ-Library
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
arXiv1 repoarXiv:2409.10103
speaker_disentangled_hubert
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
arXiv1 repoarXiv:2409.10173
jina-embeddings-v3
MusicLIME: Explainable Multimodal Music Understanding
arXiv1 repoarXiv:2409.10496
bach-or-bot
MotIF: Motion Instruction Fine-tuning
arXiv1 repoarXiv:2409.10683
motif
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
arXiv1 repoarXiv:2409.11239
KUDGE
Towards Fair RAG: On the Impact of Fair Ranking in Retrieval-Augmented Generation
arXiv1 repoarXiv:2409.11598
Fair-RAG
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
arXiv1 repoarXiv:2409.12117
nemo-nano-codec-22khz-1.89kbps-21.5fps
arXiv:2409.12147
arXiv1 repoarXiv:2409.12147
MAgICoRE
Large Language Models are Strong Audio-Visual Speech Recognition Learners
arXiv1 repoarXiv:2409.12319
Llama-AVSR
FlexiTex: Enhancing Texture Generation via Visual Guidance
arXiv1 repoarXiv:2409.12431
FlexiSyncMVD
AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework
arXiv1 repoarXiv:2409.12466
AudioEditor
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
arXiv1 repoarXiv:2409.13317
JMedBench
Higher-Order Message Passing for Glycan Representation Learning
arXiv1 repoarXiv:2409.13467
GIFFLAR
Logically Consistent Language Models via Neuro-Symbolic Integration
arXiv1 repoarXiv:2409.13724
loco-llm
Language agents achieve superhuman synthesis of scientific knowledge
arXiv1 repoarXiv:2409.13740
paper-qa
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
arXiv1 repoarXiv:2409.14074
MultiMed
PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects
arXiv1 repoarXiv:2409.14331
PISR
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
arXiv1 repoarXiv:2409.14337
MobileViews
arXiv:2409.14507
arXiv1 repoarXiv:2409.14507
SAEBench
Inference-Friendly Models With MixAttention
arXiv1 repoarXiv:2409.15012
LCKV
RAMBO: Enhancing RAG-based Repository-Level Method Body Completion
arXiv1 repoarXiv:2409.15204
RAMBO
Making s-wave superconductors topological with magnetic field
arXiv1 repoarXiv:2409.15266
SphereForge
OmniBench: Towards The Future of Universal Omni-Language Models
arXiv1 repoarXiv:2409.15272
OmniInstruct_v1
Parse Trees Guided LLM Prompt Compression
arXiv1 repoarXiv:2409.15395
Prompt-Compression
Making Text Embedders Few-Shot Learners
arXiv1 repoarXiv:2409.15700
bge-en-icl
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
arXiv1 repoarXiv:2409.16191
HelloBench
Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts
arXiv1 repoarXiv:2409.17106
CADFusion
Internalizing ASR with Implicit Chain of Thought for Efficient Speech-to-Speech Conversational LLM
arXiv1 repoarXiv:2409.17353
SpeechLLM
Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography
arXiv1 repoarXiv:2409.18119
MaMA
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
arXiv1 repoarXiv:2409.18121
rsrd
LangSAMP: Language-Script Aware Multilingual Pretraining
arXiv1 repoarXiv:2409.18199
LangSAMP
MinerU: An Open-Source Solution for Precise Document Content Extraction
arXiv1 repoarXiv:2409.18839
MinerU
FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark
arXiv1 repoarXiv:2409.19014
FLEX
HybridFlow: A Flexible and Efficient RLHF Framework
arXiv1 repoarXiv:2409.19256
verl-pipeline
On the Nonlinear Excitation of Phononic Frequency Combs in Molecules
arXiv1 repoarXiv:2409.19607
Tiny-R2
Old Optimizer, New Norm: An Anthology
arXiv1 repoarXiv:2409.20325
modded-nanogpt
The Perfect Blend: Redefining RLHF with Mixture of Judges
arXiv1 repoarXiv:2409.20370
open-perfectblend
A Hitchhikers Guide to Fine-Grained Face Forgery Detection Using Common Sense Reasoning
arXiv1 repoarXiv:2410.00485
HitchhikersGuide
From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging
arXiv1 repoarXiv:2410.01215
ComfyUI-LoRA-Optimizer
Enhancement of superconductivity coexisting with charge density wave in lattice expanded $\textrm{NbTe}_2$
arXiv1 repoarXiv:2410.01247
Titan-Memory
HelpSteer2-Preference: Complementing Ratings with Preferences
arXiv1 repoarXiv:2410.01257
Llama-3_3-Nemotron-Super-49B-GenRM
Endless Jailbreaks with Bijection Learning
arXiv1 repoarXiv:2410.01294
GA
CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
arXiv1 repoarXiv:2410.01400
CrowdCounter
SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios
arXiv1 repoarXiv:2410.01481
SonicSim
shapiq: Shapley Interactions for Machine Learning
arXiv1 repoarXiv:2410.01649
shapiq
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo
arXiv1 repoarXiv:2410.01920
TSMC4MATH
DeepProtein: Deep Learning Library and Benchmark for Protein Sequence Learning
arXiv1 repoarXiv:2410.02023
DeepProtein
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
arXiv1 repoarXiv:2410.02163
adversarial_decoding
Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
arXiv1 repoarXiv:2410.02197
general-preference-model
Contextual Document Embeddings
arXiv1 repoarXiv:2410.02525
cde-small-v1
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
arXiv1 repoarXiv:2410.02694
HELMET
LLaVA-Video: Video Instruction Tuning With Synthetic Data
arXiv1 repoarXiv:2410.02713
LLaVA-Video-178K
ToolGen: Unified Tool Retrieval and Calling via Generation
arXiv1 repoarXiv:2410.03439
ToolGen-Datasets
RAFT: Realistic Attacks to Fool Text Detectors
arXiv1 repoarXiv:2410.03658
RAFT
Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-Review
arXiv1 repoarXiv:2410.03663
Learn-from-Committee
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
arXiv1 repoarXiv:2410.03825
MonST3R_PO-TA-S-W_ViTLarge_BaseDecoder_512_dpt
Learning Code Preference via Synthetic Evolution
arXiv1 repoarXiv:2410.03837
llm-code-preference
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
arXiv1 repoarXiv:2410.03960
Llama-3.1-SwiftKV-8B-Instruct
A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages
arXiv1 repoarXiv:2410.03981
agentty
IV-Mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video Synthesis
arXiv1 repoarXiv:2410.04171
IV-mixed-Sampler
Casablanca: Data and Models for Multidialectal Arabic Speech Recognition
arXiv1 repoarXiv:2410.04527
Casablanca
CAR: Controllable Autoregressive Modeling for Visual Generation
arXiv1 repoarXiv:2410.04671
CAR
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
arXiv1 repoarXiv:2410.04727
ForgettingCurve
arXiv:2410.05102
arXiv1 repoarXiv:2410.05102
HEBO
GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting
arXiv1 repoarXiv:2410.05259
GS-VTON
AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
arXiv1 repoarXiv:2410.05295
GA
Image Watermarks are Removable Using Controllable Regeneration from Clean Noise
arXiv1 repoarXiv:2410.05470
watermarks-remover
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
arXiv1 repoarXiv:2410.06244
story-iter
TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
arXiv1 repoarXiv:2410.06511
torchtitan
InstantIR: Blind Image Restoration with Instant Generative Reference
arXiv1 repoarXiv:2410.06551
InstantIR
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
arXiv1 repoarXiv:2410.06672
circuit_backup
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
arXiv1 repoarXiv:2410.07095
aideml
VHELM: A Holistic Evaluation of Vision Language Models
arXiv1 repoarXiv:2410.07112
helm
IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking
arXiv1 repoarXiv:2410.07295
itergen
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
arXiv1 repoarXiv:2410.08020
TTFT-SIFT
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
arXiv1 repoarXiv:2410.08102
DataFlow
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation
arXiv1 repoarXiv:2410.08371
DAM
Bilinear MLPs enable weight-based mechanistic interpretability
arXiv1 repoarXiv:2410.08417
bilinear-decomposition
arXiv:2410.08709
arXiv1 repoarXiv:2410.08709
di4c
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
arXiv1 repoarXiv:2410.08792
SeeDo
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
arXiv1 repoarXiv:2410.08847
unintentional-unalignment
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
arXiv1 repoarXiv:2410.09024
felonybench
When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning
arXiv1 repoarXiv:2410.09132
MAGB
Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization
arXiv1 repoarXiv:2410.09302
verl-pipeline
Skipping Computations in Multimodal LLMs
arXiv1 repoarXiv:2410.09454
ima-lmms
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
arXiv1 repoarXiv:2410.09732
LOKI
Text4Seg: Reimagining Image Segmentation as Text Generation
arXiv1 repoarXiv:2410.09855
Text4Seg
ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains
arXiv1 repoarXiv:2410.09870
ChroKnowBench
arXiv:2410.09893
arXiv1 repoarXiv:2410.09893
RMB-Reward-Model-Benchmark
MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
arXiv1 repoarXiv:2410.10122
MuseTalk
Fed-pilot: Optimizing LoRA Allocation for Efficient Federated Fine-Tuning with Heterogeneous Clients
arXiv1 repoarXiv:2410.10200
Fed-PLoRA
KBLaM: Knowledge Base augmented Language Model
arXiv1 repoarXiv:2410.10450
KBLaM
Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification
arXiv1 repoarXiv:2410.10573
VLSA
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
arXiv1 repoarXiv:2410.10700
SafeMTData
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
arXiv1 repoarXiv:2410.10812
hart
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free
arXiv1 repoarXiv:2410.10814
MoE-Embedding
Depth Any Video with Scalable Synthetic Data
arXiv1 repoarXiv:2410.10815
DepthAnyVideo
When Does Perceptual Alignment Benefit Vision Representations?
arXiv1 repoarXiv:2410.10817
dreamsim
arXiv:2410.10819
arXiv1 repoarXiv:2410.10819
Block-Sparse-Attention
Liger Kernel: Efficient Triton Kernels for LLM Training
arXiv1 repoarXiv:2410.10989
Liger-Kernel
arXiv:2410.11163
arXiv1 repoarXiv:2410.11163
model_swarm
Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling
arXiv1 repoarXiv:2410.11236
Ctrl-U
Improving Long-Text Alignment for Text-to-Image Diffusion Models
arXiv1 repoarXiv:2410.11817
LongAlign
MoH: Multi-Head Attention as Mixture-of-Head Attention
arXiv1 repoarXiv:2410.11842
Chat-UniVi
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
arXiv1 repoarXiv:2410.12266
AudioLCM
Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
arXiv1 repoarXiv:2410.12323
Less-is-More
Towards Neural Scaling Laws for Time Series Foundation Models
arXiv1 repoarXiv:2410.12360
time-moe
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
arXiv1 repoarXiv:2410.12628
DocLayout-YOLO
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
arXiv1 repoarXiv:2410.12694
MMMM
Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media
arXiv1 repoarXiv:2410.12791
turftopic
Enterprise Benchmarks for Large Language Model Evaluation
arXiv1 repoarXiv:2410.12857
helm
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
arXiv1 repoarXiv:2410.13360
RAP-MLLM
MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures
arXiv1 repoarXiv:2410.13754
MixEval
FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model
arXiv1 repoarXiv:2410.13925
FiT-diffusers
arXiv:2410.13928
arXiv1 repoarXiv:2410.13928
SAEBench
HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation
arXiv1 repoarXiv:2410.14324
PlanGen
SNAC: Multi-Scale Neural Audio Codec
arXiv1 repoarXiv:2410.14411
snac
A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference
arXiv1 repoarXiv:2410.14442
LCKV
DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents
arXiv1 repoarXiv:2410.14803
carl_distrl
A Multimodal Vision Foundation Model for Clinical Dermatology
arXiv1 repoarXiv:2410.15038
PanDerm
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
arXiv1 repoarXiv:2410.15553
smoltalk2
Moonshine: Speech Recognition for Live Transcription and Voice Commands
arXiv1 repoarXiv:2410.15608
moonshine
TimeMixer++: A General Time Series Pattern Machine for Universal Predictive Analysis
arXiv1 repoarXiv:2410.16032
time-moe
1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs
arXiv1 repoarXiv:2410.16144
BitNet
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
arXiv1 repoarXiv:2410.16184
Llama-3_3-Nemotron-Super-49B-GenRM
Improve Vision Language Model Chain-of-thought Reasoning
arXiv1 repoarXiv:2410.16198
LLaVA-Hound-DPO
Beyond Browsing: API-Based Web Agents
arXiv1 repoarXiv:2410.16464
guaca
TIPS: Text-Image Pretraining with Spatial awareness
arXiv1 repoarXiv:2410.16512
tips
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
arXiv1 repoarXiv:2410.16770
scene-language
VoiceBench: Benchmarking LLM-Based Voice Assistants
arXiv1 repoarXiv:2410.17196
voicebench
Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss
arXiv1 repoarXiv:2410.17243
VideoLLaMA2
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
arXiv1 repoarXiv:2410.17250
JMMMU
Altogether: Image Captioning via Re-aligning Alt-text
arXiv1 repoarXiv:2410.17251
MetaCLIP
Scalable Influence and Fact Tracing for Large Language Model Pretraining
arXiv1 repoarXiv:2410.17413
bergson
ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents
arXiv1 repoarXiv:2410.17657
MING
arXiv:2410.17736
arXiv1 repoarXiv:2410.17736
llama2.mojo
Value Residual Learning
arXiv1 repoarXiv:2410.17897
modded-nanogpt
FreeVS: Generative View Synthesis on Free Driving Trajectory
arXiv1 repoarXiv:2410.18079
FreeVS
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
arXiv1 repoarXiv:2410.18194
ZIP-FIT
DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation
arXiv1 repoarXiv:2410.18666
DreamClear
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
arXiv1 repoarXiv:2410.18697
prometheus-eval
arXiv:2410.18745
arXiv1 repoarXiv:2410.18745
STRING
arXiv:2410.18798
arXiv1 repoarXiv:2410.18798
CharXiv
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
arXiv1 repoarXiv:2410.18962
GST
arXiv:2410.19278
arXiv1 repoarXiv:2410.19278
SAEBench
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction
arXiv1 repoarXiv:2410.19452
NeuroClips
CoqPilot, a plugin for LLM-based generation of proofs
arXiv1 repoarXiv:2410.19605
coqpilot
RARe: Retrieval Augmented Retrieval with In-Context Examples
arXiv1 repoarXiv:2410.20088
RARe
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
arXiv1 repoarXiv:2410.20285
moatless-tools
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
arXiv1 repoarXiv:2410.20526
circuit_backup
PaPaGei: Open Foundation Models for Optical Physiological Signals
arXiv1 repoarXiv:2410.20542
papagei-foundation-model
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
arXiv1 repoarXiv:2410.20672
mixture_of_recursions
Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation
arXiv1 repoarXiv:2410.20724
SubgraphRAG
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
arXiv1 repoarXiv:2410.20941
BLEUless_DocMT
Flaming-hot Initiation with Regular Execution Sampling for Large Language Models
arXiv1 repoarXiv:2410.21236
verl-pipeline
arXiv:2410.21357
arXiv1 repoarXiv:2410.21357
Energy-Diffusion-LLM
Are Decoder-Only Large Language Models the Silver Bullet for Code Search?
arXiv1 repoarXiv:2410.22240
DecoderLLMs-CodeSearch
Online Detection of LLM-Generated Texts via Sequential Hypothesis Testing by Betting
arXiv1 repoarXiv:2410.22318
online-llm-detection
arXiv:2410.22376
arXiv1 repoarXiv:2410.22376
Rare-to-Frequent
Image2Struct: Benchmarking Structure Extraction for Vision-Language Models
arXiv1 repoarXiv:2410.22456
helm
FlowDCN: Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution
arXiv1 repoarXiv:2410.22655
FlowDCN
Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval
arXiv1 repoarXiv:2410.23041
Role-Playing-LLM-Megumin
Controlling Language and Diffusion Models by Transporting Activations
arXiv1 repoarXiv:2410.23054
ml-lineas
On Memorization of Large Language Models in Logical Reasoning
arXiv1 repoarXiv:2410.23123
knights-and-knaves
Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms
arXiv1 repoarXiv:2410.23144
PD12M
$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources
arXiv1 repoarXiv:2410.23261
academic-pretraining
VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration
arXiv1 repoarXiv:2410.23317
GUI-KV
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
arXiv1 repoarXiv:2410.23825
GlotCC-V1
Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models
arXiv1 repoarXiv:2410.23841
InfoSearch
SelfCodeAlign: Self-Alignment for Code Generation
arXiv1 repoarXiv:2410.24198
starcoder2-instruct-15b-v0.1
DELTA: Dense Efficient Long-range 3D Tracking for any video
arXiv1 repoarXiv:2410.24211
DELTA_densetrack3d
Randomized Autoregressive Visual Generation
arXiv1 repoarXiv:2411.00776
clustermark_1d-tokenizer
CycleResearcher: Improving Automated Research via Automated Review
arXiv1 repoarXiv:2411.00816
Researcher
SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
arXiv1 repoarXiv:2411.01106
VisR-Bench
Sample-Efficient Alignment for LLMs
arXiv1 repoarXiv:2411.01493
oat
How Far is Video Generation from World Model: A Physical Law Perspective
arXiv1 repoarXiv:2411.02385
phyworld
AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
arXiv1 repoarXiv:2411.02394
autovfx
Training-free Regional Prompting for Diffusion Transformers
arXiv1 repoarXiv:2411.02395
Regional-Prompting-FLUX
SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models
arXiv1 repoarXiv:2411.02433
SLED
Dr. SoW: Density Ratio of Strong-over-weak LLMs for Reducing the Cost of Human Annotation in Preference Tuning
arXiv1 repoarXiv:2411.02481
reward_hub
ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy
arXiv1 repoarXiv:2411.02572
rxrx3-core
On the Loss of Context-awareness in General Instruction Fine-tuning
arXiv1 repoarXiv:2411.02688
context_awareness
ADOPT: Modified Adam Can Converge with Any $β_2$ with the Optimal Rate
arXiv1 repoarXiv:2411.02853
adopt
SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents
arXiv1 repoarXiv:2411.03284
rightmind
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
arXiv1 repoarXiv:2411.03628
StreamingBench
A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning
arXiv1 repoarXiv:2411.04105
prop-logic-transformer-circuit
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
arXiv1 repoarXiv:2411.04118
eval-medical-dapt
Vision Language Models are In-Context Value Learners
arXiv1 repoarXiv:2411.04549
reward-scope
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
arXiv1 repoarXiv:2411.04709
TIP-I2V
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
arXiv1 repoarXiv:2411.04923
VideoGLaMM
BitNet a4.8: 4-bit Activations for 1-bit LLMs
arXiv1 repoarXiv:2411.04965
BitNet
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
arXiv1 repoarXiv:2411.04983
stable-worldmodel
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
arXiv1 repoarXiv:2411.05195
CLIP-Embeds
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
arXiv1 repoarXiv:2411.05222
rlt
Using Language Models to Disambiguate Lexical Choices in Translation
arXiv1 repoarXiv:2411.05781
Lex-Rules
GFT: Graph Foundation Model with Transferable Tree Vocabulary
arXiv1 repoarXiv:2411.06070
GFT
M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
arXiv1 repoarXiv:2411.06176
multimodal-docs-public
ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?
arXiv1 repoarXiv:2411.06469
ClinicalBench
arXiv:2411.07186
arXiv1 repoarXiv:2411.07186
NatureLM-audio
InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance
arXiv1 repoarXiv:2411.07795
InvisMark
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
arXiv1 repoarXiv:2411.07975
Janus
Large Language Models Can Self-Improve in Long-context Reasoning
arXiv1 repoarXiv:2411.08147
SEALONG
The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models
arXiv1 repoarXiv:2411.08870
eval-medical-dapt
Reliable Decision Making via Calibration Oriented Retrieval Augmented Generation
arXiv1 repoarXiv:2411.08891
calibrag
"Should I Give Up Now?" Investigating LLM Pitfalls in Software Engineering
arXiv1 repoarXiv:2411.09916
unlazy
EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge
arXiv1 repoarXiv:2411.10224
MLRG
Does Prompt Formatting Have Any Impact on LLM Performance?
arXiv1 repoarXiv:2411.10541
opendataloader-bench
arXiv:2411.10557
arXiv1 repoarXiv:2411.10557
MLAN
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer
arXiv1 repoarXiv:2411.10781
test-time-scaling
SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization
arXiv1 repoarXiv:2411.10958
SageAttention
Scalable Autoregressive Monocular Depth Estimation
arXiv1 repoarXiv:2411.11361
VAR
Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model
arXiv1 repoarXiv:2411.12783
Med-2E3
Stylecodes: Encoding Stylistic Information For Image Generation
arXiv1 repoarXiv:2411.12811
stylecodes
VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
arXiv1 repoarXiv:2411.12915
VLM
Veryl: A New Hardware Description Language as an Altarnative to SystemVerilog
arXiv1 repoarXiv:2411.12983
veryl
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
arXiv1 repoarXiv:2411.13503
VBench
Find Any Part in 3D
arXiv1 repoarXiv:2411.13550
Find3D
Hymba: A Hybrid-head Architecture for Small Language Models
arXiv1 repoarXiv:2411.13676
hymba
Novel View Extrapolation with Video Diffusion Priors
arXiv1 repoarXiv:2411.14208
ViewExtrapolator
arXiv:2411.14280
arXiv1 repoarXiv:2411.14280
EasyHOI
Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI
arXiv1 repoarXiv:2411.14299
Masala-CHAI
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
arXiv1 repoarXiv:2411.14347
Rex-Omni
Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction
arXiv1 repoarXiv:2411.14384
Open-OmniVCus
FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data
arXiv1 repoarXiv:2411.14717
FedMLLM
OminiControl: Minimal and Universal Control for Diffusion Transformer
arXiv1 repoarXiv:2411.15098
Subjects200K
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
arXiv1 repoarXiv:2411.15100
xgrammar
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
arXiv1 repoarXiv:2411.15114
aideml
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
arXiv1 repoarXiv:2411.15115
VideoRepair
Tulu 3: Pushing Frontiers in Open Language Model Post-Training
arXiv1 repoarXiv:2411.15124
OLMoE-1B-7B-0125-Instruct
VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing
arXiv1 repoarXiv:2411.15260
VIVID-10M
Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation Knowledge
arXiv1 repoarXiv:2411.15277
FreeCure
Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT
arXiv1 repoarXiv:2411.15418
panspecies-dti
Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation
arXiv1 repoarXiv:2411.16185
Fancy123
Functionality understanding and segmentation in 3D scenes
arXiv1 repoarXiv:2411.16310
fun3du
Preference Optimization for Reasoning with Pseudo Feedback
arXiv1 repoarXiv:2411.16345
EMPO
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
arXiv1 repoarXiv:2411.16602
Chat2SVG
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
arXiv1 repoarXiv:2411.16721
ASTRA
arXiv:2411.16778
arXiv1 repoarXiv:2411.16778
uMedGround
Controllable Human Image Generation with Personalized Multi-Garments
arXiv1 repoarXiv:2411.16801
BootComp
Probing the limitations of multimodal language models for chemistry and materials research
arXiv1 repoarXiv:2411.16955
chembench
LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization
arXiv1 repoarXiv:2411.17178
VAR
CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning
arXiv1 repoarXiv:2411.17426
TransArch
WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
arXiv1 repoarXiv:2411.17459
Open-Sora-Plan
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
arXiv1 repoarXiv:2411.17465
ShowUI-desktop
CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models
arXiv1 repoarXiv:2411.18145
CHOICE
arXiv:2411.18301
arXiv1 repoarXiv:2411.18301
LaRender
ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
arXiv1 repoarXiv:2411.18363
Rex-Omni
FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving
arXiv1 repoarXiv:2411.18424
lightllm
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
arXiv1 repoarXiv:2411.18673
ac3d
Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits
arXiv1 repoarXiv:2411.18704
open-value
arXiv:2411.18895
arXiv1 repoarXiv:2411.18895
SAEBench
Puzzle: Distillation-Based NAS for Inference-Optimized LLMs
arXiv1 repoarXiv:2411.19146
Llama-3_3-Nemotron-Super-49B-v1_5
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos
arXiv1 repoarXiv:2411.19167
ObjectForesight-HOT3D-DiT
Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
arXiv1 repoarXiv:2412.00127
Orthus-7B-base
ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain
arXiv1 repoarXiv:2412.00532
mteb-1.34.14
arXiv:2412.00568
arXiv1 repoarXiv:2412.00568
the_well
Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics
arXiv1 repoarXiv:2412.00651
UMPIRE
Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild
arXiv1 repoarXiv:2412.00811
Vid-Morp
WAFFLE: Multimodal Floorplan Understanding in the Wild
arXiv1 repoarXiv:2412.00955
WAFFLE
INTELLECT-1 Technical Report
arXiv1 repoarXiv:2412.01152
prime-diloco
Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement
arXiv1 repoarXiv:2412.01282
CVPR2025_Align-KD
Free Process Rewards without Process Labels
arXiv1 repoarXiv:2412.01981
PRIME
Progress-Aware Video Frame Captioning
arXiv1 repoarXiv:2412.02071
ProgCaptioner
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
arXiv1 repoarXiv:2412.02210
CC-OCR
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention
arXiv1 repoarXiv:2412.02259
VideoGen-of-Thought
Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach
arXiv1 repoarXiv:2412.03017
OSEDiff
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
arXiv1 repoarXiv:2412.03069
TokenFlow
AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
arXiv1 repoarXiv:2412.03248
AIM
Imagine360: Immersive 360 Video Generation from Perspective Anchor
arXiv1 repoarXiv:2412.03552
Imagine360
PaliGemma 2: A Family of Versatile VLMs for Transfer
arXiv1 repoarXiv:2412.03555
gemma.cpp
Reducing Tool Hallucination via Reliability Alignment
arXiv1 repoarXiv:2412.04141
ToolHallucination
AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic
arXiv1 repoarXiv:2412.04193
al-qasida
Densing Law of LLMs
arXiv1 repoarXiv:2412.04315
Ultra-FineWeb
Liquid: Language Models are Scalable and Unified Multi-modal Generators
arXiv1 repoarXiv:2412.04332
Monkey
Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering
arXiv1 repoarXiv:2412.04459
svraster
NVILA: Efficient Frontier Visual Language Models
arXiv1 repoarXiv:2412.04468
VILA
SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models
arXiv1 repoarXiv:2412.04852
SleeperMark
A Practical Examination of AI-Generated Text Detectors for Large Language Models
arXiv1 repoarXiv:2412.05139
llm-detector-eval
Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies
arXiv1 repoarXiv:2412.05155
Multimodal-Fact-Checking-with-Vision-Language-Models
DreamColour: Controllable Video Colour Editing without Training
arXiv1 repoarXiv:2412.05180
sketchdeco-code
UniScene: Unified Occupancy-centric Driving Scene Generation
arXiv1 repoarXiv:2412.05435
UniScene
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
arXiv1 repoarXiv:2412.05818
SILMM-Implementation
Text-to-3D Generation by 2D Editing
arXiv1 repoarXiv:2412.05929
GE3D
Training Large Language Models to Reason in a Continuous Latent Space
arXiv1 repoarXiv:2412.06769
Mirage
Maya: An Instruction Finetuned Multilingual Multimodal Model
arXiv1 repoarXiv:2412.07112
maya
Hierarchical Split Federated Learning: Convergence Analysis and System Optimization
arXiv1 repoarXiv:2412.07197
SplitFM
ObjCtrl-2.5D: Training-free Object Control with Camera Poses
arXiv1 repoarXiv:2412.07721
ObjCtrl-2.5D
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
arXiv1 repoarXiv:2412.07755
Mirage
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
arXiv1 repoarXiv:2412.07760
SynCamVideo-Dataset
NLPineers@ NLU of Devanagari Script Languages 2025: Hate Speech Detection using Ensembling of BERT-based models
arXiv1 repoarXiv:2412.08163
NLPineers
StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements
arXiv1 repoarXiv:2412.08503
ComfyUI-StyleStudio
Large Concept Models: Language Modeling in a Sentence Representation Space
arXiv1 repoarXiv:2412.08821
fairseq2
FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction
arXiv1 repoarXiv:2412.09573
FreeSplatter
Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion
arXiv1 repoarXiv:2412.09593
Neural-LightRig
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
arXiv1 repoarXiv:2412.09818
MERaLiON-AudioLLM-Whisper-SEA-LION
BrushEdit: All-In-One Image Inpainting and Editing
arXiv1 repoarXiv:2412.10316
BrushEdit
SCBench: A KV Cache-Centric Analysis of Long-Context Methods
arXiv1 repoarXiv:2412.10319
SCBench
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
arXiv1 repoarXiv:2412.10321
jailbreak-objectives
Generative AI in Medicine
arXiv1 repoarXiv:2412.10337
llava-rad
EvalGIM: A Library for Evaluating Generative Image Models
arXiv1 repoarXiv:2412.10604
EvalGIM
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
arXiv1 repoarXiv:2412.10704
VisDoM
Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection
arXiv1 repoarXiv:2412.11506
glimpse
ColorFlow: Retrieval-Augmented Image Sequence Colorization
arXiv1 repoarXiv:2412.11815
ColorFlow
PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting
arXiv1 repoarXiv:2412.12096
PanSplat
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation
arXiv1 repoarXiv:2412.12559
EXIT
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
arXiv1 repoarXiv:2412.12785
Visual-Region
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
arXiv1 repoarXiv:2412.12932
Mirage
Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction
arXiv1 repoarXiv:2412.13110
gec-attribute
ROMAS: A Role-Based Multi-Agent System for Database monitoring and Planning
arXiv1 repoarXiv:2412.13520
DB-GPT
Clio: Privacy-Preserving Insights into Real-World AI Use
arXiv1 repoarXiv:2412.13678
enabling-independent-research
Open Universal Arabic ASR Leaderboard
arXiv1 repoarXiv:2412.13788
open_universal_arabic_asr_leaderboard
Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models
arXiv1 repoarXiv:2412.14133
PopVQA
Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
arXiv1 repoarXiv:2412.14642
BioMedGPT-Mol
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
arXiv1 repoarXiv:2412.14711
zen5
Scylla: Translating an Applicative Subset of C to Safe Rust
arXiv1 repoarXiv:2412.15042
libcrux
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
arXiv1 repoarXiv:2412.15194
MMLU-CF
arXiv:2412.15206
arXiv1 repoarXiv:2412.15206
AutoTrust
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
arXiv1 repoarXiv:2412.15208
OpenEMMA
EnvGS: Modeling View-Dependent Appearance with Environment Gaussian
arXiv1 repoarXiv:2412.15215
EnvGS
Building an Explainable Graph-based Biomedical Paper Recommendation System (Technical Report)
arXiv1 repoarXiv:2412.15229
NarrativeRecommender
Dimension Reduction with Locally Adjusted Graphs
arXiv1 repoarXiv:2412.15426
PaCMAP
Insights into resource utilization of code small language models serving with runtime engines and execution providers
arXiv1 repoarXiv:2412.15441
energy-ml-serving
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation
arXiv1 repoarXiv:2412.15529
XRAG
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
arXiv1 repoarXiv:2412.15649
SLAM-LLM
WebLLM: A High-Performance In-Browser LLM Inference Engine
arXiv1 repoarXiv:2412.15803
web-llm
arXiv:2412.16117
arXiv1 repoarXiv:2412.16117
Mate
Aria-UI: Visual Grounding for GUI Instructions
arXiv1 repoarXiv:2412.16256
Aria-UI_Data
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment
arXiv1 repoarXiv:2412.16334
dinov3
RCAEval: A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data
arXiv1 repoarXiv:2412.17015
kibana
A Reality Check on Context Utilisation for Retrieval-Augmented Generation
arXiv1 repoarXiv:2412.17031
druid
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
arXiv1 repoarXiv:2412.17667
versa
VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
arXiv1 repoarXiv:2412.18194
vlabench_primitive_ft_dataset
Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations
arXiv1 repoarXiv:2412.18955
GDRetriever
Jasper and Stella: distillation of SOTA embedding models
arXiv1 repoarXiv:2412.19048
RAG-Retrieval
RAG with Differential Privacy
arXiv1 repoarXiv:2412.19291
dp-rag
An Engorgio Prompt Makes Large Language Model Babble on
arXiv1 repoarXiv:2412.19394
Engorgio-prompt
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis
arXiv1 repoarXiv:2412.19723
SeeClick
Fortran2CPP: Automating Fortran-to-C++ Translation using LLMs via Multi-Turn Dialogue and Dual-Agent Integration
arXiv1 repoarXiv:2412.19770
Fortran2Cpp
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
arXiv1 repoarXiv:2412.20070
Med-MAT
Navigating Image Restoration with VAR's Distribution Alignment Prior
arXiv1 repoarXiv:2412.21063
VAR
Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation
arXiv1 repoarXiv:2412.21117
Prometheus
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
arXiv1 repoarXiv:2412.21187
DeepEnlighten
Titans: Learning to Memorize at Test Time
arXiv1 repoarXiv:2501.00663
Titan-Memory
AutoPresent: Designing Structured Visuals from Scratch
arXiv1 repoarXiv:2501.00912
AutoPresent
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
arXiv1 repoarXiv:2501.01005
flashinfer
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
arXiv1 repoarXiv:2501.01034
Multitask-National-Speech-Corpus-v1
SVFR: A Unified Framework for Generalized Video Face Restoration
arXiv1 repoarXiv:2501.01235
SVFR
LEO-Split: A Semi-Supervised Split Learning Framework over LEO Satellite Networks
arXiv1 repoarXiv:2501.01293
SplitFM
arXiv:2501.02531
arXiv1 repoarXiv:2501.02531
awesome-ai-sre
REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
arXiv1 repoarXiv:2501.03262
DeepEnlighten
The Multiple Equal-Difference Structure of Cyclotomic Cosets
arXiv1 repoarXiv:2501.03516
ru-promptriever
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
arXiv1 repoarXiv:2501.03936
PPTAgent
RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark
arXiv1 repoarXiv:2501.04440
sardet_100k
SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
arXiv1 repoarXiv:2501.04689
stable-point-aware-3d
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
arXiv1 repoarXiv:2501.04765
HDM-xut-340M-anime
Reproducing HotFlip for Corpus Poisoning Attacks in Dense Retrieval
arXiv1 repoarXiv:2501.04802
hotflip_corpus_poisoning
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
arXiv1 repoarXiv:2501.05510
OVO-Bench
Approximate well-balanced WENO finite difference schemes using a global-flux quadrature method with multi-step ODE integrator weights
arXiv1 repoarXiv:2501.06155
Sonus-Lab
A General Framework for Inference-time Scaling and Steering of Diffusion Models
arXiv1 repoarXiv:2501.06848
Fk-Diffusion-Steering
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
arXiv1 repoarXiv:2501.07171
open-pmc-18m
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
arXiv1 repoarXiv:2501.07525
RadAlign
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
arXiv1 repoarXiv:2501.07730
clustermark_1d-tokenizer
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
arXiv1 repoarXiv:2501.07888
tarsier
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
arXiv1 repoarXiv:2501.08248
ml-icr2
MiniMax-01: Scaling Foundation Models with Lightning Attention
arXiv1 repoarXiv:2501.08313
MMLongBench-Doc
arXiv:2501.08325
arXiv1 repoarXiv:2501.08325
GameFactory-Dataset
Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise
arXiv1 repoarXiv:2501.08331
Go-with-the-Flow
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation
arXiv1 repoarXiv:2501.08549
VRS-HQ
GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge
arXiv1 repoarXiv:2501.08913
raid
FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training
arXiv1 repoarXiv:2501.09213
Med-R1
Foundations of Large Language Models
arXiv1 repoarXiv:2501.09223
Medical-Assistant
Vision-Language Models Do Not Understand Negation
arXiv1 repoarXiv:2501.09425
negbench
Revealing the $χ_{\rm eff}$-$q$ Correlation among Coalescing Binary Black Holes and Tentative Evidence for AGN-driven Hierarchical Mergers
arXiv1 repoarXiv:2501.09495
Titan-Memory
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
arXiv1 repoarXiv:2501.09502
ViSpeak
Enhancing the De-identification of Personally Identifiable Information in Educational Data
arXiv1 repoarXiv:2501.09765
PrivacyAI
AI-Generated Music Detection and its Challenges
arXiv1 repoarXiv:2501.10111
deepfake-detector
ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
arXiv1 repoarXiv:2501.10132
ComplexFuncBench
MechIR: A Mechanistic Interpretability Framework for Information Retrieval
arXiv1 repoarXiv:2501.10165
MechIR
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
arXiv1 repoarXiv:2501.10970
mmar-freeform
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
arXiv1 repoarXiv:2501.11325
CatV2TON
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
arXiv1 repoarXiv:2501.11733
MobileAgent
Parallel Sequence Modeling via Generalized Spatial Propagation Network
arXiv1 repoarXiv:2501.12381
GSPN
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
arXiv1 repoarXiv:2501.12570
O1-Pruner
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
arXiv1 repoarXiv:2501.12835
AdaRAGUE
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
arXiv1 repoarXiv:2501.12895
TPO
Episodic Memories Generation and Evaluation Benchmark for Large Language Models
arXiv1 repoarXiv:2501.13121
StepDeepResearch
Design of Bayesian Clinical Trials with Clustered Data
arXiv1 repoarXiv:2501.13218
Titan-Memory
Improving Video Generation with Human Feedback
arXiv1 repoarXiv:2501.13918
VideoReward
The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities
arXiv1 repoarXiv:2501.13921
BreezyVoice
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
arXiv1 repoarXiv:2501.13926
Image-Generation-CoT
Longitudinal Abuse and Sentiment Analysis of Hollywood Movie Dialogues using Language Models
arXiv1 repoarXiv:2501.13948
sentimentanalysis-Hollywood
Constructive Ordinal Exponentiation
arXiv1 repoarXiv:2501.14542
TypeTopology
MatAnyone: Stable Video Matting with Consistent Memory Propagation
arXiv1 repoarXiv:2501.14677
MatAnyone
CodeMonkeys: Scaling Test-Time Compute for Software Engineering
arXiv1 repoarXiv:2501.14723
codemonkeys
MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models
arXiv1 repoarXiv:2501.15000
opendataloader-bench
Overview of the Amphion Toolkit (v0.2)
arXiv1 repoarXiv:2501.15442
Vevo1.5
Distributional Surgery for Language Model Activations
arXiv1 repoarXiv:2501.15758
OT-Intervention
SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model
arXiv1 repoarXiv:2501.15830
SpatialVLA
PISCO: Pretty Simple Compression for Retrieval-Augmented Generation
arXiv1 repoarXiv:2501.16075
pisco
A foundation model for human-AI collaboration in medical literature mining
arXiv1 repoarXiv:2501.16255
DeepRetrieval
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
arXiv1 repoarXiv:2501.16411
PhysBench
360Brew: A Decoder-only Foundation Model for Personalized Ranking and Recommendation
arXiv1 repoarXiv:2501.16450
linkedin-skills
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
arXiv1 repoarXiv:2501.16937
TAID
AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
arXiv1 repoarXiv:2501.17148
MidSteer
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
arXiv1 repoarXiv:2501.17161
SFTvsRL_Data
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
arXiv1 repoarXiv:2501.17433
Virus
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
arXiv1 repoarXiv:2501.17811
Janus-Pro-7B
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
arXiv1 repoarXiv:2501.18324
VDAct
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
arXiv1 repoarXiv:2501.18427
SANA1.5_1.6B_1024px
ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation
arXiv1 repoarXiv:2501.18460
ExeCoder
Track-On: Transformer-based Online Point Tracking with Memory
arXiv1 repoarXiv:2501.18487
track_on
R.I.P.: Better Models by Survival of the Fittest Prompts
arXiv1 repoarXiv:2501.18578
fairseq2
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
arXiv1 repoarXiv:2501.18585
unlazy
arXiv:2501.18593
arXiv1 repoarXiv:2501.18593
Jazz
Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models
arXiv1 repoarXiv:2501.18877
des
Visual Autoregressive Modeling for Image Super-Resolution
arXiv1 repoarXiv:2501.18993
VARSR
Enabling Autonomic Microservice Management through Self-Learning Agents
arXiv1 repoarXiv:2501.19056
ACV
Scalable-Softmax Is Superior for Attention
arXiv1 repoarXiv:2501.19399
Devstral-Small-2-24B-Instruct-2512
Low-Rank Adapting Models for Sparse Autoencoders
arXiv1 repoarXiv:2501.19406
sae_kl_finetune
arXiv:2502.00055
arXiv1 repoarXiv:2502.00055
awesome-ai-sre
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
arXiv1 repoarXiv:2502.00203
Llama-3_3-Nemotron-Super-49B-v1_5
RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
arXiv1 repoarXiv:2502.00392
PhysicalAI-VANTAGE-Bench
arXiv:2502.00640
arXiv1 repoarXiv:2502.00640
collabllm
COVE: COntext and VEracity prediction for out-of-context images
arXiv1 repoarXiv:2502.01194
5pils
On Almost Surely Safe Alignment of Large Language Models at Inference-Time
arXiv1 repoarXiv:2502.01208
inf-guard
Preference Leakage: A Contamination Problem in LLM-as-a-judge
arXiv1 repoarXiv:2502.01534
Llama-Krikri-8B-Instruct-GGUF
Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding
arXiv1 repoarXiv:2502.01563
Rope_with_LLM
SliderSpace: Decomposing the Visual Capabilities of Diffusion Models
arXiv1 repoarXiv:2502.01639
sliderspace
arXiv:2502.01651
arXiv1 repoarXiv:2502.01651
llama2.mojo
Layer by Layer: Uncovering Hidden Representations in Language Models
arXiv1 repoarXiv:2502.02013
information_flow
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
arXiv1 repoarXiv:2502.02589
coconut_pancap
Reconstructing 3D Flow from 2D Data with Diffusion Transformer
arXiv1 repoarXiv:2502.02593
graphify
arXiv:2502.02716
arXiv1 repoarXiv:2502.02716
drowse
arXiv:2502.03052
arXiv1 repoarXiv:2502.03052
dlm-jailbreak-transfer
CARROT: A Cost Aware Rate Optimal Router
arXiv1 repoarXiv:2502.03261
router
High-Fidelity Simultaneous Speech-To-Speech Translation
arXiv1 repoarXiv:2502.03382
tts-1.6b-en_fr
Pre-training Epidemic Time Series Forecasters with Compartmental Prototypes
arXiv1 repoarXiv:2502.03393
CAPE
Do Large Language Model Benchmarks Test Reliability?
arXiv1 repoarXiv:2502.03461
mmlu-redux
Efficient Image Restoration via Latent Consistency Flow Matching
arXiv1 repoarXiv:2502.03500
ELIR
DynVFX: Augmenting Real Videos with Dynamic Content
arXiv1 repoarXiv:2502.03621
dynvfx
arXiv:2502.03979
arXiv1 repoarXiv:2502.03979
Music2Emotion
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
arXiv1 repoarXiv:2502.04320
ConceptAttention
When One LLM Drools, Multi-LLM Collaboration Rules
arXiv1 repoarXiv:2502.04506
model_collaboration
Fast Video Generation with Sliding Tile Attention
arXiv1 repoarXiv:2502.04507
FastVideo
Towards Cost-Effective Reward Guided Text Generation
arXiv1 repoarXiv:2502.04517
FaRMA
arXiv:2502.04522
arXiv1 repoarXiv:2502.04522
improvnet
Sparsity-Based Interpolation of External, Internal and Swap Regret
arXiv1 repoarXiv:2502.04543
Tiny-R2
EigenLoRAx: Recycling Adapters to Find Principal Subspaces for Resource-Efficient Adaptation and Inference
arXiv1 repoarXiv:2502.04700
EigenLoRA
Chest X-ray Foundation Model with Global and Local Representations Integration
arXiv1 repoarXiv:2502.05142
CheXFound
Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency
arXiv1 repoarXiv:2502.05317
mlx-vit-tune
APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding
arXiv1 repoarXiv:2502.05431
APE
Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models
arXiv1 repoarXiv:2502.05945
targeted_intervention
VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer
arXiv1 repoarXiv:2502.05979
Omni-Effects
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
arXiv1 repoarXiv:2502.06737
VersaPRM
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
arXiv1 repoarXiv:2502.06773
OpenRLHF
Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content
arXiv1 repoarXiv:2502.07138
Video-vs-Meme-Hate
GENERator: A Long-Context Generative Genomic Foundation Model
arXiv1 repoarXiv:2502.07272
carbon-pretraining-corpus
Enhance-A-Video: Better Generated Video for Free
arXiv1 repoarXiv:2502.07508
Enhance-A-Video
Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting
arXiv1 repoarXiv:2502.07608
time2lang
TransMLA: Multi-Head Latent Attention Is All You Need
arXiv1 repoarXiv:2502.07864
TransArch
Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
arXiv1 repoarXiv:2502.07963
MedLitSpin
Training Sparse Mixture Of Experts Text Embedding Models
arXiv1 repoarXiv:2502.07972
nomic-embed-text-v2-moe
Light-A-Video: Training-free Video Relighting via Progressive Light Fusion
arXiv1 repoarXiv:2502.08590
Light-A-Video
Harnessing Vision Models for Time Series Analysis: A Survey
arXiv1 repoarXiv:2502.08869
TS-RAG
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
arXiv1 repoarXiv:2502.08910
hip-attention
CRANE: Reasoning with constrained LLM generation
arXiv1 repoarXiv:2502.09061
CRANE
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
arXiv1 repoarXiv:2502.09674
LRP-eXplains-Transformers
Precise Parameter Localization for Textual Generation in Diffusion Models
arXiv1 repoarXiv:2502.09935
t2i-text-localization
STAR: Spectral Truncation and Rescale for Model Merging
arXiv1 repoarXiv:2502.10339
ComfyUI-LoRA-Optimizer
ReStyle3D: Scene-Level Appearance Transfer with Semantic Correspondences
arXiv1 repoarXiv:2502.10377
ReStyle3D
arXiv:2502.10385
arXiv1 repoarXiv:2502.10385
hashing-baseline
Region-Adaptive Sampling for Diffusion Transformers
arXiv1 repoarXiv:2502.10389
RAS
I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models
arXiv1 repoarXiv:2502.10458
ThinkDiff
D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security
arXiv1 repoarXiv:2502.10931
awesome-cybersecurity-agentic-ai
Investigating Language Preference of Multilingual RAG Systems
arXiv1 repoarXiv:2502.11175
LanguagePreference
MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation
arXiv1 repoarXiv:2502.11234
maskflow
MARS: Mesh AutoRegressive Model for 3D Shape Detailization
arXiv1 repoarXiv:2502.11390
VAR
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
arXiv1 repoarXiv:2502.11494
DART
Continuous Diffusion Model for Language Modeling
arXiv1 repoarXiv:2502.11564
dlms-sinks
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
arXiv1 repoarXiv:2502.11598
watermarks-remover
HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims
arXiv1 repoarXiv:2502.11753
5pils
ChordFormer: A Conformer-Based Architecture for Large-Vocabulary Audio Chord Recognition
arXiv1 repoarXiv:2502.11840
SheetSage2
JoLT: Joint Probabilistic Predictions on Tabular Data Using LLMs
arXiv1 repoarXiv:2502.11877
jolt
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
arXiv1 repoarXiv:2502.11880
BitNet
A-MEM: Agentic Memory for LLM Agents
arXiv1 repoarXiv:2502.12110
ai-memory
Idiosyncrasies in Large Language Models
arXiv1 repoarXiv:2502.12150
llm-idiosyncrasies
Diffusion Models without Classifier-free Guidance
arXiv1 repoarXiv:2502.12154
classifier-free-guidance-pytorch
Independence Tests for Language Models
arXiv1 repoarXiv:2502.12292
model-tracing
YOLOv12: Attention-Centric Real-Time Object Detectors
arXiv1 repoarXiv:2502.12524
DINOV3-YOLOV12
CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image
arXiv1 repoarXiv:2502.12894
CAST
SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation
arXiv1 repoarXiv:2502.13128
SongGen
AIDE: AI-Driven Exploration in the Space of Code
arXiv1 repoarXiv:2502.13138
aideml
REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation
arXiv1 repoarXiv:2502.13270
REALTALK
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
arXiv1 repoarXiv:2502.13383
DataFlow
GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization
arXiv1 repoarXiv:2502.13398
BioMedGPT-Mol
Medical Image Classification with KAN-Integrated Transformers and Dilated Neighborhood Attention
arXiv1 repoarXiv:2502.13693
medmnistc-api
GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking
arXiv1 repoarXiv:2502.13766
gimmick
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
arXiv1 repoarXiv:2502.13989
erase-eval
Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
arXiv1 repoarXiv:2502.14044
FairLLaVA
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
arXiv1 repoarXiv:2502.14282
MobileAgent
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
arXiv1 repoarXiv:2502.14420
ChatVLA_public
Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis
arXiv1 repoarXiv:2502.14767
tree-of-debate
Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
arXiv1 repoarXiv:2502.14768
OpenRLHF
arXiv:2502.14866
arXiv1 repoarXiv:2502.14866
Block-Sparse-Attention
PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths
arXiv1 repoarXiv:2502.14902
sweet-search
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
arXiv1 repoarXiv:2502.14925
CodePromptZip-Token-Pruning
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models
arXiv1 repoarXiv:2502.15451
minimind
KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation
arXiv1 repoarXiv:2502.15602
kadtk
Learning to Reason from Feedback at Test-Time
arXiv1 repoarXiv:2502.15771
FTTT
arXiv:2502.15814
arXiv1 repoarXiv:2502.15814
slam_scaled
A Close Look at Decomposition-based XAI-Methods for Transformer Language Models
arXiv1 repoarXiv:2502.15886
LRP-eXplains-Transformers
arXiv:2502.15964
arXiv1 repoarXiv:2502.15964
minions
Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare
arXiv1 repoarXiv:2502.16051
mentat
arXiv:2502.16681
arXiv1 repoarXiv:2502.16681
SAEBench
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
arXiv1 repoarXiv:2502.17173
CheemsRM
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
arXiv1 repoarXiv:2502.17254
reinforce-attacks-llms
Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts
arXiv1 repoarXiv:2502.17297
RARE
Delta Decompression for MoE-based LLMs Compression
arXiv1 repoarXiv:2502.17298
D2MoE
On Relation-Specific Neurons in Large Language Models
arXiv1 repoarXiv:2502.17355
relation-specific-neurons
ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis
arXiv1 repoarXiv:2502.17475
minimind
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
arXiv1 repoarXiv:2502.18017
ViDoSeek
TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning
arXiv1 repoarXiv:2502.18431
SmallPlan
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
arXiv1 repoarXiv:2502.18460
dpr-scale
SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation
arXiv1 repoarXiv:2502.18793
SolEval
FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting
arXiv1 repoarXiv:2502.18834
FinTSB
LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm
arXiv1 repoarXiv:2502.19103
LongEval
CritiQ: Mining Data Quality Criteria from Human Preferences
arXiv1 repoarXiv:2502.19279
CritiQ
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
arXiv1 repoarXiv:2502.19400
TheoremExplainAgent
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
arXiv1 repoarXiv:2502.19412
helm
Stay Focused: Problem Drift in Multi-Agent Debate
arXiv1 repoarXiv:2502.19559
rightmind
Self-rewarding correction for mathematical reasoning
arXiv1 repoarXiv:2502.19613
verl-pipeline
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
arXiv1 repoarXiv:2502.19634
MedVLM-R1
Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation
arXiv1 repoarXiv:2502.20056
MLRG
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts
arXiv1 repoarXiv:2502.20395
R2-T2
Protecting multimodal large language models against misleading visualizations
arXiv1 repoarXiv:2502.20503
acl2026-misleading-visualizations
Autoregressive Medical Image Segmentation via Next-Scale Mask Prediction
arXiv1 repoarXiv:2502.20784
VAR
Adaptive Keyframe Sampling for Long Video Understanding
arXiv1 repoarXiv:2502.21271
CRAFT
Decoupling Content and Expression: Two-Dimensional Detection of AI-Generated Text
arXiv1 repoarXiv:2503.00258
truth-mirror
Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models
arXiv1 repoarXiv:2503.00743
ScoreRS
DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting
arXiv1 repoarXiv:2503.00784
DuoDecoding
TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop
arXiv1 repoarXiv:2503.01013
TS-RAG
DLF: Extreme Image Compression with Dual-generative Latent Fusion
arXiv1 repoarXiv:2503.01428
Dual-generative-Latent-Fusion
Effective High-order Graph Representation Learning for Credit Card Fraud Detection
arXiv1 repoarXiv:2503.01556
antifraud
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
arXiv1 repoarXiv:2503.01623
localmod
Detecting Stylistic Fingerprints of Large Language Models
arXiv1 repoarXiv:2503.01659
humanizer
Elliptic Loss Regularization
arXiv1 repoarXiv:2503.02138
Qwen3.6-VL-REAP-26B-A3B
Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression
arXiv1 repoarXiv:2503.02812
qfilters
Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation
arXiv1 repoarXiv:2503.03492
FindTrack
Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models
arXiv1 repoarXiv:2503.03669
parlant
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
arXiv1 repoarXiv:2503.04647
Implicit-Cross-Lingual-Rewarding
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
arXiv1 repoarXiv:2503.04691
MedRBench
Generating Millions Of Lean Theorems With Proofs By Exploring State Transition Graphs
arXiv1 repoarXiv:2503.04772
re-rl
FedMABench: Benchmarking Mobile Agents on Decentralized Heterogeneous User Data
arXiv1 repoarXiv:2503.05143
FedMABench
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
arXiv1 repoarXiv:2503.05255
CMMCoT
R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
arXiv1 repoarXiv:2503.05592
SWE-Master
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills
arXiv1 repoarXiv:2503.05641
zen5
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
arXiv1 repoarXiv:2503.05683
WikiBigEdit
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
arXiv1 repoarXiv:2503.05840
transformer-tricks
Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning
arXiv1 repoarXiv:2503.06034
llm-rankers
X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation
arXiv1 repoarXiv:2503.06134
X2I
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
arXiv1 repoarXiv:2503.06749
Vision-R1
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
arXiv1 repoarXiv:2503.06923
ToCa
Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition
arXiv1 repoarXiv:2503.06984
MelQCD-main
PE3R: Perception-Efficient 3D Reconstruction
arXiv1 repoarXiv:2503.07507
PE3R
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
arXiv1 repoarXiv:2503.07649
TS-RAG
OminiControl2: Efficient Conditioning for Diffusion Transformers
arXiv1 repoarXiv:2503.08280
OminiControl
Referring to Any Person
arXiv1 repoarXiv:2503.08507
Rex-Omni
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
arXiv1 repoarXiv:2503.08679
CoTFaithChecker
Seal Your Backdoor with Variational Defense
arXiv1 repoarXiv:2503.08829
VIBE
Optimal Control of Medical Drug in a Nonlocal Model of Solid Tumor Growth
arXiv1 repoarXiv:2503.09208
sarashina2.2-ocr
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
arXiv1 repoarXiv:2503.09402
VLog
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
arXiv1 repoarXiv:2503.09595
pisa-experiments
CASteer: Cross-Attention Steering for Controllable Concept Erasure
arXiv1 repoarXiv:2503.09630
CASteer
V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video
arXiv1 repoarXiv:2503.09631
V2M4
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
arXiv1 repoarXiv:2503.09780
ai-agent-privacy
Faster Inference of LLMs using FP8 on the Intel Gaudi
arXiv1 repoarXiv:2503.09975
neural-compressor
MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion
arXiv1 repoarXiv:2503.10289
MaterialMVP
Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
arXiv1 repoarXiv:2503.10460
Light-R1
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1
arXiv1 repoarXiv:2503.10635
M-Attack_AdvSamples
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
arXiv1 repoarXiv:2503.10673
ZeroSumEval
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
arXiv1 repoarXiv:2503.10679
ml-lineas
Neighboring Autoregressive Modeling for Efficient Visual Generation
arXiv1 repoarXiv:2503.10696
NAR
FlowTok: Flowing Seamlessly Across Text and Image Tokens
arXiv1 repoarXiv:2503.10772
clustermark_1d-tokenizer
TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools
arXiv1 repoarXiv:2503.10970
ToolUniverse
MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens
arXiv1 repoarXiv:2503.11315
MMS-LLaMA
Safe-VAR: Safe Visual Autoregressive Model for Text-to-Image Generative Watermarking
arXiv1 repoarXiv:2503.11324
VAR
Efficient Distributed MLLM Training with Cornstarch
arXiv1 repoarXiv:2503.11367
Cornstarch
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
arXiv1 repoarXiv:2503.11647
MultiCamVideo-Dataset
VGGT: Visual Geometry Grounded Transformer
arXiv1 repoarXiv:2503.11651
VGGT-1B
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
arXiv1 repoarXiv:2503.11832
Unlearn-Trace
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
arXiv1 repoarXiv:2503.12559
ACL25-AdaReTaKe
Reliable and Efficient Amortized Model-based Evaluation
arXiv1 repoarXiv:2503.13335
helm
xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference
arXiv1 repoarXiv:2503.13427
xlstm
Why Do Multi-Agent LLM Systems Fail?
arXiv1 repoarXiv:2503.13657
prisma
Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
arXiv1 repoarXiv:2503.13939
Med-R1
MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
arXiv1 repoarXiv:2503.13964
Mdocagent-dataset
Inference-Time Intervention in Large Language Models for Reliable Requirement Verification
arXiv1 repoarXiv:2503.14130
targeted_intervention
MoonCast: High-Quality Zero-Shot Podcast Generation
arXiv1 repoarXiv:2503.14345
MoonCast
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
arXiv1 repoarXiv:2503.14350
VEGGIE-VidEdit
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
arXiv1 repoarXiv:2503.14492
cosmos-transfer1
Measuring AI Ability to Complete Long Software Tasks
arXiv1 repoarXiv:2503.14499
unlazy
MusicInfuser: Making Video Diffusion Listen and Dance
arXiv1 repoarXiv:2503.14505
MusicInfuser
VisNumBench: Evaluating Number Sense of Multimodal Large Language Models
arXiv1 repoarXiv:2503.14939
mllm_number_sense
Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator
arXiv1 repoarXiv:2503.15457
test-time-scaling
Object-Spatial Programming
arXiv1 repoarXiv:2503.15812
jac
Survey on Evaluation of LLM-based Agents
arXiv1 repoarXiv:2503.16416
StepDeepResearch
XAttention: Block Sparse Attention with Antidiagonal Scoring
arXiv1 repoarXiv:2503.16428
RetrievalAttention
Sonata: Self-Supervised Learning of Reliable Point Representations
arXiv1 repoarXiv:2503.16429
sonata
arXiv:2503.16611
arXiv1 repoarXiv:2503.16611
PanoramaGenInpaint
Aligning Text-to-Music Evaluation with Human Preferences
arXiv1 repoarXiv:2503.16669
TuneJury
Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need
arXiv1 repoarXiv:2503.17272
sae_kl_finetune
Won: Establishing Best Practices for Korean Financial NLP
arXiv1 repoarXiv:2503.17963
Won-Instruct
PolarFree: Polarization-based Reflection-free Imaging
arXiv1 repoarXiv:2503.18055
PolarFree
Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures
arXiv1 repoarXiv:2503.18565
distil_xlstm
Any6D: Model-free 6D Pose Estimation of Novel Objects
arXiv1 repoarXiv:2503.18673
Any6D
CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models
arXiv1 repoarXiv:2503.18886
CFG-Zero-Star
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
arXiv1 repoarXiv:2503.18943
ml-unigen
Aether: Geometric-Aware Unified World Modeling
arXiv1 repoarXiv:2503.18945
AetherV1
Understanding and Improving Information Preservation in Prompt Compression for LLMs
arXiv1 repoarXiv:2503.19114
information-preservation-in-prompt-compression
Scaling Down Text Encoders of Text-to-Image Diffusion Models
arXiv1 repoarXiv:2503.19897
DistillT5
Scaling Vision Pre-Training to 4K Resolution
arXiv1 repoarXiv:2503.19903
Robopoint_Humble
Audio-centric Video Understanding Benchmark without Text Shortcut
arXiv1 repoarXiv:2503.19951
AVUTBenchmark
RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy
arXiv1 repoarXiv:2503.20158
rxrx3-core
UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models
arXiv1 repoarXiv:2503.20504
VASE
Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models
arXiv1 repoarXiv:2503.20752
RoboBrain2.5
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields
arXiv1 repoarXiv:2503.20776
Feature4X
BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology
arXiv1 repoarXiv:2503.20880
BioX-CPath
ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation
arXiv1 repoarXiv:2503.21729
ReaRAG-20k
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
arXiv1 repoarXiv:2503.21755
VBench
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
arXiv1 repoarXiv:2503.21758
Lumina-Image-2.0
PharMolixFM: All-Atom Foundation Models for Molecular Modeling and Generation
arXiv1 repoarXiv:2503.21788
OpenBioMed
WMCopier: Forging Invisible Image Watermarks on Arbitrary Images
arXiv1 repoarXiv:2503.22330
WMCopier
Text-Only Data Synthesis for Vision Language Model Training
arXiv1 repoarXiv:2503.22655
Modality_Gap_Theory
From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D
arXiv1 repoarXiv:2503.22976
SPAR-Bench
Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
arXiv1 repoarXiv:2503.23278
awesome-mcp-security
VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
arXiv1 repoarXiv:2503.23368
VLIPP
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
arXiv1 repoarXiv:2503.23907
HumanAesExpert
AI2Agent: An End-to-End Framework for Deploying AI Projects as Autonomous Agents
arXiv1 repoarXiv:2503.23948
ai2apps
A Multi-Stage Auto-Context Deep Learning Framework for Tissue and Nuclei Segmentation and Classification in H&E-Stained Histological Images of Advanced Melanoma
arXiv1 repoarXiv:2503.23958
PumaSubmit
Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
arXiv1 repoarXiv:2503.24290
Open-Reasoner-Zero
Easi3R: Estimating Disentangled Motion from DUSt3R Without Training
arXiv1 repoarXiv:2503.24391
Easi3R
Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge
arXiv1 repoarXiv:2504.00042
knowledge-gap
Universal Zero-shot Embedding Inversion
arXiv1 repoarXiv:2504.00147
adversarial_decoding
An Illusion of Progress? Assessing the Current State of Web Agents
arXiv1 repoarXiv:2504.01382
UI-TARS
Diffusion-Guided Gaussian Splatting for Large-Scale Unconstrained 3D Reconstruction and Novel View Synthesis
arXiv1 repoarXiv:2504.01960
SphereForge
T*: Re-thinking Temporal Search for Long-Form Video Understanding
arXiv1 repoarXiv:2504.02259
LongVideoHaystack
Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation
arXiv1 repoarXiv:2504.02438
ViLAMP-llava-qwen_sig
Exploration-Driven Generative Interactive Environments
arXiv1 repoarXiv:2504.02515
stable-retro
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
arXiv1 repoarXiv:2504.02904
post-training-mechanistic-analysis
Noiser: Bounded Input Perturbations for Attributing Large Language Models
arXiv1 repoarXiv:2504.02911
Noiser
MedSAM2: Segment Anything in 3D Medical Images and Videos
arXiv1 repoarXiv:2504.03600
MedSAM2
MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits
arXiv1 repoarXiv:2504.03767
awesome-mcp-security
Clinical ModernBERT: An efficient and long context encoder for biomedical text
arXiv1 repoarXiv:2504.03964
DiagnosisCoding
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
arXiv1 repoarXiv:2504.04423
UniToken
VSLAM-LAB: A Comprehensive Framework for Visual SLAM Methods and Datasets
arXiv1 repoarXiv:2504.04457
VSLAM-LAB
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
arXiv1 repoarXiv:2504.04715
llm-api-audit
FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
arXiv1 repoarXiv:2504.04842
FantasyTalking
M-Prometheus: A Suite of Open Multilingual LLM Judges
arXiv1 repoarXiv:2504.04953
prometheus-eval
SmolVLM: Redefining small and efficient multimodal models
arXiv1 repoarXiv:2504.05299
SmolVLM2-2.2B-Instruct
Gaussian Mixture Flow Matching Models
arXiv1 repoarXiv:2504.05304
LakonLab
STAGE: Stemmed Accompaniment Generation through Prefix-Based Conditioning
arXiv1 repoarXiv:2504.05690
stage
Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
arXiv1 repoarXiv:2504.05812
EMPO
HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
arXiv1 repoarXiv:2504.05897
HybriMoE
AEGIS: Human Attention-based Explainable Guidance for Intelligent Vehicle Systems
arXiv1 repoarXiv:2504.05950
AEGIS
CAI: An Open, Bug Bounty-Ready Cybersecurity AI
arXiv1 repoarXiv:2504.06017
awesome-cybersecurity-agentic-ai
QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform
arXiv1 repoarXiv:2504.06136
qgen-studio
Large language models as uncertainty-calibrated optimizers for experimental discovery
arXiv1 repoarXiv:2504.06265
sego
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
arXiv1 repoarXiv:2504.06792
EASYEP
OSCAR: Online Soft Compression And Reranking
arXiv1 repoarXiv:2504.07109
pisco
Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction
arXiv1 repoarXiv:2504.07375
UniHand
Malware analysis assisted by AI with R2AI
arXiv1 repoarXiv:2504.07574
r2ai
BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation
arXiv1 repoarXiv:2504.07955
BoxDreamer
Detect Anything 3D in the Wild
arXiv1 repoarXiv:2504.07958
DetAny3D
Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction
arXiv1 repoarXiv:2504.07961
Geo4D
The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
arXiv1 repoarXiv:2504.08066
aideml
Out of Style: RAG's Fragility to Linguistic Variation
arXiv1 repoarXiv:2504.08231
RAG-fragility-to-linguistic-variation
Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies
arXiv1 repoarXiv:2504.08623
awesome-mcp-security
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
arXiv1 repoarXiv:2504.08850
SpecEE
LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping
arXiv1 repoarXiv:2504.08902
LatentGenerativeAnamorphoses
PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
arXiv1 repoarXiv:2504.08966
PACT
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
arXiv1 repoarXiv:2504.09710
DUMP
Reasoning Models Can Be Effective Without Thinking
arXiv1 repoarXiv:2504.09858
CoDE-Stop
Guiding Reasoning in Small Language Models with LLM Assistance
arXiv1 repoarXiv:2504.09923
SMART
Aligning Anime Video Generation with Human Feedback
arXiv1 repoarXiv:2504.10044
Index-anisora
RealHarm: A Collection of Real-World Language Model Application Failures
arXiv1 repoarXiv:2504.10277
realharm
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
arXiv1 repoarXiv:2504.10445
RealWebAssist
Efficient Process Reward Model Training via Active Learning
arXiv1 repoarXiv:2504.10559
ActivePRM
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
arXiv1 repoarXiv:2504.10766
Gradient_Unified
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer
arXiv1 repoarXiv:2504.11289
UniAnimate-DiT
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
arXiv1 repoarXiv:2504.11456
DeepMath-103K
Activated LoRA: Fine-tuned LLMs for Intrinsics
arXiv1 repoarXiv:2504.12397
granitelib-rag-r1.0
One Model to Rig Them All: Diverse Skeleton Rigging with UniRig
arXiv1 repoarXiv:2504.12451
UniRig
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
arXiv1 repoarXiv:2504.12562
ZeroSumEval
Chinese-Vicuna: A Chinese Instruction-following Llama-based Model
arXiv1 repoarXiv:2504.12737
Chinese-Vicuna
MCP Guardian: A Security-First Layer for Safeguarding MCP-Based AI System
arXiv1 repoarXiv:2504.12757
awesome-mcp-security
Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts
arXiv1 repoarXiv:2504.12782
MACE
SoK: Security of EMV Contactless Payment Systems
arXiv1 repoarXiv:2504.12812
awesome-connected-things-sec
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
arXiv1 repoarXiv:2504.13059
RoboTwin
Long Range Navigator (LRN): Extending robot planning horizons beyond metric maps
arXiv1 repoarXiv:2504.13149
nebula2-wildos
Long-context Non-factoid Question Answering in Indic Languages
arXiv1 repoarXiv:2504.13615
IndicGenQA
Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation
arXiv1 repoarXiv:2504.14011
fashion-rag
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
arXiv1 repoarXiv:2504.14225
PersonaMem-v2
LoRe: Personalizing LLMs via Low-Rank Reward Modeling
arXiv1 repoarXiv:2504.14439
LoRe
BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation
arXiv1 repoarXiv:2504.14538
BookWorld
TAPIP3D: Tracking Any Point in Persistent 3D Geometry
arXiv1 repoarXiv:2504.14717
tapip3d
arXiv:2504.15071
arXiv1 repoarXiv:2504.15071
aria-medium-embedding
EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models
arXiv1 repoarXiv:2504.15133
WikiBigEdit
Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
arXiv1 repoarXiv:2504.15275
PURE
Event2Vec: Processing Neuromorphic Events Directly by Representations in Vector Space
arXiv1 repoarXiv:2504.15371
repro-event2vec-neuromorphic-events-vector-space
VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation
arXiv1 repoarXiv:2504.15659
VeriCoder
Dynamic Early Exit in Reasoning Models
arXiv1 repoarXiv:2504.15895
CoDE-Stop
LLMSR@XLLM25: Less is More: Enhancing Structured Multi-Agent Reasoning via Quality-Guided Distillation
arXiv1 repoarXiv:2504.16408
Less-is-More
MAGIC: Near-Optimal Data Attribution for Deep Learning
arXiv1 repoarXiv:2504.16430
bergson
Enhancing LLM-Based Agents via Global Planning and Hierarchical Execution
arXiv1 repoarXiv:2504.16563
agentic-awesome-skills
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
arXiv1 repoarXiv:2504.16727
Visual-Variations-Robustness
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
arXiv1 repoarXiv:2504.17207
APC-VLM
Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense Retrieval
arXiv1 repoarXiv:2504.17884
unsupervised_corpus_poisoning
Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
arXiv1 repoarXiv:2504.17950
mindcraft
Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface
arXiv1 repoarXiv:2504.18430
IRON
Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
arXiv1 repoarXiv:2504.19413
khms-memory
semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
arXiv1 repoarXiv:2504.19867
Semi-PD
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
arXiv1 repoarXiv:2504.19956
awesome-cybersecurity-agentic-ai
Simplified and Secure MCP Gateways for Enterprise AI Integration
arXiv1 repoarXiv:2504.19997
awesome-mcp-security
Learning Streaming Video Representation via Multitask Training
arXiv1 repoarXiv:2504.20041
streamformer-timesformer
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
arXiv1 repoarXiv:2504.20734
UniversalRAG
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
arXiv1 repoarXiv:2504.20938
circuit_backup
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math
arXiv1 repoarXiv:2504.21233
Phi-4-mini-reasoning
LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving
arXiv1 repoarXiv:2505.00284
LightEMMA
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
arXiv1 repoarXiv:2505.00703
Image-Generation-CoT
FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
arXiv1 repoarXiv:2505.01322
FreeInsert
Practical Efficiency of Muon for Pretraining
arXiv1 repoarXiv:2505.02222
Muon
Connecting Independently Trained Modes via Layer-Wise Connectivity
arXiv1 repoarXiv:2505.02604
repro-connecting-independently-trained-modes-layer-wise-low-loss-paths
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference
arXiv1 repoarXiv:2505.02922
RetrievalAttention
UCSC at SemEval-2025 Task 3: Context, Models and Prompt Optimization for Automated Hallucination Detection in LLM Output
arXiv1 repoarXiv:2505.03030
semeval-2025-task3
RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation
arXiv1 repoarXiv:2505.03275
gbrain
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
arXiv1 repoarXiv:2505.03673
RoboOS
CoCoB: Adaptive Collaborative Combinatorial Bandits for Online Recommendation
arXiv1 repoarXiv:2505.03840
graphify
Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
arXiv1 repoarXiv:2505.04021
prism-research
arXiv:2505.04080
arXiv1 repoarXiv:2505.04080
llama2.mojo
Multimodal Integrated Knowledge Transfer to Large Language Models through Preference Optimization with Biomedical Applications
arXiv1 repoarXiv:2505.05736
MINT-LLM
Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
arXiv1 repoarXiv:2505.06356
maya
Embedding Atlas: Low-Friction, Interactive Embedding Visualization
arXiv1 repoarXiv:2505.06386
embedding-atlas
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
arXiv1 repoarXiv:2505.07293
attention-influence
FLUXSynID: A Framework for Identity-Controlled Synthetic Face Generation with Document and Live Images
arXiv1 repoarXiv:2505.07530
FLUXSynID
RAI: Flexible Agent Framework for Embodied AI
arXiv1 repoarXiv:2505.07532
rai
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
arXiv1 repoarXiv:2505.07879
OMGM
Behind Maya: Building a Multilingual Vision Language Model
arXiv1 repoarXiv:2505.08910
maya
MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment
arXiv1 repoarXiv:2505.09372
MAKE
Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
arXiv1 repoarXiv:2505.09439
Omni-R1
A large-scale evaluation of commonsense knowledge in humans and large language models
arXiv1 repoarXiv:2505.10309
commonsense-llm-eval
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
arXiv1 repoarXiv:2505.10483
UniEval
MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly
arXiv1 repoarXiv:2505.10610
MMLongBench
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
arXiv1 repoarXiv:2505.10696
TartanGround
Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
arXiv1 repoarXiv:2505.10844
figures4papers
Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
arXiv1 repoarXiv:2505.11029
cr4wm
DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
arXiv1 repoarXiv:2505.11196
DiCo
Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models
arXiv1 repoarXiv:2505.11245
Diffusion-NPO
TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
arXiv1 repoarXiv:2505.11329
tokenweave
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
arXiv1 repoarXiv:2505.11475
Llama-3_3-Nemotron-Super-49B-GenRM
SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training
arXiv1 repoarXiv:2505.11594
SageAttention
Self-NPO: Data-Free Diffusion Model Enhancement via Truncated Diffusion Fine-Tuning
arXiv1 repoarXiv:2505.11777
Diffusion-NPO
TinyRS-R1: Compact Multimodal Language Model for Remote Sensing
arXiv1 repoarXiv:2505.12099
TinyRS
Estimation of Treatment Harm Rate via Partitioning
arXiv1 repoarXiv:2505.12209
zen5
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
arXiv1 repoarXiv:2505.12370
ScreenSpot-Pro-GUI-Grounding
Harnessing the Universal Geometry of Embeddings
arXiv1 repoarXiv:2505.12540
vec2vec
Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment
arXiv1 repoarXiv:2505.12669
t2m-inferalign
arXiv:2505.12674
arXiv1 repoarXiv:2505.12674
ml-sid-dit
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
arXiv1 repoarXiv:2505.13031
MindOmni
Thinkless: LLM Learns When to Think
arXiv1 repoarXiv:2505.13379
Thinkless-1.5B-Warmup
VSA: Faster Video Diffusion with Trainable Sparse Attention
arXiv1 repoarXiv:2505.13389
FastVideo
Krikri: Advancing Open Large Language Models for Greek
arXiv1 repoarXiv:2505.13772
Llama-Krikri-8B-Instruct-GGUF
Let's Verify Math Questions Step by Step
arXiv1 repoarXiv:2505.13903
DataFlow
MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem
arXiv1 repoarXiv:2505.14148
LLM-MM-Agent
Capturing the Effects of Quantization on Trojans in Code LLMs
arXiv1 repoarXiv:2505.14200
quantized-code-llm-security
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection
arXiv1 repoarXiv:2505.14318
Radar
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
arXiv1 repoarXiv:2505.14362
DeepEyes-7B
PRL: Prompts from Reinforcement Learning
arXiv1 repoarXiv:2505.14412
PRL-Prompts-from-Reinforcement-Learning
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
arXiv1 repoarXiv:2505.14454
VidCom2
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
arXiv1 repoarXiv:2505.14640
VideoChat-Flash
Quartet: Native FP4 Training Can Be Optimal for Large Language Models
arXiv1 repoarXiv:2505.14669
lectures
Reward Reasoning Model
arXiv1 repoarXiv:2505.14674
RRM-7B
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
arXiv1 repoarXiv:2505.14682
ml-unigen
Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing
arXiv1 repoarXiv:2505.14881
TrafficComposer
PiFlow: Principle-Aware Scientific Discovery with Multi-Agent Collaboration
arXiv1 repoarXiv:2505.15047
pigollum
lmgame-Bench: How Good are LLMs at Playing Games?
arXiv1 repoarXiv:2505.15146
GRL
R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization
arXiv1 repoarXiv:2505.15155
qlib
Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control
arXiv1 repoarXiv:2505.15304
sqil
MIRB: Mathematical Information Retrieval Benchmark
arXiv1 repoarXiv:2505.15585
mirb
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
arXiv1 repoarXiv:2505.15778
Soft-Thinking
arXiv:2505.16239
arXiv1 repoarXiv:2505.16239
DOVE
arXiv:2505.16369
arXiv1 repoarXiv:2505.16369
xares
LaViDa: A Large Diffusion Language Model for Multimodal Understanding
arXiv1 repoarXiv:2505.16839
LaViDa
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
arXiv1 repoarXiv:2505.16967
rlhn-680K
X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs
arXiv1 repoarXiv:2505.16997
rightmind
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
arXiv1 repoarXiv:2505.17017
Image-Generation-CoT
LLM Agents for Interactive Exploration of Historical Cadastre Data: Framework and Application to Venice
arXiv1 repoarXiv:2505.17148
venice-agents
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
arXiv1 repoarXiv:2505.17568
JALMBench
arXiv:2505.17598
arXiv1 repoarXiv:2505.17598
dlm-jailbreak-transfer
VIBE: Vector Index Benchmark for Embeddings
arXiv1 repoarXiv:2505.17810
vibe
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
arXiv1 repoarXiv:2505.17862
Daily-Omni
WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions
arXiv1 repoarXiv:2505.18151
WonderPlay
arXiv:2505.18186
arXiv1 repoarXiv:2505.18186
musicdiscovery
Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens
arXiv1 repoarXiv:2505.18237
CoDE-Stop
Dynamic Risk Assessments for Offensive Cybersecurity Agents
arXiv1 repoarXiv:2505.18384
awesome-cybersecurity-agentic-ai
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
arXiv1 repoarXiv:2505.18561
CoT-RVS
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
arXiv1 repoarXiv:2505.19037
Speech-IFEval
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
arXiv1 repoarXiv:2505.19075
UniR
Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data
arXiv1 repoarXiv:2505.19274
cadet-embed-base-v1
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation
arXiv1 repoarXiv:2505.19462
T5Gemma-TTS
Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks
arXiv1 repoarXiv:2505.19472
Parallel-Hybrid-Model
Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection
arXiv1 repoarXiv:2505.19475
VDS-TTT
STRAP: Spatio-Temporal Pattern Retrieval for Out-of-Distribution Generalization
arXiv1 repoarXiv:2505.19547
A2TTA
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
arXiv1 repoarXiv:2505.19586
TailorKV
Multi-Agent Collaboration via Evolving Orchestration
arXiv1 repoarXiv:2505.19591
ChatDev
ErpGS: Equirectangular Image Rendering enhanced with 3D Gaussian Regularization
arXiv1 repoarXiv:2505.19883
SphereForge
EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
arXiv1 repoarXiv:2505.20033
Voice-Acting-Pipeline
Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers
arXiv1 repoarXiv:2505.20128
SearchLM
syftr: Pareto-Optimal Generative AI
arXiv1 repoarXiv:2505.20266
syftr
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
arXiv1 repoarXiv:2505.20279
VLM-3R
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
arXiv1 repoarXiv:2505.20322
data_for_STA
SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
arXiv1 repoarXiv:2505.20411
SWE-rebench
DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data
arXiv1 repoarXiv:2505.20460
EmbodiedGen
RefAV: Towards Planning-Centric Scenario Mining
arXiv1 repoarXiv:2505.20981
RefAV
Who Reasons in the Large Language Models?
arXiv1 repoarXiv:2505.20993
SpaceOm
SageAttention2++: A More Efficient Implementation of SageAttention2
arXiv1 repoarXiv:2505.21136
SageAttention
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
arXiv1 repoarXiv:2505.21374
Video-Holmes
QuARI: Query Adaptive Retrieval Improvement
arXiv1 repoarXiv:2505.21647
QuARI
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
arXiv1 repoarXiv:2505.21906
ChatVLA_public
TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation
arXiv1 repoarXiv:2505.22176
table-metric-study
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
arXiv1 repoarXiv:2505.22179
FR-Spec
Inference-Time Scaling of Discrete Diffusion Models via Importance Weighting and Optimal Proposal Design
arXiv1 repoarXiv:2505.22524
smc_ddm_iclr
Thinking with Generated Images
arXiv1 repoarXiv:2505.22525
thinking-with-generated-images
A Tool for Generating Exceptional Behavior Tests With Large Language Models
arXiv1 repoarXiv:2505.22818
exLong
RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation
arXiv1 repoarXiv:2505.22846
coqpilot
From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrieval
arXiv1 repoarXiv:2505.23059
SMR
Speeding up Model Loading with fastsafetensors
arXiv1 repoarXiv:2505.23072
fastsafetensors
Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning
arXiv1 repoarXiv:2505.23195
Prune-then-Finetune
Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization
arXiv1 repoarXiv:2505.23387
Venus
KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
arXiv1 repoarXiv:2505.23416
KVzip
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
arXiv1 repoarXiv:2505.23606
Muddit
Inference-time Scaling of Diffusion Models through Classical Search
arXiv1 repoarXiv:2505.23614
Diffusion-inference-scaling
AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora
arXiv1 repoarXiv:2505.23628
AutoSchemaKG
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
arXiv1 repoarXiv:2505.23656
VideoREPA
EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast
arXiv1 repoarXiv:2505.23732
EmotionRankCLAP
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
arXiv1 repoarXiv:2505.23885
owl
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
arXiv1 repoarXiv:2505.24063
TCM-Ladder
Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning
arXiv1 repoarXiv:2505.24478
cognee
arXiv:2505.24685
arXiv1 repoarXiv:2505.24685
misp-galaxy
LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text
arXiv1 repoarXiv:2505.24826
minimind
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
arXiv1 repoarXiv:2505.24873
minimax-remover
ACE-Step: A Step Towards Music Generation Foundation Model
arXiv1 repoarXiv:2506.00045
ACE-Step
Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis
arXiv1 repoarXiv:2506.00433
LatentWaveletDiffusion
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
arXiv1 repoarXiv:2506.00519
CausalAbstain
Learning with Calibration: Exploring Test-Time Computing of Spatio-Temporal Forecasting
arXiv1 repoarXiv:2506.00635
A2TTA
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
arXiv1 repoarXiv:2506.00975
NTPP
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
arXiv1 repoarXiv:2506.01192
GigaAM
DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing
arXiv1 repoarXiv:2506.01430
FlowEdit
Policy as Code, Policy as Type
arXiv1 repoarXiv:2506.01446
extensible-mcp
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
arXiv1 repoarXiv:2506.01844
smolvla_base
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
arXiv1 repoarXiv:2506.01953
Fast-in-Slow
Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem
arXiv1 repoarXiv:2506.02040
awesome-mcp-security
Answer Convergence as a Signal for Early Stopping in Reasoning
arXiv1 repoarXiv:2506.02536
CoDE-Stop
Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
arXiv1 repoarXiv:2506.02557
KUEA
Solving Inverse Problems with FLAIR
arXiv1 repoarXiv:2506.02680
FLAIR
Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
arXiv1 repoarXiv:2506.02738
open-pmc-18m
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
arXiv1 repoarXiv:2506.03096
fuselip
Native-Resolution Image Synthesis
arXiv1 repoarXiv:2506.03131
NiT-diffusers
Test-Time Scaling of Diffusion Models via Noise Trajectory Search
arXiv1 repoarXiv:2506.03164
diffusion-tts
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
arXiv1 repoarXiv:2506.03355
LEAF
Seed-Coder: Let the Code Model Curate Data for Itself
arXiv1 repoarXiv:2506.03524
swallow-code-v2
MiMo-VL Technical Report
arXiv1 repoarXiv:2506.03569
MiMo-VL-7B-SFT-2508
Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
arXiv1 repoarXiv:2506.03610
Orak
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
arXiv1 repoarXiv:2506.03828
AssetOpsBench
Pre$^3$: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation
arXiv1 repoarXiv:2506.03887
lightllm
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
arXiv1 repoarXiv:2506.04034
Rex-Omni
EuroLLM-9B: Technical Report
arXiv1 repoarXiv:2506.04079
AMALIA-9B-0626-DPO
OpenThoughts: Data Recipes for Reasoning Models
arXiv1 repoarXiv:2506.04178
OpenThoughts-114k
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
arXiv1 repoarXiv:2506.04207
Revisual-R1
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
arXiv1 repoarXiv:2506.04225
HunyuanWorld-Voyager
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
arXiv1 repoarXiv:2506.04363
WorldPrediction
Identity Testing for Circuits with Exponentiation Gates
arXiv1 repoarXiv:2506.04529
mirage
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
arXiv1 repoarXiv:2506.04598
CLIP_benchmark
Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
arXiv1 repoarXiv:2506.04614
MobileAgent
ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning
arXiv1 repoarXiv:2506.04941
ArtVIP
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
arXiv1 repoarXiv:2506.05140
AudioLens
arXiv:2506.05301
arXiv1 repoarXiv:2506.05301
SeedVR2-3B
arXiv:2506.05414
arXiv1 repoarXiv:2506.05414
SAVVY-Bench
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
arXiv1 repoarXiv:2506.05551
MLLM-Semantic-Hallucination
BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading
arXiv1 repoarXiv:2506.06271
becominglit
Memory OS of AI Agent
arXiv1 repoarXiv:2506.06326
MemoryOS
Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict
arXiv1 repoarXiv:2506.06485
LLM-KnowledgeConflict-TaskMatters
Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
arXiv1 repoarXiv:2506.07001
Adversarial-Paraphrasing
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
arXiv1 repoarXiv:2506.07468
selfplay-redteaming
R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation
arXiv1 repoarXiv:2506.07826
R3D2
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
arXiv1 repoarXiv:2506.08009
Self-Forcing
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
arXiv1 repoarXiv:2506.08052
recogdrive_tome
LEANN: A Low-Storage Vector Index
arXiv1 repoarXiv:2506.08276
LEANN
arXiv:2506.08641
arXiv1 repoarXiv:2506.08641
TiViT
Edit Flows: Flow Matching with Edit Operations
arXiv1 repoarXiv:2506.09018
dllm
CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
arXiv1 repoarXiv:2506.09109
CAIRE
Seedance 1.0: Exploring the Boundaries of Video Generation Models
arXiv1 repoarXiv:2506.09113
Awesome-AITools
TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval
arXiv1 repoarXiv:2506.09114
TRACE-TimeseriesRAG-Dataset
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
arXiv1 repoarXiv:2506.09513
ReasonMed
MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions
arXiv1 repoarXiv:2506.09556
medusa
Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks
arXiv1 repoarXiv:2506.09593
medmnistc-api
Query-Level Uncertainty in Large Language Models
arXiv1 repoarXiv:2506.09669
query_level_uncertainty
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
arXiv1 repoarXiv:2506.09736
Vision-Matters
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
arXiv1 repoarXiv:2506.09985
vjepa2
A quantum semantic framework for natural language processing
arXiv1 repoarXiv:2506.10077
npcpy
SoK: Evaluating Jailbreak Guardrails for Large Language Models
arXiv1 repoarXiv:2506.10597
JailbreakGuardrailBenchmark
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
arXiv1 repoarXiv:2506.10600
EmbodiedGen
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
arXiv1 repoarXiv:2506.10622
sdialog
Skillful joint probabilistic weather forecasting from marginals
arXiv1 repoarXiv:2506.10772
weathernext
CyclicReflex: Improving Reasoning Models via Cyclical Reflection Token Scheduling
arXiv1 repoarXiv:2506.11077
CyclicReflex
Fed-HeLLo: Efficient Federated Foundation Model Fine-Tuning with Heterogeneous LoRA Allocation
arXiv1 repoarXiv:2506.12213
Fed-PLoRA
Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech
arXiv1 repoarXiv:2506.12311
Phonikud-yi
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
arXiv1 repoarXiv:2506.12336
Trust-videoLLMs
Extracting Composition-Dependent Diffusion Coefficients Over a Very Large Composition Range in NiCoFeCrMn High Entropy Alloy Following Strategic Design of Diffusion Couples and Physics Informed Neural Network Numerical Method
arXiv1 repoarXiv:2506.12345
unlimited-ocr-benchmark
FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
arXiv1 repoarXiv:2506.12494
FlexRAG
Scaling Test-time Compute for LLM Agents
arXiv1 repoarXiv:2506.12928
OAgents
MAMMA: Markerless & Automatic Multi-Person Motion Action Capture
arXiv1 repoarXiv:2506.13040
mamma
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
arXiv1 repoarXiv:2506.13053
ZipVoice
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy
arXiv1 repoarXiv:2506.13284
AceReason-1.1-SFT
BUT System for the MLC-SLM Challenge
arXiv1 repoarXiv:2506.13414
DiCoW_v3_2
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
arXiv1 repoarXiv:2506.13585
MiniMax-AI.github.io
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
arXiv1 repoarXiv:2506.14003
Unlearn-Trace
SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement
arXiv1 repoarXiv:2506.14035
SimpleDoc
Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection
arXiv1 repoarXiv:2506.14473
RAM-APL
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
arXiv1 repoarXiv:2506.14512
SpaceQwen2.5-VL-3B-Instruct
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
arXiv1 repoarXiv:2506.14965
Reasoning360
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
arXiv1 repoarXiv:2506.15655
sweet-search
OAgents: An Empirical Study of Building Effective Agents
arXiv1 repoarXiv:2506.15741
OAgents
FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
arXiv1 repoarXiv:2506.15742
kontext-bench
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning
arXiv1 repoarXiv:2506.16499
aideml
Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models
arXiv1 repoarXiv:2506.16853
RATTPO
arXiv:2506.17055
arXiv1 repoarXiv:2506.17055
FM-music-tagging
Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025
arXiv1 repoarXiv:2506.17077
WhisperLiveKit
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
arXiv1 repoarXiv:2506.17337
MedVLMBench
Distilling On-device Language Models for Robot Planning with Minimal Human Intervention
arXiv1 repoarXiv:2506.17486
PRISM
Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster
arXiv1 repoarXiv:2506.18034
LLM4Seg
TAB: Unified Benchmarking of Time Series Anomaly Detection Methods
arXiv1 repoarXiv:2506.18046
TAB
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
arXiv1 repoarXiv:2506.18421
finepdfs
USAD: Universal Speech and Audio Representation via Distillation
arXiv1 repoarXiv:2506.18843
usad
Benchmarking Music Generation Models and Metrics via Human Preference Studies
arXiv1 repoarXiv:2506.19085
TuneJury
Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study
arXiv1 repoarXiv:2506.19794
DataMind
MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
arXiv1 repoarXiv:2506.19835
MAM
Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
arXiv1 repoarXiv:2506.20251
Qresafe
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
arXiv1 repoarXiv:2506.20331
Biomed-Enriched
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
arXiv1 repoarXiv:2506.20920
fineweb-2
Little by Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts
arXiv1 repoarXiv:2506.21035
repro-little-by-little-continual-learning-via-incremental-mixture-of-rank-1-associative-memory-e
PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
arXiv1 repoarXiv:2506.21076
Hunyuan3D-Omni
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
arXiv1 repoarXiv:2506.21272
FairyGen
Bridging Offline and Online Reinforcement Learning for LLMs
arXiv1 repoarXiv:2506.21495
fairseq2
Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
arXiv1 repoarXiv:2506.21509
DLC
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
arXiv1 repoarXiv:2506.21605
Membench
The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
arXiv1 repoarXiv:2506.22419
aideml
arXiv:2506.22557
arXiv1 repoarXiv:2506.22557
dlm-jailbreak-transfer
Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
arXiv1 repoarXiv:2506.22685
Embedding-Adjustment
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
arXiv1 repoarXiv:2506.22900
MOTOR
arXiv:2506.23329
arXiv1 repoarXiv:2506.23329
IR3D-bench
Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop
arXiv1 repoarXiv:2506.23351
RoboTwin
PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions
arXiv1 repoarXiv:2506.23440
PathDiff
ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
arXiv1 repoarXiv:2507.00898
ONLY
Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
arXiv1 repoarXiv:2507.01297
compactds-retrieval
arXiv:2507.01663
arXiv1 repoarXiv:2507.01663
TransferQueue
RoboBrain 2.0 Technical Report
arXiv1 repoarXiv:2507.02029
RoboSpatial-Eval
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
arXiv1 repoarXiv:2507.02259
MemAgent
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
arXiv1 repoarXiv:2507.02554
aideml
AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
arXiv1 repoarXiv:2507.02664
AIGI-Holmes
VLAI: A RoBERTa-Based Model for Automated Vulnerability Severity Classification
arXiv1 repoarXiv:2507.03607
vulnerability-severity-classification-chinese-macbert-base
Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models
arXiv1 repoarXiv:2507.03916
SwanLab
SeqTex: Generate Mesh Textures in Video Sequence
arXiv1 repoarXiv:2507.04285
ComfyUI-SeqTex
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
arXiv1 repoarXiv:2507.04447
DreamVLA
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
arXiv1 repoarXiv:2507.04504
dLLM-CtrlGen
any4: Learned 4-bit Numeric Representation for LLMs
arXiv1 repoarXiv:2507.04610
any4
Spatio-Temporal LLM: Reasoning about Environments and Actions
arXiv1 repoarXiv:2507.05258
Spatio-Temporal-LLM
Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
arXiv1 repoarXiv:2507.05513
Eagle
An autonomous agent for auditing and improving the reliability of clinical AI models
arXiv1 repoarXiv:2507.05755
medmnistc-api
LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
arXiv1 repoarXiv:2507.06272
Monkey
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
arXiv1 repoarXiv:2507.06607
Phi-4-mini-flash-reasoning
Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
arXiv1 repoarXiv:2507.07095
MotionHub
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
arXiv1 repoarXiv:2507.07610
Spatial-Visualization-Benchmark
THUNDER: Tile-level Histopathology image UNDERstanding benchmark
arXiv1 repoarXiv:2507.07860
thunder
Predicting and generating antibiotics against future pathogens with ApexOracle
arXiv1 repoarXiv:2507.07862
ApexOracle
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
arXiv1 repoarXiv:2507.08306
M2-Reasoning
InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
arXiv1 repoarXiv:2507.08416
InstaScene
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
arXiv1 repoarXiv:2507.08771
FR-Spec
From One to More: Contextual Part Latents for 3D Generation
arXiv1 repoarXiv:2507.08772
partverse
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models
arXiv1 repoarXiv:2507.08983
trojanclimb
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
arXiv1 repoarXiv:2507.09116
WenetSpeech-Chuan
arXiv:2507.09264
arXiv1 repoarXiv:2507.09264
walrus
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
arXiv1 repoarXiv:2507.09318
ZipVoice
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
arXiv1 repoarXiv:2507.11097
DIJA
EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes
arXiv1 repoarXiv:2507.11407
EXAONE-4.0
Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?
arXiv1 repoarXiv:2507.11569
Foundation-based-reg
General Modular Harness for LLM Agents in Multi-Turn Gaming Environments
arXiv1 repoarXiv:2507.11633
stable-retro
A Survey of Deep Learning for Geometry Problem Solving
arXiv1 repoarXiv:2507.11936
VisNumBench
Kevin: Multi-Turn RL for Generating CUDA Kernels
arXiv1 repoarXiv:2507.11948
TritonForge
Characterizing State Space Model and Hybrid Language Model Performance with Long Context
arXiv1 repoarXiv:2507.12442
SSM-Scope
SpatialTrackerV2: 3D Point Tracking Made Easy
arXiv1 repoarXiv:2507.12462
SpaTrackerV2
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
arXiv1 repoarXiv:2507.12705
AudioJudge
DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
arXiv1 repoarXiv:2507.12796
DeQA-Doc
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
arXiv1 repoarXiv:2507.13985
DreamScene
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
arXiv1 repoarXiv:2507.15028
video-tt
OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities
arXiv1 repoarXiv:2507.15085
OCRGenBench
Solving Formal Math Problems by Decomposition and Iterative Reflection
arXiv1 repoarXiv:2507.15225
Seed-Prover
MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations
arXiv1 repoarXiv:2507.15255
GEM
RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records
arXiv1 repoarXiv:2507.15867
RDMA
SDBench: A Comprehensive Benchmark Suite for Speaker Diarization
arXiv1 repoarXiv:2507.16136
OpenBench
Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Control
arXiv1 repoarXiv:2507.16645
Morpheus-Software
AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation
arXiv1 repoarXiv:2507.16940
AURA
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
arXiv1 repoarXiv:2507.17520
VLA_Instruction_Tuning
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
arXiv1 repoarXiv:2507.17527
UniSS
WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training
arXiv1 repoarXiv:2507.17634
Ling-1T
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
arXiv1 repoarXiv:2507.19002
banana100-additional-iqa-models
RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
arXiv1 repoarXiv:2507.19280
RemoteReasoner
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
arXiv1 repoarXiv:2507.19634
MCIF
Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training
arXiv1 repoarXiv:2507.20291
OSEDiff
Meta CLIP 2: A Worldwide Scaling Recipe
arXiv1 repoarXiv:2507.22062
MetaCLIP
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
arXiv1 repoarXiv:2507.22953
CADS-dataset
Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
arXiv1 repoarXiv:2507.23159
Full-Duplex-Bench
Unveiling Super Experts in Mixture-of-Experts Large Language Models
arXiv1 repoarXiv:2507.23279
reap
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding
arXiv1 repoarXiv:2507.23478
3D-R1
Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
arXiv1 repoarXiv:2507.23726
Seed-Prover
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
arXiv1 repoarXiv:2507.23751
synthetic-data
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
arXiv1 repoarXiv:2507.23779
Phi-Ground
From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media
arXiv1 repoarXiv:2508.00497
SocialAlign
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
arXiv1 repoarXiv:2508.00912
Reasoning_Length_Prediction
SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy
arXiv1 repoarXiv:2508.01188
SwanLab
OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets
arXiv1 repoarXiv:2508.01630
healthadvocate
Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time
arXiv1 repoarXiv:2508.02037
STIM
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
arXiv1 repoarXiv:2508.02317
VeOmni
Efficient Agents: Building Effective Agents While Reducing Cost
arXiv1 repoarXiv:2508.02694
OAgents
AgentSight: System-Level Observability for AI Agents Using eBPF
arXiv1 repoarXiv:2508.02736
agentsight
CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data
arXiv1 repoarXiv:2508.02879
CauKer2M
LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
arXiv1 repoarXiv:2508.03440
Soft-Thinking
Unravelling the Probabilistic Forest: Arbitrage in Prediction Markets
arXiv1 repoarXiv:2508.03474
CloddsBot
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
arXiv1 repoarXiv:2508.03680
agent-lightning
HPSv3: Towards Wide-Spectrum Human Preference Score
arXiv1 repoarXiv:2508.03789
banana100-additional-iqa-models
TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
arXiv1 repoarXiv:2508.04324
EditScore
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
arXiv1 repoarXiv:2508.04796
tokenizer-flores-validation
Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment
arXiv1 repoarXiv:2508.04865
MultiPL-E
R-Zero: Self-Evolving Reasoning LLM from Zero Data
arXiv1 repoarXiv:2508.05004
R-Zero
Reasoning through Exploration: A Reinforcement Learning Framework for Robust Function Calling
arXiv1 repoarXiv:2508.05118
AWorld
CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
arXiv1 repoarXiv:2508.05242
SwanLab
Learning to Reason for Factuality
arXiv1 repoarXiv:2508.05618
fairseq2
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
arXiv1 repoarXiv:2508.05835
nemo-nano-codec-22khz-1.89kbps-21.5fps
More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment
arXiv1 repoarXiv:2508.06036
MER2025-MRAC25
arXiv:2508.06098
arXiv1 repoarXiv:2508.06098
Resonate
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
arXiv1 repoarXiv:2508.07493
VisR-Bench
arXiv:2508.07647
arXiv1 repoarXiv:2508.07647
LaRender
Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL
arXiv1 repoarXiv:2508.07976
ASearcher-train-data
Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing
arXiv1 repoarXiv:2508.09192
d3LLM
MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
arXiv1 repoarXiv:2508.09779
MoIIE
Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld
arXiv1 repoarXiv:2508.09889
AWorld
MOC: Meta-Optimized Classifier for Few-Shot Whole Slide Image Classification
arXiv1 repoarXiv:2508.09967
MOC
TexVerse: A Universe of 3D Objects with High-Resolution Textures
arXiv1 repoarXiv:2508.10868
TexVerse
VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip
arXiv1 repoarXiv:2508.10931
VSF
arXiv:2508.11131
arXiv1 repoarXiv:2508.11131
Block-Sparse-Attention
TinyTim: A Family of Language Models for Divergent Generation
arXiv1 repoarXiv:2508.11607
npcpy
QuarkMed Medical Foundation Model Technical Report
arXiv1 repoarXiv:2508.11894
MedXpertQA
Improving Densification in 3D Gaussian Splatting for High-Fidelity Rendering
arXiv1 repoarXiv:2508.12313
SphereForge
MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security
arXiv1 repoarXiv:2508.12538
awesome-mcp-security
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
arXiv1 repoarXiv:2508.12631
router
Cryfish: On deep audio analysis with Large Language Models
arXiv1 repoarXiv:2508.12666
VoicePersonification
arXiv:2508.13009
arXiv1 repoarXiv:2508.13009
Matrix-Game-2.0
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
arXiv1 repoarXiv:2508.13141
unlazy
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
arXiv1 repoarXiv:2508.14444
NVIDIA-Nemotron-Nano-9B-v2
Mobile-Agent-v3: Fundamental Agents for GUI Automation
arXiv1 repoarXiv:2508.15144
MobileAgent
Intern-S1: A Scientific Multimodal Foundation Model
arXiv1 repoarXiv:2508.15763
Intern-S1-Pro
TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference
arXiv1 repoarXiv:2508.15881
TransArch
Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
arXiv1 repoarXiv:2508.15904
PathPT
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
arXiv1 repoarXiv:2508.16929
circuit_backup
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
arXiv1 repoarXiv:2508.17623
emo-reasoning
The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
arXiv1 repoarXiv:2508.17627
CoDE-Stop
ST-Raptor: LLM-Powered Semi-Structured Table Question Answering
arXiv1 repoarXiv:2508.18190
ST-Raptor
Wan-S2V: Audio-Driven Cinematic Video Generation
arXiv1 repoarXiv:2508.18621
Wan2.2-S2V-14B
MobileCLIP2: Improving Multi-Modal Reinforced Training
arXiv1 repoarXiv:2508.20691
ml-mobileclip
rStar2-Agent: Agentic Reasoning Technical Report
arXiv1 repoarXiv:2508.20722
Open-AgentRL
On the Theoretical Limitations of Embedding-Based Retrieval
arXiv1 repoarXiv:2508.21038
vestige
Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
arXiv1 repoarXiv:2508.21048
Veritas
AHELM: A Holistic Evaluation of Audio-Language Models
arXiv1 repoarXiv:2508.21376
helm
MobiAgent: A Systematic Framework for Customizable Mobile Agents
arXiv1 repoarXiv:2509.00531
MobiAgent
CCE: Confidence-Consistency Evaluation for Time Series Anomaly Detection
arXiv1 repoarXiv:2509.01098
CCE
Reinforced Visual Perception with Tools
arXiv1 repoarXiv:2509.01656
REVPT-data
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
arXiv1 repoarXiv:2509.02020
FireRedTTS2
arXiv:2509.02398
arXiv1 repoarXiv:2509.02398
Resonate
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
arXiv1 repoarXiv:2509.02492
GRAM-RR-TrainingData
Jointly Reinforcing Diversity and Quality in Language Model Generations
arXiv1 repoarXiv:2509.02534
darling
Planning with Reasoning using Vision Language World Model
arXiv1 repoarXiv:2509.02722
cr4wm
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
arXiv1 repoarXiv:2509.04474
SpecTTS-Bench
Why Language Models Hallucinate
arXiv1 repoarXiv:2509.04664
generative-ai
Hunyuan-MT Technical Report
arXiv1 repoarXiv:2509.05209
Hunyuan-MT-7B-GGUF
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
arXiv1 repoarXiv:2509.06155
Verse-Bench
Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
arXiv1 repoarXiv:2509.06321
Text4Seg
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
arXiv1 repoarXiv:2509.06861
unlazy
Interleaving Reasoning for Better Text-to-Image Generation
arXiv1 repoarXiv:2509.06945
Vision-R1
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
arXiv1 repoarXiv:2509.06980
RL-Factory
veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
arXiv1 repoarXiv:2509.07003
veScale
arXiv:2509.07447
arXiv1 repoarXiv:2509.07447
trajgaze
Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates
arXiv1 repoarXiv:2509.09550
neucodec
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives
arXiv1 repoarXiv:2509.09838
stable-retro
Dynamic Vulnerability Patching for Heterogeneous Embedded Systems Using Stack Frame Reconstruction
arXiv1 repoarXiv:2509.10213
awesome-connected-things-sec
GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography
arXiv1 repoarXiv:2509.10344
GLAM
Towards Understanding Visual Grounding in Visual Language Models
arXiv1 repoarXiv:2509.10345
gui-agent
MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models
arXiv1 repoarXiv:2509.10569
watermarks-remover
Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning
arXiv1 repoarXiv:2509.11420
TradingAgents
UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
arXiv1 repoarXiv:2509.11543
MobileAgent
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
arXiv1 repoarXiv:2509.11552
hichunk
NeuroStrike: Neuron-Level Attacks on Aligned LLMs
arXiv1 repoarXiv:2509.11864
JailbreakLab
MMORE: Massive Multimodal Open RAG & Extraction
arXiv1 repoarXiv:2509.11937
mmore
RailSafeNet: Visual Scene Understanding for Tram Safety
arXiv1 repoarXiv:2509.12125
RailSafeNet
Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
arXiv1 repoarXiv:2509.12546
Agent4FaceForgery
Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
arXiv1 repoarXiv:2509.12883
lego-edit
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
arXiv1 repoarXiv:2509.13282
ChartGaze
Do Activation Verbalization Methods Convey Privileged Information?
arXiv1 repoarXiv:2509.13316
verb_faithfulness
OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
arXiv1 repoarXiv:2509.13347
VeOmni
BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching
arXiv1 repoarXiv:2509.13789
ltx2-vidgen-skill
LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
arXiv1 repoarXiv:2509.14252
mlx-tune
arXiv:2509.14427
arXiv1 repoarXiv:2509.14427
hashing-baseline
Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis
arXiv1 repoarXiv:2509.14579
X-Voice
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
arXiv1 repoarXiv:2509.15221
ScaleCUA-Data
DiffusionNFT: Online Diffusion Reinforcement with Forward Process
arXiv1 repoarXiv:2509.16117
GRPO
Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
arXiv1 repoarXiv:2509.16696
decoding_uncertainty
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
arXiv1 repoarXiv:2509.16972
q-frame
Evolution of Concepts in Language Model Pre-Training
arXiv1 repoarXiv:2509.17196
circuit_backup
AI Pangaea: Unifying Intelligence Islands for Adapting Myriad Tasks
arXiv1 repoarXiv:2509.17460
awdemos
Qwen3-Omni Technical Report
arXiv1 repoarXiv:2509.17765
Qwen3-Omni
WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing
arXiv1 repoarXiv:2509.18004
WSChuan-Train
GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction
arXiv1 repoarXiv:2509.18090
GeoSVR
The Illusion of Readiness in Health AI
arXiv1 repoarXiv:2509.18234
PeruMedQA
Track-On2: Enhancing Online Point Tracking with Memory
arXiv1 repoarXiv:2509.19115
track_on
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
arXiv1 repoarXiv:2509.19244
LaViDa
Formal Verification of Minimax Algorithms
arXiv1 repoarXiv:2509.20138
sunfish
Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs
arXiv1 repoarXiv:2509.20208
blendsql
Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing
arXiv1 repoarXiv:2509.20336
GraphGhost
EmbeddingGemma: Powerful and Lightweight Text Representations
arXiv1 repoarXiv:2509.20354
embeddinggemma-300m
Revisiting Data Challenges of Computational Pathology: A Pack-based Multiple Instance Learning Training Framework
arXiv1 repoarXiv:2509.20923
CPathPatchFeature
Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
arXiv1 repoarXiv:2509.21072
AWorld
NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
arXiv1 repoarXiv:2509.21309
NewtonGen
d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation
arXiv1 repoarXiv:2509.21474
d2
Self-Speculative Biased Decoding for Faster Re-Translation
arXiv1 repoarXiv:2509.21740
NoLanguageLeftWaiting
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
arXiv1 repoarXiv:2509.22167
Semantic-DACVAE-Japanese-32dim
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
arXiv1 repoarXiv:2509.22186
MinerU
Enabling Approximate Joint Sampling in Diffusion LMs
arXiv1 repoarXiv:2509.22738
ParallelBench
A Capacity-Based Rationale for Multi-Head Attention
arXiv1 repoarXiv:2509.22840
repro-a-capacity-based-rationale-for-multi-head-attention
Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
arXiv1 repoarXiv:2509.23050
understanding_lp
ToolUniverse: An open platform for democratizing AI scientists
arXiv1 repoarXiv:2509.23426
ToolUniverse
EfficientMIL: Efficient Linear-Complexity MIL Method for WSI Classification
arXiv1 repoarXiv:2509.23640
EfficientMIL
LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
arXiv1 repoarXiv:2509.23661
LLaVA-OneVision-1.5-4B-Instruct
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
arXiv1 repoarXiv:2509.23866
dart-gui
RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization
arXiv1 repoarXiv:2509.23991
SphereForge
SparseD: Sparse Attention for Diffusion Language Models
arXiv1 repoarXiv:2509.24014
SparseD
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
arXiv1 repoarXiv:2509.24391
UniFlow-Audio
Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting
arXiv1 repoarXiv:2509.24789
tsfmx
Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws
arXiv1 repoarXiv:2509.24914
repro-single-head-attention-inductive-bias-spectral-highdim
Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
arXiv1 repoarXiv:2509.25050
GRPO
Scaling Generalist Data-Analytic Agents
arXiv1 repoarXiv:2509.25084
DataMind
arXiv:2509.25127
arXiv1 repoarXiv:2509.25127
ml-sid-dit
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
arXiv1 repoarXiv:2509.25146
fast-feature-fields
DepthLM: Metric Depth From Vision Language Models
arXiv1 repoarXiv:2509.25413
angleLLM
A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
arXiv1 repoarXiv:2509.25609
abxlab
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
arXiv1 repoarXiv:2509.25848
VAPO
IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
arXiv1 repoarXiv:2509.26231
IMG-Multimodal-Diffusion-Alignment
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
arXiv1 repoarXiv:2509.26391
MotionRAG
SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From
arXiv1 repoarXiv:2509.26404
SeedPrints
fev-bench: A Realistic Benchmark for Time Series Forecasting
arXiv1 repoarXiv:2509.26468
fev
dParallel: Learnable Parallel Decoding for dLLMs
arXiv1 repoarXiv:2509.26488
d3LLM
Entropy After </Think> for reasoning model early exiting
arXiv1 repoarXiv:2509.26522
CoDE-Stop
OceanGym: A Benchmark Environment for Underwater Embodied Agents
arXiv1 repoarXiv:2509.26536
OceanGPT
LongCodeZip: Compress Long Context for Code Language Models
arXiv1 repoarXiv:2510.00446
LongCodeZip
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
arXiv1 repoarXiv:2510.00536
GUI-KV
On Predictability of Reinforcement Learning Dynamics for Large Language Models
arXiv1 repoarXiv:2510.00553
ReproAlphaRL
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
arXiv1 repoarXiv:2510.01010
banana100-additional-iqa-models
Aristotle: IMO-level Automated Theorem Proving
arXiv1 repoarXiv:2510.01346
open-atp
GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
arXiv1 repoarXiv:2510.02186
GeoPurify
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
arXiv1 repoarXiv:2510.02315
FOCUS
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
arXiv1 repoarXiv:2510.02609
RedCodeAgent
To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration
arXiv1 repoarXiv:2510.02676
ecf8
Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models
arXiv1 repoarXiv:2510.02880
MaskGRPO
Product-Quantised Image Representation for High-Quality Image Synthesis
arXiv1 repoarXiv:2510.03191
WeavePrompt
Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
arXiv1 repoarXiv:2510.03342
RoboSpatial-Eval
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
arXiv1 repoarXiv:2510.04213
wespeaker
RAP: 3D Rasterization Augmented End-to-End Planning
arXiv1 repoarXiv:2510.04333
RAP_ckpts
When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA
arXiv1 repoarXiv:2510.04849
PsiloQA
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
arXiv1 repoarXiv:2510.04885
rl-injector
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
arXiv1 repoarXiv:2510.05038
test-time-hybrid-retrieval
Agentic Misalignment: How LLMs Could Be Insider Threats
arXiv1 repoarXiv:2510.05179
ogham-mcp
NorMuon: Making Muon more efficient and scalable
arXiv1 repoarXiv:2510.05491
modded-nanogpt
TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
arXiv1 repoarXiv:2510.06063
TelecomTS
TokenChain: A Discrete Speech Chain via Semantic Token Modeling
arXiv1 repoarXiv:2510.06201
TokenChain
Conditional Denoising Diffusion Model-Based Robust MR Image Reconstruction from Highly Undersampled Data
arXiv1 repoarXiv:2510.06335
WeavePrompt
StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance
arXiv1 repoarXiv:2510.06827
StyleKeeper
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
arXiv1 repoarXiv:2510.07143
DART
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
arXiv1 repoarXiv:2510.07838
Full-Duplex-Bench
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
arXiv1 repoarXiv:2510.07962
LightReasoner
Reinforcing Diffusion Models by Direct Group Preference Optimization
arXiv1 repoarXiv:2510.08425
GRPO
ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
arXiv1 repoarXiv:2510.08457
Revisual-R1
ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation
arXiv1 repoarXiv:2510.08551
ARTDECO
GraphGhost: Tracing Structures Behind Large Language Models
arXiv1 repoarXiv:2510.08613
GraphGhost
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
arXiv1 repoarXiv:2510.08713
UniWM_Dataset
When to Reason: Semantic Router for vLLM
arXiv1 repoarXiv:2510.08731
semantic-router
Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
arXiv1 repoarXiv:2510.08807
humanoid-everyday
Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
arXiv1 repoarXiv:2510.09012
ARsample
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
arXiv1 repoarXiv:2510.09023
Palisade
SynthID-Image: Image watermarking at internet scale
arXiv1 repoarXiv:2510.09263
gemini-watermark-and-synthid-remover
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
arXiv1 repoarXiv:2510.10125
emboviz
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
arXiv1 repoarXiv:2510.10726
HunyuanWorld-Mirror
FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
arXiv1 repoarXiv:2510.10785
FAC-FACodec
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
arXiv1 repoarXiv:2510.11098
VCB-Bench
ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
arXiv1 repoarXiv:2510.11498
ArtifactsBenchmark
PhySIC: Physically Plausible 3D Human-Scene Interaction and Contact from a Single Image
arXiv1 repoarXiv:2510.11649
Phy-SIC
Diffusion Transformers with Representation Autoencoders
arXiv1 repoarXiv:2510.11690
RAE
Scaling Language-Centric Omnimodal Representation Learning
arXiv1 repoarXiv:2510.11693
LCO-Embedding
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
arXiv1 repoarXiv:2510.11696
QeRL
Demystifying Reinforcement Learning in Agentic Reasoning
arXiv1 repoarXiv:2510.11701
Open-AgentRL
Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics
arXiv1 repoarXiv:2510.12787
lean-lsp-mcp
Detect Anything via Next Point Prediction
arXiv1 repoarXiv:2510.12798
LocateAnything-3B
RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge
arXiv1 repoarXiv:2510.13590
post-graph-rag
The Art of Scaling Reinforcement Learning Compute for LLMs
arXiv1 repoarXiv:2510.13786
OpenRLHF
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
arXiv1 repoarXiv:2510.13795
Honey-Data-15M
LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization
arXiv1 repoarXiv:2510.13907
prompt-ops
xLLM Technical Report
arXiv1 repoarXiv:2510.14686
xllm
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
arXiv1 repoarXiv:2510.14969
UI-Simulator
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
arXiv1 repoarXiv:2510.14979
NEO
RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation
arXiv1 repoarXiv:2510.15362
rankseg
Attention Sinks in Diffusion Language Models
arXiv1 repoarXiv:2510.15731
dlms-sinks
BLIP3o-NEXT: Next Frontier of Native Image Generation
arXiv1 repoarXiv:2510.15857
BLIP3o
How Good Are LLMs at Processing Tool Outputs?
arXiv1 repoarXiv:2510.15955
toolJSONprocessing
Demystifying Transition Matching: When and Why It Can Beat Flow Matching
arXiv1 repoarXiv:2510.17991
TransitionFlowMatching
ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
arXiv1 repoarXiv:2510.18795
ProCLIP
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
arXiv1 repoarXiv:2510.18817
FailureSensorIQ
Search Self-play: Pushing the Frontier of Agent Capability without Supervision
arXiv1 repoarXiv:2510.18821
SSP
LightMem: Lightweight and Efficient Memory-Augmented Generation
arXiv1 repoarXiv:2510.18866
LightMem
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
arXiv1 repoarXiv:2510.18874
retaining-by-doing
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
arXiv1 repoarXiv:2510.18876
Grasp-Any-Region
Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues
arXiv1 repoarXiv:2510.19028
SCRIPTS
Addressing the Depth-of-Field Constraint: A New Paradigm for High Resolution Multi-Focus Image Fusion
arXiv1 repoarXiv:2510.19581
VAEAEDOF
Learning Affordances at Inference-Time for Vision-Language-Action Models
arXiv1 repoarXiv:2510.19752
roboeval
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement
arXiv1 repoarXiv:2510.20441
unified-audio
From Masks to Worlds: A Hitchhiker's Guide to World Models
arXiv1 repoarXiv:2510.20668
Lumina-DiMOO
AlphaFlow: Understanding and Improving MeanFlow Models
arXiv1 repoarXiv:2510.20771
alphaflow
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
arXiv1 repoarXiv:2510.20822
HoloCine
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
arXiv1 repoarXiv:2510.21311
Fines
The Principles of Diffusion Models
arXiv1 repoarXiv:2510.21890
PyTorchTutorial
Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy
arXiv1 repoarXiv:2510.22215
ViMDoc
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
arXiv1 repoarXiv:2510.22603
Llama-AVSR
Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem
arXiv1 repoarXiv:2510.22876
spec_dec
Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
arXiv1 repoarXiv:2510.22954
srt-hivemind
ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
arXiv1 repoarXiv:2510.23306
ReconViaGen
Dexbotic: Open-Source Vision-Language-Action Toolbox
arXiv1 repoarXiv:2510.23511
dexbotic
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
arXiv1 repoarXiv:2510.23538
ArtifactsBenchmark
MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
arXiv1 repoarXiv:2510.23727
MUStReason
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
arXiv1 repoarXiv:2510.24563
OSWorld-MCP
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
arXiv1 repoarXiv:2510.24693
STAR-Bench
Generative View Stitching
arXiv1 repoarXiv:2510.24718
generative_view_stitching
RNAGenScape: Property-Guided, Optimized Generation of mRNA Sequences with Manifold Langevin Dynamics
arXiv1 repoarXiv:2510.24736
figures4papers
Formalization of Auslander--Buchsbaum--Serre criterion in Lean4
arXiv1 repoarXiv:2510.24818
FLT
DINO-YOLO: Self-Supervised Pre-training for Data-Efficient Object Detection in Civil Engineering Applications
arXiv1 repoarXiv:2510.25140
DINOV3-YOLOV12
RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models
arXiv1 repoarXiv:2510.25257
RT-DETRv4
ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
arXiv1 repoarXiv:2510.25668
ALDEN
The Wiegold problem and free products of left-orderable groups
arXiv1 repoarXiv:2510.26073
superhuman
UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens
arXiv1 repoarXiv:2510.26372
unified-audio
Kimi Linear: An Expressive, Efficient Attention Architecture
arXiv1 repoarXiv:2510.26692
GatedDeltaNet-2
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
arXiv1 repoarXiv:2510.26835
semantic-router
The Denario project: Deep knowledge AI agents for scientific discovery
arXiv1 repoarXiv:2510.26887
Denario
E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources
arXiv1 repoarXiv:2510.27135
Nitro-E
Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs
arXiv1 repoarXiv:2510.27246
ogham-mcp
World Simulation with Video Foundation Models for Physical AI
arXiv1 repoarXiv:2511.00062
Cosmos-Predict2.5-14B
CompAgent: An Agentic Framework for Visual Compliance Verification
arXiv1 repoarXiv:2511.00171
compagent
ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
arXiv1 repoarXiv:2511.00511
OpenS2V-Nexus
Erasing 'Ugly' from the Internet: Propagation of the Beauty Myth in Text-Image Models
arXiv1 repoarXiv:2511.00749
BeautyStandards
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
arXiv1 repoarXiv:2511.01090
FineWeb2-Ro
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
arXiv1 repoarXiv:2511.01261
speech_drame
UniREditBench: A Unified Reasoning-based Image Editing Benchmark
arXiv1 repoarXiv:2511.01295
UniREdit-Data-100K
arXiv:2511.01833
arXiv1 repoarXiv:2511.01833
evalscope
iFlyBot-VLA Technical Report
arXiv1 repoarXiv:2511.01914
iFlyBot-VLA
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
arXiv1 repoarXiv:2511.02802
TabTune
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
arXiv1 repoarXiv:2511.02817
oolong
FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels
arXiv1 repoarXiv:2511.02872
open-atp
audio2chart: End to End Audio Transcription into playable Guitar Hero charts
arXiv1 repoarXiv:2511.03337
audio2chart
NVIDIA Nemotron Nano V2 VL
arXiv1 repoarXiv:2511.03929
Eagle
Submanifold Sparse Convolutional Networks for Automated 3D Segmentation of Kidneys and Kidney Tumours in Computed Tomography
arXiv1 repoarXiv:2511.04334
ai_cancer_research
V-Thinker: Interactive Thinking with Images
arXiv1 repoarXiv:2511.04460
V-Thinker
Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
arXiv1 repoarXiv:2511.04555
Evo-1
Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
arXiv1 repoarXiv:2511.04655
VSI-Bench
arXiv:2511.05171
arXiv1 repoarXiv:2511.05171
naturelm-audio
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
arXiv1 repoarXiv:2511.05275
TwinVLA
CGCE: Classifier-Guided Concept Erasure in Generative Models
arXiv1 repoarXiv:2511.05865
CGCE
The Station: An Open-World Environment for AI-Driven Discovery
arXiv1 repoarXiv:2511.06309
station
DIMO: Diverse 3D Motion Generation for Arbitrary Objects
arXiv1 repoarXiv:2511.07409
DIMO
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
arXiv1 repoarXiv:2511.08379
som-refusal-directions
Structured RAG for Answering Aggregative Questions
arXiv1 repoarXiv:2511.08505
aggregative_questions
LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
arXiv1 repoarXiv:2511.08544
mlx-tune
TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
arXiv1 repoarXiv:2511.08667
TabPFN
RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
arXiv1 repoarXiv:2511.09554
rf-detr
Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages
arXiv1 repoarXiv:2511.09690
fairseq2
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
arXiv1 repoarXiv:2511.09715
SliderEdit
Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
arXiv1 repoarXiv:2511.10037
Beyond-React
ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
arXiv1 repoarXiv:2511.10240
ProgRAG
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
arXiv1 repoarXiv:2511.10645
qwen38-27b-exl3
Depth Anything 3: Recovering the Visual Space from Any Views
arXiv1 repoarXiv:2511.10647
DA3-BASE
Effective Brascamp-Lieb inequalities
arXiv1 repoarXiv:2511.11091
numina-lean-agent
Moirai 2.0: When Less Is More for Time Series Forecasting
arXiv1 repoarXiv:2511.11698
moirai-2.0-R-small
Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views
arXiv1 repoarXiv:2511.12878
UniHand
PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
arXiv1 repoarXiv:2511.12998
PerTouch
Distribution Matching Distillation Meets Reinforcement Learning
arXiv1 repoarXiv:2511.13649
Z-Image-Turbo
Scaling Spatial Intelligence with Multimodal Foundation Models
arXiv1 repoarXiv:2511.13719
SenseNova-SI-1.1-Qwen2.5-VL-7B
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
arXiv1 repoarXiv:2511.14582
OmniZip
UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning
arXiv1 repoarXiv:2511.14760
ml-unigen
IPR-1: Interactive Physical Reasoner
arXiv1 repoarXiv:2511.15407
stable-retro
GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
arXiv1 repoarXiv:2511.15658
geo-bench
First Frame Is the Place to Go for Video Content Customization
arXiv1 repoarXiv:2511.15700
FFGO-Video-Customization
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
arXiv1 repoarXiv:2511.15738
Weave
Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
arXiv1 repoarXiv:2511.16043
Agent0
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
arXiv1 repoarXiv:2511.16108
SkyRL
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
arXiv1 repoarXiv:2511.16449
VLA-Pruner
Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
arXiv1 repoarXiv:2511.16664
NVIDIA-Nemotron-3-Nano-4B-GGUF
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
arXiv1 repoarXiv:2511.16757
UTS
Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
arXiv1 repoarXiv:2511.17209
SPECTRE
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
arXiv1 repoarXiv:2511.17366
RoboBrain_Dex
RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
arXiv1 repoarXiv:2511.17441
RoboCOIN
Planning with Sketch-Guided Verification for Physics-Aware Video Generation
arXiv1 repoarXiv:2511.17450
SketchVerify
ROVER: Regulator-Driven Robust Temporal Verification of Black-Box Robot Policies
arXiv1 repoarXiv:2511.17781
stable-retro
NeAR: Coupled Neural Asset-Renderer Stack
arXiv1 repoarXiv:2511.18600
NeAR
arXiv:2511.18870
arXiv1 repoarXiv:2511.18870
HunyuanVideo-1.5
Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
arXiv1 repoarXiv:2511.18890
Nemotron-Flash-1B
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
arXiv1 repoarXiv:2511.19558
spqr
RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
arXiv1 repoarXiv:2511.19704
RADIO
Leveraging Foundation Models for Histological Grading in Cutaneous Squamous Cell Carcinoma using PathFMTools
arXiv1 repoarXiv:2511.19751
PathFMTools
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
arXiv1 repoarXiv:2511.19861
Giga-World-1-Toydata
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
arXiv1 repoarXiv:2511.19900
Agent0
Boosting Reasoning in Large Multimodal Models via Activation Replay
arXiv1 repoarXiv:2511.19972
replay
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
arXiv1 repoarXiv:2511.20022
WaymoQA
PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images
arXiv1 repoarXiv:2511.20068
prada
NVIDIA Nemotron Parse 1.1
arXiv1 repoarXiv:2511.20478
NVIDIA-Nemotron-Parse-v1.1
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
arXiv1 repoarXiv:2511.20597
guaca
PixelDiT: Pixel Diffusion Transformers for Image Generation
arXiv1 repoarXiv:2511.20645
PixelDiT-ImageNet
FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain
arXiv1 repoarXiv:2511.21113
FaithFusion
When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
arXiv1 repoarXiv:2511.21192
UPA-RFAS
Exploring Fusion Strategies for Multimodal Vision-Language Systems
arXiv1 repoarXiv:2511.21889
Multimodal-Fusion-Strategies
Geometrically-Constrained Agent for Spatial Reasoning
arXiv1 repoarXiv:2511.22659
gca
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield
arXiv1 repoarXiv:2511.22677
Z-Image-Turbo
Test-time scaling of diffusions with flow maps
arXiv1 repoarXiv:2511.22688
UniGenBench
Captain Safari: A World Engine with Pose-Aligned 3D Memory
arXiv1 repoarXiv:2511.22815
Captain-Safari
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
arXiv1 repoarXiv:2511.23269
OctoMed-7B
TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
arXiv1 repoarXiv:2512.01248
TRivia-3B
EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans
arXiv1 repoarXiv:2512.01340
LongCat-Video-Avatar
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
arXiv1 repoarXiv:2512.01797
karma-electric-llama31-8b
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
arXiv1 repoarXiv:2512.02014
tuna-2
CLEF: Clinically-Guided Contrastive Learning for Electrocardiogram Foundation Models
arXiv1 repoarXiv:2512.02180
ecg-foundation-model
Guided Self-Evolving LLMs with Minimal Human Supervision
arXiv1 repoarXiv:2512.02472
R-Zero
dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
arXiv1 repoarXiv:2512.02498
dots.ocr
Spatially-Grounded Document Retrieval via Patch-to-Region Relevance Propagation
arXiv1 repoarXiv:2512.02660
Snappy
A Hierarchical Tree-based approach for creating Configurable and Static Deep Research Agent (Static-DRA)
arXiv1 repoarXiv:2512.03887
workers-personal-agent
OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference
arXiv1 repoarXiv:2512.03927
colibri
BioMedGPT-Mol: Multi-task Learning for Molecular Understanding and Generation
arXiv1 repoarXiv:2512.04629
BioMedGPT-Mol
YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases
arXiv1 repoarXiv:2512.04793
YingMusic-SVC
Intrinsically Interpretable Attention via Sparse Post-Training
arXiv1 repoarXiv:2512.05865
CLT-Forge
LightSearcher: Efficient DeepSearch via Experiential Memory
arXiv1 repoarXiv:2512.06653
MemoryOS
Group Representational Position Encoding
arXiv1 repoarXiv:2512.07805
GRAPE
OmniPSD: Layered PSD Generation with Diffusion Transformer
arXiv1 repoarXiv:2512.09247
OmniPSD
Hierarchy-Aware Multimodal Unlearning for Medical AI
arXiv1 repoarXiv:2512.09867
MedForget
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
arXiv1 repoarXiv:2512.10284
motionedit
Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
arXiv1 repoarXiv:2512.10691
RadVLM-GRPO
PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction
arXiv1 repoarXiv:2512.10888
granite-4.0-3b-vision
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
arXiv1 repoarXiv:2512.10927
FoundationMotion
Position: Universal Aesthetic Alignment Narrows Artistic Expression
arXiv1 repoarXiv:2512.11883
icml2026_position
MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models
arXiv1 repoarXiv:2512.12121
MixtureKit
arXiv:2512.12218
arXiv1 repoarXiv:2512.12218
journey-before-destination
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
arXiv1 repoarXiv:2512.12772
JointAVBench
DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
arXiv1 repoarXiv:2512.12799
DrivePI
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
arXiv1 repoarXiv:2512.13278
Open-AgentRL
RecTok: Reconstruction Distillation along Rectified Flow
arXiv1 repoarXiv:2512.13421
RecTok
Adapting MLLMs for Nuanced Video Retrieval
arXiv1 repoarXiv:2512.13511
TARA
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
arXiv1 repoarXiv:2512.14008
LaViDa
GLM-TTS Technical Report
arXiv1 repoarXiv:2512.14291
GLM-TTS
Native and Compact Structured Latents for 3D Generation
arXiv1 repoarXiv:2512.14692
TRELLIS.2-4B
FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
arXiv1 repoarXiv:2512.15420
flowbind
Corrective Diffusion Language Models
arXiv1 repoarXiv:2512.15596
ParallelBench
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
arXiv1 repoarXiv:2512.15713
DiffusionVL
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
arXiv1 repoarXiv:2512.16676
DataFlow
arXiv:2512.16899
arXiv1 repoarXiv:2512.16899
EditReward
Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
arXiv1 repoarXiv:2512.16913
SphereForge
StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors
arXiv1 repoarXiv:2512.16915
StereoPilot
Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience
arXiv1 repoarXiv:2512.17260
Seed-Prover
ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration
arXiv1 repoarXiv:2512.17298
ProCache
Xiaomi MiMo-VL-Miloco Technical Report
arXiv1 repoarXiv:2512.17436
xiaomi-mimo-vl-miloco
PathBench-MIL: A Comprehensive AutoML and Benchmarking Framework for Multiple Instance Learning in Histopathology
arXiv1 repoarXiv:2512.17517
PathBench-MIL
The HydroGym Reinforcement Learning Platform for Fluid Dynamics
arXiv1 repoarXiv:2512.17534
hydrogym
Dexterous World Models
arXiv1 repoarXiv:2512.17907
dwm
arXiv:2512.17909
arXiv1 repoarXiv:2512.17909
EditReward
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
arXiv1 repoarXiv:2512.19433
Lumina-DiMOO
How Much 3D Do Video Foundation Models Encode?
arXiv1 repoarXiv:2512.19949
VidFM3D
MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization
arXiv1 repoarXiv:2512.20135
SwanLab
QuarkAudio Technical Report
arXiv1 repoarXiv:2512.20151
unified-audio
Step-DeepResearch Technical Report
arXiv1 repoarXiv:2512.20491
StepDeepResearch
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
arXiv1 repoarXiv:2512.20573
specdiff_aoi
Quantifying Laziness, Decoding Suboptimality, and Context Degradation in Large Language Models
arXiv1 repoarXiv:2512.20662
unlazy
How important is Recall for Measuring Retrieval Quality?
arXiv1 repoarXiv:2512.20854
retrieval-response
NVIDIA Nemotron 3: Efficient and Open Intelligence
arXiv1 repoarXiv:2512.20856
NVIDIA-Nemotron-3-Nano-4B-GGUF
Streaming Video Instruction Tuning
arXiv1 repoarXiv:2512.21334
Streamo
AstraNav-Memory: Contexts Compression for Long Memory
arXiv1 repoarXiv:2512.21627
AstraNav-Memory
AstraNav-World: World Model for Foresight Control and Consistency
arXiv1 repoarXiv:2512.21714
AstraNav-World
Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
arXiv1 repoarXiv:2512.21857
ADT-Tree
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs
arXiv1 repoarXiv:2512.22219
mirage
RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
arXiv1 repoarXiv:2512.22560
Mooncake
Visual Autoregressive Modelling for Monocular Depth Estimation
arXiv1 repoarXiv:2512.22653
VAR-Depth
Reverse Personalization
arXiv1 repoarXiv:2512.22984
reverse-personalization
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
arXiv1 repoarXiv:2512.23365
spatial_mosaic_vqa
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
arXiv1 repoarXiv:2512.23578
SLM-Style-Amnesia
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
arXiv1 repoarXiv:2512.24618
Youtu-LLM-2B
mHC: Manifold-Constrained Hyper-Connections
arXiv1 repoarXiv:2512.24880
maxtext
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction
arXiv1 repoarXiv:2512.25073
MVGenMaster
STELLAR: A Search-Based Testing Framework for Large Language Model Applications
arXiv1 repoarXiv:2601.00497
STELLAR
RoboReward: General-Purpose Vision-Language Reward Models for Robotics
arXiv1 repoarXiv:2601.00675
reward-scope
Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
arXiv1 repoarXiv:2601.00753
circuit-breaker
CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
arXiv1 repoarXiv:2601.01874
cogflow_code
360-GeoGS: Geometrically Consistent Feed-Forward 3D Gaussian Splatting Reconstruction for 360 Images
arXiv1 repoarXiv:2601.02102
SphereForge
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
arXiv1 repoarXiv:2601.02356
talk2move
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
arXiv1 repoarXiv:2601.02970
ReASC
AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation
arXiv1 repoarXiv:2601.03191
anatomix
IndexTTS 2.5 Technical Report
arXiv1 repoarXiv:2601.03888
index-tts
PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography
arXiv1 repoarXiv:2601.03993
PosterVerse
Re-Rankers as Relevance Judges
arXiv1 repoarXiv:2601.04455
reranker-as-judge
SmartSearch: Process Reward-Guided Query Refinement for Search Agents
arXiv1 repoarXiv:2601.04888
SmartSearch
Atlas 2 -- Foundation models for clinical deployment
arXiv1 repoarXiv:2601.05148
PathoROB
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
arXiv1 repoarXiv:2601.05241
RoboVIP_VDM
LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
arXiv1 repoarXiv:2601.05248
last0
Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
arXiv1 repoarXiv:2601.05251
Mesh4D
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
arXiv1 repoarXiv:2601.05635
SoE
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
arXiv1 repoarXiv:2601.05679
reasoning-probing
LayerGS: Decomposition and Inpainting of Layered 3D Human Avatars via 2D Gaussian Splatting
arXiv1 repoarXiv:2601.05853
LayerGS
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
arXiv1 repoarXiv:2601.06803
laser
Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition
arXiv1 repoarXiv:2601.06972
categorize-early-asr
Solar Open Technical Report
arXiv1 repoarXiv:2601.07022
Solar-Open-100B
Dr. Zero: Self-Evolving Search Agents without Training Data
arXiv1 repoarXiv:2601.07055
drzero
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
arXiv1 repoarXiv:2601.07372
maxtext
MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head
arXiv1 repoarXiv:2601.07832
MHLA
Training Free Zero-Shot Visual Anomaly Localization via Diffusion Inversion
arXiv1 repoarXiv:2601.08022
DIVAD
Ministral 3
arXiv1 repoarXiv:2601.08584
Ministral-3-8B-Instruct-2512-GGUF
TranslateGemma Technical Report
arXiv1 repoarXiv:2601.09012
trans-gemma
A.X K1 Technical Report
arXiv1 repoarXiv:2601.09200
A.X-K1
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
arXiv1 repoarXiv:2601.09385
SLAM-LLM
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
arXiv1 repoarXiv:2601.10338
Palisade
Reasoning Models Generate Societies of Thought
arXiv1 repoarXiv:2601.10825
h2aichat
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
arXiv1 repoarXiv:2601.11141
FlashLabs-Chroma
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
arXiv1 repoarXiv:2601.11404
AI539_NLP
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
arXiv1 repoarXiv:2601.11522
UniX
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
arXiv1 repoarXiv:2601.12294
ToolPRMBench
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
arXiv1 repoarXiv:2601.12626
linear-mech-vlms
Think3D: Thinking with Space for Spatial Reasoning
arXiv1 repoarXiv:2601.13029
spagent
Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration
arXiv1 repoarXiv:2601.14235
Denario
RoboBrain 2.5: Depth in Sight, Time in Mind
arXiv1 repoarXiv:2601.14352
RoboBrain2.5-4B
Next Generation Active Learning: Mixture of LLMs in the Loop
arXiv1 repoarXiv:2601.15773
MoLLIA
Endless Terminals: Scaling RL Environments for Terminal Agents
arXiv1 repoarXiv:2601.16443
SkyRL
OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding
arXiv1 repoarXiv:2601.16538
online-spatial-intelligence
HapticMatch: An Exploration for Generative Material Haptic Simulation and Interaction
arXiv1 repoarXiv:2601.16639
GelSight_Texture_Dataset
VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding
arXiv1 repoarXiv:2601.17868
VidLaDA
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
arXiv1 repoarXiv:2601.18137
deepclause-sdk
SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction
arXiv1 repoarXiv:2601.18537
minimind
Arithmetic volumes of moduli stacks of Shtukas
arXiv1 repoarXiv:2601.18557
superhuman
A Hybrid Discriminative and Generative System for Universal Speech Enhancement
arXiv1 repoarXiv:2601.19113
unified-audio
Native LLM and MLLM Inference at Scale on Apple Silicon
arXiv1 repoarXiv:2601.19139
vllm-ios
Self-Distillation Enables Continual Learning
arXiv1 repoarXiv:2601.19897
OpenClaw-RL
CPiRi: Channel Permutation-Invariant Relational Interaction for Multivariate Time Series Forecasting
arXiv1 repoarXiv:2601.20318
CPiRi
AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
arXiv1 repoarXiv:2601.20524
AnomalyVFM
DeepSeek-OCR 2: Visual Causal Flow
arXiv1 repoarXiv:2601.20552
DeepSeek-OCR-2
Reinforcement Learning via Self-Distillation
arXiv1 repoarXiv:2601.20802
OpenClaw-RL
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
arXiv1 repoarXiv:2601.20992
asr_eval
AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions
arXiv1 repoarXiv:2601.21116
keep-the-why
MoCo: A One-Stop Shop for Model Collaboration Research
arXiv1 repoarXiv:2601.21257
model_collaboration
Irrationality of rapidly converging series: a problem of Erdős and Graham
arXiv1 repoarXiv:2601.21442
superhuman
OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models
arXiv1 repoarXiv:2601.21639
OCRVerse
Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units
arXiv1 repoarXiv:2601.21996
Mechanistic-Data-Attribution
Where Do the Joules Go? Diagnosing Inference Energy Consumption
arXiv1 repoarXiv:2601.22076
zeus
Semi-Autonomous Mathematics Discovery with Gemini: A Case Study on the Erdős Problems
arXiv1 repoarXiv:2601.22401
superhuman
dgMARK: Decoding-Guided Watermarking for Diffusion Language Models
arXiv1 repoarXiv:2601.22985
dgmark-watermarking
Strongly Polynomial Time Complexity of Policy Iteration for $L_\infty$ Robust MDPs
arXiv1 repoarXiv:2601.23229
superhuman
Eigenweights for arithmetic Hirzebruch Proportionality
arXiv1 repoarXiv:2601.23245
superhuman
Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models
arXiv1 repoarXiv:2602.00217
figures4papers
VoxServe: Streaming-Centric Serving System for Speech Language Models
arXiv1 repoarXiv:2602.00269
vox-serve
From Observations to States: Latent Time Series Forecasting
arXiv1 repoarXiv:2602.00297
LatentTSF
DecompressionLM: Deterministic, Diagnostic, and Zero-Shot Concept Graph Extraction from Language Models
arXiv1 repoarXiv:2602.00377
decompressionlm
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds
arXiv1 repoarXiv:2602.00807
repro-any3d-vla-enhancing-vla-robustness-via-diverse-point-clouds
arXiv:2602.01382
arXiv1 repoarXiv:2602.01382
EditReward
Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models
arXiv1 repoarXiv:2602.01849
self-rewarding-smc
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
arXiv1 repoarXiv:2602.02227
LatentMorph
Kimi K2.5: Visual Agentic Intelligence
arXiv1 repoarXiv:2602.02276
Kimi-K2.5
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
arXiv1 repoarXiv:2602.02431
repro-full-batch-gd-outperforms-one-pass-sgd-sample-complexity-separation
Lower bounds for multivariate independence polynomials and their generalisations
arXiv1 repoarXiv:2602.02450
superhuman
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
arXiv1 repoarXiv:2602.02475
AgentRx
RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System
arXiv1 repoarXiv:2602.02488
Open-AgentRL
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
arXiv1 repoarXiv:2602.02493
PixelGen-diffusers
Norm Anchors Make Model Edits Last
arXiv1 repoarXiv:2602.02543
ComfyUI-LoRA-Optimizer
Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis
arXiv1 repoarXiv:2602.03139
anima-turbo-4step
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection
arXiv1 repoarXiv:2602.03216
repro-token-sparse-attention-long-context-interleaved-selection
SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training
arXiv1 repoarXiv:2602.03411
SWE-World
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
arXiv1 repoarXiv:2602.03733
RegionReasoner
HY3D-Bench: Generation of 3D Assets
arXiv1 repoarXiv:2602.03907
HY3D-Bench
arXiv:2602.04315
arXiv1 repoarXiv:2602.04315
GeneralVLA
VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image
arXiv1 repoarXiv:2602.04349
VecSet-Edit
EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL
arXiv1 repoarXiv:2602.04417
ema-pg
Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging
arXiv1 repoarXiv:2602.04805
SkinTokens
PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
arXiv1 repoarXiv:2602.04876
PerpetualWonder
The Single-Multi Evolution Loop for Self-Improving Model Collaboration Systems
arXiv1 repoarXiv:2602.05182
model_collaboration
Learning to Inject: Automated Prompt Injection via Reinforcement Learning
arXiv1 repoarXiv:2602.05746
AutoInject
PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
arXiv1 repoarXiv:2602.06053
personaplex
Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making
arXiv1 repoarXiv:2602.06570
Baichuan-M3-235B
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models
arXiv1 repoarXiv:2602.07026
Modality_Gap_Theory
DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation
arXiv1 repoarXiv:2602.07371
DeepAnalyze
When Is Enough Not Enough? Illusory Completion in Search Agents
arXiv1 repoarXiv:2602.07549
illusory_completion
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
arXiv1 repoarXiv:2602.07605
Finedefics_ICLR2025
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
arXiv1 repoarXiv:2602.08387
modalities
Beyond Transcripts: A Renewed Perspective on Audio Chaptering
arXiv1 repoarXiv:2602.08979
ytseg
Atlas: Enabling Cross-Vendor Authentication for IoT
arXiv1 repoarXiv:2602.09263
awesome-connected-things-sec
UniVTAC: A Unified Simulation Platform for Visuo-Tactile Manipulation Data Generation, Learning, and Benchmarking
arXiv1 repoarXiv:2602.10093
UniVTAC
A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology
arXiv1 repoarXiv:2602.10624
DermFM-Zero
Training and Benchmarking Code Generation for Physics-Inspired Animations
arXiv1 repoarXiv:2602.10840
AgentFly
Embedding Inversion via Conditional Masked Diffusion Language Models
arXiv1 repoarXiv:2602.11047
embedding-inversion-demo
It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks
arXiv1 repoarXiv:2602.12147
TIME
dVoting: Fast Voting for dLLMs
arXiv1 repoarXiv:2602.12153
dVoting
arXiv:2602.12322
arXiv1 repoarXiv:2602.12322
foreact
Flow-Factory: A Unified Framework for Reinforcement Learning in Flow-Matching Models
arXiv1 repoarXiv:2602.12529
GRPO
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
arXiv1 repoarXiv:2602.12705
MedXpertQA
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
arXiv1 repoarXiv:2602.12806
rat-bench
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
arXiv1 repoarXiv:2602.13692
ThunderAgent
Towards Spatial Transcriptomics-driven Pathology Foundation Models
arXiv1 repoarXiv:2602.14177
SEAL
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
arXiv1 repoarXiv:2602.14486
Aristotelian
Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions
arXiv1 repoarXiv:2602.14878
tool-definition-quality-score
EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
arXiv1 repoarXiv:2602.15031
EditCtrl
ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT
arXiv1 repoarXiv:2602.15720
repro-toast-token-channel-selection-and-structured-pruning-for-efficient-vit
SAM 3D Body: Robust Full-Body Human Mesh Recovery
arXiv1 repoarXiv:2602.15989
sam-3d-body
Fast KV Compaction via Attention Matching
arXiv1 repoarXiv:2602.16284
attention-matching-rl
MerLean: An Agentic Framework for Autoformalization in Quantum Computation
arXiv1 repoarXiv:2602.16554
lean-lsp-mcp
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
arXiv1 repoarXiv:2602.16609
ColBERT-Zero
Towards a Science of AI Agent Reliability
arXiv1 repoarXiv:2602.16666
harness-evals
DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
arXiv1 repoarXiv:2602.16742
DeepVision-103K
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
arXiv1 repoarXiv:2602.16855
MobileAgent
Heterogeneous Federated Fine-Tuning with Parallel One-Rank Adaptation
arXiv1 repoarXiv:2602.16936
Fed-PLoRA
Arcee Trinity Large Technical Report
arXiv1 repoarXiv:2602.17004
donutloop-genesis
M2F: Automated Formalization of Mathematical Literature at Scale
arXiv1 repoarXiv:2602.17016
lean-lsp-mcp
PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing
arXiv1 repoarXiv:2602.17033
PartRAG
Leveraging Contrastive Learning for a Similarity-Guided Tampered Document Data Generation Pipeline
arXiv1 repoarXiv:2602.17322
RealText-V2-Syn25k
arXiv:2602.17868
arXiv1 repoarXiv:2602.17868
MantisV2Experiments
Improving Topic Modeling by Distilling Soft Labels from Language Models
arXiv1 repoarXiv:2602.17907
MCompassRAG
GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation
arXiv1 repoarXiv:2602.18164
grand_tour_box
Improving Sampling for Masked Diffusion Models via Information Gain
arXiv1 repoarXiv:2602.18176
Information-Gain-Sampler
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
arXiv1 repoarXiv:2602.18424
CapNav
From Docs to Descriptions: Smell-Aware Evaluation of MCP Server Descriptions
arXiv1 repoarXiv:2602.18914
tool-definition-quality-score
TimeRadar: A Domain-Rotatable Foundation Model for Time Series Anomaly Detection
arXiv1 repoarXiv:2602.19068
TimeRadar
Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling
arXiv1 repoarXiv:2602.19089
ani3dhuman
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
arXiv1 repoarXiv:2602.19127
DataFlow
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
arXiv1 repoarXiv:2602.19161
flash-vaed
SkillOrchestra: Learning to Route Agents via Skill Transfer
arXiv1 repoarXiv:2602.19672
SkillOrchestra
Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation
arXiv1 repoarXiv:2602.19778
ChordMiniApp
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
arXiv1 repoarXiv:2602.20577
MVLAD-AD
On Data Engineering for Scaling LLM Terminal Capabilities
arXiv1 repoarXiv:2602.21193
Nemotron-Terminal-Corpus
CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis
arXiv1 repoarXiv:2602.21637
CARE
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
arXiv1 repoarXiv:2602.21646
smt-9b-hf
Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models
arXiv1 repoarXiv:2602.22246
Diffusion_Self_Purification
veScale-FSDP: Flexible and High-Performance FSDP at Scale
arXiv1 repoarXiv:2602.22437
veScale
Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language
arXiv1 repoarXiv:2602.23201
Generalized-Neural-Memory
LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding
arXiv1 repoarXiv:2602.23881
SpecForge
Enhancing Spatial Understanding in Image Generation via Reward Modeling
arXiv1 repoarXiv:2602.24233
UniGenBench
Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
arXiv1 repoarXiv:2603.00431
Finedefics_ICLR2025
DreamWorld: Unified World Modeling in Video Generation
arXiv1 repoarXiv:2603.00466
VideoREPA
Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning
arXiv1 repoarXiv:2603.00667
HistoSelect
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
arXiv1 repoarXiv:2603.01205
CoSMo3D
Constructive and Predicative Locale Theory in Univalent Foundations
arXiv1 repoarXiv:2603.01308
TypeTopology
SimRecon: SimReady Compositional Scene Reconstruction from Real Videos
arXiv1 repoarXiv:2603.02133
SimRecon
Adaptive Personalized Federated Learning via Multi-task Averaging of Kernel Mean Embeddings
arXiv1 repoarXiv:2603.02233
repro-adaptive-personalized-fl-multi-task-averaging-kernel-mean-embedding
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
arXiv1 repoarXiv:2603.02277
felonybench
Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels
arXiv1 repoarXiv:2603.02573
Track4World
Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
arXiv1 repoarXiv:2603.02641
RE-USE
SorryDB: Can AI Provers Complete Real-World Lean Theorems?
arXiv1 repoarXiv:2603.02668
SorryDB
EvoSkill: Automated Skill Discovery for Multi-Agent Systems
arXiv1 repoarXiv:2603.02766
EvoSkill
NeuroSkill(tm): Proactive Real-Time Agentic System Capable of Modeling Human State of Mind
arXiv1 repoarXiv:2603.03212
neuroloop
PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters
arXiv1 repoarXiv:2603.04165
PlaneCycle
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild
arXiv1 repoarXiv:2603.04205
Real5-OmniDocBench
Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset
arXiv1 repoarXiv:2603.04745
FLIR-IISR
BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning
arXiv1 repoarXiv:2603.04918
BandPO
LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution
arXiv1 repoarXiv:2603.05947
LucidFlux
arXiv:2603.06007
arXiv1 repoarXiv:2603.06007
MASFactory
CHMv2: Improvements in Global Canopy Height Mapping using DINOv3
arXiv1 repoarXiv:2603.06382
dinov3
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
arXiv1 repoarXiv:2603.06688
Narrative-Weaver
WaDi: Weight Direction-aware Distillation for One-step Image Synthesis
arXiv1 repoarXiv:2603.08258
WaDi
How Far Can Unsupervised RLVR Scale LLM Training?
arXiv1 repoarXiv:2603.08660
TTRL
SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing
arXiv1 repoarXiv:2603.08982
Sparse-VideoGen
arXiv:2603.09877
arXiv1 repoarXiv:2603.09877
GenEditEvalKit
OpenClaw-RL: Train Any Agent Simply by Talking
arXiv1 repoarXiv:2603.10165
OpenClaw-RL
GLM-OCR Technical Report
arXiv1 repoarXiv:2603.10910
GLM-OCR
Detect Anything in Real Time: From Single-Prompt Segmentation to Multi-Class Detection
arXiv1 repoarXiv:2603.11441
DART
ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping
arXiv1 repoarXiv:2603.11841
wespeaker
LMEB: Long-horizon Memory Embedding Benchmark
arXiv1 repoarXiv:2603.12572
seahorse
When Drafts Evolve: Speculative Decoding Meets Online Learning
arXiv1 repoarXiv:2603.12617
OnlineSPEC
DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
arXiv1 repoarXiv:2603.12996
ParallelBench
PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
arXiv1 repoarXiv:2603.13026
PISmith
Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach
arXiv1 repoarXiv:2603.13056
CVPRW-26
Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents
arXiv1 repoarXiv:2603.13151
Awesome-OpenClaw
Diffusion Reinforcement Learning via Centered Reward Distillation
arXiv1 repoarXiv:2603.14128
GRPO
MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models
arXiv1 repoarXiv:2603.16077
mdm-prime-v2
Omnilingual MT: Machine Translation for 1,600 Languages
arXiv1 repoarXiv:2603.16309
pearmut
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
arXiv1 repoarXiv:2603.17187
MetaClaw
Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions
arXiv1 repoarXiv:2603.17522
sloptotal
CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
arXiv1 repoarXiv:2603.17829
SkyRL
Versatile Editing of Video Content, Actions, and Dynamics without Training
arXiv1 repoarXiv:2603.17989
FlowEdit
VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events
arXiv1 repoarXiv:2603.18178
Cosmos-Reason2-32B
NymeriaPlus: Enriching Nymeria Dataset with Additional Annotations and Data
arXiv1 repoarXiv:2603.18496
nymeria_dataset
Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds
arXiv1 repoarXiv:2603.18532
EmbodiedGen
The Simplicity of the Hodge Bundle
arXiv1 repoarXiv:2603.19052
superhuman
arXiv:2603.19637
arXiv1 repoarXiv:2603.19637
UniBioTransfer
MOSS-TTSD: Text to Spoken Dialogue Generation
arXiv1 repoarXiv:2603.19739
MOSS-TTS
Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams
arXiv1 repoarXiv:2603.20380
npcpy
The production of meaning in the processing of natural language
arXiv1 repoarXiv:2603.20381
npcpy
RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization
arXiv1 repoarXiv:2603.20527
repro-rmnp-row-momentum-normalized-preconditioning-for-scalable-matrix-based-optimization
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
arXiv1 repoarXiv:2603.21014
CLT-Forge
Mechanisms of Introspective Awareness
arXiv1 repoarXiv:2603.21396
introspection-mechanisms
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
arXiv1 repoarXiv:2603.21693
CEBaG
Efficient Universal Perception Encoder
arXiv1 repoarXiv:2603.22387
EUPE
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
arXiv1 repoarXiv:2603.22728
usad
MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
arXiv1 repoarXiv:2603.23067
MLLM-HWSI
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
arXiv1 repoarXiv:2603.23114
contextual_moralchoice
Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training
arXiv1 repoarXiv:2603.23885
HunyuanOCR
MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation
arXiv1 repoarXiv:2603.23896
HunyuanOCR
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
arXiv1 repoarXiv:2603.25040
Intern-S1-Pro
S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation
arXiv1 repoarXiv:2603.25702
S2D2
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
arXiv1 repoarXiv:2603.25750
sommelier
Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
arXiv1 repoarXiv:2603.25767
UTS
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
arXiv1 repoarXiv:2603.26661
GaussianGPT
TAPS: Task Aware Proposal Distributions for Speculative Sampling
arXiv1 repoarXiv:2603.27027
TAPS-Datasets
MOOZY: A Patient-First Foundation Model for Computational Pathology
arXiv1 repoarXiv:2603.27048
MOOZY
Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP
arXiv1 repoarXiv:2603.27277
codebase-memory-mcp
Falcon Perception
arXiv1 repoarXiv:2603.27365
Falcon-OCR
Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development
arXiv1 repoarXiv:2603.27460
Project-Imaging-X
Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM
arXiv1 repoarXiv:2603.27507
Chat-Scene
V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models
arXiv1 repoarXiv:2603.27650
VidCom2
ITQ3_S: High-Fidelity 3-bit LLM Inference via Interleaved Ternary Quantization with Rotation-Domain Smoothing
arXiv1 repoarXiv:2603.27914
turboquant-vllm
JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding
arXiv1 repoarXiv:2603.27942
jawildtext
Meta-Harness: End-to-End Optimization of Model Harnesses
arXiv1 repoarXiv:2603.28052
meta-harness
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
arXiv1 repoarXiv:2603.28086
MOSS-TTS
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
arXiv1 repoarXiv:2603.28458
TransArch
Rethinking Language Model Scaling under Transferable Hypersphere Optimization
arXiv1 repoarXiv:2603.28743
ohara
Gen-Searcher: Reinforcing Agentic Search for Image Generation
arXiv1 repoarXiv:2603.28767
KnowGen-Bench
MemFactory: Unified Inference & Training Framework for Agent Memory
arXiv1 repoarXiv:2603.29493
SwanLab
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
arXiv1 repoarXiv:2604.00688
OmniVoice
CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery
arXiv1 repoarXiv:2604.01658
time-series
T5Gemma-TTS Technical Report
arXiv1 repoarXiv:2604.01760
T5Gemma-TTS
Lifting Unlabeled Internet-level Data for 3D Scene Understanding
arXiv1 repoarXiv:2604.01907
SceneVersepp
RGB-Pointmap Pretraining for Unified 3D Scene Understanding
arXiv1 repoarXiv:2604.02546
UniScene3D
AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
arXiv1 repoarXiv:2604.02947
Internal-Safety-Collapse
StoryScope: Investigating idiosyncrasies in AI fiction
arXiv1 repoarXiv:2604.03136
storyscope
VOSR: A Vision-Only Generative Model for Image Super-Resolution
arXiv1 repoarXiv:2604.03225
OSEDiff
Version Control System for Data with MatrixOne
arXiv1 repoarXiv:2604.03927
matrixone
AURA: Always-On Understanding and Real-Time Assistance via Video Streams
arXiv1 repoarXiv:2604.04184
AURA
3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
arXiv1 repoarXiv:2604.04406
EmbodiedGen
MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale
arXiv1 repoarXiv:2604.04771
MinerU
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
arXiv1 repoarXiv:2604.04847
Full-Duplex-Bench
TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
arXiv1 repoarXiv:2604.04921
triattention-main
Early Stopping for Large Reasoning Models via Confidence Dynamics
arXiv1 repoarXiv:2604.04930
CoDE-Stop
StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
arXiv1 repoarXiv:2604.05014
starVLA
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
arXiv1 repoarXiv:2604.05018
academic-research-skills
This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA
arXiv1 repoarXiv:2604.05051
LLMHealthFramingEffect
TRACE: Capability-Targeted Agentic Training
arXiv1 repoarXiv:2604.05336
TRACE
WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering
arXiv1 repoarXiv:2604.05818
WikiSeeker
CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
arXiv1 repoarXiv:2604.05821
CLEAR
From Debate to Decision: Conformal Social Choice for Safe Multi-Agent Deliberation
arXiv1 repoarXiv:2604.07667
rightmind
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
arXiv1 repoarXiv:2604.08377
SkillClaw
PIArena: A Platform for Prompt Injection Evaluation
arXiv1 repoarXiv:2604.08499
PIArena
RewardFlow: Generate Images by Optimizing What You Reward
arXiv1 repoarXiv:2604.08536
RewardFlow
Enhancing LLM Problem Solving via Tutor-Student Multi-Agent Interaction
arXiv1 repoarXiv:2604.08931
rightmind
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
arXiv1 repoarXiv:2604.09057
Tora
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
arXiv1 repoarXiv:2604.09344
Sidon
Conflicts Make Large Reasoning Models Vulnerable to Attacks
arXiv1 repoarXiv:2604.09750
ConflictHarm
Pioneer Agent: Continual Improvement of Small Language Models in Production
arXiv1 repoarXiv:2604.09791
gliner2-privacy-filter-PII-multi
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
arXiv1 repoarXiv:2604.09860
RoboLab
When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling
arXiv1 repoarXiv:2604.10739
unlazy
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
arXiv1 repoarXiv:2604.10788
TInR
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
arXiv1 repoarXiv:2604.12012
tips
Nucleus-Image: Sparse MoE for Image Generation
arXiv1 repoarXiv:2604.12163
Nucleus-Image
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
arXiv1 repoarXiv:2604.14268
HY-World-2.0
VoxSafeBench: Not Just What Is Said, but Who, How, and Where
arXiv1 repoarXiv:2604.14548
VoxSafeBench
Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models
arXiv1 repoarXiv:2604.14591
MLN
ProtoTTA: Prototype-Guided Test-Time Adaptation
arXiv1 repoarXiv:2604.15494
ProtoTTA
DIRT: Database-Integrated Random Testing
arXiv1 repoarXiv:2604.16373
turso
Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy
arXiv1 repoarXiv:2604.16923
LAPD
RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation
arXiv1 repoarXiv:2604.17243
RemoteShield
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
arXiv1 repoarXiv:2604.18556
Qwen3.8-27B-GSQ-RCO-GGUF
Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs
arXiv1 repoarXiv:2604.18576
time-series
Debating the Unspoken: Role-Anchored Multi-Agent Reasoning for Half-Truth Detection
arXiv1 repoarXiv:2604.19005
rightmind
arXiv:2604.19417
arXiv1 repoarXiv:2604.19417
MERTools
GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction
arXiv1 repoarXiv:2604.19624
graft
Semantic-Fast-SAM: Efficient Semantic Segmenter
arXiv1 repoarXiv:2604.20169
Semantic-Fast-SAM
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
arXiv1 repoarXiv:2604.21072
BloomBee
DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline
arXiv1 repoarXiv:2604.21507
diarizen-tutorial
StructMem: Structured Memory for Long-Horizon Behavior in LLMs
arXiv1 repoarXiv:2604.21748
LightMem
Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents
arXiv1 repoarXiv:2604.22085
memanto
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
arXiv1 repoarXiv:2604.22328
Energy_Benchmark_TSFM_pub
MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models
arXiv1 repoarXiv:2604.23321
MMEB-V3
Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior in Generative Retrieval
arXiv1 repoarXiv:2604.23396
lost-in-decoding
IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance
arXiv1 repoarXiv:2604.23446
AssetOpsBench
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis
arXiv1 repoarXiv:2604.24198
DataMind
arXiv:2604.24622
arXiv1 repoarXiv:2604.24622
CF-VLA
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
arXiv1 repoarXiv:2604.24763
tuna-2
MAIC-UI: Making Interactive Courseware with Generative UI
arXiv1 repoarXiv:2604.25806
MAIC-UI
DeepTutor: Towards Agentic Personalized Tutoring
arXiv1 repoarXiv:2604.26962
DeepTutor
MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons
arXiv1 repoarXiv:2604.28130
MoCapAnythingV2-weights
Model Compression with Exact Budget Constraints via Riemannian Manifolds
arXiv1 repoarXiv:2605.00649
Qwen3.8-27B-GSQ-RCO-GGUF
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
arXiv1 repoarXiv:2605.00658
UniVidX
Co-Generative De Novo Functional Protein Design
arXiv1 repoarXiv:2605.00948
OpenBioMed
Shao: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
arXiv1 repoarXiv:2605.01790
Shao
Stochastic Sparse Attention for Memory-Bound Inference
arXiv1 repoarXiv:2605.01910
repro-stochastic-sparse-attention-for-memory-bound-inference
Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning
arXiv1 repoarXiv:2605.02263
Block-R1
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
arXiv1 repoarXiv:2605.02757
Seeing-Realism-from-Simulation
Compositional Neural-Cyber-Physical System Verification in the Interactive Theorem Prover of Your Choice
arXiv1 repoarXiv:2605.02790
analysis
Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based Sanitization
arXiv1 repoarXiv:2605.02977
contrastive-privacy
Ortho-Hydra: Orthogonalized Experts for DiT LoRA
arXiv1 repoarXiv:2605.03252
anima_lora
Segmenting Human-LLM Co-authored Text via Change Point Detection
arXiv1 repoarXiv:2605.03723
DetectLLMSegmentation
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
arXiv1 repoarXiv:2605.03937
minimind-o
arXiv:2605.05115
arXiv1 repoarXiv:2605.05115
drowse
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents
arXiv1 repoarXiv:2605.05185
OpenSearch-VL
MidSteer: Optimal Affine Framework for Steering Generative Models
arXiv1 repoarXiv:2605.05220
MidSteer
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
arXiv1 repoarXiv:2605.05611
X-Voice
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
arXiv1 repoarXiv:2605.06058
CoExVQA
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
arXiv1 repoarXiv:2605.06376
anima-turbo-4step
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
arXiv1 repoarXiv:2605.06388
semantic-wm
SparseForge: Efficient Semi-Structured LLM Sparsification via Annealing of Hessian-Guided Soft-Mask
arXiv1 repoarXiv:2605.06402
SparseForge
Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models
arXiv1 repoarXiv:2605.06510
tfmlens
MIND: Monge Inception Distance for Generative Models Evaluation
arXiv1 repoarXiv:2605.06797
RAEv2
Narrow Secret Loyalty Dodges Black-Box Audits
arXiv1 repoarXiv:2605.06846
whitebox-affordance-ladder
Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment
arXiv1 repoarXiv:2605.06885
Open-dLLM
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
arXiv1 repoarXiv:2605.07363
TransArch
LLM hallucinations in the wild: Large-scale evidence from non-existent citations
arXiv1 repoarXiv:2605.07723
academic-research-skills
Benchmarking Foundation Models for Renal Lesion Stratification in CT
arXiv1 repoarXiv:2605.07749
RenalVision
Anisotropic Modality Align
arXiv1 repoarXiv:2605.07825
Modality_Gap_Theory
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
arXiv1 repoarXiv:2605.08029
ml-starflow
Normalizing Trajectory Models
arXiv1 repoarXiv:2605.08078
ml-starflow
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
arXiv1 repoarXiv:2605.08366
big-pickle-swe-atlas
arXiv:2605.08703
arXiv1 repoarXiv:2605.08703
EditReward
Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal
arXiv1 repoarXiv:2605.09203
watermarks-remover
Attention Drift: What Autoregressive Speculative Decoding Models Learn
arXiv1 repoarXiv:2605.09992
Attention-Drift
Continual Harness: Online Adaptation for Self-Improving Foundation Agents
arXiv1 repoarXiv:2605.09998
prime-agent
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
arXiv1 repoarXiv:2605.10912
WildClawBench
ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
arXiv1 repoarXiv:2605.11086
exploitgym
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models
arXiv1 repoarXiv:2605.11726
Block-R1
Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters
arXiv1 repoarXiv:2605.11960
HunyuanOCR
L2P: Unlocking Latent Potential for Pixel Generation
arXiv1 repoarXiv:2605.12013
L2P
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
arXiv1 repoarXiv:2605.12487
task-aware-embedding-refinement
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
arXiv1 repoarXiv:2605.12495
AlphaGRPO
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
arXiv1 repoarXiv:2605.12500
NEO
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
arXiv1 repoarXiv:2605.12501
Phi-Ground-Any
Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing
arXiv1 repoarXiv:2605.12876
icml26-hybrid-randomized-smoothing
Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics
arXiv1 repoarXiv:2605.13171
formal-conjectures
LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics
arXiv1 repoarXiv:2605.13412
RAB-Cred
arXiv:2605.13782
arXiv1 repoarXiv:2605.13782
LMPath
TabPFN-3: Technical Report
arXiv1 repoarXiv:2605.13986
TabPFN
SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks
arXiv1 repoarXiv:2605.14051
AssetOpsBench
Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization
arXiv1 repoarXiv:2605.14373
repro-turning-stale-gradients-into-stable-gradients-coherent-coordinate-descent
Nexus : An Agentic Framework for Time Series Forecasting
arXiv1 repoarXiv:2605.14389
time-series
DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making
arXiv1 repoarXiv:2605.14403
DermAgent
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
arXiv1 repoarXiv:2605.15250
TransArch
Tweedie's Formula and Score-Driven Updating
arXiv1 repoarXiv:2605.15902
skaters
Generative 3D Gaussians with Learned Density Control
arXiv1 repoarXiv:2605.16355
TripoSplat
TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT
arXiv1 repoarXiv:2605.16572
TriALS
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
arXiv1 repoarXiv:2605.16842
Lumina-DiMOO
Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers
arXiv1 repoarXiv:2605.16941
WINO-DLLM
AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs
arXiv1 repoarXiv:2605.17535
rebuild-dossier
Stable Audio 3
arXiv1 repoarXiv:2605.17991
stable-audio-3
Improved Baselines with Representation Autoencoders
arXiv1 repoarXiv:2605.18324
RAEv2
SAME: A Semantically-Aligned Music Autoencoder
arXiv1 repoarXiv:2605.18613
same-l-decoder-lora
Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees
arXiv1 repoarXiv:2605.18654
sqlite-predict
D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting
arXiv1 repoarXiv:2605.18810
SpecForge
arXiv:2605.19130
arXiv1 repoarXiv:2605.19130
egobabyvlm
PhyWorld: Physics-Faithful World Model for Video Generation
arXiv1 repoarXiv:2605.19242
PhyWorld
Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR
arXiv1 repoarXiv:2605.19282
Pion
Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance
arXiv1 repoarXiv:2605.19628
understanding-wacky-weights
Stage-adaptive Token Selection for Efficient Omni-modal LLMs
arXiv1 repoarXiv:2605.20035
SEATS
Toto 2.0: Time Series Forecasting Enters the Scaling Era
arXiv1 repoarXiv:2605.20119
toto
Tiny-Engram: Trigger-Indexed Concept Tables for Generative Vision
arXiv1 repoarXiv:2605.20309
TinyEngram
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
arXiv1 repoarXiv:2605.20630
AssetOpsBench
Latent-space Attacks for Refusal Evasion in Language Models
arXiv1 repoarXiv:2605.21706
latent-evasion
Swift Sampling: Selecting Temporal Surprises via Taylor Series
arXiv1 repoarXiv:2605.22678
SwiftSampling
Advancing Mathematics Research with AI-Driven Formal Proof Search
arXiv1 repoarXiv:2605.22763
formal-conjectures
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
arXiv1 repoarXiv:2605.22791
GatedDeltaNet-2
Cambrian-P: Pose-Grounded Video Understanding
arXiv1 repoarXiv:2605.22819
cambrian-p
FastKernels: Benchmarking GPU Kernel Generation in Production
arXiv1 repoarXiv:2605.23215
fastkernels
From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills
arXiv1 repoarXiv:2605.23899
darwin-skill
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
arXiv1 repoarXiv:2605.23908
picbreeder-vlm-archive
Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content
arXiv1 repoarXiv:2605.24421
gaslit-aisoc
ECHO: Terminal Agents Learn World Models for Free
arXiv1 repoarXiv:2605.24517
SkyRL
Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance
arXiv1 repoarXiv:2605.24953
AssetOpsBench
Leveraging Gauge Freedom for Learning Non-Gradient Population Dynamics of Stochastic Systems
arXiv1 repoarXiv:2605.25107
repro-leveraging-gauge-freedom-for-non-gradient-population-dynamics
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning
arXiv1 repoarXiv:2605.26292
BiomedCoOp
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
arXiv1 repoarXiv:2605.26874
AssetOpsBench
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
arXiv1 repoarXiv:2605.27062
sopro
Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?
arXiv1 repoarXiv:2605.27881
SwanLab
ABot-OCR Technical Report
arXiv1 repoarXiv:2605.27978
ocr
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
arXiv1 repoarXiv:2605.28091
Qwen-Image-Flash
Rethinking Memory as Continuously Evolving Connectivity
arXiv1 repoarXiv:2605.28773
LightMem
From Pixels to Words -- Towards Native One-Vision Models at Scale
arXiv1 repoarXiv:2605.28820
NEO
Draft-OPD: On-Policy Distillation for Speculative Draft Models
arXiv1 repoarXiv:2605.29343
Draft-OPD
Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting
arXiv1 repoarXiv:2605.29401
time-series
DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?
arXiv1 repoarXiv:2605.29615
layoutlens
VikingMem: A Memory Base Management System for Stateful LLM-based Applications
arXiv1 repoarXiv:2605.29640
OpenViking
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
arXiv1 repoarXiv:2605.29659
opir-multilang-onnx
Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
arXiv1 repoarXiv:2605.29707
SpecForge
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis
arXiv1 repoarXiv:2605.30434
DataMind
APE: Agentic Prompt Enhancer for Image Generation and Editing
arXiv1 repoarXiv:2606.00204
UniGenBench
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
arXiv1 repoarXiv:2606.00206
2x-3090-GA102-300-A1-sglang-inference
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
arXiv1 repoarXiv:2606.01348
HunyuanOCR
UniD$^3$: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning
arXiv1 repoarXiv:2606.01394
UniD3
Bridging the Last Mile of Time Series Forecasting with LLM Agents
arXiv1 repoarXiv:2606.02497
time-series
Cosmos 3: Omnimodal World Models for Physical AI
arXiv1 repoarXiv:2606.02800
UniGenBench
From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting
arXiv1 repoarXiv:2606.03097
time-series
LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks
arXiv1 repoarXiv:2606.03303
superhuman
KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks
arXiv1 repoarXiv:2606.03458
KVarN
AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task
arXiv1 repoarXiv:2606.03967
WhisperLiveKit
Self-Distilled Policy Gradient
arXiv1 repoarXiv:2606.04036
halo
Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have
arXiv1 repoarXiv:2606.05107
dinov3
GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
arXiv1 repoarXiv:2606.05160
PhysicalAI-Robotics-Locomanipulation-GRAIL
AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents
arXiv1 repoarXiv:2606.05597
webgym
T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation
arXiv1 repoarXiv:2606.05975
T-FunS3D
Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios
arXiv1 repoarXiv:2606.06177
ouvia
RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
arXiv1 repoarXiv:2606.06256
RedKnot
Unsupervised Skill Discovery for Agentic Data Analysis
arXiv1 repoarXiv:2606.06416
DataMind
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
arXiv1 repoarXiv:2606.06444
usad
VoxCPM2 Technical Report
arXiv1 repoarXiv:2606.06928
VoxCPM
arXiv:2606.09516
arXiv1 repoarXiv:2606.09516
SwiftVR
End-to-End Context Compression at Scale
arXiv1 repoarXiv:2606.09659
LCLM
Collaborative Human-Agent Protocol (CHAP)
arXiv1 repoarXiv:2606.09751
chap
Rethinking the Divergence Regularization in LLM RL
arXiv1 repoarXiv:2606.09821
UniRL
Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
arXiv1 repoarXiv:2606.10968
UniRL
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
arXiv1 repoarXiv:2606.11025
UniRL
M*: A Modular, Extensible, Serving System for Multimodal Models
arXiv1 repoarXiv:2606.12688
mstar
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
arXiv1 repoarXiv:2606.13652
world-tracing-demo
FastContext: Training Efficient Repository Explorer for Coding Agents
arXiv1 repoarXiv:2606.14066
caveman
Classifying by Proxy: Explainable and Reproducible Ensemble of Proxy Tasks for Child Sexual Abuse Imagery Classification
arXiv1 repoarXiv:2606.15993
EISPCSAI
Scaling Human and G2P Supervision for Robust Phonetic Transcription
arXiv1 repoarXiv:2606.16019
ML
Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents
arXiv1 repoarXiv:2606.16038
Open-SWE-Traces
arXiv:2606.17404
arXiv1 repoarXiv:2606.17404
ELSA
MagicSim: A Unified Infrastructure for Executable Embodied Interaction
arXiv1 repoarXiv:2606.17511
robodojo_long
Vision-language models for chest radiography do not always need the image
arXiv1 repoarXiv:2606.17710
causal
HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning
arXiv1 repoarXiv:2606.17833
HumanoidArena
Uncertainty Quantification for Flow-Based Vision-Language-Action Models
arXiv1 repoarXiv:2606.18043
uq_vla
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection
arXiv1 repoarXiv:2606.19881
REDACT-PII-Benchmark
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
arXiv1 repoarXiv:2606.20177
NeFo
TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation
arXiv1 repoarXiv:2606.20709
TeleStyleV2
Tmax: A simple recipe for terminal agents
arXiv1 repoarXiv:2606.23321
tmax
RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild
arXiv1 repoarXiv:2606.23344
PP-DocLayoutV3_safetensors
DiffusionBench: On Holistic Evaluation of Diffusion Transformers
arXiv1 repoarXiv:2606.24888
diffusion-bench
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
arXiv1 repoarXiv:2606.24893
agentodyssey
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents
arXiv1 repoarXiv:2606.25556
SwanLab
Towards an Interactive Evidence-RAG Peer-Review Workspace for the Journal of Digital History
arXiv1 repoarXiv:2606.25837
journal-of-digital-history-evidence-rag
Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs
arXiv1 repoarXiv:2606.26387
figures4papers
NaviCache: Test-Time Self-Calibration Caching for Video Generation
arXiv1 repoarXiv:2606.26795
HunyuanVideo
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
arXiv1 repoarXiv:2606.26947
DyRef
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts
arXiv1 repoarXiv:2606.27881
temporal-ner
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
arXiv1 repoarXiv:2606.28128
PhysisForcing
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
arXiv1 repoarXiv:2606.28276
SimFoundry
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
arXiv1 repoarXiv:2606.29354
LSF_MDia
StrucTab: A Structured Optimization Framework for Table Parsing
arXiv1 repoarXiv:2606.29905
HunyuanOCR
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks
arXiv1 repoarXiv:2606.31074
triospect
DA-Studio: An Agentic System for End-to-End Data Analysis
arXiv1 repoarXiv:2606.31423
DeepAnalyze
Large Databases Need Small, Open-Weight Language Models
arXiv1 repoarXiv:2606.31808
blendsql
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
arXiv1 repoarXiv:2606.32028
DVG-WM
EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
arXiv1 repoarXiv:2607.01147
EquiSteer
TiRex-2: Generalizing TiRex to Multivariate Data and Streaming
arXiv1 repoarXiv:2607.01204
tirex
AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
arXiv1 repoarXiv:2607.02269
AnyGroundBench
GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training
arXiv1 repoarXiv:2607.02486
Geomix
arXiv:2607.03502
arXiv1 repoarXiv:2607.03502
Filler_Token
arXiv:2607.03788
arXiv1 repoarXiv:2607.03788
tensor-train
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
arXiv1 repoarXiv:2607.04438
ppt-master
Unified Audio Intelligence Without Regressing on Text Intelligence
arXiv1 repoarXiv:2607.05196
Nemotron-Labs-Audex-30B-A3B
Vision Pretraining for Dense Spatial Perception
arXiv1 repoarXiv:2607.05247
lingbot-vision
LLM-as-a-Verifier: A General-Purpose Verification Framework
arXiv1 repoarXiv:2607.05391
dsh-plugin-llm-verifier
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
arXiv1 repoarXiv:2607.06514
stable-retro
EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI
arXiv1 repoarXiv:2607.07459
EmbodiedGen
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
arXiv1 repoarXiv:2607.08646
UltraX-Preview
StereoSplat+: Feed-Forward Stereo Gaussian Splatting with Diffusion-Assisted Progressive Inference
arXiv1 repoarXiv:2607.08808
StereoSplat_Plus
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
arXiv1 repoarXiv:2607.09324
crowdsourced-data-tools
Index SLM Technical Report
arXiv1 repoarXiv:2607.09885
Index-1.9B
GigaAM Multilingual: Foundation Model for Underrepresented Languages
arXiv1 repoarXiv:2607.10371
GigaAM
SETA: Scaling Environments for Terminal Agents
arXiv1 repoarXiv:2607.10891
seta
Higher-Order Cell Tracking Transformer
arXiv1 repoarXiv:2607.11754
rlx-models
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
arXiv1 repoarXiv:2607.11886
AlphaGRPO
Let RGB Be the Language of Vision
arXiv1 repoarXiv:2607.12450
RINO
Self-Improvements in Modern Agentic Systems: A Survey
arXiv1 repoarXiv:2607.13104
academic-research-skills
UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets
arXiv1 repoarXiv:2607.13586
UniPhys-Bench
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
arXiv1 repoarXiv:2607.14846
asr-benchmark-optimization
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
arXiv1 repoarXiv:2607.18218
prov-gigapath
Patch Policy: Efficient Embodied Control via Dense Visual Representations
arXiv1 repoarXiv:2607.18236
patch_policy
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
arXiv1 repoarXiv:2607.19191
ABot-World-Explorer-500h
PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
arXiv1 repoarXiv:2607.20064
PRO-LONG
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
arXiv1 repoarXiv:2607.20709
labs-OO-Agents
arXiv:2607.21557
arXiv1 repoarXiv:2607.21557
OpenForge-RL
RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
arXiv1 repoarXiv:2607.21927
riskernel
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model
arXiv1 repoarXiv:2607.22083
Nanbeige4.2-3B
Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
arXiv1 repoarXiv:2607.22563
AssetOpsBench
Modeling Local Exploit Hazard - A Bayesian Framework for Quantifying Exploit Risk and Operational Efficiency
arXiv1 repoarXiv:2607.24618
researches
Specula: Scaling formal specifications for autonomous model checking of system code
arXiv1 repoarXiv:2607.25333
Specula
Coevolution of epidemic dynamics and network topology driven by disease fatality and waning immunity
arXiv1 repoarXiv:2607.25475
sci-ssci-skills
A Distributional Robustness Margin For Pathology Foundation Models
arXiv1 repoarXiv:2607.25497
croma
Revisiting the Algebraic Foundation of Relational Data
arXiv1 repoarXiv:2607.26356
prela
Contrastive ESA: Human Evaluation of Multiple Translations at Once
arXiv1 repoarXiv:2607.26640
pearmut
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
arXiv1 repoarXiv:2607.26754
stable-retro
Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences
arXiv1 repoarXiv:2607.26973
sat-bundleadjust
VETO: Towards Protecting Images From Frontier AI Editing
arXiv1 repoarXiv:2607.27292
VetoBench
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
arXiv1 repoarXiv:2607.28625
ACE-Data-0
TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving
arXiv1 repoarXiv:2607.29678
toktier
DiffusionGemma Technical Report
arXiv1 repoarXiv:2608.00146
awesome-gemma
TEngineDB-V: An OLAP-Native Vector Search System for Large-$k$ Workloads at Tencent
arXiv1 repoarXiv:2608.00650
parqdb
FATE: Frame-Level Audio-Visual Temporal Embedding
arXiv1 repoarXiv:2608.01310
FATE
arXiv:2608.01507
arXiv1 repoarXiv:2608.01507
ripwire
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
arXiv1 repoarXiv:2608.03403
ReMe
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
arXiv1 repoarXiv:2608.03979
Vision-DeepResearch
When does training on downscaled images yield the same gradients?
arXiv1 repoarXiv:2608.04448
anima-turbo-4step
K-EXAONE 2.0 Technical Report
arXiv1 repoarXiv:2608.04505
K-EXAONE-2.0-750B-A37B
TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
arXiv1 repoarXiv:2608.06223
raggy
Retrofitting Linear Attention into Diffusion Language Models
arXiv1 repoarXiv:2608.06628
LLaDA-Hybrid
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
arXiv1 repoarXiv:2608.06867
xRouteBench
Vernata: Self-Supervised Learning of LiDAR Point Representations
arXiv1 repoarXiv:2608.06919
vernata
The Spectral Neuron
arXiv1 repoarXiv:2608.08003
spectral_neuron_paper
Motif 3: Technical Report
arXiv1 repoarXiv:2608.09119
Motif-3
Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning
arXiv1 repoarXiv:2608.10438
continuous-interaction-diffusion
AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations
arXiv1 repoarXiv:2608.11123
AlbumentationsX
HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing
arXiv1 repoarXiv:2608.12122
HandEdit
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
arXiv1 repoarXiv:2608.13546
Evoke
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
arXiv1 repoarXiv:2608.13966
Qwen3.8-27B-QUASAR-NVFP4
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
arXiv1 repoarXiv:2608.14577
Internal-Safety-Collapse
Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI
arXiv1 repoarXiv:2608.16319
relarena
Le Critique: Privileged Value Functions for LLM Reinforcement Learning
arXiv1 repoarXiv:2608.16739
prime-values
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
arXiv1 repoarXiv:2608.16812
ConceptEdit-12M
Agent Lightning v1.0: Towards Harnessed Agentic RL
arXiv1 repoarXiv:2608.17528
agent-lightning
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability
arXiv1 repoarXiv:2608.19338
observerbench
PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks
arXiv1 repoarXiv:2608.21381
PersonaMem-v2
Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
arXiv1 repoarXiv:2608.22602
accel-sim-framework
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
arXiv1 repoarXiv:2608.23041
AutoSaddler
Prime Agent: A Self-Improving RLM Harness
arXiv1 repoarXiv:2608.23552
prime-agent
Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal
arXiv1 repoarXiv:2608.23616
rebuild-dossier
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
arXiv1 repoarXiv:2608.23691
station
Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
arXiv1 repoarXiv:2608.23873
semantic-overlays
pigzpp: Fast, Parallel, Portable Compression for the Whole Stack
arXiv1 repoarXiv:2608.24153
pigzpp
Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
arXiv1 repoarXiv:2608.24188
paritok-4b-v1
Meta$^n$: Recursive Self-Improvement through Emergent Depth
arXiv1 repoarXiv:2608.24735
meta-n
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
arXiv1 repoarXiv:2608.24845
BVD
post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding at Synthesis
arXiv1 repoarXiv:2608.24921
post-graph-rag
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
arXiv1 repoarXiv:2608.25218
turnbench
Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images
arXiv1 repoarXiv:2608.29348
TotalSegmentator
SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature
arXiv1 repoarXiv:2608.30214
Spark-234K
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
arXiv1 repoarXiv:2609.00111
Qwen-Drive-1.0-4B
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
arXiv1 repoarXiv:2609.00551
LightMem
Bandits in Prod: Hyperparameter Optimization at Inference Time
arXiv1 repoarXiv:2609.01335
IMABO
H3-World: Turning Language Understanding into World Control
arXiv1 repoarXiv:2609.01560
H3-World
VibeVoice-ASR-Streaming Technical Report
arXiv1 repoarXiv:2609.02812
VibeVoice-ASR-Streaming-7B
Last Translation Benchmark
arXiv1 repoarXiv:2609.04173
last-translation-benchmark
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
arXiv1 repoarXiv:2609.04196
Puffin-16M
Search in Power-Law Networks
arXiv1 repoarXiv:cs/0103016
littleballoffur
Adaptive Interaction Using the Adaptive Agent Oriented Software Architecture (AAOSA)
arXiv1 repoarXiv:cs/9812015
neuro-san-studio
Google AI beats top human players at strategy game StarCraft II
Nature1 repoNature:d41586-019-03298-6
learn-claude-code
Human-level control through deep reinforcement learning
Nature1 repoNature:nature14236
learn-claude-code
The 4D nucleome project
Nature1 repoNature:nature23884
alphagenome_research
Nature:nbt
Nature1 repoNature:nbt
nextflow
Nature:ng
Nature1 repoNature:ng
Histopathology_Benchmark
Circularly and elliptically polarized light under water and the Umov effect
Nature1 repoNature:s41377-019-0143-0
Awesome-Polarization
Nature-inspired chiral metasurfaces for circular polarization detection and full-Stokes polarimetric measurements
Nature1 repoNature:s41377-019-0184-4
Awesome-Polarization
Flavonoid intake is associated with lower mortality in the Danish Diet Cancer and Health Cohort
Nature1 repoNature:s41467-019-11622-x
HowToLiveLonger
The 4D Nucleome Data Portal as a resource for searching and visualizing curated nucleomics data
Nature1 repoNature:s41467-022-29697-4
alphagenome_research
Out-of-distribution generalization for learning quantum dynamics
Nature1 repoNature:s41467-023-39381-w
Awesome-Out-Of-Distribution-Detection
Augmenting interpretable models with large language models during training
Nature1 repoNature:s41467-023-43713-1
imodelsX
Sopa: a technology-invariant pipeline for analyses of image-based spatial omics
Nature1 repoNature:s41467-024-48981-z
sopa
Towards building multilingual language model for medicine
Nature1 repoNature:s41467-024-52417-z
MMedLM
Quantifying the reasoning abilities of LLMs on clinical cases
Nature1 repoNature:s41467-025-64769-1
MedRBench
A generalizable pathology foundation model using a unified knowledge distillation pretraining framework
Nature1 repoNature:s41551-025-01488-4
GPFM
A new era in functional genomics screens
Nature1 repoNature:s41576-021-00409-w
awesome-CRISPR
Highly accurate protein structure prediction with AlphaFold
Nature1 repoNature:s41586-021-03819-2
paper-reading
Highly accurate protein structure prediction for the human proteome
Nature1 repoNature:s41586-021-03828-1
alphafold3
Advancing mathematics by guiding human intuition with AI
Nature1 repoNature:s41586-021-04086-x
paper-reading
Solving olympiad geometry without human demonstrations
Nature1 repoNature:s41586-023-06747-5
superhuman
Accurate structure prediction of biomolecular interactions with AlphaFold 3
Nature1 repoNature:s41586-024-07487-w
alphafold3
A multimodal generative AI copilot for human pathology
Nature1 repoNature:s41586-024-07618-3
TITAN
A Drosophila computational brain model reveals sensorimotor processing
Nature1 repoNature:s41586-024-07763-9
fruit-fly-simulation
Larger and more instructable language models become less reliable
Nature1 repoNature:s41586-024-07930-y
llm-reliability
Scalable watermarking for identifying large language model outputs
Nature1 repoNature:s41586-024-08025-4
watermarks-remover
Functional evaluation and clinical classification of BRCA2 variants
Nature1 repoNature:s41586-024-08388-8
carbon
Mastering diverse control tasks through world models
Nature1 repoNature:s41586-025-08744-2
UI-TARS
Eye structure shapes neuron function in Drosophila motion vision
Nature1 repoNature:s41586-025-09276-5
fly
Advancing regulatory variant effect prediction with AlphaGenome
Nature1 repoNature:s41586-025-10014-0
alphagenome
Operational tropical cyclone forecasting with AI
Nature1 repoNature:s41586-026-10953-2
weathernext
A DNA language model based on multispecies alignment predicts the effects of genome-wide variants
Nature1 repoNature:s41587-024-02511-w
clinvar-vep-final
A guide to deep learning in healthcare
Nature1 repoNature:s41591-018-0316-z
Awesome-AI4Med
Evaluation and accurate diagnoses of pediatric diseases using artificial intelligence
Nature1 repoNature:s41591-018-0335-9
Awesome-AI4Med
Magnitude, demographics and dynamics of the effect of the first wave of the COVID-19 pandemic on all-cause mortality in 21 industrialized countries
Nature1 repoNature:s41591-020-1112-0
HowToLiveLonger
Optimized glycemic control of type 2 diabetes with reinforcement learning: a proof-of-concept trial
Nature1 repoNature:s41591-023-02552-9
DiagnosisCoding
Transparent medical image AI via an image–text foundation model grounded in medical literature
Nature1 repoNature:s41591-024-02887-x
DermFM-Zero
A multimodal vision foundation model for clinical dermatology
Nature1 repoNature:s41591-025-03747-y
PanDerm
Holistic evaluation of large language models for medical tasks with MedHELM
Nature1 repoNature:s41591-025-04151-2
helm
nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation
Nature1 repoNature:s41592-020-01008-z
TriALS
Alevin-fry unlocks rapid, accurate and memory-frugal quantification of single-cell RNA-seq data
Nature1 repoNature:s41592-022-01408-3
alevin-fry
Cellpose 2.0: how to train your own model
Nature1 repoNature:s41592-022-01663-4
cellpose
A cellular segmentation algorithm with fast customization
Nature1 repoNature:s41592-022-01664-3
cellpose
Cellpose3: one-click image restoration for improved cellular segmentation
Nature1 repoNature:s41592-025-02595-5
cellpose
A visual–omics foundation model to bridge histopathology with spatial transcriptomics
Nature1 repoNature:s41592-025-02707-1
AtlasPatch
Novae: a graph-based foundation model for spatial transcriptomics data
Nature1 repoNature:s41592-025-02899-6
sopa
High-parameter spatial multi-omics through histology-anchored integration
Nature1 repoNature:s41592-025-02926-6
SpatialEx
The ENERTALK dataset, 15 Hz electricity consumption data from 22 houses in Korea
Nature1 repoNature:s41597-019-0212-5
awesome-nilm
MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports
Nature1 repoNature:s41597-019-0322-0
rad-dino
A voltage and current measurement dataset for plug load appliance identification in households
Nature1 repoNature:s41597-020-0389-7
awesome-nilm
Kvasir-Capsule, a video capsule endoscopy dataset
Nature1 repoNature:s41597-021-00920-z
Endo-FM
The IDEAL household energy dataset, electricity, gas, contextual sensor data and survey data for 255 UK homes
Nature1 repoNature:s41597-021-00921-y
awesome-nilm
PAPILA: Dataset with fundus images and clinical data of both eyes of the same patient for glaucoma assessment
Nature1 repoNature:s41597-022-01388-1
FairMedFM
The heterogeneous pharmacological medical biochemical network PharMeBINet
Nature1 repoNature:s41597-022-01510-3
pykeen
BRAX, Brazilian labeled chest x-ray dataset
Nature1 repoNature:s41597-022-01608-8
rad-dino
Solar and wind power data from the Chinese State Grid Renewable Energy Generation Forecasting Competition
Nature1 repoNature:s41597-022-01696-6
Energy-EVA
MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification
Nature1 repoNature:s41597-022-01721-8
medmnistc-api
Three Dimensional Lower Extremity Musculoskeletal Geometry of the Visible Human Female and Male
Nature1 repoNature:s41597-022-01905-2
Awesome-Medical-Dataset
A dataset for medical instructional video classification and question answering
Nature1 repoNature:s41597-023-02036-y
Awesome-AI4Med
A Spitzoid Tumor dataset with clinical metadata and Whole Slide Images for Deep Learning models
Nature1 repoNature:s41597-023-02585-2
Awesome-Medical-Dataset
Open Fundus Photograph Dataset with Pathologic Myopia Recognition and Anatomical Structure Annotation
Nature1 repoNature:s41597-024-02911-2
Awesome-Medical-Dataset
MAPLES-DR: MESSIDOR Anatomical and Pathological Labels for Explainable Screening of Diabetic Retinopathy
Nature1 repoNature:s41597-024-03739-6
Awesome-Medical-Dataset
AMD-SD: An Optical Coherence Tomography Image Dataset for wet AMD Lesions Segmentation
Nature1 repoNature:s41597-024-03844-6
Awesome-Medical-Dataset
FARFUM-RoP, A dataset for computer-aided detection of Retinopathy of Prematurity
Nature1 repoNature:s41597-024-03897-7
Awesome-Medical-Dataset
An open-access lumbosacral spine MRI dataset with enhanced spinal nerve root structure resolution
Nature1 repoNature:s41597-024-03919-4
Awesome-Medical-Dataset
LungHist700: A dataset of histological images for deep learning in pulmonary pathology
Nature1 repoNature:s41597-024-03944-3
Awesome-Medical-Dataset
Dual Leap Motion Controller 2: A Robust Dataset for Multi-view Hand Pose Recognition
Nature1 repoNature:s41597-024-03968-9
Awesome-Medical-Dataset
Investigating the Quality of DermaMNIST and Fitzpatrick17k Dermatological Image Datasets
Nature1 repoNature:s41597-025-04382-5
medmnistc-api
The “Podcast” ECoG dataset for modeling neural activity during natural language comprehension
Nature1 repoNature:s41597-025-05462-2
automated-brain-explanations
A Custom Annotated Dataset for Segmentation of Pulmonary Veins, Arteries, and Airways
Nature1 repoNature:s41597-025-06074-6
TotalSegmentator
Learning a Health Knowledge Graph from Electronic Medical Records
Nature1 repoNature:s41598-017-05778-z
Awesome-AI4Med
Automated Gleason grading of prostate cancer tissue microarrays via deep learning
Nature1 repoNature:s41598-018-30535-1
Histopathology_Benchmark
Pessimism is associated with greater all-cause and cardiovascular mortality, but optimism is not protective
Nature1 repoNature:s41598-020-69388-y
HowToLiveLonger
A large dataset of white blood cells containing cell locations and types, along with segmented nuclei and cytoplasm
Nature1 repoNature:s41598-021-04426-x
DinoBloom
Bioacoustic classification of avian calls from raw sound waveforms with an open-source deep learning architecture
Nature1 repoNature:s41598-021-95076-6
NIPS4Bplus
iCodon customizes gene expression based on the codon composition
Nature1 repoNature:s41598-022-15526-7
CodonBERT
Contrastive language and vision learning of general fashion concepts
Nature1 repoNature:s41598-022-23052-9
Cool-GenAI-Fashion-Papers
Extracting and visualizing hidden activations and computational graphs of PyTorch models with TorchLens
Nature1 repoNature:s41598-023-40807-0
torchlens
PetBERT: automated ICD-11 syndromic disease coding for outbreak detection in first opinion veterinary electronic health records
Nature1 repoNature:s41598-023-45155-7
DiagnosisCoding
Global birdsong embeddings enable superior transfer learning for bioacoustic classification
Nature1 repoNature:s41598-023-49989-z
perch
Data-driven blood glucose level prediction in type 1 diabetes: a comprehensive comparative analysis
Nature1 repoNature:s41598-024-70277-x
nocturnal-hypo-gly-prob-forecast
A multiscale model for multivariate time series forecasting
Nature1 repoNature:s41598-024-82417-4
Time-Series-Library
Multiple model visual feature embedding and selection method for an efficient ocular disease classification
Nature1 repoNature:s41598-024-84922-y
Project-Imaging-X
Asymmetric dual pathway fusion for histology driven spatial transcriptomic prediction
Nature1 repoNature:s41598-026-63426-x
STP-Bench
VetTag: improving automated veterinary diagnosis coding via large-scale language modeling
Nature1 repoNature:s41746-019-0113-1
DiagnosisCoding
A large language model for electronic health records
Nature1 repoNature:s41746-022-00742-2
DiagnosisCoding
PatchSorter: a high throughput deep learning digital pathology tool for object labeling
Nature1 repoNature:s41746-024-01150-4
PatchSorter
Towards evaluating and building versatile large language models for medicine
Nature1 repoNature:s41746-024-01390-4
MedS-Ins
HONeYBEE: enabling scalable multimodal AI in oncology through foundation model-driven embeddings
Nature1 repoNature:s41746-025-02003-4
HoneyBee
Embeddings of clinical codes enable knowledge-grounded AI in medicine
Nature1 repoNature:s41746-026-02664-9
ClinVec
Deep learning models for predicting RNA degradation via dual crowdsourcing
Nature1 repoNature:s42256-022-00571-8
CodonBERT
Evaluation of post-hoc interpretability methods in time-series classification
Nature1 repoNature:s42256-023-00620-w
Awesome-Time-Series-Explainability
Parameter-efficient fine-tuning of large-scale pre-trained language models
Nature1 repoNature:s42256-023-00626-4
LLMSurvey
Defending ChatGPT against jailbreak attack via self-reminders
Nature1 repoNature:s42256-023-00765-8
llm-jailbreaking-defense
Exploring scalable medical image encoders beyond text supervision
Nature1 repoNature:s42256-024-00965-w
rad-dino
ImmunoStruct enables multimodal deep learning for immunogenicity prediction
Nature1 repoNature:s42256-025-01163-y
figures4papers
Large language models as uncertainty-calibrated optimizers for experimental discovery
Nature1 repoNature:s42256-026-01283-z
gollum
High-content CRISPR screening
Nature1 repoNature:s43586-021-00093-4
awesome-CRISPR
On-chip solar power source for self-powered smart microsensors in bulk CMOS process
Nature1 repoNature:s44172-025-00358-w
solarcircuits
PodGPT: an audio-augmented large language model for research and education
Nature1 repoNature:s44385-025-00022-0
PodGPT
Nature:sdata20157
Nature1 repoNature:sdata20157
awesome-nilm
Nature:sdata2018251
Nature1 repoNature:sdata2018251
RefinedVision
Multi-class texture analysis in colorectal cancer histology
Nature1 repoNature:srep27988
Histopathology_Benchmark