Papers

7,217 papers (arXiv, Nature, ACM) referenced from a repo's own README — the strongest signal that the project implements or was directly influenced by it.

Quantifying the Carbon Emissions of Machine Learning

arXiv159 repos

arXiv:1910.09700

stable-diffusion-v1-4, timesfm-tourism-monthly, EleutherAI_pythia-6.9b-deduped__sft__tldr, SmolLM3-3B-QAT-Baseline-Q, KD-Tinker, MMed-Llama3.1-70B, roberta-large-mnli, distilbert-base-uncased-distilled-squad, stable-diffusion-v-1-4-original, mistral-europe_culture, mistral-northamerica_culture, mistral-news_l, Llama3-v2-iterative-DPO-iter3, Llama3-v2-iterative-DPO-iter1, mistral-reddit_l, mistral-reddit_c, Llama3-v2-iterative-DPO-iter2, mistral-africa_culture, mistral-asia_culture, mistral-southamerica_culture, mistral-news_c, mistral-news_r, mistral-reddit_r, SongGen_mixed_pro, SongGen_interleaving_A_V, flan-t5-xl, FlexCAD, stable-diffusion-x4-upscaler, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, coreml-stable-diffusion-2-1-base, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-2-base-palettized, coreml-stable-diffusion-2-1-base-palettized, coreml-stable-diffusion-2-base, stable-diffusion-2-1-base, higgs-audio-v2-tokenizer, QwQ-32B-Preview-Pruned, Marco-o1-7B-Pruned, dalle-mega, dalle-mini, flan-t5-large, sup-simcse-roberta-large, sup-simcse-bert-large-uncased, hana-alpha44, hana-small_MNIST-MODERNBERT, hana-alpha30, hana-alpha32, hana-alpha14_cifar10-128_TS-1000_1000e, hana-small_alpha8_TS-1000_400e, hana-alpha17, hana-small_alpha7, hana-small_MNIST-5e, hana-alpha45, hana-alpha42, hana-alpha39, hana-alpha38, hana-alpha31, hana-alpha29, hana-small_MNIST-MODERNBERT-3e, hana-small_alpha10, hana-alpha11_cifar10_TS-1000_500e, hana-alpha35, hana-alpha12_cifar10_TS-1000_500e, hana-small_MNIST-BATCHED, hana-alpha22, hana-alpha37, mistral-7b-hermes-rm-skywork, mistral-7b-ppo-hermes-v0.3, llama3-8b-final-ppo-v0.3, mistral-7b-hermes-dpo-v0.2, mistral-7b-hermes-cdpo-v0.2, mistral-7b-ppo-clean-hermes, prometheus-bgb-8x7b-v2.0, speecht5_tts, flan-t5-base, Memalpha-4B, stable-diffusion-2-1, llava-llama-3-8b-hqedit, ecot-openvla-7b-oxe, Mistral-7B-Instruct-v0.2-gad-bv4nogram3-merged, Mistral-7B-Instruct-v0.2-gad-slianogram3-merged, Mistral-7B-Instruct-v0.2-gad-cp8-merged, woodsoloadd_codeonly_rt_add2, wood_v2_sftr1, wood_v2_sftr4_filt, llama2-mha-from-scratch-ckpt-125M, stable-diffusion-inpainting, rad-dino, LangSAMP, GPT-J-6B-Skein, EmoWhisper-AnS-Small-v0.1, k2-v1, Llasa-1B-GRPO-2000, llama-2-7b-chat-it, llama-2-7b-chat-zh, Mistral-7B-Instruct-v0.2-es, Mixtral-8x7B-Instruct-nl, Llama-3-8B-Instruct-nl, Llama-3-8B-Instruct-it, Llama-3-8B-Instruct-es, Llama-3-8B-Instruct-fr, Llama-3-8B-Instruct-de, Llama-3-8B-Instruct-pt, Llama-3-8B-Instruct-ru, Llama-3-8B-Instruct-hi, llama-2-7b-chat-nl, llama-2-7b-chat-fr, llama-2-7b-chat-de, llama-2-7b-chat-tr, llama-2-7b-chat-bn, llama-2-7b-chat-es, llama-2-7b-chat-ru, Mistral-7B-Instruct-v0.2-nl, Mistral-7B-Instruct-v0.2-de, llama-2-13b-chat-nl, llama-2-13b-chat-es, llama-2-13b-chat-fr, llama-2-7b-chat-pt, llama-2-7b-chat-hi, llama-2-7b-chat-pl-polish-polski, llama-2-7b-chat-eu, llava-mlan-llama2-7b, llava-mlan-vicuna-7b, llava-mlan-v-llama2-7b, llava-mlan-v-vicuna-7b, quant_deepseekmath, flan-t5-small, Sentiment-Reasoning-Eng, t5-large-encoder-only-bf16, MalayaLLM-Paligemma-VQA-3B-Adapters, distilgpt2, untoken-v1, untoken-v2, bert-base-turkish-sentiment-analysis, LLaVA_MORE-llama_3_1-reasoning-finetuning, LLaVA_MORE-gemma_2_9b-siglip2-finetuning, LLaVA_MORE-gemma_2_9b-finetuning, ke-t5-base, bert-base, gemma-2b-quote-generation-82000, stable-diffusion-safety-checker, NatureLM-audio, CADFusion, Rationale_predictor, chatbot-qa-path, X-VLA-libero-object-peft, X-VLA-simpler-widowx-peft, X-VLA-libero-spatial-peft, X-VLA-libero-long-peft, X-VLA-libero-goal-peft, PredEx_Llama-2-7B_Pred-Exp_Instruction-Tuned, PredEx_Llama-2-7B_Pred-Exp, show-o-512x512, show-o-w-clip-vit-512x512, show-o-512x512-wo-llava-tuning, UniRL, stable-diffusion-v1.5-webnn

LLaMA: Open and Efficient Foundation Language Models

arXiv72 repos

arXiv:2302.13971

baize-lora-7B, ik_llama.cpp, llama.cpp, stanford_alpaca, api-for-open-llm, OpenOrca, HuatuoGPT, PMC-LLaMA, RAIN, llama.cpp, open_llama_7b_v2, llama.cpp.qwen2.5vl, vicuna-7B-1.1-HF, GPTQ-for-LLaMa, openalpaca_7b_700bt_preview, openalpaca_3b_600bt_preview, Llama-Chinese, atomic-llama-cpp-turboquant, Llama-Chinese, Llama-X, Chinese-Vicuna, vicuna-13b-delta-v1.1, vicuna-13B-1.1-HF, vicuna-7b-delta-v1.1, alpaca-7b-reproduced, beaver-7b-v2.0-cost, beaver-7b-v1.0-cost, beaver-7b-unified-cost, beaver-7b-v3.0-cost, beaver-7b-v3.0-reward, beaver-7b-v1.0, beaver-7b-v3.0, beaver-7b-v2.0-reward, beaver-7b-v2.0, beaver-7b-v1.0-reward, beaver-7b-unified-reward, baize-lora-30B, baize-healthcare-lora-7b, baize-lora-13B, mpt-1b-redpajama-200b, mpt-1b-redpajama-200b-dolly, vicuna-7b-delta-v0, ClinicalNLP_PMCVQA, RRHF, calm2-7b-chat, calm2-7B-chat-GPTQ, llama2_chinese, JobList, llama.cpp, Chinese-Vicuna, vicuna-13b-v1.3, KnowLM, beaver-dam-7b, llmtools, open_llama_7b, GPlatty-30B, Platypus-30B, SuperPlatty-30B, llama-chat, llama.cpp, vicuna-7b-v1.3, llmtools, koalpaca, OpenOrca-KO, kwater, KOR-OpenOrca-Platypus, OpenOrca-gugugo-ko, medalpaca, llama2.zig, llama2.scala, llama2.java, llama2.cs

Llama 2: Open Foundation and Fine-Tuned Chat Models

arXiv54 repos

arXiv:2307.09288

StableBeluga2, RAIN, eCeLLM, Llama-Chinese, Llama-Chinese, Chinese-Llama-2-7b, heron-chat-git-Llama-2-7b-v0, heron-preliminary-git-Llama-2-70b-v0, heron-chat-git-ELYZA-fast-7b-v0, ELYZA-japanese-Llama-2-7b-fast, Llama-2-7B-Chat-GGUF, Llama-2-70B-chat-AWQ, Llama-2-13B-Chat-GGUF, ELYZA-japanese-Llama-2-7b-fast-instruct, llama2_chinese, Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-GPTQ, speechless-llama2-13b, Speechless-Llama2-13B-GGML, speechless-llama2-dolphin-orca-platypus-13b, Speechless-Llama2-13B-GPTQ, speechless-llama2-hermes-orca-platypus-wizardlm-13b, Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-GGUF, Speechless-Llama2-Hermes-Orca-Platypus-WizardLM-13B-AWQ, speechless-llama2-hermes-orca-platypus-13b, speechless-llama2-luban-orca-platypus-13b, Speechless-Llama2-13B-GGUF, ELYZA-japanese-CodeLlama-7b-instruct, Llama-2-7B-Chat-GGML, Llama-2-13B-Chat-GGML, llama-2-7b-chat, Llama-2-7b-Chat-GPTQ, llama-2-13b-chat, Explore_llamav2_with_TGI, codellama-13b-chat, Platypus2-70B, Platypus2-70B-instruct, Camel-Platypus2-13B, Platypus2-13B, Platypus-70B-adapters, Camel-Platypus2-70B, Stable-Platypus2-13B, Platypus2-7B, Platypus-7B-adapters, Platypus-13B-adapters, corningQA-llama2-13b-chat, genai-llama2-infer-vertex, EMGerman-LLama-Streamlit-Chatbot, ContraDecode, KO-Platypus2-7B-ex, Llama-2-ko-7b-Chat, StableBeluga-7B, Llama-2-7B-GPTQ, Llama-2-13B-Chat-GPTQ, llama2.zig

LoRA: Low-Rank Adaptation of Large Language Models

arXiv52 repos

arXiv:2106.09685

prompt-api, YiVal, Qwen, writing-assistance-apis, stanford_alpaca, MindSpeed-MM, deepseek-r1-finetune, Continual-NExT, RWKV-LM-LoRA, SanAssist, Llama-Chinese, Llama-Chinese, Chinese-Vicuna, quickllm, LoRA, h2o-llmstudio, adapt_med_seg, vigogne, BLOOM-LORA, colpali, colpali-v1.2, colpali-v1.1, colpali-v1.3, colqwen2-v1.0, colqwen2.5-v0.2, colqwen2.5-v0.1, colSmol-500M, colSmol-256M, colqwen2-v0.1, llama2_chinese, NTU_ADL_Team11_Final, Bloom-Lora, Chinese-Vicuna, llmtools, Platypus-30B, Platypus, CharForge, robopoint-v1-vicuna-v1.5-13b-lora, robopoint-v1-llama-2-13b-lora, robopoint-v1-llama-2-7b-lora, robopoint-v1-vicuna-v1.5-7b-lora, lit-llama, llmtools, MoA, Role-Playing-LLM-Megumin, medalpaca-lora-7b-8bit, medalpaca-lora-30b-8bit, medalpaca-lora-13b-8bit, level3_nlp_finalproject-nlp-12, topxgen, AXRU, ruvector

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

arXiv46 repos

arXiv:2404.16821

InternVL2-8B-MPO, InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-2B, InternVL3_5-4B-HF, InternVL3_5-8B-HF, InternVL3_5-4B, InternVL-Chat-V1-5, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL-Chat-V1-2-Plus, InternVL3_5-8B, InternVL-Chat-V1-2, InternVL3_5-38B-HF, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-30B-A3B-HF, MMPR-v1.1, InternVL3_5-1B, InternVL-14B-224px, InternVL-Chat-V1-5-AWQ, InternVL2_5-1B, InternVL3_5-14B, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL-Chat-V1-1, MMPR, InternVL2_5-78B, InternVL3_5-1B-HF, InternVL3_5-241B-A28B-HF, InternVL2-1B, InternVL2_5-8B, Mini-InternVL-Chat-4B-V1-5, InternVL2_5-4B, InternVL2_5-2B, Mini-InternVL-Chat-2B-V1-5, InternVL2-8B, InternVL2-4B, InternVL2-2B, InternVL3-38B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, detikzify-v2-8b, InternVL2_5-26B, InternVL3-1B-hf

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

arXiv45 repos

arXiv:2312.14238

InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-2B, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, InternVL-Chat-V1-5, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL-Chat-V1-2-Plus, InternVL3_5-8B, InternVL-Chat-V1-2, InternVL3_5-8B-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL3_5-14B-HF, InternVL3_5-30B-A3B-HF, MMPR-v1.1, InternVL2_5-78B, InternVL2-8B-MPO, InternVL-Chat-V1-5-AWQ, InternVL2_5-1B, InternVL3_5-14B, InternVL3_5-2B-HF, MMPR, InternVL-Chat-V1-1, InternVL-14B-224px, InternVL3_5-1B, InternVL3_5-1B-HF, InternVL3_5-38B, InternVL3_5-30B-A3B, InternVL3_5-241B-A28B-HF, InternVL2-1B, InternVL2_5-4B, InternVL2_5-2B, Mini-InternVL-Chat-2B-V1-5, InternVL2-8B, InternVL2-4B, InternVL2-2B, Mini-InternVL-Chat-4B-V1-5, InternVL2_5-8B, InternVL3-38B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, InternVL2_5-26B, InternVL3-1B-hf

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

arXiv45 repos

arXiv:2412.05271

InternVL3_5-1B-HF, InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL-Chat-V1-5, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL-Chat-V1-2-Plus, InternVL3_5-8B, InternVL-Chat-V1-2, InternVL3_5-2B, InternVL3_5-8B-HF, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-30B-A3B-HF, InternVL3_5-4B-HF, MMPR-v1.1, InternVL3_5-1B, InternVL-14B-224px, InternVL-Chat-V1-5-AWQ, InternVL3_5-14B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL3_5-241B-A28B-HF, InternVL-Chat-V1-1, InternVL2_5-1B, InternVL2_5-78B, InternVL3_5-30B-A3B, InternVL2-1B, InternVL2_5-8B, Mini-InternVL-Chat-4B-V1-5, InternVL2_5-4B, InternVL2_5-2B, Mini-InternVL-Chat-2B-V1-5, InternVL2-8B, InternVL2-4B, InternVL2-2B, InternVL3-38B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, InternVL2_5-26B, OpenVid-1M, OpenVid, InternVL3-1B-hf

High-Resolution Image Synthesis with Latent Diffusion Models

arXiv42 repos

arXiv:2112.10752

stable-diffusion, stable-diffusion-xl-base-1.0, stable-diffusion-v1-4, stable-diffusion-xl-1.0-inpainting-0.1, stable-diffusion-v-1-4-original, Diff-Pruning, visual-chatgpt-zh, stable-diffusion-x4-upscaler, PixArt-LCM-XL-2-1024-MS, PixArt-XL-2-256x256, PixArt-XL-2-1024-MS, PixArt-XL-2-512x512, Taiyi-Stable-Diffusion-1B-Chinese-EN-v0.1, Taiyi-Stable-Diffusion-1B-Chinese-v0.1, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, coreml-stable-diffusion-2-base-palettized, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-2-1-base, coreml-stable-diffusion-xl-base-with-refiner, PixArt-Sigma-XL-2-512-MS, coreml-stable-diffusion-2-1-base-palettized, coreml-stable-diffusion-xl-base-ios, ldm-celebahq-256, coreml-stable-diffusion-2-base, stable-diffusion-2-1-base, cool-japan-diffusion-for-learning-2-0, picasso-diffusion-1-1, diff-mining, stable-diffusion-2-1, stable-diffusion-xl-refiner-1.0, sdxl-vae, Stable-Diffusion-FineTuned-zh-v1, Stable-Diffusion-FineTuned-zh-v2, Stable-Diffusion-FineTuned-zh-v0, stable-diffusion-inpainting, stable-diffusion-uncrop, ldm3d-4c, Cover-Generator, stable-diffusion-v1.5-webnn

Building a Large Japanese Web Corpus for Large Language Models

arXiv42 repos

arXiv:2404.17733

Swallow-70b-hf, Swallow-70b-NVE-hf, Swallow-7b-NVE-hf, Swallow-MS-7b-v0.1, Swallow-7b-NVE-instruct-hf, Llama-3-Swallow-70B-v0.1, Swallow-70b-instruct-hf, Swallow-7b-plus-hf, Swallow-70b-NVE-instruct-hf, Llama-3.1-Swallow-8B-v0.1, Swallow-7b-hf, Swallow-MX-8x7b-NVE-v0.1, Llama-3-Swallow-8B-v0.1, Gemma-2-Llama-Swallow-9b-pt-v0.1, Llama-3.1-Swallow-70B-v0.1, Swallow-7b-instruct-hf, Llama-3.1-Swallow-8B-v0.2, Llama-3.3-Swallow-70B-v0.4, Swallow-13b-hf, Swallow-13b-instruct-hf, Swallow-13b-NVE-hf, Gemma-2-Llama-Swallow-2b-pt-v0.1, Gemma-2-Llama-Swallow-27b-pt-v0.1, Qwen3-Swallow-32B-SFT-v0.2, Qwen3-Swallow-30B-A3B-CPT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2, Llama-3.1-Swallow-8B-v0.5, Qwen3-Swallow-32B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2, GPT-OSS-Swallow-20B-RL-v0.1, GPT-OSS-Swallow-120B-SFT-v0.1, Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-20B-SFT-v0.1, Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-8B-SFT-v0.2, GPT-OSS-Swallow-120B-RL-v0.1-MXFP4, Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-20B-RL-v0.1-MXFP4, Qwen3-Swallow-8B-CPT-v0.2, Qwen3-Swallow-30B-A3B-SFT-v0.2, GPT-OSS-Swallow-120B-RL-v0.1, Qwen3-Swallow-32B-CPT-v0.2

Qwen3 Technical Report

arXiv41 repos

arXiv:2505.09388

Qwen3-30B-A3B, Qwen3-32B, Qwen3-235B-A22B, Qwen3-Coder-30B-A3B-Instruct, Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-4B-Thinking-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-Coder-480B-A35B-Instruct, Qwen3-Coder-30B-A3B-Instruct-FP8, Qwen3-Coder-480B-A35B-Instruct-FP8, Qwen3-0.6B-GPTQ-Int8, Qwen3-4B-AWQ, gliner-stream-pii-v1.0, qwen3-8b-base, SDAR-8B-Chat-b16, SDAR-8B-Chat-b32, SDAR-8B-Chat-b64, SDAR, SDAR-30B-A3B-Sci, SDAR-4B-Chat, SDAR-8B-Chat, SDAR-1.7B-Chat, SDAR-30B-A3B-Chat, Qwen3-Next-80B-A3B-Instruct, Qwen3-0.6B, Qwen3-Swallow-8B-SFT-v0.2, Qwen3-Swallow-32B-SFT-v0.2, Qwen3-Swallow-30B-A3B-CPT-v0.2, Medical-Qwen3-Swallow-32B, Qwen3-Swallow-30B-A3B-SFT-v0.2, Qwen3-Swallow-32B-CPT-v0.2, Qwen3-Swallow-32B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4, Medical-Qwen3-Swallow-30B-A3B, Medical-Qwen3-Swallow-8B, Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-8B-CPT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2

YaRN: Efficient Context Window Extension of Large Language Models

arXiv40 repos

arXiv:2309.00071

Yarn-Mistral-7b-64k, Qwen2.5-VL, Qwen3-VL, QwQ, Qwen3-30B-A3B, Qwen2.5-72B-Instruct, Qwen2.5-32B-Instruct, Qwen3-32B, Qwen3-235B-A22B, Qwen2.5-Coder-32B-Instruct, QwQ-32B, Qwen2.5-Coder-14B-Instruct, Qwen3-Coder-Next-GGUF, Qwen2.5-VL-72B-Instruct, Llama-3-8B-Instruct-Gradient-1048k, Qwen2-VL, Qwen2.5-VL-3B-Instruct-GGUF, Qwen2-72B-Instruct, Qwen3-4B-AWQ, Qwen2.5-14B-Instruct, Llama-3-70B-Instruct-Gradient-1048k, Llama-3-70B-Instruct-Gradient-262k, Llama-3-8B-Instruct-262k, Llama-3-8B-Instruct-Gradient-4194k, Qwen2-57B-A14B-Instruct, Chinese-LLaMA-Alpaca-2, Qwen2.5-VL-7B-Instruct-GGUF, Yarn-Llama-2-13b-128k, Yarn-Llama-2-7b-128k, yarn, Yarn-Llama-2-13b-64k, Yarn-Llama-2-70b-32k, Yarn-Llama-2-7b-64k, Yarn-Solar-10b-64k, Yarn-Mistral-7b-128k, Yarn-Solar-10b-32k, Qwen3-Next-80B-A3B-Instruct, Qwen2.5-14B-Instruct-AWQ, Qwen2.5-VL-7B-Instruct-GGUF, Qwen2.5-Coder-7B-Instruct

Learning Transferable Visual Models From Natural Language Supervision

arXiv36 repos

arXiv:2103.00020

CLIP, clip-vit-base-patch32, stable-diffusion-v1-4, clip-ViT-B-32, stable-diffusion-v-1-4-original, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, contrastors, Semantic-Segment-Anything, contrastors, tf-transformers, clip-italian, clip-italian, clip-image-search, BiomedCoOp, clip-vit-base-patch16, LAVIS, ml-mobileclip, stable-diffusion-inpainting, visual-spatial-reasoning, CLIP-Chinese, Multilingual-CLIP, turkish-clip, vilmedic, prismatic-vlms, BiomedVLP-CXR-BERT-specialized, vit_large_patch14_clip_224.openai, stable-diffusion-safety-checker, TiViT, ru-clip, ru-clip, vit-dog, align-text-encoders, stable-diffusion-v1.5-webnn

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

arXiv30 repos

arXiv:1810.04805

bert-base-uncased, bert, bert-base-NER, bert-base-chinese, bert-base-cased, bert-base-multilingual-cased, DiagnosisCoding, EMS-Pipeline, tf-transformers, legalbert-large-1.7M-1, legalbert-large-1.7M-2, europeana-bert, dllm, MCBG, Malware_survey_my_experiments, CodeXGLUE, Vorbereitung, Bert-Implementation-From-Scratch, Word-Embeddings-Repository-for-Turkish, BertWithPretrained, BertWithPretrained, bangla-bert-base, bangla-bert, kcbert-base, KcBERT, BERTify, IndicAbusive, rudetoxifier, ASR-Knowledge-Transferring, caption

Rewriting Pre-Training Data Boosts LLM Performance in Math and Code

arXiv30 repos

arXiv:2505.02881

Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0002500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0012500, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0005000, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0002500, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0002500, Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0010000, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0007500, Llama-3.1-8B-code-ablation-exp3-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0007500, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0012500, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0010000, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0010000, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0002500, Llama-3.3-Swallow-70B-v0.4, Llama-3.1-8B-code-ablation-exp2-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0002500, Llama-3.1-8B-math-ablation-exp2-LR2.5e-5-WD0.1-iter0005000, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0007500, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0010000, swallow-code, swallow-code-v2, Llama-3.1-8B-math-ablation-exp1-LR2.5e-5-WD0.1-iter0012500, Llama-3.1-Swallow-8B-v0.5, Llama-3.1-8B-code-ablation-exp1-LR2.5e-5-MINLR2.5E-6-WD0.1-iter0012500, swallow-math, swallow-math-v2, swallow-code-math

On the Dangers of Stochastic Parrots

ACM28 repos

ACM:3442188.3445922

hacker-laws, roberta-large-mnli, distilbert-base-uncased-distilled-squad, transfo-xl-wt103, bert-base-chinese, idefics-80b-instruct, sup-simcse-roberta-large, sup-simcse-bert-large-uncased, geolm-base-toponym-recognition, cosmo-xl, idefics2-8b, legal-xlm-longformer-base, legal-xlm-roberta-large, legal-xlm-roberta-base, idefics-9b-instruct, distilgpt2, legal-english-roberta-base, legal-croatian-roberta-base, legal-swiss-roberta-large, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large, bert-base-uncased-hatexplain-rationale-two, ke-t5-base, bert-base, stable-diffusion-safety-checker, opus-mt-ru-en

Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

arXiv28 repos

arXiv:1910.10683

ChatLM-mini-Chinese, t5x, chronos-t5-base, c4, SwissArmyTransformer, t5-v1_1-xxl, chronos-bolt-base, chronos-bolt-tiny, chronos-bolt-mini, chronos-bolt-small, chronos-t5-tiny, chronos-t5-small, chronos-t5-mini, chronos-t5-large, google_t5-v1_1-xxl_encoderonly, Glot500, tk-instruct-11b-def-pos, Tk-Instruct, tf-transformers, turkish-bert, tk-instruct-3b-def, nano_flan_t5, text-to-text-transfer-transformer, tk-instruct-11b-def, ke-t5, level2-nlp-datacentric-nlp-03, russe_detox_2022, russe_detox_2022

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

arXiv28 repos

arXiv:2306.00978

awq-embed, lorax, ScaleLLM, nunchaku, llm-awq, AutoAWQ, VILA, efficient-transformers, InternVL-Chat-V1-5-AWQ, vllm, vllm, nunchaku, vllm-turboquant, trtllm, TensorRT-LLM, smash, dgx-spark-quantization, Qwen3-VL-2B-GRACE-W4G128-AWQ, awq4nvomni, llm-awq, lmdeploy-v100, vllm-old, vllm_amd_sleep, vllm, higgs-audio-vllm, lmdeploy-dev, cosyvoice3-inference-acceleration, lmdeploy

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

arXiv28 repos

arXiv:2411.10442

InternVL3_5-30B-A3B-HF, InternVL3_5-1B-HF, InternVL3-2B, xtuner, InternVL, MMPR-Tiny, MMPR-v1.2, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL3_5-8B, InternVL3_5-8B-HF, InternVL3_5-2B, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, MMPR-v1.1, InternVL2-8B-MPO, InternVL3_5-1B, InternVL3_5-14B, MMPR, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL3_5-241B-A28B-HF, InternVL3-38B, InternVL3-1B-hf

DFlash: Block Diffusion for Flash Speculative Decoding

arXiv27 repos

arXiv:2602.06036

Muse-Glimmer-30B, SpecForge, dflash, Qwen3-4B-DFlash-b16, DeepSpec, tpu-spec-decode, dflash-b4d1109c, dflash_benchmark, Kimi-K2.5-DFlash, Qwen3.6-27B-DFlash, Qwen3.6-35B-A3B-DFlash, gemma-4-26B-A4B-it-DFlash, gemma-4-31B-it-DFlash, MiniMax-M2.5-DFlash, Qwen3.5-27B-DFlash, Qwen3.5-122B-A10B-DFlash, Qwen3.5-35B-A3B-DFlash, gpt-oss-20b-DFlash, gpt-oss-120b-DFlash, Qwen3-Coder-30B-A3B-DFlash, Qwen3.5-9B-DFlash, Qwen3.5-4B-DFlash, Qwen3-8B-DFlash-b16, Qwen3-Coder-Next-DFlash, LLaMA3.1-8B-Instruct-DFlash-UltraChat, dflasher, dflash_eval

Attention Is All You Need

arXiv26 repos

arXiv:1706.03762

speech-emotion-recognition, atomic-agents, pi-web-agent, bert, graphify, NeMo-Agent-Toolkit-Examples, nature-skills, prompt-engineering, onnxruntime-training-examples, minGPT, document-to-podcast, dissertation-project, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, Speech-Transformer, nmt, nugie-jax-nemotron-3-nano, tiny-vllm, canary-1b, diar_sortformer_4spk-v1, transformer-tricks, indonesian-language-models, LaTeX-OCR, NAVER_AIRUSH_Grammar_Error_Correction, llama2.zig, llama2.go

ColPali: Efficient Document Retrieval with Vision Language Models

arXiv26 repos

arXiv:2407.01449

vidore-benchmark, colpali, colpali, colpali-v1.2, colpali-v1.1, colpali-v1.3, tomoro-colqwen3-embed-4b, colqwen2-v1.0, colqwen2.5-v0.1, colqwen2.5-v0.2, colSmol-256M, colSmol-500M, colpali, colpali, colqwen2-v0.1, colpali_train_set, SampadaEmbed, vllm-factory, esg_reports_v2, biomedical_lectures_v2, economics_reports_v2, esg_reports_human_labeled_v2, multi-modal-rag-with-colpali, colpali_t, colpali-tutorial, Colpali-Reproducibility

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

arXiv25 repos

arXiv:2406.12793

GLM-Z1-32B-0414, glm-4-9b-chat, GLM-4, glm-4-9b, glm-4-9b-chat-1m, GLM-Z1-32B-0414, glm-4-9b-chat-hf, glm-4-9b-chat-1m-hf, glm-4v-9b, glm-4v-9b, chatglm2-6b-32k, chatglm2-6b, chatglm-6b-int8, chatglm3-6b-32k, chatglm-6b, chatglm2-6b-int4, chatglm-6b-int4, GLM-Z1-9B-0414, chatglm3-6b-32k, glm-4-9b-chat-1m, chatglm3-6b-base, GLM-4, chatglm-6b-int8, chatglm-6b-int4, visualglm-6b

Code Llama: Open Foundation Models for Code

arXiv24 repos

arXiv:2308.12950

CodeLlama-7b-Instruct-hf, speechless-codellama-34b-v2.0, speechless-codellama-34b-v2.0-GGUF, speechless-codellama-34b-v2.0-GPTQ, speechless-codellama-34b-v1.0, speechless-codellama-34b-v2.0-AWQ, speechless-codellama-orca-13b, speechless-codellama-airoboros-orca-platypus-13b, speechless-codellama-dolphin-orca-platypus-34b, speechless-codellama-platypus-13b, CodeLlama-7b-hf, ELYZA-japanese-CodeLlama-7b-instruct, CodeLlama-7B-Instruct-GPTQ, CodeLlama-7B-GGUF, CodeLlama-7b-Python-hf, CodeLlama-34b-Python-hf, CodeLlama-13b-hf, CodeLlama-13b-Python-hf, CodeLlama-13b-Instruct-hf, CodeLlama-34b-hf, CodeLlama-34b-Instruct-hf, CodeLlama-70b-hf, CodeLlama-70b-Python-hf, CodeLlama-70b-Instruct-hf

EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test

arXiv24 repos

arXiv:2503.01840

SpecForge, EAGLE-llama2-chat-7B, EAGLE-LLaMA3-Instruct-70B, EAGLE-Qwen2-7B-Instruct, DeepSpec, EAGLE, EAGLE3-Vicuna1.3-13B, EAGLE3-LLaMA3.1-Instruct-8B, EAGLE3-LLaMA3.3-Instruct-70B, EAGLE3-DeepSeek-R1-Distill-LLaMA-8B, EAGLE-llama2-chat-13B, EAGLE-Vicuna-7B-v1.3, GLM-4.7-Flash-Eagle3, EAGLE-Vicuna-33B-v1.3, EAGLE-Vicuna-13B-v1.3, EAGLE-LLaMA3-Instruct-8B, EAGLE-llama2-chat-70B, EAGLE-mixtral-instruct-8x7B, EAGLE-Qwen2-72B-Instruct, EAGLE-LLaMA3.1-Instruct-8B, tpu-spec-decode, quantized-eagle, EAGLE, Enhanced-Eagle

Language Models are Few-Shot Learners

arXiv23 repos

arXiv:2005.14165

awesome-prompt-engineering, ik_llama.cpp, llama.cpp, lm-evaluation-harness, llm.c, llm.c, mmlu, llama.cpp, promptsource, falcon-40b, falcon-7b, falcon-rw-1b, llama.cpp.qwen2.5vl, atomic-llama-cpp-turboquant, lm-evaluation-harness, opt-125m, falcon-40b-instruct, opt-13b, opt-2.7b, llama.cpp, opt-350m, relbert, llama.cpp

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

arXiv23 repos

arXiv:2210.17323

deepcompressor, ChatGLM-6B, glq, ChatGLM-6B, text-generation-inference, lorax, ScaleLLM, llm-awq, gptq, PodGPT, efficient-transformers, vllm, vllm, vllm-turboquant, smash, GPTQ-for-LLaMa, awq-embed, awq4nvomni, llm-awq, vllm-old, vllm_amd_sleep, vllm, higgs-audio-vllm

MedVLThinker: Simple Baselines for Multimodal Medical Reasoning

arXiv23 repos

arXiv:2508.02669

MedVLSynther-7B-RL_5K_internvl-glm, MedVLSynther-7B-RL_5K_no-verify, MedVLThinker-3B-SFT_5K, MedVLThinker, MedVLSynther-3B-RL_1K, MedVLSynther-3B-RL_2K, MedVLSynther-3B-RL_10K, MedVLSynther-3B-RL_5K_qwen-glm, MedVLSynther-3B-RL_5K_glm-glm, MedVLSynther-3B-RL_5K, MedVLSynther-3B-RL_5K_internvl-glm, MedVLSynther-3B-RL_13K, MedVLSynther-7B-RL_2K, MedVLSynther-7B-RL_10K, MedVLSynther-7B-RL_5K_qwen-glm, MedVLSynther-7B-RL_1K, MedVLSynther-3B-RL_5K_no-verify, MedVLSynther-7B-RL_5K, MedVLSynther-3B-RL_5K_PMC-style, MedVLSynther-7B-RL_13K, MedVLSynther-7B-RL_5K_glm-glm, MedVLSynther-7B-RL_5K_PMC-style, MedVLThinker-7B-SFT_5K

Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs

arXiv23 repos

arXiv:2510.25867

MedVLSynther-7B-RL_5K_internvl-glm, MedVLSynther-7B-RL_5K_no-verify, MedVLSynther, MedVLSynther-3B-RL_1K, MedVLSynther-3B-RL_2K, MedVLSynther-3B-RL_10K, MedVLSynther-3B-RL_5K_qwen-glm, MedVLSynther-3B-RL_5K_glm-glm, MedVLSynther-3B-RL_5K, MedVLSynther-3B-RL_5K_internvl-glm, MedVLSynther-3B-RL_13K, MedVLSynther-7B-RL_2K, MedVLSynther-7B-RL_10K, MedVLSynther-7B-RL_5K_qwen-glm, MedVLSynther-7B-RL_1K, MedVLSynther-3B-RL_5K_no-verify, MedVLSynther-7B-RL_5K, MedVLSynther-3B-RL_5K_PMC-style, MedVLSynther-7B-RL_13K, MedVLSynther-7B-RL_5K_glm-glm, MedVLThinker-3B-SFT_5K, MedVLSynther-7B-RL_5K_PMC-style, MedVLThinker-7B-SFT_5K

Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance

arXiv22 repos

arXiv:2410.16261

InternVL, InternVL-Chat-V1-5, InternVL-Chat-V1-2-Plus, InternVL-Chat-V1-2, InternVL2_5-1B, InternVL-Chat-V1-1, InternVL-Chat-V1-5-AWQ, InternVL-14B-224px, InternVL2_5-78B, InternVL2-1B, InternVL2_5-8B, InternVL2_5-4B, Mini-InternVL-Chat-4B-V1-5, Mini-InternVL-Chat-2B-V1-5, InternVL2_5-2B, InternVL2-8B, InternVL2-2B, InternVL2-4B, InternViT-300M-448px, InternVL2-Llama3-76B, InternVL2-26B, InternVL2_5-26B

TimeLMs: Diachronic Language Models from Twitter

arXiv21 repos

arXiv:2202.03829

tweeteval, twitter-roberta-base-sentiment-latest, twitter-roberta-base-2021-124m, timelms, twitter-roberta-base-2019-90m, twitter-roberta-base-dec2020, twitter-roberta-base-jun2021, twitter-roberta-base-mar2020, twitter-roberta-base-jun2020, twitter-roberta-base-sep2020, twitter-roberta-base-mar2021, twitter-roberta-base-dec2021, twitter-roberta-base-jun2022, twitter-roberta-base-sep2021, twitter-roberta-base-mar2022, twitter-roberta-base-mar2022-15M-incr, twitter-roberta-base-jun2022-15M-incr, twitter-roberta-base-2022-154m, twitter-roberta-base-sep2022, twitter-roberta-large-2022-154m, tgbot-hate-speech

WizardLM: Empowering large pre-trained language models to follow complex instructions

arXiv21 repos

arXiv:2304.12244

WizardLM_evol_instruct_70k, WizardCoder-15B-V1.0, WizardLM-70B-V1.0, EasyInstruct, WizardMath-7B-V1.1, WizardLM-70B-V1.0, WizardMath-70B-V1.0, WizardCoder-Python-13B-V1.0, WizardCoder-15B-V1.0, WizardMath-7B-V1.0, WizardLM-13B-V1.0, WizardCoder-Python-34B-V1.0, WizardLM_evol_instruct_V2_196k, WizardMath-7B-V1.1, WizardCoder-33B-V1.1, WizardLM-13B-V1.2, WizardLM_evol_instruct_70k, slm-innovator-lab, evolve-instruct, ko-instruction-dataset, Ko.WizardLM_evol_instruct_V2_196k

EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

arXiv21 repos

arXiv:2401.15077

EAGLE-llama2-chat-7B, EAGLE, EAGLE3-LLaMA3.3-Instruct-70B, EAGLE3-Vicuna1.3-13B, EAGLE3-LLaMA3.1-Instruct-8B, EAGLE3-DeepSeek-R1-Distill-LLaMA-8B, EAGLE-llama2-chat-13B, EAGLE-LLaMA3-Instruct-70B, EAGLE-Qwen2-7B-Instruct, EAGLE-Vicuna-7B-v1.3, EAGLE-Vicuna-33B-v1.3, EAGLE-Vicuna-13B-v1.3, EAGLE-llama2-chat-70B, EAGLE-mixtral-instruct-8x7B, EAGLE-LLaMA3-Instruct-8B, EAGLE-Qwen2-72B-Instruct, EAGLE-LLaMA3.1-Instruct-8B, quantized-eagle, EAGLE, Enhanced-Eagle, mlx-flash

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

arXiv21 repos

arXiv:2402.03300

minimind, tunix, MedVLM-R1, AntAngelMed, AntAngelMed, Qwen2.5-3B-GRPO-MATH-1EPOCH, Qwen2.5-1.5B-GRPO-MATH-1EPOCH, async-grpo, xtuner, deepseek-math-7b-rl, deepseek-math-7b-base, MM-EUREKA, CPGD-7B, MedicalGPT, detikzify-v2.5-8b, SkinTokens, SkinTokens, marcello, rl-unsloth, chart-rvr-3b, chart-rvr-hard-3b

Chronos: Learning the Language of Time Series

arXiv21 repos

arXiv:2403.07815

chronos-t5-base, sundial-base-128m, uni2ts, chronos-forecasting, chronos-t5-large, chronos-bolt-base, chronos-bolt-tiny, chronos-bolt-mini, chronos-2, chronos-bolt-small, chronos-t5-tiny, chronos-t5-mini, moirai-2.0-R-small, chronos-t5-small, stock_forecasting, Samay, A2TTA, moment, patchtst-fm-r1, granite-timeseries-patchtst-fm-r1, uni2ts

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

arXiv21 repos

arXiv:2504.10479

InternVL3_5-30B-A3B-HF, InternVL3_5-1B-HF, InternVL3-2B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-4B, InternVL3_5-38B-HF, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL3_5-8B, InternVL3_5-8B-HF, InternVL3_5-2B, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, InternVL3_5-14B, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-38B, InternVL3_5-241B-A28B-HF, InternVL3_5-1B, InternVL3-38B, InternVL3-1B-hf

Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

arXiv21 repos

arXiv:2508.01691

voxlect-mandarin-cantonese-dialect-whisper-small, voxlect-indic-lid-whisper-small, voxlect-german-dialect-whisper-small, voxlect-german-dialect-whisper-large-v3, voxlect-spanish-dialect-mms-lid-256, voxlect-english-dialect-mms-lid-256, voxlect-indic-lid-whisper-large-v3, voxlect-french-dialect-mms-lid-256, voxlect-thai-dialect-mms-lid-256, voxlect-spanish-dialect-whisper-small, voxlect-mandarin-cantonese-dialect-mms-lid-256, voxlect-english-dialect-whisper-large-v3, voxlect-indic-lid-mms-lid-256, voxlect-spanish-dialect-whisper-large-v3, voxlect-german-dialect-mms-lid-256, voxlect-thai-dialect-whisper-small, voxlect-thai-dialect-whisper-large-v3, voxlect-mandarin-cantonese-dialect-whisper-large-v3, voxlect-french-dialect-whisper-small, voxlect-english-dialect-whisper-small, voxlect-french-dialect-whisper-large-v3

ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT

arXiv20 repos

arXiv:2004.12832

RAG-RLRC-LaySum, colqwen2.5-v0.2, docs-reference, bge-m3, sweet-search, colpali, colpali-v1.2, colpali-v1.1, colpali-v1.3, colqwen2-v1.0, colqwen2.5-v0.1, colSmol-500M, colSmol-256M, colqwen2-v0.1, plaidrepro, colbertv2.0, ColBERT, odqa_baseline_code, jina-colbert-v1-en, KolBERT

Efficient Memory Management for Large Language Model Serving with PagedAttention

arXiv20 repos

arXiv:2309.06180

vllm, flash-attention, vllm, flash-attention, lorax, vllm, vllm, vllm, vllm-turboquant, vllm, tiny-vllm, vllm-old, vllm_amd_sleep, vllm, higgs-audio-vllm, vllm, flash-attention, flash-attention, KsanaLLM, vllm

Qwen Technical Report

arXiv20 repos

arXiv:2309.16609

Qwen, Qwen-7B-Chat, Qwen-14B-Chat, Qwen1.5-0.5B, Qwen1.5-7B-Chat, Qwen1.5-7B, Qwen1.5-14B, Qwen1.5-14B-Chat, Qwen1.5-0.5B-Chat, Qwen1.5-1.8B, Qwen1.5-1.8B-Chat, Qwen-7B, Qwen-1_8B, CodeQwen1.5-7B-Chat, Qwen1.5-110B-Chat, Qwen1.5-32B-Chat, Qwen1.5-4B-Chat, Qwen1.5-72B, Qwen1.5-32B-Chat-GPTQ-Int4, Qwen1.5-4B

EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation

arXiv20 repos

arXiv:2405.18991

EasyAnimate, EasyAnimateV5.1-7b-zh, EasyAnimateV5.1-12b-zh, EasyAnimateV5.1-7b-zh-Control, EasyAnimateV3-XL-2-InP-512x512, EasyAnimateV2-XL-2-512x512, EasyAnimateV2-XL-2-768x768, EasyAnimateV4-XL-2-InP, EasyAnimateV5-7b-zh, EasyAnimateV5.1-12b-zh-InP, EasyAnimateV5.1-7b-zh-diffusers, EasyAnimateV5.1-12b-zh-Control-Camera, EasyAnimateV5.1-7b-zh-Control-Camera, EasyAnimateV3-XL-2-InP-960x960, EasyAnimateV5.1-7b-zh-InP, EasyAnimateV3-XL-2-InP-768x768, EasyAnimateV5.1-12b-zh-Control, EasyAnimateV5-12b-zh, EasyAnimateV5-7b-zh-InP, EasyAnimateV5-12b-zh-Control

EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

arXiv20 repos

arXiv:2406.16858

EAGLE-llama2-chat-7B, EAGLE-Qwen2-7B-Instruct, EAGLE, EAGLE3-Vicuna1.3-13B, EAGLE3-LLaMA3.1-Instruct-8B, EAGLE3-LLaMA3.3-Instruct-70B, EAGLE3-DeepSeek-R1-Distill-LLaMA-8B, EAGLE-llama2-chat-13B, EAGLE-LLaMA3-Instruct-70B, EAGLE-Vicuna-7B-v1.3, EAGLE-Vicuna-33B-v1.3, EAGLE-Vicuna-13B-v1.3, EAGLE-llama2-chat-70B, EAGLE-mixtral-instruct-8x7B, EAGLE-LLaMA3-Instruct-8B, EAGLE-Qwen2-72B-Instruct, EAGLE-LLaMA3.1-Instruct-8B, quantized-eagle, EAGLE, Enhanced-Eagle

Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

arXiv20 repos

arXiv:2412.02595

NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-3-Nano-4B-GGUF, Qwen3-Swallow-32B-SFT-v0.2, Qwen3-Swallow-30B-A3B-CPT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2, Qwen3-Swallow-30B-A3B-SFT-v0.2, Qwen3-Swallow-32B-CPT-v0.2, Qwen3-Swallow-32B-RL-v0.2, Qwen3-Swallow-8B-RL-v0.2, GPT-OSS-Swallow-120B-RL-v0.1, GPT-OSS-Swallow-20B-RL-v0.1, GPT-OSS-Swallow-20B-SFT-v0.1, GPT-OSS-Swallow-120B-SFT-v0.1, Qwen3-Swallow-8B-RL-v0.2-AWQ-INT4, Qwen3-Swallow-8B-SFT-v0.2, Qwen3-Swallow-30B-A3B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-120B-RL-v0.1-MXFP4, Qwen3-Swallow-32B-RL-v0.2-AWQ-INT4, GPT-OSS-Swallow-20B-RL-v0.1-MXFP4, Qwen3-Swallow-8B-CPT-v0.2

70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)

arXiv20 repos

arXiv:2504.11651

BAGEL-7B-MoT-DF11, DFloat11, Bagel-DFloat11, DeepSeek-R1-Distill-Qwen-32B-DF11, DeepSeek-R1-Distill-Qwen-14B-DF11, Llama-3.1-8B-Instruct-DF11, gemma-3-12b-it-DF11, FLUX.1-Krea-dev-DF11, FLUX.1-dev-DF11, DeepSeek-R1-Distill-Llama-8B-DF11, Wan2.1-T2V-14B-Diffusers-DF11, Qwen3-32B-DF11, Qwen3-4B-DF11, Qwen3-14B-DF11, DeepSeek-R1-Distill-Qwen-7B-DF11, Phi-4-reasoning-plus-DF11, gemma-3-4b-it-DF11, Qwen3-8B-DF11, gemma-3-27b-it-DF11, BAGEL-DFloat11-Windows

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

arXiv20 repos

arXiv:2508.18265

InternVL3_5-30B-A3B-HF, MMPR-Tiny, MMPR-v1.2, InternVL3_5-4B, InternVL3_5-GPT-OSS-20B-A4B-Preview-HF, InternVL3_5-38B-HF, InternVL3_5-GPT-OSS-20B-A4B-Preview, InternVL3_5-8B-HF, InternVL3_5-8B, InternVL3_5-2B, InternVL3_5-14B-HF, InternVL3_5-241B-A28B, InternVL3_5-4B-HF, InternVL3_5-1B, InternVL3_5-14B, InternVL3_5-30B-A3B, InternVL3_5-2B-HF, InternVL3_5-241B-A28B-HF, InternVL3_5-1B-HF, InternVL3_5-38B

DeBERTa: Decoding-enhanced BERT with Disentangled Attention

arXiv19 repos

arXiv:2006.03654

LoRA, LEXTREME, mdeberta-v3-base, DeBERTa, deberta-v3-xsmall, deberta-v2-xxlarge, deberta-large, deberta-base-mnli, deberta-v2-xlarge, deberta-v2-xxlarge-mnli, deberta-v2-xlarge-mnli, deberta-v3-small, deberta-base, deberta-xlarge, deberta-v3-base, deberta-v3-large, deberta-xlarge-mnli, deberta-large-mnli, DeBERTa_TxtClassifier

RoFormer: Enhanced Transformer with Rotary Position Embedding

arXiv19 repos

arXiv:2104.09864

falcon-40b, roformer_v2_chinese_char_base, Baichuan-13B-Chat, falcon-7b, roformer_chinese_base, Baichuan-13B-Base, nomic-bert-2048, cholesky_encoder, snowflake-arctic-embed-m-long, arctic-embed, stablecode-completion-alpha-3b-4k, mesh-transformer-jax, falcon-40b-instruct, tiny-vllm, polyglot-ko-1.3b, polyglot-ko-3.8b, polyglot-ko-5.8b, llama2.zig, llama2.go

Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

arXiv19 repos

arXiv:2406.08464

magpie, Magpie-Reasoning-150K, Magpie-Qwen2-Air-3M-v0.1, Magpie-Qwen2.5-Pro-1M-v0.1, Magpie-Llama-3.1-Pro-DPO-100K-v0.1, Magpie-Llama-3.3-Pro-1M-v0.1, Llama-3-8B-Magpie-Align-v0.2, Magpie-Pro-DPO-100K-v0.1, Llama-3-8B-Magpie-Align-v0.3, Magpie-Air-DPO-100K-v0.1, Magpie-Llama-3.1-Pro-1M-v0.1, Llama-3-8B-Magpie-Align-v0.1, Magpie-Qwen2-Pro-1M-v0.1, Magpie-Qwen2-Pro-200K-Chinese, aurora-m2, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-70B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.2, swallow-gemma-magpie-v0.1

Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data

arXiv19 repos

arXiv:2504.09895

zephyr-7b-alpha-conf-sft, RefAlign, NQ-Subset-500, Llama-3.3-70B-Inst-awq_ultrafeedback_1in3, llama3-ultrafeedback-bertscore-bart-large-mnli, alpaca-7b-ref-meteor, Llama-2-13b-hf-conf-refalign, Mistral-7B-v0.1-conf-sft, Llama-3.3-70B-Inst-awq_SafeRLHF, alpaca-7b-ref-bertscore, Llama-2-7b-hf-conf-refalign, Llama-2-7b-hf-conf-sft, Mistral-7B-v0.1-conf-refalign, Llama-2-13b-hf-conf-sft, zephyr-7b-alpha-conf-refalign, Mistral-7B-Instruct-v0.2-ref-simpo, Mistral-7B-Instruct-v0.2-refalign, Llama-3-8B-Instruct-ref-simpo, Llama-3-8B-Instruct-refalign

Scaling Instruction-Finetuned Language Models

arXiv18 repos

arXiv:2210.11416

flan-t5-xl, flan-t5-large, pdfai-back, flan-t5-base, optimized-parler-tts, flan-alpaca-gpt4-xl, flan-alpaca, flan-alpaca-large, flan-alpaca-base, flan-alpaca-xl, flan-alpaca-xxl, flan-gpt4all-xl, flan-sharegpt-xl, blip2-flan-t5-xl, blip2-flan-t5-xxl, flan-t5-small, flan-ul2-dolly, flan-ul2-dolly-lora

Robust Speech Recognition via Large-Scale Weak Supervision

arXiv18 repos

arXiv:2212.04356

whisper, whisper-large-v3, whisper-large-v3-turbo, whisper-medium, whisper-tiny, kotoba-whisper-v2.0, kotoba-whisper-v1.0, dissertation-project, whisper-jax, whisper-tiny.en, whisper-base, whisper-large, whisper-small.en, whisper-small, whisper-medium.en, whisperspeech, final-project-level3-nlp-01, whisper-base-webnn

WizardCoder: Empowering Code Large Language Models with Evol-Instruct

arXiv18 repos

arXiv:2306.08568

evalplus, WizardLM_evol_instruct_70k, WizardCoder-15B-V1.0, WizardLM-70B-V1.0, WizardMath-7B-V1.1, WizardLM-70B-V1.0, WizardMath-7B-V1.0, WizardLM-13B-V1.2, WizardCoder-Python-13B-V1.0, WizardCoder-15B-V1.0, WizardMath-70B-V1.0, WizardLM-13B-V1.0, WizardCoder-Python-34B-V1.0, WizardMath-7B-V1.1, WizardLM_evol_instruct_70k, WizardLM_evol_instruct_V2_196k, WizardCoder-33B-V1.1, Ko.WizardLM_evol_instruct_V2_196k

Improved Baselines with Visual Instruction Tuning

arXiv18 repos

arXiv:2310.03744

LLaVA, minimind-v, M4-Instruct-Data, llava-v1.6-34b-hf, llava-v1.6-vicuna-13b-hf, llava-v1.6-vicuna-7b-hf, llava-v1.6-mistral-7b-hf, libra-llava-rad, llava-rad, table-llava-v1.5-13b, table-llava-v1.5-7b, table-llava-v1.5-7b-hf, llava-bench-in-the-wild, llava-1.5-665k-instructions, TriPlaneLLaVA, LLaVA-toy, multi_token, TinyLLava

LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

arXiv18 repos

arXiv:2404.05961

DermL2V-tmp, llm2vec, DermL2V-training, LLM2Vec-Llama-2-7b-chat-hf-mntp, LLM2Vec-Mistral-7B-Instruct-v2-unsup-simcse, LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervised, LLM2Vec-Llama-2-7b-chat-hf-mntp-supervised, LLM2Vec-Llama-2-7b-chat-hf-unsup-simcse, LLM2Vec-Sheared-LLaMA-mntp-supervised, LLM2Vec-Meta-Llama-3-8B-Instruct-mntp, LLM2Vec-Mistral-7B-Instruct-v2-mntp, LLM2Vec-Sheared-LLaMA-unsup-simcse, LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-unsup-simcse, LLM2Vec-Sheared-LLaMA-mntp, LLM2Vec-Mistral-7B-Instruct-v2-mntp-supervised, NetTAG, llm-idiosyncrasies, bach-or-bot

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

arXiv18 repos

arXiv:2406.07476

VideoLLaMA2, VideoLLaMA2.1-7B-16F-Base, Multi-Source-Video-Captioning, VideoLLaMA2.1-7B-AV, VideoLLaMA2-72B, VideoLLaMA2.1-7B-16F, VideoLLaMA2-72B-Base, VideoLLaMA2-7B, VideoLLaMA2-8x7B-Base, VideoLLaMA2-8x7B, VideoLLaMA2-7B-16F, VideoLLaMA2-7B-Base, VideoLLaMA2-7B-16F-Base, VideoLLaMA3-7B-Image, VideoLLaMA3-2B-Image, AVProunRLForVideoLLaMa2, MM-PreTrain, JavisUnd-Eval

AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model

arXiv18 repos

arXiv:2606.19100

AMALIA, amalia-vl-eval, DocVQA-PT, MME-PT, MMMU-Pro-PT, AI2D-PT, TextVQA-PT, ChartQA-PT, OCRBench-PT, InfographicVQA-PT, POPE-PT, COCO-Caption2017-PT, MMStar-PT, MATH-Vision-PT, MMMU-PT, SEED-Bench-PT, RealWorldQA-PT, EmbSpatial-Bench-PT

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

arXiv17 repos

arXiv:2010.11929

vision_transformer, tf-transformers, hls-foundation-os, coyo-dataset, coyo-700m, coyo-labeled-300m, vit-l16-coyo-labeled-300m-i1k384, vit-l16-coyo-labeled-300m-i1k512, vit-l16-coyo-labeled-300m, vilmedic, vit-base-patch16-224-in21k, vit-large-patch16-224-in21k, camie-tagger-v2, nsfw_image_detection, LaTeX-OCR, vit-base-patch32-384, vit-dog

CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

arXiv17 repos

arXiv:2203.13474

sven_modified, sven, codegen-16B-nl, CodeGen, MOSS, moss-moon-003-sft-int4, moss-moon-003-sft-plugin-int4, moss-moon-003-base, moss-moon-003-sft-plugin-int8, moss-moon-003-sft, moss-moon-003-sft-int8, moss-moon-003-sft-plugin, MOSS, codegen-6B-mono, jaxformer, codegen-16B-mono, diff-codegen-6b-v2

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

arXiv17 repos

arXiv:2205.14135

graphify, flash-attention, web-stable-diffusion, flash-attention, GPT-2, falcon-40b, falcon-rw-1b, falcon-7b, Chinese-CLIP, falcon-40b-instruct, mpt-1b-redpajama-200b, mpt-1b-redpajama-200b-dolly, nano_flan_t5, speechless-starcoder2-15b, flash-attention, flash-attention, Block-Sparse-Attention

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

arXiv17 repos

arXiv:2301.12597

Make-It-3D, heron-chat-blip-ja-stablelm-base-7b-v1-llava-620k, heron-chat-blip-ja-stablelm-base-7b-v1, heron-chat-blip-ja-stablelm-base-7b-v0, LAVIS, mBLIP, mblip-mt0-xl, mblip-bloomz-7b, blip2-flan-t5-xl, vlis, blip2-flan-t5-xxl, vqazero, Zero-and-Few-Shot-Visual-Question-Answering, VLSA, visualglm-6b, VisualGLM-6B, VisualGLM-6B

Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

arXiv17 repos

arXiv:2306.02858

Video-LLaMA-Series, VideoLLaMA2, VideoLLaMA2.1-7B-16F-Base, Multi-Source-Video-Captioning, VideoLLaMA2.1-7B-AV, VideoLLaMA2-72B, VideoLLaMA2.1-7B-16F, VideoLLaMA2-72B-Base, VideoLLaMA2-7B, VideoLLaMA2-8x7B-Base, VideoLLaMA2-8x7B, VideoLLaMA2-7B-16F, VideoLLaMA2-7B-Base, VideoLLaMA2-7B-16F-Base, VideoLLaMA3-7B-Image, VideoLLaMA3-2B-Image, AVProunRLForVideoLLaMa2

WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

arXiv17 repos

arXiv:2308.09583

WizardLM_evol_instruct_70k, WizardCoder-15B-V1.0, WizardLM-70B-V1.0, WizardMath-7B-V1.1, WizardLM-70B-V1.0, WizardLM-13B-V1.2, WizardCoder-15B-V1.0, WizardMath-70B-V1.0, WizardLM-13B-V1.0, WizardCoder-Python-34B-V1.0, WizardLM_evol_instruct_V2_196k, WizardCoder-Python-13B-V1.0, WizardMath-7B-V1.1, WizardCoder-33B-V1.1, WizardMath-7B-V1.0, WizardLM_evol_instruct_70k, Ko.WizardLM_evol_instruct_V2_196k

VCR: A Task for Pixel-Level Complex Reasoning in Vision Language Models via Restoring Occluded Text

arXiv17 repos

arXiv:2406.06462

VCR-wiki-en-easy, VCR, VCR-wiki-en-easy-test-100, VCR-wiki-zh-easy-test-500, VCR-wiki-en-easy-test-500, VCR-wiki-en-hard-test-500, VCR-wiki-en-hard-test-100, VCR-wiki-zh-hard-test-100, VCR-wiki-zh-easy-test-100, VCR-wiki-zh-hard-test, VCR-wiki-en-easy-test, VCR-wiki-zh-hard-test-500, VCR-wiki-zh-easy-test, VCR-wiki-zh-hard, VCR-wiki-en-hard-test, VCR-wiki-en-hard, VCR-wiki-zh-easy

Rank1: Test-Time Compute for Reranking in Information Retrieval

arXiv17 repos

arXiv:2502.18418

rank1-7b, rank1, rank1-1.5b, rank1-llama3-8b-awq, rank1-R1-MSMARCO, rank1-mistral-2501-24b, rank1-mistral-2501-24b-awq, rank1-training-data, rank1-14b-awq, rank1-Run-Files, rank1-32b-awq, rank1-7b-awq, rank1-3b, rank1-14b, rank1-0.5b, rank1-32b, rank1-llama3-8b

Perception Encoder: The best visual embeddings are not at the output of the network

arXiv17 repos

arXiv:2504.13181

Muse-Glimmer-30B, perception_models, PE-Video, PE-Core-T16-384, PE-Core-S16-384, PE-Core-B16-224, PE-Core-L14-336, PE-Core-G14-448, PE-Lang-L14-448, PE-Lang-L14-448-Tiling, PE-Lang-G14-448, PE-Spatial-B16-512, PE-Spatial-L14-448, PE-Lang-G14-448-Tiling, PE-Spatial-G14-448, PE-Spatial-T16-512, PE-Spatial-S16-512

Seq vs Seq: An Open Suite of Paired Encoders and Decoders

arXiv17 repos

arXiv:2507.11412

ettin-encoder-150m, ettin-encoder-400m, ettin-encoder-32m, ettin-encoder-68m, ettin-encoder-vs-decoder, ettin-decoder-1b, ettin-decoder-17m, ettin-decoder-400m, ettin-decoder-32m, ettin-encoder-17m, ettin-checkpoints, ettin-encoder-1b, ettin-decay-data, ettin-pretraining-data, ettin-decoder-68m, ettin-decoder-150m, ettin-extension-data

The Pile: An 800GB Dataset of Diverse Text for Language Modeling

arXiv16 repos

arXiv:2101.00027

GLM, falcon-40b, minipile, falcon-7b, pile-uncopyrighted, pythia-12b, pythia-1.4b, pythia-410m, LLaMA-MiLe-Loss, BiLLa, gpt-neo-2.7B, GPT-J-6B-Janeway, GPT-J-6B-Shinen, pythia-2.8b, diff-codegen-6b-v2, gpt-neo-1.3B

Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

arXiv16 repos

arXiv:2306.05685

synthetic_text_to_sql, mt_bench_prompts, Multi-Modality-Arena, vicuna-7B-1.1-HF, mt_bench_human_judgments, vicuna-13b-delta-v1.1, vicuna-13B-1.1-HF, vicuna-7b-delta-v1.1, vicuna-7b-delta-v0, llm-jailbreaking-defense, stablelm-zephyr-3b, vicuna-13b-v1.3, llm-router, vicuna-7b-v1.3, KoMT-Bench, KoMT-Bench

BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset

arXiv16 repos

arXiv:2307.04657

beaver-7b-v2.0-cost, beaver-7b-unified-cost, PKU-SafeRLHF-10K, beaver-7b-v3.0-cost, beaver-7b-v3.0-reward, beaver-7b-v1.0, beaver-7b-v1.0-cost, beaver-7b-v3.0, beaver-7b-v2.0, beaver-7b-v1.0-reward, beaver-7b-unified-reward, beaver-7b-v2.0-reward, PKU-SafeRLHF-30K, beavertails, BeaverTails, BeaverTails-Evaluation

RoBERTa: A Robustly Optimized BERT Pretraining Approach

arXiv15 repos

arXiv:1907.11692

roberta-large-mnli, EasyNLP, roberta-base, DiagnosisCoding, tf-transformers, legalbert-large-1.7M-1, legalbert-large-1.7M-2, LoRA, GPT_Ranker, roberta_toxicity_classifier, Reddit-Sports-Sentiment-Analysis, roberta-large, deid_roberta_i2b2, ehr_deidentification, caption

SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers

arXiv15 repos

arXiv:2105.15203

segformer-b0-finetuned-ade-512-512, SegFormer, segformer-b2-finetuned-ade-512-512, segformer-b0-finetuned-cityscapes-1024-1024, Semantic-Segment-Anything, segformer-b5-finetuned-cityscapes-1024-1024, segformer-b5-finetuned-ade-640-640, segformer-b2-finetuned-cityscapes-1024-1024, segformer_b2_clothes, segformer-tf-transformers, SegFormer-Training_From-Scratch_vs_HuggingFace, EdgeSeg, segformer-tf-transformers, segformer-b3-fashion, segformer_b3_clothes

Crosslingual Generalization through Multitask Finetuning

arXiv15 repos

arXiv:2211.01786

awesome-totally-open-chatgpt, bloomz, mt0-base, xmtf, xP3megds, xwinograd, xP3, xP3x, bloomz-560m, bloomz-3b, mt0-xl, bloomz-1b1, bloomz-7b1-mt, bloomz-1b7, bloomz-7b1

Direct Preference Optimization: Your Language Model is Secretly a Reward Model

arXiv15 repos

arXiv:2305.18290

BELLE, tunix, zephyr-7b-alpha, zephyr-7b-beta, Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, Reinforcement-Learning-Full-Pipeline, shisa-7b-v1, alpaca_farm, MedicalGPT, stablelm-zephyr-3b, sacpo, p-sacpo, dpo-arithmo-mistral-7B, llm_factuality_tuning

arXiv:2308.07317

arXiv15 repos

arXiv:2308.07317

Jellyfish-13B, Platypus2-70B, Platypus-13B-adapters, Platypus2-70B-instruct, Camel-Platypus2-13B, Stable-Platypus2-13B, Platypus2-13B, Platypus-70B-adapters, Camel-Platypus2-70B, Platypus2-7B, Platypus-7B-adapters, Platypus-30B, Platypus, KOpen-platypus, KO-Platypus2-7B-ex

StarVector: Generating Scalable Vector Graphics Code from Images and Text

arXiv15 repos

arXiv:2312.11556

vtracer, star-vector, starvector-1b-im2svg, starvector-8b-im2svg, svg-stack, text2svg-stack, svg-stack-simple, svg-diagrams, svg-fonts, svg-emoji, svg-fonts-simple, svg-emoji-simple, FIGR-SVG, svg-icons, svg-icons-simple

xLAM: A Family of Large Action Models to Empower AI Agent Systems

arXiv15 repos

arXiv:2409.03215

xLAM-7b-r, xLAM, xLAM-1b-fc-r, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, xLAM-2-1b-fc-r, xLAM-v0.1-r, xLAM-2-3b-fc-r-gguf, xLAM-2-32b-fc-r, xLAM-8x22b-r, xLAM-7b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, xLAM-8x7b-r, Llama-xLAM-2-8b-fc-r, xLAM-2-3b-fc-r

Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

arXiv15 repos

arXiv:2409.12191

Qwen2.5-VL, Qwen3-VL, Qwen2-VL-7B-Instruct, Qwen2.5-VL-72B-Instruct, Qwen2-VL, Qwen2.5-VL-3B-Instruct-GGUF, Baichuan-Omni-1.5, UGround-V1-7B, UGround-V1-2B, UGround-V1-72B, Qwen2.5-VL-7B-Instruct-GGUF, Qwen2-VL-7B-Instruct, qwen2vit600m, Qwen2.5-VL-7B-Instruct-GGUF, ChatVLA_public

Self-Instruct: Aligning Language Models with Self-Generated Instructions

arXiv14 repos

arXiv:2212.10560

stanford_alpaca, camel, alpaca_eval, self-instruct-seed, COIG, COIG, Chinese-Vicuna, moss-002-sft-data, EasyInstruct, vigogne, openchat, Chinese-Vicuna, koalpaca, kwater

Adding Conditional Control to Text-to-Image Diffusion Models

arXiv14 repos

arXiv:2302.05543

ControlNet, visual-chatgpt-zh, MistoLine, sd-controlnet-canny, control_v11p_sd15_openpose, controlnet-canny-sdxl-1.0, BDM1.0, ControlNet_AnimalPose, controlnet-openpose-sdxl-1.0, controlnet-scribble-sdxl-1.0, sd-controlnet-seg, control_v11f1p_sd15_depth, control_v11p_sd15_normalbae, sd-controlnet-scribble

Sigmoid Loss for Language Image Pre-Training

arXiv14 repos

arXiv:2303.15343

siglip-so400m-patch14-384, siglip2-giant-opt-patch16-384, siglip2-so400m-patch16-512, siglip2-base-patch16-224, siglip2-base-patch16-naflex, open-value, siglip-base-patch16-224, ml-mobileclip, CCD, siglip2-so400m-patch16-naflex, prismatic-vlms, vit_base_patch16_siglip_512.v2_webli, siglip2-so400m-patch14-384, siglip2-so400m-patch14-224

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

arXiv14 repos

arXiv:2309.14509

VeOmni, MindSpeed-MM, Wan2.2-T2V-A14B-Diffusers, Wan2.1-VACE-14B, EasyContext, Wan2.2, Wan2.1, Wan2.2-T2V-A14B, Wan2.1-VACE-1.3B, Wan2.1-FLF2V-14B-720P, maxdiffusion, OpenDiT, Wan2.2-Lightning, Wan2.2-Lightning

Simple and Effective Masked Diffusion Language Models

arXiv14 repos

arXiv:2406.07524

mdlm-owt, dllm, SDAR, dLLM-RL, FreeDave, mdlm, bd3lm, BDM, bd3lms, BD_DNA, bd3lm-owt-block_size1024-pretrain, ar-noeos-owt, diffusion-ebm, Open-dLLM

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

arXiv14 repos

arXiv:2408.01800

OmniLMM, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB, MiniCPM-V-4.6-Thinking-GPTQ, MiniCPM-o-4_5-gguf, MiniCPM-o-4_5-AWQ, MiniCPM-o-4_5

ActionStudio: A Lightweight Framework for Data and Training of Large Action Models

arXiv14 repos

arXiv:2503.22673

xLAM-7b-r, xLAM, xLAM-1b-fc-r, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, Llama-xLAM-2-8b-fc-r, xLAM-2-1b-fc-r, xLAM-2-3b-fc-r, xLAM-2-32b-fc-r, xLAM-8x22b-r, xLAM-7b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, xLAM-8x7b-r, xLAM-2-3b-fc-r-gguf

Helios: Real Real-Time Long Video Generation Model

arXiv14 repos

arXiv:2603.04379

Open-Sora-Plan, MagicTime, ChronoMagic-Bench, ConsisID, Helios, Helios-14B-RealTime, Helios-14B-RealTime-AOTI, Helios-Base, HeliosBench-Weights, Helios-Mid, Helios-Distilled, OpenS2V-Nexus, helios, tele_Gen

Measuring Massive Multitask Language Understanding

arXiv13 repos

arXiv:2009.03300

XVERSE-7B, XuanYuan, mmlu, PodGPT, mmlu-redux, Baichuan2, Qwen-14B-Chat, Qwen-7B-Chat, Baichuan-13B-Base, Baichuan-13B-Chat, Qwen-1_8B, JMMLU, mmlu_ru

BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models

arXiv13 repos

arXiv:2104.08663

beir, SFR-Embedding-Mistral, multilingual-e5-base, multilingual-e5-large, e5-large-unsupervised, e5-small-unsupervised, e5-base-unsupervised, e5-mistral-7b-instruct, multilingual-e5-large-instruct, e5-large-v2, gte-multilingual-base, RAG, RAGatouille

SimCSE: Simple Contrastive Learning of Sentence Embeddings

arXiv13 repos

arXiv:2104.08821

DermL2V-training, DermL2V-tmp, llm2vec, SimCSE, sup-simcse-roberta-large, sup-simcse-bert-large-uncased, Thai-Sentence-Vector-Benchmark, simcse-model-roberta-base-thai, simcse-model-XLMR, simcse-model-m-bert-thai-cased, simcse-model-wangchanberta, simcse-model-phayathaibert, simcse-model-distil-m-bert

Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

arXiv13 repos

arXiv:2108.12409

bloom, bloom-optimizer-states, falcon-rw-1b, Baichuan-13B-Base, Baichuan-13B-Chat, bloom-560m, bloom-1b7, mpt-1b-redpajama-200b, mpt-1b-redpajama-200b-dolly, bloom-7b1, bloom-1b1, bloom-3b, jina-colbert-v1-en

OCR-free Document Understanding Transformer

arXiv13 repos

arXiv:2111.15664

donut, donut-base-finetuned-cord-v2, donut-base-finetuned-cord-v1, donut-base-finetuned-cord-v1-2560, donut-base-finetuned-zhtrainticket, donut-base-finetuned-rvlcdip, donut-base-finetuned-docvqa, donut-base, donut-proto, aipeaks-pipeline-workshop, donut, donut1, donut-master

Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition

arXiv13 repos

arXiv:2305.05084

parakeet-rnnt-1.1b, parakeet-tdt_ctc-1.1b, parakeet-ctc-0.6b, parakeet-tdt_ctc-110m, parakeet-tdt-1.1b, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, parakeet-rnnt-0.6b, SLU_pipeline, stt_ar_fastconformer_hybrid_large_pcd_v1.0, canary-1b, diar_sortformer_4spk-v1, stt_ru_fastconformer_hybrid_large_pc

MultiLegalPile: A 689GB Multilingual Legal Corpus

arXiv13 repos

arXiv:2306.02069

LEXTREME, Multi_Legal_Pile, legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, MultilingualLegalLMPretraining, legal-swiss-roberta-large, legal-english-roberta-base, legal-croatian-roberta-base, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

arXiv13 repos

arXiv:2306.02707

StableBeluga2, OpenOrca, orca_dpo_pairs, orca_mini_7B-GPTQ, neural-chat-7b-v3-1, vigogne, Mistral-7B-OpenOrca, Jellyfish-13B, OpenOrca-KO, KOR-OpenOrca-Platypus, OpenOrca-gugugo-ko, orca_mini_3B-GGML, StableBeluga-7B

PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

arXiv13 repos

arXiv:2310.00426

PixArt-alpha, PixArt-alpha, PixArt-LCM-XL-2-1024-MS, PixArt-alpha, PixArt-XL-2-512x512, PixArt-LCM, PixArt-XL-2-256x256, PixArt-XL-2-1024-MS, PixArt-Sigma-XL-2-512-MS, Open-Sora-Plan-v1.2.0, pixeart, PixArt-alpha, PixArt-LCM

Safe RLHF: Safe Reinforcement Learning from Human Feedback

arXiv13 repos

arXiv:2310.12773

alpaca-7b-reproduced, safe-rlhf, beaver-7b-v2.0-cost, beaver-7b-v3.0-cost, beaver-7b-v3.0-reward, beaver-7b-v1.0, beaver-7b-v1.0-cost, beaver-7b-unified-cost, beaver-7b-v3.0, beaver-7b-v2.0, beaver-7b-v1.0-reward, beaver-7b-unified-reward, beaver-7b-v2.0-reward

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

arXiv13 repos

arXiv:2311.10122

LanguageBind, MoE-LLaVA, Video-LLaVA, MoE-LLaVA-Qwen-1.8B-4e, MoE-LLaVA-Phi2-2.7B-4e, MoE-LLaVA-Phi2-2.7B-4e-384, MoE-LLaVA-StableLM-1.6B-4e-384, MoE-LLaVA-StableLM-Pretrain, MoE-LLaVA-Phi2-Pretrain, MoE-LLaVA-StableLM-1.6B-4e, MoE-LLaVA-Qwen-Pretrain, MoE-LLaVA-Phi2-384-Pretrain, LLMBind

Efficient Multimodal Learning from Data-centric Perspective

arXiv13 repos

arXiv:2402.11530

Bunny-v1_1-data, Bunny, Bunny-Llama-3-8B-V, Bunny-v1_0-3B-zh, Bunny-v1_0-4B-gguf, Bunny-v1_0-4B, Bunny-v1_1-4B, Bunny-v1_1-Llama-3-8B-V, Bunny-Llama-3-8B-V-gguf, Bunny-v1_0-data, Bunny-v1_0-2B-zh, bunny-phi-2-siglip-lora, Bunny-v1_0-3B

AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning

arXiv13 repos

arXiv:2402.15506

xLAM-7b-r, xLAM, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, xLAM-8x22b-r, xLAM-2-3b-fc-r-gguf, xLAM-2-32b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, xLAM-8x7b-r, Llama-xLAM-2-8b-fc-r, xLAM-v0.1-r, xLAM-2-1b-fc-r, xLAM-2-3b-fc-r

ViTamin: Designing Scalable Vision Models in the Vision-Language Era

arXiv13 repos

arXiv:2404.02132

ViTamin, ViTamin-XL-384px, ViTamin-L-256px, ViTamin-L-224px, ViTamin-L-336px, ViTamin-L-384px, ViTamin-L2-384px, ViTamin-L2-256px, ViTamin-XL-256px-s13B, ViTamin-L2-224px, ViTamin-L2-336px, ViTamin-XL-336px, ViTamin-XL-256px

YOLOv10: Real-Time End-to-End Object Detection

arXiv13 repos

arXiv:2405.14458

yolov10, yolov10s, YOLOv10, Yolov10, yolov10n, yolov10m, yolov10b, yolov10l, yolov10x, yolov10, spectra, yolo10-train, CalEstimator

arXiv:2407.21783

arXiv13 repos

arXiv:2407.21783

STRING, Llama-3.1-Swallow-70B-v0.1, Llama-3.1-Swallow-8B-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-70B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.2, Llama-3.1-Swallow-8B-v0.2, Llama-3.1-Swallow-70B-Instruct-v0.3, Llama-3.3-Swallow-70B-v0.4, Llama-3.3-Swallow-70B-Instruct-v0.4, Llama-3.1-Swallow-8B-Instruct-v0.3, Llama-3.1-Swallow-8B-Instruct-v0.5, Llama-3.1-Swallow-8B-v0.5

Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

arXiv13 repos

arXiv:2503.09573

ECHO, dllm, LightningRL, SDAR, dLLM-RL, FreeDave, bd3lm, BDM, bd3lms, BD_DNA, bd3lm-owt-block_size1024-pretrain, sedd-noeos-owt, ar-noeos-owt

Wan: Open and Advanced Large-Scale Video Generative Models

arXiv13 repos

arXiv:2503.20314

Wan2.2-I2V-A14B-Diffusers, Wan2.2-Animate-14B, Wan2.2-TI2V-5B-Diffusers, Wan2.1-VACE-14B, Wan2.2-T2V-A14B-Diffusers, Wan2.2-S2V-14B, Wan2.2, Wan2.1, Wan2.2-T2V-A14B, Wan2.2-TI2V-5B, Wan2.2-I2V-A14B, Wan2.1-VACE-1.3B, Wan2.1-FLF2V-14B-720P

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

arXiv12 repos

arXiv:1910.01108

distilbert-base-uncased, scifact, distilbert-base-uncased-distilled-squad, DistilBERT-BERT_WNLI_NER, berts, turkish-bert, distilbert-base-turkish-cased, europeana-bert, swift-coreml-transformers, DistilBERT, LEXTREME, distilgpt2

Data Bootstrapping Approaches to Improve Low Resource Abusive Language Detection for Indic Languages

arXiv12 repos

arXiv:2204.12543

english-abusive-MuRIL, IndicAbusive, malayalam-codemixed-abusive-MuRIL, bengali-abusive-MuRIL, tamil-codemixed-abusive-MuRIL, kannada-codemixed-abusive-MuRIL, hindi-abusive-MuRIL, marathi-codemixed-abusive-MuRIL, urdu-codemixed-abusive-MuRIL, indic-abusive-allInOne-MuRIL, hindi-codemixed-abusive-MuRIL, urdu-abusive-MuRIL

MTEB: Massive Text Embedding Benchmark

arXiv12 repos

arXiv:2210.07316

mteb, SFR-Embedding-Mistral, multilingual-e5-base, multilingual-e5-large, e5-large-unsupervised, e5-small-unsupervised, e5-base-unsupervised, e5-mistral-7b-instruct, multilingual-e5-large-instruct, e5-large-v2, mteb-1.34.14, ru_sci_bench_mteb

Segment Anything

arXiv12 repos

arXiv:2304.02643

geti-instant-learn, sd-webui-inpaint-anything, sam-vit-large, Depth-Estimation, sd-webui-inpaint-anything, MobileSAM, Flow-Inference-Time-Scaling, SEED-Data-Edit-Part2-3, SEED-Data-Edit, 3D-LLM, medsam-vit-base, SAMReg

Universal and Transferable Adversarial Attacks on Aligned Language Models

arXiv12 repos

arXiv:2307.15043

OBLITERATUS, JBB-Behaviors, llm-attacks, jailbreakbench, RAIN, nanoGCG, Magic_Words, certified-llm-safety, Jailbreak_LLM, GA, llm-jailbreaking-defense, LLMart

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

arXiv12 repos

arXiv:2308.12966

Qwen2-VL-7B-Instruct, Qwen2.5-VL-72B-Instruct, Qwen2.5-VL-3B-Instruct-GGUF, UGround-V1-7B, UGround-V1-2B, UGround-V1-72B, Qwen2.5-VL-7B-Instruct-GGUF, Qwen-VL, Qwen2-VL-7B-Instruct, Qwen-VL-Chat, CoBSAT, Qwen2.5-VL-7B-Instruct-GGUF

Mistral 7B

arXiv12 repos

arXiv:2310.06825

Mistral-7B-v0.1, Mistral-7B-Instruct-v0.2, eCeLLM, Mistral-7B-Instruct-v0.1, SFR-Embedding-Mistral, ROOT-RAG, Mixtral-8x7B-Instruct-v0.1, speechless-mistral-six-in-one-7b, speechless-mistral-dolphin-orca-platypus-samantha-7B-GGUF, speechless-mistral-dolphin-orca-platypus-samantha-7b, speechless-mistral-dolphin-orca-platypus-samantha-7B-GPTQ, speechless-mistral-dolphin-orca-platypus-samantha-7B-AWQ

CogVLM: Visual Expert for Pretrained Language Models

arXiv12 repos

arXiv:2311.03079

glm-4v-9b, glm-4v-9b, cogvlm2-llama3-chat-19B, cogagent-chat-hf, cogvlm-chat-hf, cogvlm-base-490-hf, cogvlm-grounding-base-hf, cogvlm-grounding-generalist-hf, cogagent-vqa-hf, cogvlm-base-224-hf, cogagent-9b-20241220, visualglm-6b

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

arXiv12 repos

arXiv:2401.15947

LanguageBind, MoE-LLaVA, Video-LLaVA, MoE-LLaVA-Phi2-2.7B-4e, MoE-LLaVA-Qwen-1.8B-4e, MoE-LLaVA-Phi2-2.7B-4e-384, MoE-LLaVA-StableLM-1.6B-4e-384, MoE-LLaVA-Phi2-Pretrain, MoE-LLaVA-StableLM-1.6B-4e, MoE-LLaVA-Phi2-384-Pretrain, MoE-LLaVA-StableLM-Pretrain, MoE-LLaVA-Qwen-Pretrain

Yi: Open Foundation Models by 01.AI

arXiv12 repos

arXiv:2403.04652

Yi, Yi-1.5-9B-Chat, Yi-1.5-34B-Chat, Yi-1.5-6B-Chat, Yi-34B-Chat, Yi-1.5-6B, Yi-1.5-9B-Chat-16K, Yi-1.5-9B, Yi-1.5-9B-32K, Yi-VL-6B, Yi-VL-34B, llm_project

InternLM2 Technical Report

arXiv12 repos

arXiv:2403.17297

internlm2_5-7b-chat-1m, internlm2-20b, internlm3-8b-instruct, internlm2-chat-7b, internlm2_5-7b-chat, internlm2-chat-1_8b, internlm2_5-7b, internlm2-7b, internlm2-chat-20b, Bespoke-MiniCheck-7B, internlm2-7b-reward, internlm2-20b-reward

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

arXiv12 repos

arXiv:2403.18814

MGM, MGM-2B, MGM-7B, MGM-13B, MGM-8B, MGM-8x7B, MGM-34B, MGM-7B-HD, MGM-13B-HD, MGM-8B-HD, MGM-8x7B-HD, MGM-34B-HD

Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities

arXiv12 repos

arXiv:2404.17790

Swallow-70b-hf, Swallow-70b-NVE-hf, Swallow-13b-NVE-hf, Swallow-7b-NVE-hf, Swallow-7b-NVE-instruct-hf, Swallow-70b-instruct-hf, Swallow-70b-NVE-instruct-hf, Swallow-7b-plus-hf, Swallow-7b-hf, Swallow-13b-instruct-hf, Swallow-7b-instruct-hf, Swallow-13b-hf

Qwen2 Technical Report

arXiv12 repos

arXiv:2407.10671

evalplus, Qwen2.5-72B-Instruct, Qwen2.5-32B-Instruct, Qwen2.5-Coder-32B-Instruct, Qwen2.5-Coder-14B-Instruct, Qwen2.5-3B-Instruct, Qwen2.5-0.5B-Instruct, Qwen2.5-3B, Qwen2.5-14B-Instruct, Qwen2.5-14B-Instruct-AWQ, Qwen2.5-Coder-7B-Instruct, Qwen2.5-Coder-0.5B

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

arXiv12 repos

arXiv:2502.14786

siglip2-giant-opt-patch16-384, siglip2-so400m-patch16-512, siglip2-base-patch16-224, siglip2-base-patch16-naflex, pd12m, siglip2-so400m-patch16-naflex, vit_base_patch16_siglip_512.v2_webli, RAE, siglip2-so400m-patch14-384, piedomains-image, TiViT, siglip2-so400m-patch14-224

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

arXiv12 repos

arXiv:2503.14734

Eagle, physical-ai-studio, GR00T-N1.6-G1-PnPAppleToPlate, GR00T-N1.6-DROID, GR00T-N1.6-bridge, GR00T-N1.6-fractal, GR00T-N1.6-BEHAVIOR1k, GR00T-N1.7-SimplerEnv-Fractal, GR00T-N1.7-DROID, GR00T-N1.7-SimplerEnv-Bridge, GR00T-N1.7-LIBERO, GR00T-N1.7-3B

FG-CLIP: Fine-Grained Visual and Textual Alignment

arXiv12 repos

arXiv:2505.05071

fg-clip-base, FG-CLIP, fg-clip2-large, fg-clip2-base, DCI-CN, fg-clip-large, BoxClass-CN, DOCCI-CN, fg-clip2-so400m, FineHARD, LIT-CN, vit-dog

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

arXiv12 repos

arXiv:2508.06471

GLM-4.7, GLM-4.7-Flash, GLM-4.5, GLM-4.5-FP8, GLM-4.7-FP8, GLM-4.5-Air, GLM-4.6, GLM-4.6-FP8, GLM-4.5, GLM-4.5-Air-FP8, GLM-4.5-Base, GLM-4.5-Air-Base

BYOL: Bring Your Own Language Into LLMs

arXiv12 repos

arXiv:2601.10804

byol, Global-MMLU-Lite, byol-nya-12b-cpt, byol-mri-1b-cpt, byol-nya-12b-merged, byol-mri-4b-cpt, byol-mri-4b-merged, byol-nya-4b-cpt, byol-nya-1b-cpt, byol-mri-12b-cpt, byol-mri-12b-merged, byol-nya-4b-merged

RLDX-1 Technical Report

arXiv12 repos

arXiv:2605.03269

RLDX-1, RLDX-1-PT-IMG, RLDX-1-MT-DROID, RLDX-1-PT, RLDX-1-MT-ALLEX, RLDX-1-FT-LIBERO, RLDX-1-FT-GR1, RLDX-1-FT-SIMPLER-GOOGLE, RLDX-1-FT-ROBOCASA, RLDX-1-FT-SIMPLER-WIDOWX, RLDX-1-FT-RC365, RLDX-FineAct

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

arXiv12 repos

arXiv:2605.08985

OmniLMM, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB, MiniCPM-V-4.6-Thinking-GPTQ, LLaVA-UHD-v4

Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

arXiv11 repos

arXiv:1908.10084

paraphrase-multilingual-mpnet-base-v2, SimCSE, LateOn-Code, LateOn-Code-edge, distiluse-base-multilingual-cased-v2, splade-ecommerce-esci, SauerkrautLM-Multi-Reason-ModernColBERT, ColBERT-Zero, langcache-embed-v1, langcache-embed-v2, average_word_embeddings_glove.6B.300d

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

arXiv11 repos

arXiv:1909.08053

awesome-gpu-engineering, bloom, MindSpeed-MM, bloom-optimizer-states, bloom-560m, awsome-llm-papers, bloom-1b7, mesh-transformer-jax, bloom-7b1, bloom-1b1, bloom-3b

GLM: General Language Model Pretraining with Autoregressive Blank Infilling

arXiv11 repos

arXiv:2103.10360

GLM, SwissArmyTransformer, chatglm2-6b-32k, chatglm2-6b, chatglm-6b, chatglm2-6b-int4, chatglm3-6b-32k, chatglm3-6b-32k, chatglm3-6b-base, visualglm-6b, KoGLM

LexGLUE: A Benchmark Dataset for Legal Language Understanding in English

arXiv11 repos

arXiv:2110.00976

legalbert-large-1.7M-1, legalbert-large-1.7M-2, legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, legal-english-roberta-base, legal-swiss-roberta-large, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large

Training language models to follow instructions with human feedback

arXiv11 repos

arXiv:2203.02155

ik_llama.cpp, llama.cpp, Open-Assistant, databricks-dolly-15k, no_robots, llama.cpp, llama.cpp.qwen2.5vl, atomic-llama-cpp-turboquant, alpaca_farm, llama.cpp, llama.cpp

BigVGAN: A Universal Neural Vocoder with Large-Scale Training

arXiv11 repos

arXiv:2206.04658

HierSpeechpp, bigvgan_v2_24khz_100band_256x, BigVGAN, bigvgan_v2_44khz_128band_512x, bigvgan_base_24khz_100band, bigvgan_22khz_80band, bigvgan_v2_22khz_80band_fmax8k_256x, bigvgan_v2_22khz_80band_256x, bigvgan_24khz_100band, bigvgan_v2_44khz_128band_256x, bigvgan_base_22khz_80band

LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain

arXiv11 repos

arXiv:2301.13126

LEXTREME, legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, lextreme, legal-swiss-roberta-large, legal-english-roberta-base, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large

DINOv2: Learning Robust Visual Features without Supervision

arXiv11 repos

arXiv:2304.07193

geti-instant-learn, rf-detr, dinov2-small, dinov2-large, dinov2-giant, hashing-baseline, prismatic-vlms, RAE, CVProject, DinoBloom, TiViT

Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

arXiv11 repos

arXiv:2305.14233

ultrachat_200k, zephyr-7b-alpha, zephyr-7b-beta, Llama-3-8B-Instruct-Gradient-1048k, Llama-3-70B-Instruct-Gradient-262k, Llama-3-70B-Instruct-Gradient-1048k, Llama-3-8B-Instruct-262k, Llama-3-8B-Instruct-Gradient-4194k, llm-finetuning, ultrachat, Instruct-SkillMix-SDD

QLoRA: Efficient Finetuning of Quantized LLMs

arXiv11 repos

arXiv:2305.14314

Qwen, qlora, text-generation-inference, smash, GPTQ-for-LLaMa, vigogne, llmtools, llama2-70b, llmtools, level3_nlp_finalproject-nlp-12, ainn-final

Unified Training of Universal Time Series Forecasting Transformers

arXiv11 repos

arXiv:2402.02592

moirai-1.0-R-base, uni2ts, moirai-2.0-R-small, lotsa_data, stock_forecasting, Samay, A2TTA, moment, moirai-1.0-R-large, moirai-1.0-R-small, uni2ts

RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics

arXiv11 repos

arXiv:2406.10721

RoboPoint, robopoint-v1-vicuna-v1.5-13b, robopoint-v1-vicuna-v1.5-13b-lora, robopoint-v1-llama-2-13b, robopoint_data, robopoint-v1-llama-2-13b-lora, robopoint-v1-llama-2-7b-lora, robopoint-v1-vicuna-v1.5-7b-lora, where2place, Robopoint_Humble, Cosmos-Reason2-32B

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

arXiv11 repos

arXiv:2503.06520

refCOCOg_2k_840, Seg-Zero, refCOCOg_9k_840, Seg-Zero-7B, VisionReasoner_multi_object_7k_840, ReasonSeg_val, ReasonSeg_test, VisionReasoner-7B, VisionReasoner, Seg-Zero, VisionReasoner

LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning

arXiv11 repos

arXiv:2503.15621

LLaVA-MORE, LLaVA_MORE-llama_3_1-8B-pretrain, LLaVA_MORE-llama_3_1-8B-siglip-finetuning, LLaVA_MORE-llama_3_1-8B-S2-pretrain, LLaVA_MORE-llama_3_1-8B-siglip-pretrain, LLaVA_MORE-llama_3_1-8B-finetuning, LLaVA_MORE-llama_3_1-8B-S2-finetuning, LLaVA_MORE-llama_3_1-8B-S2-siglip-pretrain, LLaVA_MORE-llama_3_1-8B-S2-siglip-finetuning, LLaVA_MORE-phi_4-finetuning, LLaVA_MORE-gemma_2_9b-finetuning

VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning

arXiv11 repos

arXiv:2505.12081

Seg-Zero, VisionReasoner_multi_object_7k_840, refCOCOg_9k_840, ReasonSeg_val, VisionReasoner_multi_object_1k_840, ReasonSeg_test, VisionReasoner-7B, VisionReasoner, TaskRouter-1.5B, VisionReasoner, Seg-Zero

Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models

arXiv11 repos

arXiv:2507.17702

AntAngelMed, AntAngelMed, Ling-flash-base-2.0, Ling-mini-2.0, Ling-mini-base-2.0-5T, Ling-mini-base-2.0, Ling-mini-base-2.0-10T, Ling-mini-base-2.0-15T, Ling-mini-base-2.0-20T, Ling-flash-2.0, Ling-1T

MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

arXiv11 repos

arXiv:2509.18154

OmniLMM, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB, MiniCPM-V-4.6-Thinking-GPTQ

X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model

arXiv11 repos

arXiv:2510.10274

X-VLA-RoboTwin2, X-VLA-Google-Robot, xvla-base, X-VLA, X-VLA-Libero, X-VLA-SoftFold, X-VLA-Calvin-ABC_D, X-VLA-Pt, X-VLA-AgiWorld-Challenge, X-VLA-WidowX, X-VLA-VLABench

A visual-language foundation model for computational pathology

Nature11 repos

Nature:s41591-024-02856-4

AtlasPatch, PIANO, KEEP, KEEP, VLSA, SEAL, HistAug, histaug-conch, PathPT, Histopathology_Benchmark, dpfm_factory

A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark

arXiv10 repos

arXiv:1910.04867

CLIP-convnext_xxlarge-laion2B-s34B-b82K-augreg-soup, CLIP-ViT-H-14-laion2B-s32B-b79K, CLIP-ViT-bigG-14-laion2B-39B-b160k, CLIP-ViT-B-32-laion2B-s34B-b79K, CLIP-ViT-L-14-laion2B-s32B-b82K, CLIP-ViT-g-14-laion2B-s12B-b42K, CLIP-ViT-B-32-roberta-base-laion2B-s12B-b32k, CLIP-ViT-H-14-frozen-xlm-roberta-large-laion5B-s13B-b90k, CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft-soup, CLIP-ViT-B-16-laion2B-s34B-b88K

TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification

arXiv10 repos

arXiv:2010.12421

tweetnlp, twitter-roberta-base, twitter-roberta-base-irony, twitter-roberta-base-offensive, twitter-roberta-base-emoji, twitter-roberta-base-emotion, tweeteval, twitter-roberta-base-hate, twitter-roberta-base-sentiment, lares

Self-attention Does Not Need $O(n^2)$ Memory

arXiv10 repos

arXiv:2112.05682

LLaVA, FastChat, FastChat, Libra, multilingual_mt_bench, FastChat, TriPlaneLLaVA, hermes-llava, LLaVA-toy, TinyLLava

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

arXiv10 repos

arXiv:2205.11487

stable-diffusion, stable-diffusion-v1-4, stable-diffusion-v-1-4-original, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, stable-diffusion-inpainting, stable-diffusion-uncrop, stable-diffusion-v1.5-webnn

Matryoshka Representation Learning

arXiv10 repos

arXiv:2205.13147

RAG-Retrieval, ogham-mcp, contrastors, nomic-embed-text-v2-moe, nomic-embed-text-v1.5, UForm, contrastors, cholesky_encoder, arctic-embed, langcache-embed-v2

Classifier-Free Diffusion Guidance

arXiv10 repos

arXiv:2207.12598

stable-diffusion, stable-diffusion-v1-4, stable-diffusion-v-1-4-original, coreml-stable-diffusion-v1-5-palettized, coreml-stable-diffusion-1-4-palettized, coreml-stable-diffusion-v1-5, coreml-stable-diffusion-v1-4, stable-diffusion-inpainting, smalldiffusion, stable-diffusion-v1.5-webnn

Visual Instruction Tuning

arXiv10 repos

arXiv:2304.08485

LLaVA, minimind-v, MedAI-project, llava-implementation, nanoMFM, TriPlaneLLaVA, CoBSAT, LLaVA-toy, multi_token, TinyLLava

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

arXiv10 repos

arXiv:2306.00890

LLaVA, llava-med-v1.5-mistral-7b, LLaVA-Med, libra-llava-med-v1.5-mistral-7b, libra-llava-rad, llava-rad, TriPlaneLLaVA, LLaVA-toy, LLaDA-MedV, TinyLLava

h2oGPT: Democratizing Large Language Models

arXiv10 repos

arXiv:2306.08161

h2ogpt, fork, honogpt, OpenGPT-v1, h2ogpt, H2O_AI, h2ogpt, h2ogpt, llmbot, h2ogpt

C-Pack: Packed Resources For General Chinese Embeddings

arXiv10 repos

arXiv:2309.07597

bge-small-zh-v1.5, bge-large-zh-v1.5, bge-base-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-en-icl, bge-large-en, bge-base-en, bge-small-en, mteb-1.34.14

A decoder-only foundation model for time-series forecasting

arXiv10 repos

arXiv:2310.10688

timesfm, timesfm-3.0-pytorch, timesfm-1.0-200m, uni2ts, Samay, A2TTA, timesfm-1.0-200m-pytorch, moment, timesfm-2.5-200m-pytorch, uni2ts

HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs

arXiv10 repos

arXiv:2311.09774

HuatuoGPT2-SFT-GPT4-140K, HuatuoGPT2-13B, HuatuoGPT2-Pretraining-Instruction, HuatuoGPT2-34B, HuatuoGPT2-7B, HuatuoGPT-II, HuatuoGPT2-7B-8bits, HuatuoGPT2-7B-4bits, HuatuoGPT2-34B-8bits, HuatuoGPT2-34B-4bits

Self-Play Preference Optimization for Language Model Alignment

arXiv10 repos

arXiv:2405.00675

SPPO, Mistral7B-PairRM-SPPO-Iter1, Gemma-2-9B-It-SPPO-Iter1, Gemma-2-9B-It-SPPO-Iter3, Llama-3-Instruct-8B-SPPO-Iter1, Mistral7B-PairRM-SPPO-Iter2, Gemma-2-9B-It-SPPO-Iter2, Mistral7B-PairRM-SPPO-Iter3, Llama-3-Instruct-8B-SPPO-Iter3, Llama-3-Instruct-8B-SPPO-Iter2

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

arXiv10 repos

arXiv:2408.06072

CogVideo, CogVideoX1.5-5B, CogVideoX-2b, CogVideoX-5b-I2V, CogVideoX1.5-5B-I2V, CogVideoX-2b, CogVideoX1.5-5B, CogVideoX-5b, CogVideoX-5b-I2V, cogvlm2-llama3-caption

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

arXiv10 repos

arXiv:2409.18042

emova_speech_tokenizer_hf, EMOVA, qwen2vit600m, emova-qwen-2-5-3b-hf, emova-qwen-2-5-3b, emova-qwen-2-5-7b-hf, emova-qwen-2-5-72b, EMOVA_speech_tokenizer, emova-qwen-2-5-7b, emova-qwen-2-5-72b-hf

Open-Sora Plan: Open-Source Large Video Generation Model

arXiv10 repos

arXiv:2412.00131

Open-Sora-Plan, Open-Sora-Plan-v1.3.0, MagicTime, ChronoMagic-Bench, MoE-LLaVA, Video-LLaVA, ConsisID, OpenS2V-Nexus, UniWorld-V1, UniWorld-V1-NF4

Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model

arXiv10 repos

arXiv:2501.02790

DenseRewardRLHF-PPO, Phi-3-mini-4k-token-ppo-60k, Phi-3-mini-4k-segment-ppo-60k, meta-llama-3.1-instruct-8b-token-ppo-60k, meta-llama-3.1-instruct-8b-segment-ppo-60k, Phi-3-mini-4k-bandit-ppo-60k, meta-llama-3.1-instruct-8b-bandit-ppo-60k, rlhflow-llama-3-sft-8b-v2-segment-ppo-60k, rlhflow-llama-3-sft-8b-v2-token-ppo-60k, rlhflow-llama-3-sft-8b-v2-bandit-ppo-60k

Building Instruction-Tuning Datasets from Human-Written Instructions with Open-Weight Large Language Models

arXiv10 repos

arXiv:2503.23714

Llama-3.1-Swallow-8B-Instruct-v0.5, Llama-3.1-Swallow-70B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.1, Llama-3.1-Swallow-8B-Instruct-v0.2, Llama-3.1-Swallow-8B-Instruct-v0.3, Llama-3.1-Swallow-70B-Instruct-v0.3, Llama-3.3-Swallow-70B-Instruct-v0.4, Gemma-2-Llama-Swallow-9b-it-v0.1, Gemma-2-Llama-Swallow-27b-it-v0.1, Gemma-2-Llama-Swallow-2b-it-v0.1

APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay

arXiv10 repos

arXiv:2504.03601

xLAM, Llama-xLAM-2-70b-fc-r, xLAM-2-1b-fc-r-gguf, Llama-xLAM-2-8b-fc-r, xLAM-2-1b-fc-r, xLAM-2-3b-fc-r, xLAM-2-32b-fc-r, Llama-xLAM-2-8b-fc-r-gguf, APIGen-MT-5k, xLAM-2-3b-fc-r-gguf

REAP the Experts: Why Pruning Prevails for One-Shot MoE compression

arXiv10 repos

arXiv:2510.13999

Qwen3-Coder-30B-A3B-REAP-AWQ, llm-compressor, Qwen3-Coder-REAP-25B-A3B, Qwen3-Coder-REAP-25B-A3B-AWQ, inkling-mlx, Inkling-MLX-REAP12-4bit, Inkling-Small-MLX-REAP25-4bit, Inkling-MLX-REAP25-4bit, Inkling-MLX-REAP50-4bit, turboquant-vllm

Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning

arXiv10 repos

arXiv:2512.19687

pe-av-large, pe-av-small-16-frame, pe-av-base-16-frame, pe-av-large-16-frame, pe-av-small, pe-a-frame-small, pe-a-frame-base, pe-a-frame-large, pe-av-base, Echo-ViLD

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

arXiv10 repos

arXiv:2604.27393

MiniCPM-V-4.6-Thinking-GPTQ, MiniCPM-V-4.6, MiniCPM-V-4.6-gguf, MiniCPM-V-4.6-BNB, MiniCPM-V-4.6-GPTQ, MiniCPM-V-4.6-AWQ, MiniCPM-V-4.6-Thinking, MiniCPM-V-4.6-Thinking-gguf, MiniCPM-V-4.6-Thinking-AWQ, MiniCPM-V-4.6-Thinking-BNB

MolmoAct2: Action Reasoning Models for Real-world Deployment

arXiv10 repos

arXiv:2605.02881

MolmoAct2, MolmoAct2-Pretrain, Molmo2-ER, MolmoAct2-SO100_101, MolmoAct2-Think, MolmoAct2-BimanualYAM, MolmoAct2-Think-LIBERO, MolmoAct2-LIBERO, MolmoAct2-DROID, vla-edge

MOSS-VL Technical Report

arXiv10 repos

arXiv:2608.15045

MOSS-VL-Base-0708, MOSS-VL-Base-0408, MOSS-VL, MOSS-VL-Instruct-0708-NF4, MOSS-VL-Instruct-0708-FP8, MOSS-VL-Realtime-NF4, MOSS-VL-Realtime-FP8, MOSS-VL-Realtime, MOSS-VL-Instruct-0408, MOSS-VL-Instruct-0708

Fast Transformer Decoding: One Write-Head is All You Need

arXiv9 repos

arXiv:1911.02150

ChatGLM-6B, ChatGLM-6B, falcon-40b, chatglm2-6b-32k, chatglm2-6b, falcon-7b, chatglm2-6b-int4, falcon-40b-instruct, WebGLM

Revisiting Pre-Trained Models for Chinese Natural Language Processing

arXiv9 repos

arXiv:2004.13922

chinese-roberta-wwm-ext, chinese-xlnet-base, chinese-macbert-base, chinese-macbert-large, chinese-bert-wwm-ext, chinese-electra-base-discriminator, pycorrector, macbert4csc-base-chinese, repo-5146-pycorrector

IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding

arXiv9 repos

arXiv:2009.05387

indobert-base-p2, indobert-lite-large-p1, indobert-lite-large-p2, indobert-base-p1, indobert-large-p1, indobert-lite-base-p2, indobert-lite-base-p1, indobert-large-p2, TweetSentimentIndoBERT

Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

arXiv9 repos

arXiv:2204.07705

instructor-embedding, COIG, COIG, FLAN, tk-instruct-11b-def-pos, Tk-Instruct, instructor-embedding, tk-instruct-3b-def, tk-instruct-11b-def

GIT: A Generative Image-to-text Transformer for Vision and Language

arXiv9 repos

arXiv:2205.14100

git-large, GenerativeImage2Text, heron-chat-git-Llama-2-7b-v0, heron-preliminary-git-Llama-2-70b-v0, heron-chat-git-ja-stablelm-base-7b-v1, heron-chat-git-ja-stablelm-base-7b-v0, heron-chat-git-ELYZA-fast-7b-v0, CVProject, git-large-coco

PaLI: A Jointly-Scaled Multilingual Language-Image Model

arXiv9 repos

arXiv:2209.06794

siglip-so400m-patch14-384, siglip2-giant-opt-patch16-384, siglip2-so400m-patch16-512, siglip2-base-patch16-224, siglip2-base-patch16-naflex, siglip-base-patch16-224, siglip2-so400m-patch16-naflex, siglip2-so400m-patch14-384, siglip2-so400m-patch14-224

GLM-130B: An Open Bilingual Pre-trained Model

arXiv9 repos

arXiv:2210.02414

chatglm2-6b-32k, chatglm2-6b, chatglm2-6b-int4, chatglm3-6b-32k, chatglm-6b, chatglm3-6b-32k, chatglm3-6b-base, visualglm-6b, P-tuning-v2

ReAct: Synergizing Reasoning and Acting in Language Models

arXiv9 repos

arXiv:2210.03629

Awesome-OpenClaw, prompt-engineering, tau-bench, Qwen-7B-Chat, Qwen-14B-Chat, ReWOO, agentodyssey, AgentTuning, hacktech24_app_testing

ChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain Knowledge

arXiv9 repos

arXiv:2303.14070

medalpaca-13b, medalpaca-7b, doctorwithbloom, doctorwithbloom, doctorwithbloomz-7b1, doctorwithbloomz-7b1-mt, medalpaca-lora-30b-8bit, medalpaca-lora-7b-8bit, medalpaca-lora-13b-8bit

Exploring Human-Like Translation Strategy with Large Language Models

arXiv9 repos

arXiv:2305.04118

llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA

Iterative Translation Refinement with Large Language Models

arXiv9 repos

arXiv:2306.03856

llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA, topxgen-llama-4-scout-TEaR

One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support

arXiv9 repos

arXiv:2306.09237

legal-xlm-roberta-large, legal-xlm-longformer-base, legal-xlm-roberta-base, legal-english-roberta-base, legal-swiss-roberta-large, legal-swiss-roberta-base, legal-swiss-longformer-base, legal-portuguese-roberta-base, legal-english-roberta-large

ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

arXiv9 repos

arXiv:2310.04564

PowerInfer, ReluLLaMA-7B, ReluLLaMA-13B, prosparse-llama-2-13b, prosparse-llama-2-7b, Bamboo-base-v0_1, ReluFalcon-40B, Bamboo-DPO-v0_1, ReluLLaMA-70B

Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation

arXiv9 repos

arXiv:2312.02145

Marigold, marigold, marigold, marigold-depth-v1-1, marigold-normals-v1-1, marigold-iid-appearance-v1-1, marigold-iid-lighting-v1-1, marigold-depth-lcm-v1-0, prisma

Bilateral Reference for High-Resolution Dichotomous Image Segmentation

arXiv9 repos

arXiv:2401.03407

BiRefNet_HR, BiRefNet, FeyNobg, BiRefNet, BiRefNet_HR-matting, BiRefNet_dynamic, BiRefNet-matting, BiRefNet_lite-2K, BiRefNet-portrait

TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement

arXiv9 repos

arXiv:2402.16379

llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA, topxgen-llama-4-scout-TEaR

Evolutionary Optimization of Model Merging Recipes

arXiv9 repos

arXiv:2403.13187

evolutionary-model-merge, EvoLLM-JP-A-v1-7B, EvoLLM-JP-v1-10B, EvoLLM-JP-v1-7B, EvoVLM-JP-v1-7B, JA-VG-VQA-500, JA-VLM-Bench-In-the-Wild, python_moder_merge, adaptation-evolutionary-model-merge

OpenVLA: An Open-Source Vision-Language-Action Model

arXiv9 repos

arXiv:2406.09246

modified_libero_rlds, openvla-7b, openvla, openvla-7b-finetuned-libero-spatial, openvla-7b-finetuned-libero-object, openvla-7b-finetuned-libero-10, openvla-7b-finetuned-libero-goal, openvla-7b-prismatic, openvla-main-dhot1

Translating Step-by-Step: Decomposing the Translation Process for Improved Translation Quality of Long-Form Texts

arXiv9 repos

arXiv:2409.06790

llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA

VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

arXiv9 repos

arXiv:2410.05160

VLM2Vec-V2.0, MMEB-train, VLM2Vec-LLaVa-Next, VLM2Vec-Full, MMEB-V2, VLM2Vec-Qwen2VL-2B, MMEB-V3, VLM2Vec-Qwen2VL-7B, MMEB-eval

Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

arXiv9 repos

arXiv:2410.05243

UGround, UGround-V1-72B, UGround-V1-2B, UGround-V1-7B, llava_uground, android_world_seeact_v, Mind2Web_Live_SeeAct_V, UGround, mind2web-live-seeact-v

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

arXiv9 repos

arXiv:2501.02669

VLM_S2H, Eagle-X2-Llama3-8B-ConsecutiveTableReadout-Mix-160k, Eagle-X2-Llama3-8B-GridNavigation-AlignMixPlus-120k, Eagle-X2-Llama3-8B, Eagle-X2-Llama3-8B-GridNavigation-MixPlus-120k, Eagle-X2-Llama3-8B-VisualAnalogy-AlignMixPlus-120k, Eagle-X2-Llama3-8B-TableReadout-MixPlus-240k, Eagle-X2-Llama3-8B-VisualAnalogy-MixPlus-120k, Eagle-X2-Llama3-8B-TableReadout-AlignMixPlus-240k

Provence: efficient and robust context pruning for retrieval-augmented generation

arXiv9 repos

arXiv:2501.16214

provence-reranker-debertav3-v1, squeez, semantic-highlight-bilingual-v1, open_provence, open-provence-reranker-v1, query-context-pruner-multilingual-Qwen3-4B, open-provence-reranker-v1-gte-modernbert-base, open-provence-reranker-large-v1, open-provence-reranker-xsmall-v1

DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

arXiv9 repos

arXiv:2503.01183

DiffRhythm, DiffRhythm-base, DiffRhythm-full, DiffRhythm-vae, DiffRhythm2, Diffrhythm-linux, DiffRhythm, DiffRhythm-windows, DiffRhythm

Compositional Translation: A Novel LLM-based Approach for Low-resource Machine Translation

arXiv9 repos

arXiv:2503.04554

topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-BoA, llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-TEaR

MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning

arXiv9 repos

arXiv:2504.10160

topxgen-llama-4-scout-REFINE, topxgen-llama-4-scout-BoA, llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

arXiv9 repos

arXiv:2505.09358

Marigold, marigold, marigold-depth-lcm-v1-0, marigold-normals, marigold-iid, marigold-depth-v1-1, marigold-normals-v1-1, marigold-iid-appearance-v1-1, marigold-iid-lighting-v1-1

Learning to Reason without External Rewards

arXiv9 repos

arXiv:2505.19590

Qwen3-14B-Intuitor-MATH-1EPOCH, Qwen2.5-1.5B-Intuitor-MATH-1EPOCH, OLMo-2-7B-SFT-Intuitor-MATH-1EPOCH, Qwen2.5-3B-Intuitor-MATH-1EPOCH, Intuitor, Qwen2.5-3B-GRPO-MATH-1EPOCH, OLMo-2-7B-SFT-GRPO-MATH-1EPOCH, Qwen3-14B-GRPO-MATH-1EPOCH, Qwen2.5-1.5B-GRPO-MATH-1EPOCH

Show-o2: Improved Native Unified Multimodal Models

arXiv9 repos

arXiv:2506.15564

Show-o, show-o2-1.5B, show-o2-1.5B-w-video-und, show-o2-7B, show-o2-7B-w-video-und, show-o-512x512-wo-llava-tuning, show-o, show-o2-1.5B-HQ, show-o-w-clip-vit

gpt-oss-120b & gpt-oss-20b Model Card

arXiv9 repos

arXiv:2508.10925

gpt-oss-120b, gpt-oss-20b, GPT-OSS-Swallow-120B-RL-v0.1, GPT-OSS-Swallow-20B-RL-v0.1, GPT-OSS-Swallow-20B-SFT-v0.1, GPT-OSS-Swallow-120B-SFT-v0.1, Medical-GPT-OSS-Swallow-120B, GPT-OSS-Swallow-120B-RL-v0.1-MXFP4, GPT-OSS-Swallow-20B-RL-v0.1-MXFP4

Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding

arXiv9 repos

arXiv:2510.08668

Hulu-Med-32B, Hulu-Med-4B, Hulu-Med-7B, Hulu-Med-14B, Hulu-Med, Hulu-Med-235A22, Hulu-Med-Flash-Preview-27B, Hulu-Med-30A3, hulumed-prober

LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens

arXiv9 repos

arXiv:2510.11919

topxgen-llama-4-scout-REFINE, llm-reasoning-mt, topxgen-llama-4-scout-MAPS, topxgen-llama-4-scout-CoT, topxgen-llama-4-scout-SBYS, topxgen-llama-4-scout-TEaR, topxgen-llama-4-scout-CompTra, topxgen-llama-4-scout-Decomp, topxgen-llama-4-scout-BoA

Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens

arXiv9 repos

arXiv:2511.19418

CoVT, CoVT-Dataset, CoVT-7B-seg, CoVT-7B-seg_depth_dino_edge, CoVT-LLaVA-13B-depth, CoVT-7B-seg_depth_dino, CoVT-7B-depth, covt_data, COVT

LFM2 Technical Report

arXiv9 repos

arXiv:2511.23404

LFM2.5-Audio-1.5B, LFM2.5-350M, LFM2-VL-450M-ONNX, LFM2-2.6B-Exp-ONNX, LFM2-8B-A1B-ONNX, LFM2.5-Audio-1.5B-JP, LFM2-ColBERT-350M, kani-tts-2-en, kani-tts-2-pt

Radiology Report Generation with Layer-Wise Anatomical Attention

arXiv9 repos

arXiv:2512.16841

LAnA-v3, layer-wise-anatomical-attention, LAnA-Arxiv, LAnA, LAnA-v2, LAnA-MIMIC, LAnA-v5, LAnA-MIMIC-CHEXPERT, LAnA-v4

Data Science and Technology Towards AGI Part I: Tiered Data Management

arXiv9 repos

arXiv:2602.09003

Ultra-FineWeb-L1, UltraData-RL-2609, UltraData-SFT-Agent-2609, UltraData-Code, MiniCPM5-2B, MiniCPM5-2B-GGUF, MiniCPM5-1B-GGUF, MiniCPM5-1B, Ultra-FineWeb

MOSS-TTS Technical Report

arXiv9 repos

arXiv:2603.18090

MOSS-TTS-v1.5, MOSS-TTS-Nano, MOSS-TTS, MOSS-TTS-Local-Transformer-v1.5, MOSS-Audio-Tokenizer-Nano, MOSS-TTS-Realtime, MOSS-TTS-Norwegian-LoRA, MOSS-TTS-Nano-100M-ONNX, MOSS-Audio-Tokenizer-Nano-ONNX

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention

arXiv9 repos

arXiv:2606.07639

MOSS-VL-Base-0408, MOSS-VL-Instruct-0708-NF4, MOSS-VL-Instruct-0708-FP8, MOSS-VL-Realtime-NF4, MOSS-VL-Realtime-FP8, MOSS-VL-Realtime, MOSS-VL-Base-0708, MOSS-VL-Instruct-0408, MOSS-VL-Instruct-0708

Towards a general-purpose foundation model for computational pathology

Nature9 repos

Nature:s41591-024-02857-3

AtlasPatch, PIANO, TRIDENT, CPathPatchFeature, SEAL, HistAug, TridentEdited, histaug-uni, aegis

DCSpell: A Detector-Corrector Framework for Chinese Spelling Error Correction

ACM8 repos

ACM:3404835.3463050

macbert4mdcspell_v1, macbert4mdcspell_v2, macbert4csc_v2, macbert4mdcspell_v3, macbert4csc_v1, bert4csc_v1, Chinese-text-correction-papers, CTCResources

HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

arXiv8 repos

arXiv:2010.05646

MockingBird, nemo-nano-codec-22khz-1.89kbps-21.5fps, tts-hifigan-libritts-16kHz, fish-diffusion, tts-hifigan-ljspeech, dla-tts, hifi-gan, univnet

KLUE: Korean Language Understanding Evaluation

arXiv8 repos

arXiv:2105.09680

roberta-large, klue, pko-t5-base, bert-base, KLUE, pko-t5, pko-t5-large, pko-t5-small

SciFive: a text-to-text transformer model for biomedical literature

arXiv8 repos

arXiv:2106.03598

SciFive-large-Pubmed_PMC-MedNLI, SciFive, SciFive-base-Pubmed_PMC, SciFive-large-Pubmed_PMC, SciFive-large-Pubmed, SciFive-base-Pubmed, SciFive-base-PMC, SciFive-large-PMC

DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

arXiv8 repos

arXiv:2111.09543

DiagnosisCoding, LEXTREME, mdeberta-v3-base, DeBERTa, deberta-v3-small, deberta-v3-xsmall, deberta-v3-base, deberta-v3-large

OPT: Open Pre-trained Transformer Language Models

arXiv8 repos

arXiv:2205.01068

safe-rlhf, opt-125m, OPT-13B-Erebus, opt-13b, opt-2.7b, OPT-6B-nerys-v2, OPT-6.7B-Erebus, opt-350m

Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence

arXiv8 repos

arXiv:2209.02970

Fengshenbang-LM, Erlangshen-MegatronBert-1.3B, Taiyi-Stable-Diffusion-1B-Chinese-EN-v0.1, Taiyi-Stable-Diffusion-1B-Chinese-v0.1, Erlangshen-DeBERTa-v2-710M-Chinese, Erlangshen-DeBERTa-v2-97M-Chinese, Erlangshen-DeBERTa-v2-320M-Chinese, BDM1.0

Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation

arXiv8 repos

arXiv:2211.06687

larger_clap_music, clap-htsat-unfused, hashing-baseline, CLAP, clap-demo, larger_clap_music_and_speech, audio-embeddings, shira_audio

Scalable Diffusion Models with Transformers

arXiv8 repos

arXiv:2212.09748

DiT, nanoMFM, smalldiffusion, mdlm, bd3lms, bd3lm, BDM, BD_DNA

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning

arXiv8 repos

arXiv:2301.13688

OpenOrca, dolma, FLAN, Mistral-7B-OpenOrca, Jellyfish-13B, OpenOrca-KO, KOR-OpenOrca-Platypus, OpenOrca-gugugo-ko

T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

arXiv8 repos

arXiv:2302.08453

t2i-adapter-canny-sdxl-1.0, T2I-Adapter, T2I-Adapter, t2i-adapter-openpose-sdxl-1.0, t2i-adapter-sketch-sdxl-1.0, t2i-adapter-lineart-sdxl-1.0, t2i-adapter-depth-zoe-sdxl-1.0, t2i-adapter-depth-midas-sdxl-1.0

The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

arXiv8 repos

arXiv:2306.01116

fineweb, falcon-refinedweb, falcon-40b, falcon-rw-1b, falcon-7b, falcon-40b-instruct, SEA-PILE-v1, llm_project

INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models

arXiv8 repos

arXiv:2306.04757

InstructEvalImpact, flan-alpaca-gpt4-xl, flan-gpt4all-xl, flan-alpaca-large, flan-alpaca-base, flan-alpaca-xl, flan-alpaca-xxl, flan-sharegpt-xl

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

arXiv8 repos

arXiv:2308.01390

Otter, OpenFlamingo-9B-vitl-mpt7b, Otter, open_flamingo, OpenFlamingo-3B-vitl-mpt1b-langinstruct, OpenFlamingo-3B-vitl-mpt1b, OpenFlamingo-4B-vitl-rpj3b, OpenFlamingo-4B-vitl-rpj3b-langinstruct

Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

arXiv8 repos

arXiv:2308.09662

flan-alpaca-gpt4-xl, red-instruct, starling-7B, HarmfulQA, flan-alpaca-large, flan-alpaca-base, flan-alpaca-xl, flan-sharegpt-xl

ModelScope-Agent: Building Your Customizable Agent System with Open-source Large Language Models

arXiv8 repos

arXiv:2309.00986

ms-agent, swift, modelscope-agent, ms-swift, ms-swift, vmopd, ms, dense-retention-rl

MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

arXiv8 repos

arXiv:2309.12284

SOLAR-10.7B-Instruct-v1.0, Arithmo2-Mistral-7B, MetaMathQA, Arithmo2-Mistral-7B-adapter, Arithmo-Mistral-7B, Arithmo-Data, GSM8K_Backward, neural-chat-7b-v3-3

MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning

arXiv8 repos

arXiv:2310.03731

MathCoder, MathCodeInstruct-Plus, MathCodeInstruct, MathCoder-CL-7B, MathCodeInstruct, MathCoder-L-7B, MathCoder-L-13B, MathCoder-CL-34B

A Multi-Task Embedder For Retrieval Augmented LLMs

arXiv8 repos

arXiv:2310.07554

bge-small-zh-v1.5, bge-large-zh-v1.5, bge-base-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-large-en, bge-base-en, bge-small-en

Skywork: A More Open Bilingual Foundation Model

arXiv8 repos

arXiv:2310.19341

Skywork-13B-base, Skywork, mock_gsm8k_test, Skywork-13B-Math-8bits, Skywork-13B-Math, Skywork-13B-Base-3.1TB, ChineseDomainModelingEval, Skywork-13B-Base-8bits

Seamless: Multilingual Expressive and Streaming Speech Translation

arXiv8 repos

arXiv:2312.05187

w2v-bert-2.0, seamless_communication, seamless-streaming, seamless-m4t-medium, seamless-m4t-v2-large, seamless-m4t-large, conformer-shaw, fairseq2

AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One

arXiv8 repos

arXiv:2312.06709

RADIO, C-RADIOv3-L, C-RADIOv3-B, C-RADIOv4-SO400M, C-RADIOv4-H, C-RADIOv3-g, C-RADIOv3-H, RADIO

CogAgent: A Visual Language Model for GUI Agents

arXiv8 repos

arXiv:2312.08914

CogVLM, CogAgent, cogagent-chat-hf, cogagent-vqa-hf, CogAgent, cogagent-9b-20241220, CogVLM, CogAgent

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

arXiv8 repos

arXiv:2312.16886

MobileVLM, MobileLLaMA-2.7B-Chat, MobileLLaMA-1.4B-Chat, MobileVLM-1.7B, MobileVLM-3B, MobileLLaMA-1.4B-Base, MobileLLaMA-2.7B-Base, MobileLISA

Improving Text Embeddings with Large Language Models

arXiv8 repos

arXiv:2401.00368

DermL2V-training, Anchor-Embedding, DermL2V-tmp, llm2vec, SFR-Embedding-Mistral, tinyllama-embed, e5-mistral-7b-instruct, multilingual-e5-large-instruct

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

arXiv8 repos

arXiv:2401.06066

DeepSeek-Coder-V2-Instruct, DeepSeek-Coder-V2, DeepSeek-Coder-V2-Lite-Base, DeepSeek-Coder-V2-Lite-Instruct, DeepSeek-Coder-V2-Base, deepseek-moe-16b-chat, deepseek-moe-16b-base, nugie-jax-nemotron-3-nano

M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

arXiv8 repos

arXiv:2402.03216

bge-m3, bge-reranker-v2-m3, OKEAN, semantic-highlight-bilingual-v1, pyterrier_dr, pyterrier_dr_jpq, gte-multilingual-base, bge-reranker-v2.5-gemma2-lightweight

MOMENT: A Family of Open Time-series Foundation Models

arXiv8 repos

arXiv:2402.03885

Samay, MOMENT-1-large, moment, MOMENT-1-small, Timeseries-PILE, MOMENT-1-base, FM4Motor, TiViT

HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

arXiv8 repos

arXiv:2402.04249

OBLITERATUS, garak, JBB-Behaviors, HarmBench, HarmBench-Llama-2-13b-cls, HarmBench-Mistral-7b-val-cls, HarmBench-Llama-2-13b-cls-multimodal-behaviors, karma-electric-llama31-8b

Towards Building Multilingual Language Model for Medicine

arXiv8 repos

arXiv:2402.13963

MMedC, MMed-Llama-3-8B, MMedLM, MMedLM2, MMedBench, MMedLM2-1.8B, MMed-Llama-3-8B-EnIns, MMedLM

M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models

arXiv8 repos

arXiv:2404.00578

M3D, M3D-RefSeg, M3D-LaMed-Llama-2-7B, M3D-CLIP, M3D-VQA, M3D-Seg, RAIdio-Agents, adapt_med_seg

Towards Large-Scale Training of Pathology Foundation Models

arXiv8 repos

arXiv:2404.15217

towards_large_pathology_fms, vit_base_patch16_224.kaiko_ai_towards_large_pathology_fms, vit_small_patch16_224.kaiko_ai_towards_large_pathology_fms, vit_large_patch14_reg4_dinov2.kaiko_ai_towards_large_pathology_fms, vit_small_patch8_224.kaiko_ai_towards_large_pathology_fms, vit_base_patch8_224.kaiko_ai_towards_large_pathology_fms, Midnight, vit_large_patch14_reg4_224.kaiko_ai_towards_large_pathology_fms

RLHF Workflow: From Reward Modeling to Online RLHF

arXiv8 repos

arXiv:2405.07863

Online-RLHF, LLaMA3-SFT-v2, LLaMA3-SFT, pair-preference-model-LLaMA3-8B, Llama3-SFT-v2.0-epoch3, Llama3-SFT-v2.0-epoch1, FsfairX-LLaMA3-RM-v0.1, LLaMA3-iterative-DPO-final

Ovis: Structural Embedding Alignment for Multimodal Large Language Model

arXiv8 repos

arXiv:2405.20797

Ovis2.5-9B, Ovis2-34B, Ovis1.6-Gemma2-9B-GPTQ-Int4, Ovis1.6-Llama3.2-3B-GPTQ-Int4, Ovis1.6-Llama3.2-3B, Ovis2-34B-GPTQ-Int4, Ovis1.6-Gemma2-9B, Ovis2.5-2B

Refusal in Language Models Is Mediated by a Single Direction

arXiv8 repos

arXiv:2406.11717

ds4, OBLITERATUS, heretic, obliteratus, Heretic-Abliteration, som-refusal-directions, ds4, ds4

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

arXiv8 repos

arXiv:2406.16860

cambrian-13b, cambrian-34b, cambrian-phi3-3b, CV-Bench, cambrian-8b, Cambrian-10M, Cambrian-Alignment, phyworld

APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets

arXiv8 repos

arXiv:2406.18518

xLAM-7b-r, xLAM-1b-fc-r, xLAM-7b-fc-r, xLAM-1b-fc-r-gguf, xLAM-v0.1-r, xLAM-8x22b-r, xLAM-8x7b-r, xLAM-7b-fc-r-gguf

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

arXiv8 repos

arXiv:2406.18629

Math-Step-DPO-10K, Step-DPO, DeepSeekMath-RL-Step-DPO, Qwen2-7B-SFT-Step-DPO, Llama-3-70B-SFT-Step-DPO, Qwen2-57B-A14B-SFT-Step-DPO, Qwen1.5-32B-SFT-Step-DPO, Qwen2-72B-Instruct-Step-DPO

Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models

arXiv8 repos

arXiv:2409.11136

ru-promptriever, promptriever, RepLLaMA-reproduced, promptriever-mistral-v0.1-7b-v1, promptriever-llama3.1-8b-instruct-v1, promptriever-llama3.1-8b-v1, promptriever-llama2-7b-v1, ru-promptriever-qwen3-4b

PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation

arXiv8 repos

arXiv:2410.01680

RADIO, C-RADIOv3-L, C-RADIOv3-B, C-RADIOv4-SO400M, C-RADIOv4-H, C-RADIOv3-g, C-RADIOv3-H, RADIO

SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

arXiv8 repos

arXiv:2410.10629

nunchaku, nunchaku, Sana_1600M_1024px_diffusers, SANA, Sana_600M_512px_diffusers, Sana_1600M_512px_diffusers, Sana-fork, Sana_1600M_512px_MultiLing

Prithvi-EO-2.0: A Versatile Multi-Temporal Foundation Model for Earth Observation Applications

arXiv8 repos

arXiv:2412.02732

Prithvi-EO-2.0-NYC-Pluvial, Prithvi-EO-2.0-300M-TL-Sen1Floods11, Prithvi-EO-2.0, Prithvi-EO-2.0-300M, Prithvi-EO-2.0-600M, Prithvi-EO-2.0-300M-TL, Prithvi-EO-2.0-600M-TL, Land-Change-Detection

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

arXiv8 repos

arXiv:2501.12948

awesome-deepseek-prompts, DeepSeek-R1, DeepSeek-R1-0528, DeepSeek-R1-Distill-Llama-70B, DeepSeek-R1-Distill-Qwen-32B, VLAC, DeepSeek-R1-0528-Qwen3-8B, marcello

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

arXiv8 repos

arXiv:2502.04128

ch-tts-llasa-rl-grpo, xcodec2, LLaSA_training, X-Codec-2.0, xcodec2, Llasa_opensource_speech_data_160k_hours_tokenized, xcodec2-hf, tts

R2MED: A Benchmark for Reasoning-Driven Medical Retrieval

arXiv8 repos

arXiv:2505.14558

Bioinformatics, Biology, MedXpertQA-Exam, Medical-Sciences, MedQA-Diag, PMC-Clinical, IIYi-Clinical, PMC-Treatment

Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning

arXiv8 repos

arXiv:2505.16410

Tool-Star, Tool-Star-Qwen-3B, Tool-Star-Qwen-1.5B, Tool-Star-Qwen-0.5B, Multi-Tool-RL-10K, Tool-Star-Qwen-7B, Tool-Star-SFT-54K, Tool-Star

OmniGen2: Towards Instruction-Aligned Multimodal Generation

arXiv8 repos

arXiv:2506.18871

OmniGen2, OmniGen2, OmniGen2-EditScore7B, OmniContext, Macro-OmniGen2, X2I2, ComfyUI-OmniGen2, OmniGen2-EditScore7B-v1.1

GLiNER2: An Efficient Multi-Task Information Extraction System with Schema-Driven Interface

arXiv8 repos

arXiv:2507.18546

gliner2-base-v1, gliner2-large-v1, gliner2-multi-v1, gliner2.5-base-v1, gliner2.5-small-v1, gliner2.5-multi-v1, ArXiv-GLiNER2-Intelligence-Extractor, gliner2.5-multi-v1-onnx

RynnEC: Bringing MLLMs into Embodied World

arXiv8 repos

arXiv:2508.14160

WorldVLA, RynnVLA-002, RynnVLA-001, RynnEC, RynnEC-2B, RynnEC-Bench, RynnEC-7B, RynnVLA-002

ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine

arXiv8 repos

arXiv:2508.14706

ShizhenGPT-7B-Omni, TCM-Pretrain-Data-ShizhenGPT, ShizhenGPT-32B-VL, TCM-Instruction-Tuning-ShizhenGPT, ShizhenGPT, ShizhenGPT-7B-LLM, ShizhenGPT-7B-VL, ShizhenGPT-32B-LLM

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

arXiv8 repos

arXiv:2508.20751

UniGenBench_Leaderboard, UniGenBench, UniGenBench-EvalModel-qwen-72b-v1, UniGenBench-Eval-Images, UniGenBench-EvalModel-qwen3vl-32b-v1, UniGenBench_Leaderboard_Chinese, UniGenBench_Leaderboard_English_Long, UniGenBench_Leaderboard_Chinese_Long

CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding

arXiv8 repos

arXiv:2509.23379

CCD, libra-v1.0-7b, libra-llava-med-v1.5-mistral-7b, libra-llava-rad, libra-maira-2, CheXpert-plus-RRG, IU-Xray-RRG, Medical-CXR-VQA

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

arXiv8 repos

arXiv:2510.10921

FG-CLIP, fg-clip2-base, DCI-CN, BoxClass-CN, DOCCI-CN, LIT-CN, fg-clip2-so400m, fg-clip2-large

Uniform Discrete Diffusion with Metric Path for Video Generation

arXiv8 repos

arXiv:2510.24717

URSA, URSA-0.6B-FSQ320, URSA-0.6B-IBQ1024, URSA-1.7B-IBQ512-UDMGRPO-GenEval, URSA-1.7B-IBQ1024, URSA-1.7B-FSQ320, URSA-1.7B-IBQ512-UDMGRPO-PickScore, URSA-1.7B-IBQ512

OneThinker: All-in-one Reasoning Model for Image and Video

arXiv8 repos

arXiv:2512.03043

Video-R1, OneThinker, OneThinker-8B, OneThinker-SFT-Qwen3-8B, OneThinker-eval, OneThinker-train-data, CoLT-8B, CoLT_Train_Dataset

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

arXiv8 repos

arXiv:2512.22905

JavisDiT, JavisGPT, JavisGPT-v1.0-7B-Instruct, JavisGPT-v0.1-7B-Instruct, JavisInst-Omni, AV-FineTune, JavisUnd-Eval, MM-PreTrain

GLM-5: from Vibe Coding to Agentic Engineering

arXiv8 repos

arXiv:2602.15763

GLM-5.3-Flash, GLM-5.3, GLM-5.3-Flash-GGUF, GLM-5.2, GLM-5, GLM-5.1, GLM-5.2-EXL3-TR3-3.0bpw, GLM-5.2-FP8

Lance: Unified Multimodal Modeling by Multi-Task Synergy

arXiv8 repos

arXiv:2605.18678

lance-mlx, Lance, Lance-3B-bf16, Lance-3B-AWQ-INT4, Lance-3B-8bit, Lance-3B-Video-bf16, Lance, lance-quant

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

arXiv8 repos

arXiv:2607.13125

Boogu-Image-0.1-Edit, Boogu-Image-0.1-Turbo, Boogu-Image-0.1-Base, Boogu-Image, Boogu-Image-0.1-Edit-Turbo, Boogu-Image-0.1-Edit-fp8, Boogu-Image-0.1-Base-fp8, Boogu-Image-0.1-Turbo-fp8

A whole-slide foundation model for digital pathology from real-world data

Nature8 repos

Nature:s41586-024-07441-w

AtlasPatch, PIANO, TRIDENT, CPathPatchFeature, TridentEdited, TITAN, aegis, dpfm_factory

BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

arXiv7 repos

arXiv:1910.13461

llama-2-jax, dalle-mini, lares, GPT_Ranker, ClipSumary, QA-SLM, final-project-level3-nlp-02

wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

arXiv7 repos

arXiv:2006.11477

speech-emotion-recognition, HierSpeechpp, dissertation-project, GigaAM, wav2vec2-base, tmh, gsoc-wav2vec2

8-bit Optimizers via Block-wise Quantization

arXiv7 repos

arXiv:2110.02861

bloom, bloom-optimizer-states, bloom-560m, bloom-1b7, bloom-7b1, bloom-1b1, bloom-3b

Multitask Prompted Training Enables Zero-Shot Task Generalization

arXiv7 repos

arXiv:2110.08207

FLAN, promptsource, T0pp, P3, T0, T0_3B, art

WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing

arXiv7 repos

arXiv:2110.13900

ASR-Transcription-Router, CrisperWhisper, faster_CrisperWhisper, UniSpeech, CrisperWhisper, moshi, wavlm-large

Progressive Distillation for Fast Sampling of Diffusion Models

arXiv7 repos

arXiv:2202.00512

stable-diffusion-x4-upscaler, coreml-stable-diffusion-2-1-base, coreml-stable-diffusion-2-base-palettized, coreml-stable-diffusion-2-1-base-palettized, coreml-stable-diffusion-2-base, stable-diffusion-2-1-base, stable-diffusion-2-1

No Language Left Behind: Scaling Human-Centered Machine Translation

arXiv7 repos

arXiv:2207.04672

nllb-moe-54b, Glot500, WhisperLiveKit, flores200, SONAR, nmtscore, NoLanguageLeftWaiting

LAION-5B: An open large-scale dataset for training next generation image-text models

arXiv7 repos

arXiv:2210.08402

CLIP-convnext_xxlarge-laion2B-s34B-b82K-augreg-soup, OpenFlamingo-9B-vitl-mpt7b, OpenFlamingo-3B-vitl-mpt1b-langinstruct, OpenFlamingo-4B-vitl-rpj3b-langinstruct, OpenFlamingo-3B-vitl-mpt1b, OpenFlamingo-4B-vitl-rpj3b, CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft-soup

Fast Inference from Transformers via Speculative Decoding

arXiv7 repos

arXiv:2211.17192

LMFlow, distil-whisper, LLMSpeculativeSampling, RemoteSpeculativeDecoding, Hierarchical-Speculative-Decoding, GPTFast, NoLanguageLeftWaiting

Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

arXiv7 repos

arXiv:2303.05499

geti-instant-learn, LocateAnything-3B, GroundingDINO, GroundingDINO, Grounding_DINO_demo, DINO, Flow-Inference-Time-Scaling

Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data

arXiv7 repos

arXiv:2304.01196

falcon-40b-instruct, baize-chatbot, baize-v2-13b, baize-v2-7b, baize, vigogne, rulm

Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

arXiv7 repos

arXiv:2304.06939

OpenFlamingo-9B-vitl-mpt7b, OpenFlamingo-3B-vitl-mpt1b-langinstruct, OpenFlamingo-4B-vitl-rpj3b-langinstruct, OpenFlamingo-3B-vitl-mpt1b, OpenFlamingo-4B-vitl-rpj3b, SEED, mmc4

LIMA: Less Is More for Alignment

arXiv7 repos

arXiv:2305.11206

LimaRP, llm-finetuning, InstructionGPT-4, LLaMA2-Accessory, CoT-llama2, open-korean-instructions, ko-lima

Extending Context Window of Large Language Models via Positional Interpolation

arXiv7 repos

arXiv:2306.15595

chatglm2-6b-32k, Open-Sora-Plan, Open-Sora-Plan-v1.3.0, Chinese-LLaMA-Alpaca-2, LongQLoRA, Long-QLORA, article_gpt

OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

arXiv7 repos

arXiv:2306.16527

idefics-80b-instruct, idefics2-8b, Idefics3-8B-Llama3, idefics-9b-instruct, OBELISC, OBELICS, OBELICS

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

arXiv7 repos

arXiv:2307.01952

stable-diffusion-xl-base-1.0, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-xl-base-with-refiner, coreml-stable-diffusion-xl-base-ios, stable-diffusion-xl-refiner-1.0, cagliostro-webui, diffusion-augmentation

SA-Solver: Stochastic Adams Solver for Fast Sampling of Diffusion Models

arXiv7 repos

arXiv:2309.05019

PixArt-alpha, PixArt-XL-2-256x256, PixArt-XL-2-1024-MS, PixArt-XL-2-512x512, PixArt-Sigma-XL-2-512-MS, pixeart, PixArt-alpha

Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs

arXiv7 repos

arXiv:2309.05516

auto-round, DeepSeek-R1-int2-mixed-sym-inc, vllm, vllm-turboquant, vllm-old, vllm_amd_sleep, vllm

AnglE-optimized Text Embeddings

arXiv7 repos

arXiv:2309.12871

UAE-Large-V1, AnglE, UAE-Code-Large-V1, pubmed-angle-base-en, angle-llama-13b-nli, pubmed-angle-large-en, angle-llama-7b-nli-v2

Distil-Whisper: Robust Knowledge Distillation via Large-Scale Pseudo Labelling

arXiv7 repos

arXiv:2311.00430

kotoba-whisper-v2.0, kotoba-whisper, distil-whisper, distil-large-v2, kotoba-whisper-v1.0, distil-medium.en, distil-small.en

Nomic Embed: Training a Reproducible Long Context Text Embedder

arXiv7 repos

arXiv:2402.01613

contrastors, nomic-embed-text-v1.5, nomic-embed-text-v1-ablated, nomic-embed-text-v1-unsupervised, contrastors, ColBERT-Zero, modernbert-embed-base

Natural language guidance of high-fidelity text-to-speech with synthetic annotations

arXiv7 repos

arXiv:2402.01912

parler-tts-mini-v1, parler-tts, dataspeech, parler-tts-large-v1, libritts-r-filtered-speaker-descriptions, mls-eng-speaker-descriptions, parler-tts-vietnamese-v1-stage2

World Model on Million-Length Video And Language With Blockwise RingAttention

arXiv7 repos

arXiv:2402.08268

LWM, Llama-3-8B-Instruct-Gradient-1048k, Llama-3-70B-Instruct-Gradient-262k, Llama-3-8B-Instruct-262k, Llama-3-70B-Instruct-Gradient-1048k, Llama-3-8B-Instruct-Gradient-4194k, lwm

IEPile: Unearthing Large-Scale Schema-Based Information Extraction Corpus

arXiv7 repos

arXiv:2402.14710

OneKE, IEPile, iepie, llama2-13b-iepile-lora, OneKE, baichuan2-13b-iepile-lora, iepile

Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

arXiv7 repos

arXiv:2403.03163

Design2Code, Design2Code_human_eval_pairwise, Design2Code_human_eval_reference_vs_gpt4v, Design2Code-hf, Design2Code-18B-v0, Design2Code, Design2Code-HARD

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

arXiv7 repos

arXiv:2404.01258

train_video_and_instruction, LLaVA-Hound-DPO, LLaVA-Hound-DPO, test_video_and_instruction, LLaVA-Hound-SFT, MixEval-Archon, MixEval

MagicTime: Time-lapse Video Generation Models as Metamorphic Simulators

arXiv7 repos

arXiv:2404.05014

MagicTime, MagicTime, ChronoMagic, MagicTime, ChronoMagic-Bench, ConsisID, OpenS2V-Nexus

MANTIS: Interleaved Multi-Image Instruction Tuning

arXiv7 repos

arXiv:2405.01483

Mantis, Mantis-8B-Idefics2, Mantis-8B-clip-llama3, Mantis-8B-siglip-llama3, Mantis-Instruct, Mantis-8B-Fuyu, SpaceMantis

Arctic-Embed: Scalable, Efficient, and Accurate Text Embedding Models

arXiv7 repos

arXiv:2405.05374

cholesky_encoder, snowflake-arctic-embed-m, snowflake-arctic-embed-xs, snowflake-arctic-embed-s, snowflake-arctic-embed-l, snowflake-arctic-embed-m-long, arctic-embed

USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

arXiv7 repos

arXiv:2405.07719

xDiT, mochi-xdit, HunyuanVideo, HunyuanVideo, HunyuanVideo-I2V, HunyuanVideo-I2V, HunyuanWorld-Voyager

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

arXiv7 repos

arXiv:2405.08748

HunyuanDiT-v1.2-Diffusers, HunyuanDiT, HunyuanDiT-v1.1, HunyuanDiT-v1.2-Diffusers-Distilled, HunyuanDiT, HunyuanDiT-v1.2, HunyuanDiT

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

arXiv7 repos

arXiv:2406.09952

CLIP_COCO, CLIP_TROHN-Img, TROHN-Text, TROHN-Img, CLIP_TROHN-Text, CLIP_Detector, CLIP_TROHN-Img_Detector

ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation

arXiv7 repos

arXiv:2406.18522

MagicTime, ChronoMagic-Bench, ChronoMagic-Pro, ChronoMagic-ProH, ChronoMagic-Bench, ConsisID, OpenS2V-Nexus

Embedding And Clustering Your Data Can Improve Contrastive Pretraining

arXiv7 repos

arXiv:2407.18887

cholesky_encoder, snowflake-arctic-embed-m, snowflake-arctic-embed-xs, snowflake-arctic-embed-s, snowflake-arctic-embed-m-long, snowflake-arctic-embed-l, arctic-embed

Gemma 2: Improving Open Language Models at a Practical Size

arXiv7 repos

arXiv:2408.00118

modded-nanogpt, Gemma-2-Llama-Swallow-9b-pt-v0.1, Gemma-2-Llama-Swallow-2b-pt-v0.1, Gemma-2-Llama-Swallow-27b-it-v0.1, Gemma-2-Llama-Swallow-27b-pt-v0.1, Gemma-2-Llama-Swallow-2b-it-v0.1, Gemma-2-Llama-Swallow-9b-it-v0.1

Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

arXiv7 repos

arXiv:2408.02657

Lumina-T2X, Lumina-mGPT, Lumina-mGPT-7B-512, Lumina-mGPT-7B-768, Lumina-mGPT-7B-768-Omni, Lumina-mGPT-7B-1024, Lumina-mGPT-34B-512

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

arXiv7 repos

arXiv:2409.17146

Molmo-7B-D-0924, molmo, Molmo-7B-O-0924, MolmoE-1B-0924, pixmo-docs, CoSyn-400K, CoSyn-point

Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach

arXiv7 repos

arXiv:2410.03160

PusaV1_training, Pusa-VidGen, PusaV0.5_Training, Pusa-Wan2.2-V1, PusaV1, Pusa-V0.5, Mochi-Full-Finetuner

arXiv:2410.09724

arXiv7 repos

arXiv:2410.09724

Reward-Calibration, mistral-7b-ppo-c-hermes, llama3-8b-crm-final-v0.1, llama3-8b-final-ppo-m-v0.3, mistral-7b-ppo-m-hermes, mistral-7b-hermes-crm-skywork, llama3-8b-final-ppo-c-v0.3

How to Evaluate Reward Models for RLHF

arXiv7 repos

arXiv:2410.14872

PPE, PPE-Human-Preference-V1, PPE-MMLU-Pro-Best-of-K, PPE-MATH-Best-of-K, PPE-GPQA-Best-of-K, PPE-IFEval-Best-of-K, PPE-MBPP-Plus-Best-of-K

Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances

arXiv7 repos

arXiv:2410.18775

VINE, VINE-R-Dec, VINE-R-Enc, VINE-B-Enc, VINE-B-Dec, W-Bench, WMCopier

LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch

arXiv7 repos

arXiv:2411.11171

LLaMmlein, LLaMmlein_120M_prerelease, LLaMmlein_120M, LLaMmlein_7B, LLaMmlein_1B, LLaMmlein_1B_prerelease, LLaMmlein-Dataset

RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models

arXiv7 repos

arXiv:2412.07679

RADIO, C-RADIOv3-L, C-RADIOv3-B, C-RADIOv4-SO400M, C-RADIOv4-H, C-RADIOv3-g, C-RADIOv3-H

Multimodal Latent Language Modeling with Next-Token Diffusion

arXiv7 repos

arXiv:2412.08635

VibeVoice, VibeVoice-1.5B, VibeVoice-Realtime-0.5B, VibeVoice-7B, VibeVoice, VibeVoice-Large, VibeVoice

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

arXiv7 repos

arXiv:2412.13663

indic-modernBERT, ModernBERT-large, DiagnosisCoding, LightOnOCR-2-1B, ModernBERT, vllm-factory, LightOnOCR-2-1B-base

VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

arXiv7 repos

arXiv:2501.00574

VideoChat-Flash-Qwen2_5-7B-1M_res224, VideoChat-Flash-Qwen2-7B_res448, VideoChat-Flash-Training-Data, InternVL_2_5_HiCo_R16, VideoChat-Flash-Qwen2_5-2B_res448, VideoChat-Flash-Qwen2_5-7B_InternVideo2-1B, VideoChat-Flash-Qwen2-7B_res224

TripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models

arXiv7 repos

arXiv:2502.06608

TripoSG-scribble, TripoSG, TripoSG, 2D23D, tripoSG-pipeline, avera, TripoSG-fc5c409

Magma: A Foundation Model for Multimodal AI Agents

arXiv7 repos

arXiv:2502.13130

Magma, Magma-Mind2Web-SoM, Magma-AITW-SoM, Magma-8B, Magma-OXE-ToM, Magma-Video-ToM, Magma-820K

LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification

arXiv7 repos

arXiv:2502.17421

longspec-longchat-13b-16k, longspec-Llama-3-8B-Instruct-262k, longspec-longchat-7b-v1.5-32k, longspec-QwQ-32B-Preview, longspec-vicuna-7b-v1.5-16k, longspec-vicuna-13b-v1.5-16k, longspec-data

LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation

arXiv7 repos

arXiv:2502.20583

LiteASR, lite-whisper-large-v3-fast, lite-whisper-large-v3-turbo, lite-whisper-large-v3-turbo-fast, lite-whisper-large-v3-acc, lite-whisper-large-v3, lite-whisper-large-v3-turbo-acc

Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models

arXiv7 repos

arXiv:2503.01774

difix_ref, harmonizer, Difix3D, difix, Fixer, Difix3d-3dgs-demo, difix3d

VACE: All-in-One Video Creation and Editing

arXiv7 repos

arXiv:2503.07598

Wan2.1-VACE-14B, Wan2.1, Wan2.1-VACE-1.3B, VACE-LTX-Video-0.9, VACE-Wan2.1-1.3B-Preview, VACE-Benchmark, VACE-Annotators

OmniSVG: A Unified Scalable Vector Graphics Generation Model

arXiv7 repos

arXiv:2504.06263

OmniSVG, MMSVG-Illustration, OmniSVG1.1_4B, OmniSVG1.1_8B, OmniSVG, OmniSVG-train, omnisvg-train

Step1X-Edit: A Practical Framework for General Image Editing

arXiv7 repos

arXiv:2504.17761

GEdit-Bench, Step1X-Edit, Step1X-Edit-v1p2-preview, Step1X-Edit-v1p2, Step1X-Edit-v1p1-diffusers, Step1X-Edit, Step1X-Edit

TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

arXiv7 repos

arXiv:2504.19874

turbovec, mlx-vlm, turboquant_plus, rotorquant, quant.cpp, turboquant-vllm, sqlite-vector

General-Reasoner: Advancing LLM Reasoning Across All Domains

arXiv7 repos

arXiv:2505.14652

General-Reasoner, General-Reasoner-Qwen2.5-14B, General-Reasoner-Qwen3-14B, WebInstruct-verified, General-Reasoner-Qwen2.5-7B, General-Reasoner-Qwen3-4B, RLPR-Evaluation

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

arXiv7 repos

arXiv:2505.17012

SpaceQwen2.5-VL-3B-Instruct, SpaceOm, SpaceThinker-Qwen2.5VL-3B, SpatialScore, SpatialScore, SpatialCorpus, SpatialScore

HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

arXiv7 repos

arXiv:2505.22705

HiDream-E1, HiDream-E1-1, HiDream-I1-Full, HiDream-E1-Full, HiDream-I1, HiDream-I1-Dev, HiDream-I1-Fast

Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

arXiv7 repos

arXiv:2507.05257

Mem-alpha, memalphaljx, knowl, MemoryAgentBench, MemoryAgentBench, tmp-mem-alpha, agentic-memory

VibeVoice Technical Report

arXiv7 repos

arXiv:2508.19205

VibeVoice, VibeVoice-1.5B, VibeVoice-Realtime-0.5B, VibeVoice-7B, VibeVoice, VibeVoice-Large, VibeVoice

GUIrilla: A Scalable Framework for Automated Desktop UI Exploration

arXiv7 repos

arXiv:2510.16051

GUIrilla, GUIrilla-Trees, GUIrilla-Gold, GUIrilla-See-3B, GUIrilla-See-7B, GUIrilla-Task, GUIrilla-See-0.7B

Steering Evaluation-Aware Language Models to Act Like They Are Deployed

arXiv7 repos

arXiv:2510.20487

steering-eval-awareness-public, large-finetune, steering-eval-awareness, eval-evasion, wood_v2_sftr4_filt, steering, steering-eval-awareness-public-v2

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

arXiv7 repos

arXiv:2511.00088

Alpamayo-R1-10B, alpamayo, alpamayo_, alpamayo-recipes, alphamayo_VLA_test_webui_opti., alpamayo1.5, RiskWorld

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

arXiv7 repos

arXiv:2512.07348

MICo-150K, OmniGen2-MICo, Qwen-Image-MICo, MICo-150K, BLIP3o-Next-MICo, BAGEL-MICo, Lumina-DiMOO-MICo

LLaDA2.0: Scaling Up Diffusion Language Models to 100B

arXiv7 repos

arXiv:2512.15745

LLaDA2.0-mini, LLaDA2.0-flash-preview, LLaDA2.0-flash, LLaDA2.0-mini-preview, LLaDA2.0-flash-CAP, LLaDA2.0-mini-CAP, dllm

Creating a Hybrid Rule and Neural Network Based Semantic Tagger using Silver Standard Data: the PyMUSAS framework for Multilingual Semantic Annotation

arXiv7 repos

arXiv:2601.09648

experimental-wsd, English-USAS-Mosaico, PyMUSAS-Neural-English-Small-BEM, PyMUSAS-Neural-English-Base-BEM, PyMUSAS-Neural-Multilingual-Small-BEM, PyMUSAS-Neural-Multilingual-Base-BEM, USAS-WSD

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

arXiv7 repos

arXiv:2601.10611

molmo2, MolmoWeb-8B, MolmoWeb-Pretrained-8B, MolmoWeb-Pretrained-4B, MolmoWeb-4B, MolmoWeb-8B-Native, MolmoWeb-4B-Native

Innovator-VL: A Multimodal Large Language Model for Scientific Discovery

arXiv7 repos

arXiv:2601.19325

Innovator-VL, Innovator-VL-8B-Instruct, Innovator-VL-8B-Thinking, Innovator-VL-RL-172K, Innovator-VL-Instruct-46M, PreMidTrainVL-Qwen3Dense, PreMidTrainVL

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

arXiv7 repos

arXiv:2601.22709

LLaVA-1.5-7B-GRACE-W4G128, Qwen3-VL-2B-GRACE-W4G128-AWQ, Qwen3-VL-2B-GRACE-W4G128, Qwen3-VL-2B-GRACE-BF16, LLaVA-1.5-7B-GRACE-W4G128-AWQ, Qwen3-VL-2B-GRACE-W8G128, GRACE-VLM

ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation

arXiv7 repos

arXiv:2602.00744

acestep-v15-xl-base, Ace-Step1.5, ACE-Step-v1.5-chinese-new-year-LoRA, acestep-v15-xl-sft, acestep-v15-xl-turbo, acestep-v15-base, acestep-v15-sft

MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models

arXiv7 repos

arXiv:2602.10934

MOSS-TTS-Nano, MOSS-Audio-Tokenizer-v2, MOSS-Audio-Tokenizer-Nano, MOSS-Audio-Tokenizer, MOSS-TTS-Realtime, MOSS-Audio-Tokenizer-Nano-ONNX, MOSS-TTS-Nano-100M-ONNX

The Million-Label NER: Breaking Scale Barriers with GLiNER bi-encoder

arXiv7 repos

arXiv:2602.18487

gliner-bi-base-v2.0, gliner-linker-large-v1.0, gliner-bi-small-v2.0, gliner-bi-large-v2.0, gliner-bi-edge-v2.0, gliner-linker-base-v1.0, gliner-linker-rerank-v1.0

arXiv:2602.20903

arXiv7 repos

arXiv:2602.20903

TextPecker, TextPecker-8B-InternVL3, SD3.5M-TextPecker-SQPA, Flux.1-dev-TextPecker-SQPA, TextPecker-1.5M, QwenImage-TextPecker-SQPA, TextPecker-8B-Qwen3VL

FireRedASR2S: A State-of-the-Art Industrial-Grade All-in-One Automatic Speech Recognition System

arXiv7 repos

arXiv:2603.10420

FireRedLID, FireRedPunc, FireRedASR2-LLM, FireRedVAD, FireRedASR2S, FireRedVAD, FireRedASR2S

A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens

arXiv7 repos

arXiv:2604.04913

deltatok, deltatok-kinetics, seg-head-vspw, depth-head-kitti, rgb-head-imagenet, seg-head-cityscapes, deltaworld-kinetics

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

arXiv7 repos

arXiv:2606.25369

Joyo-Kanji-Yomi-Benchmark, joyo-kanji-yomi-benchmark, kana-whisper, JoyoKanji-Yomi-Benchmark, sarashina2.2-tts, Irodori-TTS-v4-Small, sarashina2.2-tts

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

arXiv7 repos

arXiv:2607.17423

TimeLens2-4B, TimeLens2-2B, TimeLens2-8B, TimeLens2-2B-SFT, TimeLens2-4B-SFT, TimeLens2-8B-SFT, TimeLens2-93K

A visual–language foundation model for pathology image analysis using medical Twitter

Nature7 repos

Nature:s41591-023-02504-3

PIANO, AtlasPatch, KEEP, KEEP, VLSA, PathPT, Histopathology_Benchmark

BERTScore: Evaluating Text Generation with BERT

arXiv6 repos

arXiv:1904.09675

YiVal, gec-metrics, bert_score, lares, mslr-shared-task, KoBERTScore

CodeSearchNet Challenge: Evaluating the State of Semantic Code Search

arXiv6 repos

arXiv:1909.09436

CodeBERTa-language-id, CodeBERTa-small-v1, CodeRL, codet5-large, codet5-large-ntp-py, CodeXGLUE

Root Mean Square Layer Normalization

arXiv6 repos

arXiv:1910.07467

idefics-80b-instruct, tiny-vllm, gemma-scope-2b-pt-transcoders, transformer-tricks, idefics-9b-instruct, llama2.zig

Unsupervised Cross-lingual Representation Learning at Scale

arXiv6 repos

arXiv:1911.02116

Multilingual-MiniLM-L12-H384, xlm-roberta-base, xlm-roberta-large, XLM, tydiqa-primary-task-xlm-roberta-large, caption

YOLOv4: Optimal Speed and Accuracy of Object Detection

arXiv6 repos

arXiv:2004.10934

darknet, darknetcv, tensorflow-yolov4-tflite, Complex-YOLOv4-Pytorch, darknet, pytorch-YOLOv4

Denoising Diffusion Probabilistic Models

arXiv6 repos

arXiv:2006.11239

Diff-Pruning, ddpm-ema-bedroom-256, ddpm-cifar10-32, deepfake_multiLID, DDPM_vs_DDIM, smalldiffusion

Language-agnostic BERT Sentence Embedding

arXiv6 repos

arXiv:2007.01852

Multilingual-CLIP, xnli_bn, squad_bn, russe_detox_2022, russe_detox_2022, MERA

Beyond English-Centric Multilingual Machine Translation

arXiv6 repos

arXiv:2010.11125

m2m100_1.2B, EasyNMT, m2m100_418M, Easy-Translate, ContraDecode, m2m100-12B-avg-5-ckpt

SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

arXiv6 repos

arXiv:2108.01073

stable-diffusion, stable-diffusion-xl-base-1.0, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-xl-base-with-refiner, coreml-stable-diffusion-xl-base-ios, stable-diffusion-xl-refiner-1.0

MoEfication: Transformer Feed-forward Layers are Mixtures of Experts

arXiv6 repos

arXiv:2110.01786

ReluLLaMA-7B, ReluLLaMA-13B, ReluLLaMA-70B, Bamboo-base-v0_1, ReluFalcon-40B, Bamboo-DPO-v0_1

Masked Autoencoders Are Scalable Vision Learners

arXiv6 repos

arXiv:2111.06377

mae, hls-foundation-os, TiViT, vit-mae-large, vit-mae-huge, vit-mae-base

LiT: Zero-Shot Transfer with Locked-image text Tuning

arXiv6 repos

arXiv:2111.07991

vision_transformer, CLIP_benchmark, contrastors, nomic-embed-vision-v1, nomic-embed-vision-v1.5, contrastors

ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

arXiv6 repos

arXiv:2112.01488

finetuning-bert-for-IR, plaidrepro, colbertv2.0, ColBERT, jina-colbert-v1-en, KolBERT

Datasheet for the Pile

arXiv6 repos

arXiv:2201.07311

minipile, pythia-12b, pythia-1.4b, pythia-410m, pythia-2.8b, diff-codegen-6b-v2

From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective

arXiv6 repos

arXiv:2205.04733

splade-cocondenser-ensembledistil, splade, splade-cocondenser-selfdistil, splade-ecommerce-esci, Splade_PP_en_v1, SPLADERunner

Generative Language Models for Paragraph-Level Question Generation

arXiv6 repos

arXiv:2210.03992

tweetnlp, qg_tweetqa, qag_tweetqa, t5-small-tweetqa-qa, t5-base-tweetqa-qag, t5-large-squad

SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control

arXiv6 repos

arXiv:2210.17432

ssd-lm, mdlm, bd3lm, bd3lms, BDM, BD_DNA

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

arXiv6 repos

arXiv:2211.01324

stable-diffusion-xl-base-1.0, paint-with-words-sd, coreml-stable-diffusion-xl-base, coreml-stable-diffusion-xl-base-with-refiner, coreml-stable-diffusion-xl-base-ios, stable-diffusion-xl-refiner-1.0

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

arXiv6 repos

arXiv:2211.01335

Chinese-CLIP, chinese-clip-vit-huge-patch14, chinese-clip-vit-base-patch16, chinese-clip-vit-large-patch14, chinese-clip-vit-large-patch14-336px, BDM1.0

Text Embeddings by Weakly-Supervised Contrastive Pre-training

arXiv6 repos

arXiv:2212.03533

ROOT-RAG, e5-large-unsupervised, e5-small-unsupervised, e5-base-unsupervised, e5-mistral-7b-instruct, e5-large-v2

Benchmarking Self-Supervised Learning on Diverse Pathology Datasets

arXiv6 repos

arXiv:2212.04690

resnet50.lunit_mocov2, resnet50.lunit_swav, vit_small_patch8_224.lunit_dino, resnet50.lunit_bt, vit_small_patch16_224.lunit_dino, benchmark-ssl-pathology

SantaCoder: don't reach for the stars!

arXiv6 repos

arXiv:2301.03988

sven_modified, sven, bigcode-evaluation-harness, santacoder-fim-task, MultiPL-E, MultiPL-E

BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

arXiv6 repos

arXiv:2303.00915

MediMeta-C, RobustMedCLIP, RobustMedCLIP, BiomedCoOp, libra-llava-rad, llava-rad

Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

arXiv6 repos

arXiv:2304.13705

aloha_sim_transfer_cube_human, physical-ai-studio, act_gyoza_cherry3_v2, act_gyoza_shiratama_v2, act_gyoza_pickplace_synth, act_gyoza_grape_v2

StarCoder: may the source be with you!

arXiv6 repos

arXiv:2305.06161

MultiPL-E, magicoder, Magicoder-S-DS-6.7B, Magicoder-S-CL-7B, Magicoder-DS-6.7B, Magicoder-CL-7B

InstructIE: A Bilingual Instruction-based Information Extraction Dataset

arXiv6 repos

arXiv:2305.11527

IEPile, InstructIE, llama2-13b-iepile-lora, baichuan2-13b-iepile-lora, EasyInstruct, KnowLM

Simple and Controllable Music Generation

arXiv6 repos

arXiv:2306.05284

optimized-parler-tts, whisperspeech, musicgen-medium, musicgen-melody, musicgen-melody-large, encodec_32khz

ModelScope Text-to-Video Technical Report

arXiv6 repos

arXiv:2308.06571

Text-To-Video-Finetuning, text-to-video-synthesis-colab, text-to-video-ms-1.7b, text-to-video-ms-1.7b, mcm, i2vgen-xl

A Family of Pretrained Transformer Language Models for Russian

arXiv6 repos

arXiv:2309.10931

rugpt3large_based_on_gpt2, rugpt3medium_based_on_gpt2, rugpt3small_based_on_gpt2, rugpt3large_based_on_gpt2, rugpt3medium_based_on_gpt2, rugpt3small_based_on_gpt2

UltraFeedback: Boosting Language Models with Scaled AI Feedback

arXiv6 repos

arXiv:2310.01377

zephyr-7b-alpha, zephyr-7b-beta, ultrafeedback_binarized, Llama-3-8B-Magpie-Align-v0.1, Llama-3-8B-Magpie-Align-v0.3, Llama-3-8B-Magpie-Align-v0.2

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

arXiv6 repos

arXiv:2310.01852

LanguageBind, LanguageBind, VIDAL-Depth-Thermal, MoE-LLaVA, Video-LLaVA, LLMBind

Ring Attention with Blockwise Transformers for Near-Infinite Context

arXiv6 repos

arXiv:2310.01889

MindSpeed-MM, Wan2.1-VACE-14B, EasyContext, Wan2.1, Wan2.1-VACE-1.3B, Wan2.1-FLF2V-14B-720P

Prometheus: Inducing Fine-grained Evaluation Capability in Language Models

arXiv6 repos

arXiv:2310.08491

prometheus, Feedback-Collection, prometheus-7b-v2.0, prometheus-7b-v1.0-fp16, prometheus-13b-v1.0-fp16, open-korean-instructions

DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

arXiv6 repos

arXiv:2310.12190

DynamiCrafter_1024, DynamiCrafter, DynamiCrafter_512, DynamiCrafter, DynamiCrafter_512_Interp, DynamiCrafter_pruned

Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

arXiv6 repos

arXiv:2310.16834

mdlm, bd3lm, bd3lms, BDM, BD_DNA, sedd-noeos-owt

Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch

arXiv6 repos

arXiv:2311.03099

Twin-Merging, model_merging, Mario, supermario_v2, MergeLM, ComfyUI-LoRA-Optimizer

An Efficient Self-Supervised Cross-View Training For Sentence Embedding

arXiv6 repos

arXiv:2311.03228

SCT-model-phayathaibert, SCT-KD-model-phayathaibert, SCT-model-XLMR, SCT-KD-model-wangchanberta, SCT-model-wangchanberta, SCT-KD-model-XLMR

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

arXiv6 repos

arXiv:2311.07919

LoRA-MCL, Qwen2-Audio-7B, Qwen-Audio-Chat, Qwen-Audio, Speech-IFEval, Qwen2-Audio-7B-Instruct

SegVol: Universal and Interactive Volumetric Medical Image Segmentation

arXiv6 repos

arXiv:2311.13385

M3D-Seg, SegVol, M3D-RefSeg, SegVol, DL2-group5-med-seg, adapt_med_seg

LMDrive: Closed-Loop End-to-End Driving with Large Language Models

arXiv6 repos

arXiv:2312.07488

LMDrive, LMDrive, LMDrive-vicuna-v1.5-7b-v1.0, LMDrive-llava-v1.5-7b-v1.0, LMDrive-vision-encoder-r50-v1.0, LMDrive-llama-7b-v1.0

VILA: On Pre-training for Visual Language Models

arXiv6 repos

arXiv:2312.07533

llm-awq, VILA, awq-embed, awq4nvomni, llm-awq, VILA-2.7b

SOLAR 10.7B: Scaling Large Language Models with Simple yet Effective Depth Up-Scaling

arXiv6 repos

arXiv:2312.15166

SOLAR-10.7B-v1.0, SOLAR-10.7B-Instruct-v1.0, yarn, LDCC-SOLAR-10.7B, iDUS, corningQA-solar-10.7b-v1.0

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

arXiv6 repos

arXiv:2401.12168

OpenSpaces, SpaceQwen2.5-VL-3B-Instruct, SpaceLLaVA, SpaceOm, SpaceThinker-Qwen2.5VL-3B, SpaceMantis

Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

arXiv6 repos

arXiv:2402.11746

red-instruct, CategoricalHarmfulQA, starling-7B, HarmfulQA, resta, CategoricalHarmfulQ

ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

arXiv6 repos

arXiv:2402.13516

prosparse-llama-2-13b, prosparse-llama-2-7b, prosparse-llama-2-7b-gguf, prosparse-llama-2-13b-gguf, prosparse-llama-2-13b-predictor, prosparse-llama-2-7b-predictor

Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation

arXiv6 repos

arXiv:2403.08002

libra-llava-rad, llava-rad, llava-rad, chexprompt, llava-dino, Explanability_in_VLM

Layer-Condensed KV Cache for Efficient Inference of Large Language Models

arXiv6 repos

arXiv:2405.10637

StepDeepResearch, LCKV, tinyllama-lckv-w10-100b, tinyllama-lckv-w2-100b, tinyllama-lckv-w10-ft-250b, tinyllama-lckv-w2-ft-100b

SimPO: Simple Preference Optimization with a Reference-Free Reward

arXiv6 repos

arXiv:2405.14734

SimPO, CPO_SIMPO, Reinforcement-Learning-Full-Pipeline, Llama-3-8B-Magpie-Align-v0.1, Llama-3-8B-Magpie-Align-v0.3, Llama-3-8B-Magpie-Align-v0.2

arXiv:2405.17976

arXiv6 repos

arXiv:2405.17976

Yuan2-M32, Yuan2-M32-gguf-int4, Yuan2-M32-hf-int8, Yuan2-M32-hf, Yuan2-M32-gguf, Yuan2-M32-hf-int4

LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

arXiv6 repos

arXiv:2406.05113

LlavaGuard, LlavaGuard-v1.2-0.5B-OV, LlavaGuard-v1.2-0.5B-OV-hf, LlavaGuard-v1.2-7B-OV, LlavaGuard-v1.2-7B-OV-hf, compagent

Depth Anything V2

arXiv6 repos

arXiv:2406.09414

Depth-Anything-V2-Large, Depth-Anything-V2-Small, Depth-Anything-V2-Base, coreml-depth-anything-v2-small, svraster, Depth-Anything-V2-Small-hf

Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

arXiv6 repos

arXiv:2406.12845

Magpie-Pro-DPO-100K-v0.1, Magpie-Air-DPO-100K-v0.1, Magpie-Llama-3.1-Pro-DPO-100K-v0.1, Llama-3-8B-Magpie-Align-v0.1, Llama-3-8B-Magpie-Align-v0.3, Llama-3-8B-Magpie-Align-v0.2

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

arXiv6 repos

arXiv:2406.17557

modded-nanogpt, fineweb, fineweb-edu, fineweb-2, Automodel, ml_filter

HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale

arXiv6 repos

arXiv:2406.19280

HuatuoGPT-Vision-34B, HuatuoGPT-Vision-7B, PubMedVision, Medical_Multimodal_Evaluation_Data, HuatuoGPT-Vision-7B-Qwen2.5VL, MedAI-project

MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention

arXiv6 repos

arXiv:2407.02490

LLMLingua, Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, RetrievalAttention, Block-Sparse-Attention

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

arXiv6 repos

arXiv:2407.11691

VLMEvalKit, Investigating_MultiEncoder_Redundancy, VLMEvalKit, ChatVLA_public, sa2va_eval, EAI_VLMEvalKit

OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection

arXiv6 repos

arXiv:2407.16237

origen_dataset_debug, OriGen, origen_dataset_instruction, OriGen_Fix, OriGen, origen_dataset_description

Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology

arXiv6 repos

arXiv:2408.00738

PIANO, Midnight, SEAL, HistAug, histaug-virchow2, dpfm_factory

miniCTX: Neural Theorem Proving with (Long-)Contexts

arXiv6 repos

arXiv:2408.03350

ntp-mathlib-instruct-st, miniCTX-v2, ntp-mathlib-st-deepseek-coder-1.3b, ntp-mathlib-context-deepseek-coder-1.3b, ntp-mathlib-instruct-ctx, ntp-mathlib

SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning

arXiv6 repos

arXiv:2408.05517

swift, ms-swift, ms-swift, vmopd, ms, dense-retention-rl

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

arXiv6 repos

arXiv:2408.15998

Eagle-X5-34B-Plus, Eagle-X5-13B, Eagle-X4-13B-Plus, Eagle-X4-8B-Plus, Eagle-X5-13B-Chat, Eagle-X5-7B

Diabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management

arXiv6 repos

arXiv:2409.13191

Diabetica, Diabetica-1.5B, Diabetica-SFT, Diabetica-o1, Diabetica-o1-SFT, Diabetica-7B

Moshi: a speech-text foundation model for real-time dialogue

arXiv6 repos

arXiv:2410.00037

minimind-o, tts-1.6b-en_fr, moshi-rag, personaplex, moshi, eval-moshi

Aria: An Open Multimodal Native Mixture-of-Experts Model

arXiv6 repos

arXiv:2410.05993

Aria-Chat, MMLongBench-Doc, Aria, Aria, Aria-Base-64K, Aria-Base-8K

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

arXiv6 repos

arXiv:2411.15296

lmms-eval, Thyme, Investigating_MultiEncoder_Redundancy, VLMEvalKit, Video-MME, UniG2U

Identity-Preserving Text-to-Video Generation by Frequency Decomposition

arXiv6 repos

arXiv:2411.17440

MagicTime, ChronoMagic-Bench, ConsisID, ConsisID-preview, ConsisID-preview-Data, OpenS2V-Nexus

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis

arXiv6 repos

arXiv:2411.19378

libra-v1.0-7b, libra-llava-med-v1.5-mistral-7b, libra-llava-rad, libra-v1.0-3b, libra-maira-2, Libra

Multimodal Whole Slide Foundation Model for Pathology

arXiv6 repos

arXiv:2411.19666

TRIDENT, VLSA, HistAug, TridentEdited, histaug-conch_v15, aegis

RedStone: Curating General, Code, Math, and QA Data for Large Language Models

arXiv6 repos

arXiv:2412.03398

RedStone, RedStone, RedStone-QA-mcq, RedStone-Code-python, RedStone-Math, RedStone-QA-oq

LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

arXiv6 repos

arXiv:2412.09262

LatentSync, LatentSync, LatentSync-1.6, LatentSync1.5-mac, LatentSync-1.5, screencastgen

CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion Models

arXiv6 repos

arXiv:2412.13195

CoMPaSS, CoMPaSS-FLUX.1, CoMPaSS-SD2.1, CoMPaSS-SD1.4, CoMPaSS-SD1.5, CoMPaSS-FLUX.1-dev-ComfyUI

Qwen2.5 Technical Report

arXiv6 repos

arXiv:2412.15115

QwQ-32B, d3LLM, Baichuan-Omni-1.5, smoltalk2, Baichuan-Audio, FLUXSynID

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

arXiv6 repos

arXiv:2501.07542

visual-thinker, AlphaMaze-v0.2-1.5B, Mirage, MVoT, UniWM, GoViG

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models

arXiv6 repos

arXiv:2501.14818

Eagle, Eagle2-2B, GR00T-N1.5-3B, Eagle2-1B, Eagle2-9B, llama-nemotron-embed-vl-1b-v2

Qwen2.5-1M Technical Report

arXiv6 repos

arXiv:2501.15383

Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen2.5-7B-Instruct-1M, Qwen2.5-14B-Instruct-1M, Qwen3-Next-80B-A3B-Instruct

s1: Simple test-time scaling

arXiv6 repos

arXiv:2501.19393

unlazy, s1, s1.1-32B, s1K-step-conditional-control-old, step-conditional-control-old, crrrocq

PolarQuant: Quantizing KV Caches with Polar Transformation

arXiv6 repos

arXiv:2502.02617

turboquant_plus, quant.cpp, Qwen3.5-9B-PolarQuant-MLX-4bit, Qwen3.5-9B-PolarEngine-v4, Qwen3.5-9B-PolarQuant-Q5, polarengine-vllm

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

arXiv6 repos

arXiv:2502.02737

smollm, finemath, smoltalk, SmolLM2-360M, SmolLM2-1.7B, smollm2-1.7b-instruct

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

arXiv6 repos

arXiv:2502.05512

IndexTTS-2, parrots, Index-TTS, index-tts, indexTTS2, IndexTTS

Large Language Diffusion Models

arXiv6 repos

arXiv:2502.09992

dLLM-cache, d3LLM, iLLaDA-8B-Instruct, LLaDA-MoE-7B-A1B-Base, dllm, SDAR

Qwen2.5-VL Technical Report

arXiv6 repos

arXiv:2502.13923

Qwen2.5-VL, Qwen3-VL, Qwen2-VL, ScreenSpot-Pro-GUI-Grounding, CharXiv, Fara_Test

ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation

arXiv6 repos

arXiv:2502.15543

ParamMute, ParamMute-8B-KTO, ParamMute-7B, ParamMute-8B-SFT, CoConflictQA, PIP-KAG

YuE: Scaling Open Foundation Models for Long-Form Music Generation

arXiv6 repos

arXiv:2503.08638

YuE2-3B, YuE, SheetSage2, YuE2-Vae-legacy, YuE2-Vae, WildSongBench

DAPO: An Open-Source LLM Reinforcement Learning System at Scale

arXiv6 repos

arXiv:2503.14476

ReTool, tunix, DAPO, xtuner, MM-EUREKA, VL-Rethinker

m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models

arXiv6 repos

arXiv:2504.00869

m1, m1-7B-23K, m1-7B-1K, m1-32B-1K, m23k-tokenized, m1k-tokenized

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

arXiv6 repos

arXiv:2504.05599

Skywork-R1V, Skywork-R1V2-38B, Skywork-R1V-38B-AWQ, Skywork-R1V-38B, Skywork-R1V2-38B-AWQ, sky-rv1

arXiv:2504.07962

arXiv6 repos

arXiv:2504.07962

GLUS, GLUS-S-partial, GLUS-A, GLUS-S, myGLUS, space_glus

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

arXiv6 repos

arXiv:2504.13180

perception_models, PE-Lang-L14-448, PE-Lang-G14-448, PLM-Image-Auto, PLM-Video-Human, PLM-Video-Auto

Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video

arXiv6 repos

arXiv:2504.19475

ViT-Prisma, sparse-autoencoder-clip-b-32-sae-vanilla-x64-layer-10-hook_mlp_out-l1-1e-05, ViT-Prisma, my_prisma, df-prisma, ViT-Prisma-fix

Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning

arXiv6 repos

arXiv:2505.13886

Game-RL-InternVL2.5-8B, GameQA-text, Game-RL-InternVL3-8B, Game-RL-Qwen2.5-VL-7B, GameQA-140K, GameQA-5K

AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound

arXiv6 repos

arXiv:2505.14142

AudSemThinker, audsem-simple, audsemthinker-qa, audsem, audsemthinker, audsemthinker-qa-grpo

Emerging Properties in Unified Multimodal Pretraining

arXiv6 repos

arXiv:2505.14683

Bagel, bytedance_BAGEL-7B-MoT-INT8, Macro-Bagel, Vision-R1, ComfyUI-BAGEL, Bagel

OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

arXiv6 repos

arXiv:2505.20292

MagicTime, ChronoMagic-Bench, ConsisID, OpenS2V-5M, OpenS2V-Nexus, OpenS2V-Eval

Frame In-N-Out: Unbounded Controllable Image-to-Video Generation

arXiv6 repos

arXiv:2505.21491

FrameINO, FrameINO_Wan2.2_5B_Stage2_MotionINO_v1.5, FrameINO_Wan2.2_5B_Stage2_MotionINO_v1.6, FrameINO_CogVideoX_Stage2_MotionINO_v1.0, FrameINO_CogVideoX_Stage1_Motion_v1.0, FrameINO_Wan2.2_5B_Stage1_Motion_v1.5

Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences

arXiv6 repos

arXiv:2506.02095

cyclereward, CyclePrefDB-I2T, CycleReward-I2T, CycleReward-Combo, CyclePrefDB-T2I, CycleReward-T2I

UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

arXiv6 repos

arXiv:2506.03147

ImgEdit, UniWorld, UniWorld-V1, UniWorld-V1, UniWorld-V1, UniWorld-V1-NF4

SpatialLM: Training Large Language Models for Structured Indoor Modeling

arXiv6 repos

arXiv:2506.07491

SpatialLM1.1-Qwen-0.5B, SpatialLM, SpatialLM-Dataset, SpatialLM1.1-Llama-1B, SpatialLM-TestSet, SPATIALLM

MiniCPM4: Ultra-Efficient LLMs on End Devices

arXiv6 repos

arXiv:2506.07900

MiniCPM4-8B, FR-Spec, MiniCPM5-2B, MiniCPM5-2B-GGUF, MiniCPM5-1B-GGUF, MiniCPM5-1B

Step-Audio 2 Technical Report

arXiv6 repos

arXiv:2507.16632

Step-Audio-2-mini, Step-Audio2, Step-Audio-2-mini-Think, StepEval-Audio-Toolcall, StepEval-Audio-Paralinguistic, Step-Audio-2-mini-Base

On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

arXiv6 repos

arXiv:2508.05629

swift, ms-swift, ms-swift, vmopd, ms, dense-retention-rl

Controllable Latent Space Augmentation for Digital Pathology

arXiv6 repos

arXiv:2508.14588

HistAug, histaug-conch, histaug-conch_v15, histaug-virchow2, histaug-hoptimus1, histaug-uni

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

arXiv6 repos

arXiv:2509.01986

DIM, DIM-T2I, DIM-4.6B-Edit, DIM-Edit, DIM-4.6B-T2I, DIM-4.6B-Edit-Stage1

Tequila: Trapping-free Ternary Quantization for Large Language Models

arXiv6 repos

arXiv:2509.23809

Qwen3-1.7B_eagle3, Qwen3-4B_eagle3, Qwen3-8B_eagle3, Qwen3-14B_eagle3, Qwen3-a3B_eagle3, Qwen3-32B_eagle3

SpecExit: Accelerating Large Reasoning Model via Speculative Exit

arXiv6 repos

arXiv:2509.24248

Qwen3-a3B_eagle3, Qwen3-14B_eagle3, Qwen3-1.7B_eagle3, Qwen3-4B_eagle3, Qwen3-8B_eagle3, Qwen3-32B_eagle3

Fast-dLLM v2: Efficient Block-Diffusion LLM

arXiv6 repos

arXiv:2509.26328

Fast_dLLM_v2_7B, d3LLM, Fast-dLLM, dev-dllm, relay, Fast_dLLM_v2_1.5B

SDAR: A Synergistic Diffusion-AutoRegression Paradigm for Scalable Sequence Generation

arXiv6 repos

arXiv:2510.06303

SDAR, SDAR-4B-Chat, SDAR-8B-Chat, SDAR-30B-A3B-Sci, SDAR-1.7B-Chat, SDAR-30B-A3B-Chat

Emu3.5: Native Multimodal Models are World Learners

arXiv6 repos

arXiv:2510.26583

Emu3.5-VisionTokenizer, Emu3.5, Emu3.5-Image, Emu3.5, Emu35-Comfyui-Nodes, Emu35-Image-NF4

Cambrian-S: Towards Spatial Supersensing in Video

arXiv6 repos

arXiv:2511.04670

vsi-590k, Cambrian-S-3M, Cambrian-S-7B, cambrian-s-3b, cambrian-s-0.5b, cambrian-s-1.5b

d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation

arXiv6 repos

arXiv:2601.07568

d3LLM_LLaDA, d3LLM, trajectory_data_dream_32, d3LLM_Dream_Coder, d3LLM_Dream, trajectory_data_llada_32

Qwen3-TTS Technical Report

arXiv6 repos

arXiv:2601.15621

Qwen3-TTS-12Hz-1.7B-VoiceDesign, Qwen3-TTS-12Hz-1.7B-Base, Qwen3-TTS-12Hz-0.6B-CustomVoice, Qwen3-TTS-12Hz-0.6B-Base, Qwen3-TTS-Tokenizer-12Hz, qwen3-tts

vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models

arXiv6 repos

arXiv:2602.02204

vllm-omni, vllm-GPT-SoVITS, cmu-15642, vllm-omni, vllm-omni-hunyuanimage3, vllm-omni-minicpmo45-npu

DM4CT: Benchmarking Diffusion Models for Computed Tomography Reconstruction

arXiv6 repos

arXiv:2602.18589

lodochallenge_pixel_diffusion, lodochallenge_latent_diffusion, synchrotron_pixel_diffusion, lodoind_latent_diffusion, lodoind_pixel_diffusion, synchrotron_latent_diffusion

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

arXiv6 repos

arXiv:2604.04911

SpatialEdit, SpatialEdit-500K, SpatialEdit-16B, SpatialEdit-Bench, JoyAI-Image-SpatialEdit-Bench, JoyAI-Image-SpatialEdit

ACL-Verbatim: hallucination-free question answering for research

arXiv6 repos

arXiv:2605.21102

verbatim-rag, verbatim-rag-modern-bert-v2, verbatim-spans, acl-anthology-md, acl-verbatim-spans, acl-verbatim-modernbert

Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search

arXiv6 repos

arXiv:2605.30917

v-splade, v-splade-quality, v-splade-efficient, SPLADE-mlx, v-splade-efficient-mlx, v-splade-quality-mlx

Gemma 4 Technical Report

arXiv6 repos

arXiv:2607.02770

gemma-4-31B-it-qat-q4_0-unquantized, awesome-gemma, gemma-4-31B-it, gemma-4-e4b-it, gemma-4-12B-it, gemma-4-12b-it-qat-q4_0-gguf

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

arXiv6 repos

arXiv:2607.04988

InternVLA-A1, InternVLA-A1.5-DOMINO, InternVLA-A1.5-base, InternVLA-A1.5-RoboTwin, InternVLA-A1.5-Libero, InternVLA-A-series

Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion

arXiv6 repos

arXiv:2607.25572

vulnerability-attack-technique-classification-roberta-base, vulnerability-attack-technique-biencoder, cve-attack-mapping-paper, vulnerability-attack-technique-classification-roberta-base-llm-expanded, vulnerability-attack-techniques, vulnerability-attack-techniques-llm-scaling

Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

arXiv6 repos

arXiv:2608.10932

CamSFT-4B, CamInject-4B, CamSFT-8B, CamDistill-8B, CamInject-8B, CamDistill-4B

An Efficiency Study for SPLADE Models

ACM5 repos

ACM:3477495.3531833

splade, efficient-splade-V-large-doc, efficient-splade-VI-BT-large-doc, efficient-splade-V-large-query, efficient-splade-VI-BT-large-query

Microsoft COCO: Common Objects in Context

arXiv5 repos

arXiv:1405.0312

LocateAnything-3B, mscoco-it, MSCOCO, huggingface-datasets_MSCOCO, COCO-Caption2017

Deep Residual Learning for Image Recognition

arXiv5 repos

arXiv:1512.03385

yolov5, resnet50.tv_in1k, AIGI-Holmes, EasyOCR, resnet-50

SQuAD: 100,000+ Questions for Machine Comprehension of Text

arXiv5 repos

arXiv:1606.05250

squad, squad_v2, squad_bn, t5-large-encoder-only-bf16, CoConflictQA

Verified Low-Level Programming Embedded in F*

arXiv5 repos

arXiv:1703.00053

libcrux, everest, karamel, kremlin, libcrux

The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

arXiv5 repos

arXiv:1801.03924

materialgan, LatentSync, LatentSync1.5-mac, PerceptualSimilarity, WeavePrompt

Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

arXiv5 repos

arXiv:1908.08962

bert, tapas, bert-tiny-historic-multilingual-cased, bert-mini-historic-multilingual-cased, bert-small-historic-multilingual-cased

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

arXiv5 repos

arXiv:1910.02054

ColossalAI, MindSpeed-MM, awsome-llm-papers, stablecode-completion-alpha-3b-4k, mesh-transformer-jax

GLU Variants Improve Transformer

arXiv5 repos

arXiv:2002.05202

t5-v1_1-xxl, google_t5-v1_1-xxl_encoderonly, nomic-bert-2048, rwkv, llama2.zig

FLERT: Document-Level Features for Named Entity Recognition

arXiv5 repos

arXiv:2011.06993

ner-german-large, ner-english-ontonotes-large, ner-spanish-large, flair, ner-dutch-large

HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection

arXiv5 repos

arXiv:2012.10289

bert-base-uncased-hatexplain, HateXplain, bert-base-uncased-hatexplain-rationale-two, HateXplain, DeepLearningProject

VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

arXiv5 repos

arXiv:2101.00390

voxpopuli, voxpopuli, wav2vec2-large-100k-voxpopuli, unispeech-sat-large, wavlm-large

XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond

arXiv5 repos

arXiv:2104.12250

twitter-xlm-roberta-base-sentiment, xlm-t, twitter-xlm-roberta-base, xlm-twitter-politics-sentiment, multilingual-hate-speech-robacofi

SpeechBrain: A General-Purpose Speech Toolkit

arXiv5 repos

arXiv:2106.04624

lang-id-voxlingua107-ecapa, spkrec-ecapa-voxceleb, spkrec-xvect-voxceleb, spkrec-resnet-voxceleb, spkrec-ecapa-voxceleb-mel-spec

Evaluating Large Language Models Trained on Code

arXiv5 repos

arXiv:2107.03374

human-eval, llm-humaneval-benchmarks, FTTT, openai_humaneval, research2

CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and Generation

arXiv5 repos

arXiv:2109.05729

bart-base-chinese, CPT, bart-large-chinese, cpt-base, cpt-large

Challenges in Detoxifying Language Models

arXiv5 repos

arXiv:2109.07445

fineweb, fineweb-edu, falcon-refinedweb, finepdfs, fineweb-2

Few-shot Learning with Multilingual Language Models

arXiv5 repos

arXiv:2112.10668

xglm-2.9B, xglm-564M, polyglot, mGPT, mGPT

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

arXiv5 repos

arXiv:2201.12086

Semantic-Segment-Anything, blip-image-captioning-base, LAVIS, blip-vqa-base, blip-vqa-capfilt-large

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

arXiv5 repos

arXiv:2204.05862

hh-rlhf, artgpt2tox, RAIN, RRHF, LLM-Pref-Mark-UI

InCoder: A Generative Model for Code Infilling and Synthesis

arXiv5 repos

arXiv:2204.05999

sven_modified, sven, incoder-6B, incoder, incoder-1B

Petals: Collaborative Inference and Fine-tuning of Large Models

arXiv5 repos

arXiv:2209.01188

petals, subnet-llm, petals, bloombee_add_models, petals

Zero-Shot Learners for Natural Language Understanding via a Unified Multiple Choice Perspective

arXiv5 repos

arXiv:2210.08590

Ziya-LLaMA-13B-Pretrain-v1, Fengshenbang-LM, Ziya-LLaMA-13B-v1, Ziya-BLIP2-14B-Visual-v1, Ziya-LLaMA-13B-v1.1

AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

arXiv5 repos

arXiv:2211.06679

stable-diffusion-webui, FlagAI, AltDiffusion-m9, AltDiffusion, generative-ai

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

arXiv5 repos

arXiv:2211.10438

nunchaku, nunchaku, trtllm, TensorRT-LLM, TensorRT-LLM

TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities

arXiv5 repos

arXiv:2212.06385

gpt2-chinese-cluecorpussmall, t5-small-chinese-cluecorpussmall, t5-base-chinese-cluecorpussmall, Linly, Chinese-ChatLLaMA

Precise Zero-Shot Dense Retrieval without Relevance Labels

arXiv5 repos

arXiv:2212.10496

prompt-engineering, docs-reference, llm-search, spring-ai-extensions, KoPrivateGPT

A Watermark for Large Language Models

arXiv5 repos

arXiv:2301.10226

watermarks-remover, text-generation-inference, Adversarial-Paraphrasing, impossibility-watermark, lm-watermarking

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

arXiv5 repos

arXiv:2303.16199

LLaMA-Adapter, Point-Bind_Point-LLM, LLaMA-Adapter, LLaMA2-Accessory, lit-llama

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

arXiv5 repos

arXiv:2304.01373

TinyLlama, pythia-12b, pythia-1.4b, pythia-410m, pythia-2.8b

Scaling Speech Technology to 1,000+ Languages

arXiv5 repos

arXiv:2305.13516

mms-300m, gujarati-vsr, fcbh-dataset-io, mms-tts-spa, mms-tts-eng

The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning

arXiv5 repos

arXiv:2305.14045

KoCoT_2000, CoT-Collection, CoT-Collection, Multilingual-CoT-Collection, KoCommercial-Dataset

Segment Anything in High Quality

arXiv5 repos

arXiv:2306.01567

geti-instant-learn, sd-webui-inpaint-anything, sd-webui-inpaint-anything, interior-segment-labeler, SAMReg

A Technical Report for Polyglot-Ko: Open-Source Large-Scale Korean Language Models

arXiv5 repos

arXiv:2306.02254

polyglot, polyglot-ko-1.3b, polyglot-ko-3.8b, polyglot-ko-5.8b, level3_nlp_finalproject-nlp-12

RED$^{\rm FM}$: a Filtered and Multilingual Relation Extraction Dataset

arXiv5 repos

arXiv:2306.09802

mrebel-large, SREDFM, REDFM, mdeberta-v3-base-triplet-critic-xnli, mrebel-large-32

BayLing: Bridging Cross-lingual Alignment and Instruction Following through Interactive Translation for Large Language Models

arXiv5 repos

arXiv:2306.10968

BayLing, bayling-13b-v1.1, bayling-7b-diff, bayling-13b-diff, BayLing

Cross-Lingual Cross-Age Group Adaptation for Low-Resource Elderly Speech Emotion Recognition

arXiv5 repos

arXiv:2306.14517

elderly_ser, SER-wav2vec2-large-xlsr-53-eng-zho-adults, SER-wav2vec2-large-xlsr-53-eng-zho-elderly, SER-wav2vec2-large-xlsr-53-eng-zho-all-age, YueMotion

FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

arXiv5 repos

arXiv:2307.08691

graphify, lorax, GPT-2, vigogne, Block-Sparse-Attention

AlpaGasus: Training A Better Alpaca with Fewer Data

arXiv5 repos

arXiv:2307.08701

KoRAE, KoRAE-13b, KoRAE-13b-DPO, KoRAE_filtered_12k, original-KoRAE-13b-3ep

LISA: Reasoning Segmentation via Large Language Model

arXiv5 repos

arXiv:2308.00692

LISA-Llama-3, DINO_LISA, LISA-Training, LISA, LISA

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

arXiv5 repos

arXiv:2308.01907

all-seeing, CRPE, AS-Core, AS-V2, AS-100M

Towards General Text Embeddings with Multi-stage Contrastive Learning

arXiv5 repos

arXiv:2308.03281

gte-large-zh, gte-base-zh, gte-large-en-v1.5, gte-large, gte-base

Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs

arXiv5 repos

arXiv:2308.09895

self-oss-instruct-sc2-exec-filter-50k, MultiPL-T-StarCoder2_15B, stack-dedup-python-testgen-starcoder-filter-v2, MultiPL-T-CodeLlama_34b, MultiPL-T-DeepSeekCoder_33b

How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary Detection

arXiv5 repos

arXiv:2308.13177

omchat-v2.0-13B-single-beta_hf, omchat, OmDet, OVDEval, OmAgent

Efficient Streaming Language Models with Attention Sinks

arXiv5 repos

arXiv:2309.17453

streaming-llm, KVQuant, dlms-sinks, mlx-flash, Block-Sparse-Attention

Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature

arXiv5 repos

arXiv:2310.05130

L2D, AdaDetectGPT, fast-detect-gpt, fast-detect-gpt, worker-fast-detect-gpt

Fine-Tuning LLaMA for Multi-Stage Text Retrieval

arXiv5 repos

arXiv:2310.08319

tinyllama-embed, repllama-v1-7b-lora-passage, rankllama-v1-7b-lora-passage, pyterrier_genrank, RepLLaMA-reproduced

Llemma: An Open Language Model For Mathematics

arXiv5 repos

arXiv:2310.10631

math-lm, proof-pile-2, llemma_34b, llemma_7b, llmstep

SkyMath: Technical Report

arXiv5 repos

arXiv:2310.16713

Skywork-13B-base, Skywork-13B-Math-8bits, Skywork-13B-Base-3.1TB, Skywork-13B-Base-8bits, Skywork-13B-Math

LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing

arXiv5 repos

arXiv:2311.00571

LLaVA, GLIGEN, TriPlaneLLaVA, LLaVA-toy, TinyLLava

LongQLoRA: Efficient and Effective Method to Extend Context Length of Large Language Models

arXiv5 repos

arXiv:2311.04879

LongQLoRA, LongQLoRA-Llama2-7b-8k, LongQLoRA-Vicuna-13b-8k, Long-QLORA, article_gpt

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

arXiv5 repos

arXiv:2311.05437

LLaVA, TriPlaneLLaVA, LLaVA-toy, LLaVA-Plus-Codebase, TinyLLava

Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

arXiv5 repos

arXiv:2311.06242

Florence-2-base, Florence-2-base-ft, Florence-2-large, Florence-2-large-ft, interior-segment-labeler

Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

arXiv5 repos

arXiv:2311.08046

LanguageBind, Chat-UniVi-Instruct, Chat-UniVi, Chat-UniVi-13B, Chat-UniVi

GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer

arXiv5 repos

arXiv:2311.08526

gliner_multi_pii-v1, gliner-stream-pii-v1.0, gliner_small-v2.5, gliner_small-v2.1, agentic-graphrag

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

arXiv5 repos

arXiv:2311.12793

ShareGPT4V-7B, ShareGPT4V, ShareCaptioner, ShareGPT4V-13B, ShareGPT4V

Magicoder: Empowering Code Generation with OSS-Instruct

arXiv5 repos

arXiv:2312.02120

evalplus, Magicoder-S-DS-6.7B, Magicoder-S-CL-7B, Magicoder-DS-6.7B, Magicoder-CL-7B

Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval

arXiv5 repos

arXiv:2312.15503

bge-large-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-reranker-v2-m3, bge-reranker-v2.5-gemma2-lightweight

Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

arXiv5 repos

arXiv:2401.01044

Auffusion, auffusion-full, auffusion, auffusion-full-no-adapter, MMDisCo

GPT-4V(ision) is a Generalist Web Agent, if Grounded

arXiv5 repos

arXiv:2401.01614

UGround-V1-7B, UGround-V1-2B, UGround-V1-72B, UGround, Multimodal-Mind2Web

Latte: Latent Diffusion Transformer for Video Generation

arXiv5 repos

arXiv:2401.03048

Latte-1, Latte, Latte-0, Latte, Latte-1

Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series

arXiv5 repos

arXiv:2401.03955

ttm-research-r2, granite-timeseries-ttm-r2, Samay, moment, granite-timeseries-ttm-v1

Executable Code Actions Elicit Better LLM Agents

arXiv5 repos

arXiv:2402.01030

rlm, CodeActAgent-Llama-2-7b, code-act, CodeActAgent-Mistral-7b-v0.1, CodeActAgent-Mistral-7b-v0.1.q8_0.gguf

DoRA: Weight-Decomposed Low-Rank Adaptation

arXiv5 repos

arXiv:2402.09353

hymba, ohara, llama3-chinese, llama3-chinese, Llama3-Chinese-Lora

Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models

arXiv5 repos

arXiv:2402.14207

dspy, storm, gpt-researcher, awesome-dspy, dsp

The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

arXiv5 repos

arXiv:2402.19474

all-seeing, CRPE, AS-Core, AS-V2, AS-100M

Apollo: A Lightweight Multilingual Medical LLM towards Democratizing Medical AI to 6B People

arXiv5 repos

arXiv:2403.03640

ApolloCorpus, Apollo-7B, Apollo-34B, Apollo-72B, PodGPT

DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation

arXiv5 repos

arXiv:2403.08857

HunyuanDiT, HunyuanDiT-v1.1, HunyuanDiT, HunyuanDiT-v1.2, HunyuanDiT

FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions

arXiv5 repos

arXiv:2403.15246

mteb-1.34.14, FollowIR, FollowIR-7B, FollowIR-train, ru-promptriever-qwen3-4b

Applying Guidance in a Limited Interval Improves Sample and Distribution Quality in Diffusion Models

arXiv5 repos

arXiv:2404.07724

OmniGen2, OmniGen2, OmniGen2-EditScore7B, Macro-OmniGen2, OmniGen2-EditScore7B-v1.1

TAVGBench: Benchmarking Text to Audible-Video Generation

arXiv5 repos

arXiv:2404.14381

JavisBench, JavisGPT, JavisInst-Omni, MM-PreTrain, AV-FineTune

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

arXiv5 repos

arXiv:2404.14396

SEED-Data-Edit-Part2-3, SEED-X, SEED-X-17B, SEED-Data-Edit, SEED

Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation

arXiv5 repos

arXiv:2406.02347

flash-diffusion, flash-pixart, flash-sd3, flash-sdxl, flash-sd

LADI v2: Multi-label Dataset and Classifiers for Low-Altitude Disaster Imagery

arXiv5 repos

arXiv:2406.02780

ladi-overview, LADI-v2-dataset, LADI-v2-classifier-large, LADI-v2-classifier-large-reference, LADI-v2-classifier-small

Scaling and evaluating sparse autoencoders

arXiv5 repos

arXiv:2406.04093

nanointerpret, sparsify, dictionary_learning, cli, notebooks

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

arXiv5 repos

arXiv:2406.04325

ShareGPT4Video, ShareGPT4Video, sharegpt4video-8b, ShareCaptioner-Video, stereopilot-replica-accelerate

Multimodal Table Understanding

arXiv5 repos

arXiv:2406.08100

table-llava-v1.5-13b, Table-LLaVA, MMTab, table-llava-v1.5-7b, table-llava-v1.5-7b-hf

FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

arXiv5 repos

arXiv:2407.04051

minimind-o, FunASR, SenseVoice, SenseVoice, FunASR

ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

arXiv5 repos

arXiv:2407.06135

Anole-7b-v0.1, anole, Anole-7b, UniWM, GoViG

FlashNorm: Fast Normalization for Transformers

arXiv5 repos

arXiv:2407.09577

transformer-tricks, gemma-4-E2B-FlashNorm, Llama-3.1-8B-FlashNorm, gemma-4-E2B-FlashNorm-strict, Llama-3.2-1B-FlashNorm

CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

arXiv5 repos

arXiv:2407.15886

catvton-flux, catvton-unstudio-flux, CatVTON, CatVTON, CatVTON-MaskFree

SAM 2: Segment Anything in Images and Videos

arXiv5 repos

arXiv:2408.00714

segment-anything, geti-instant-learn, sam2-hiera-large, sam2-hiera-tiny, sam2.1-hiera-large

OLMoE: Open Mixture-of-Experts Language Models

arXiv5 repos

arXiv:2409.02060

OLMoE, OLMoE-mix-0924, OLMoE-1B-7B-0924-SFT, OLMoE-1B-7B-0924-Instruct, OLMoE-1B-7B-0125-Instruct

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

arXiv5 repos

arXiv:2409.04429

vila-u, VLAC, OpenVid-1M, OpenVid, vila-u-7b-256

Block-Attention for Efficient Prefilling

arXiv5 repos

arXiv:2409.15355

Block-Attention, Tulu3-Block-FT, Tulu3-SFT, Tulu3-RAG, GraphKV

Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale

arXiv5 repos

arXiv:2409.17115

DCLM-pro, web-doc-refining-lm, math-doc-refining-lm, web-chunk-refining-lm, math-chunk-refining-lm

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

arXiv5 repos

arXiv:2410.06885

F5-TTS, F5-TTS, StyleStream, StyleStream, X-Voice

LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

arXiv5 repos

arXiv:2410.10813

hippo-memory, Awareness-Market, ogham-mcp, Titan-Memory, post-graph-rag

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

arXiv5 repos

arXiv:2410.18451

Skywork-Reward-Preference-80K-v0.2, Skywork-Reward-Llama-3.1-8B, Skywork-Reward-Gemma-2-27B, Skywork-Reward-Gemma-2-27B-v0.2, Skywork-Reward-Llama-3.1-8B-v0.2

HunyuanVideo: A Systematic Framework For Large Video Generative Models

arXiv5 repos

arXiv:2412.03603

HunyuanVideo, HunyuanVideo-PromptRewrite, HunyuanVideo, HunyuanVideo-I2V, HunyuanVideo-I2V

Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

arXiv5 repos

arXiv:2412.18319

Mulberry-SFT, Mulberry, Mulberry_llama_11b, Mulberry_qwen2vl_7b, Mulberry_llava_8b

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

arXiv5 repos

arXiv:2412.18605

Orient-Anything, Orient-Anything, OriNet, OriAnyV2_ckpt, orient-anything

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

arXiv5 repos

arXiv:2412.18925

medical-o1-reasoning-SFT, HuatuoGPT-o1-70B, HuatuoGPT-o1-7B, HuatuoGPT-o1-8B, medical-o1-verifiable-problem

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

arXiv5 repos

arXiv:2502.09927

granite-vision-3.3-2b, granite-vision-3.3-2b-GGUF, granite-vision-3.1-2b-preview, granite-vision-4.1-4b, granite-4.0-3b-vision

OmniSQL: Synthesizing High-quality Text-to-SQL Data at Scale

arXiv5 repos

arXiv:2503.02240

OmniSQL, SynSQL-2.5M, OmniSQL-14B, OmniSQL-7B, OmniSQL-32B

YOLOE: Real-Time Seeing Anything

arXiv5 repos

arXiv:2503.07465

yoloe, yolov10, yoloe, yoloe, yoloe

ViSpeak: Visual Instruction Feedback in Streaming Videos

arXiv5 repos

arXiv:2503.12769

ViSpeak, StreamingBench, StreamingBench, ViSpeak-s2, ViSpeak-s3

LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis

arXiv5 repos

arXiv:2503.21749

LeX-Art, LeX-Lumina, LeX-Bench, LeX-Data-10K, LeX-Enhancer-full

TerraMind: Large-Scale Generative Multimodality for Earth Observation

arXiv5 repos

arXiv:2504.11171

TerraMind-1.0-base, terramind, TerraMind-1.0-large, TerraMind-1.0-tiny, TerraMind-1.0-small

Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

arXiv5 repos

arXiv:2504.16656

Skywork-R1V, Skywork-R1V2-38B, Skywork-R1V-38B-AWQ, Skywork-R1V2-38B-AWQ, sky-rv1

Process Reward Models That Think

arXiv5 repos

arXiv:2504.16828

ThinkPRM, ThinkPRM-14B, ThinkPRM-7B, ThinkPRM-1.5B, thinkprm-1K-verification-cots

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

arXiv5 repos

arXiv:2505.07608

MiMo-7B-RL, MiMo-7B-Base, MiMo-7B-RL-0530, MiMo-7B-RL-Zero, MiMo-7B-SFT

MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning

arXiv5 repos

arXiv:2505.10557

MM-MathInstruct, Img2Code, MathCoder-VL-2B, FigCodifier, MathCoder-VL-8B

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

arXiv5 repos

arXiv:2506.05573

PartCrafter, PartCrafter-Scene, PartCrafter, modly-partcrafter-extension, Accelerator0701

BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation

arXiv5 repos

arXiv:2506.07530

BitVLA-CoreAI, BitVLA, bitvla-bitsiglipL-224px-bf16, bitvla-bf16, bitvla-siglipL-224px-bf16

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

arXiv5 repos

arXiv:2506.09965

ViLaSR, ViLaSR, ViLaSR, ViLaSR-cold-start, ViLaSR-data

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

arXiv5 repos

arXiv:2506.21619

Awesome-AITools, IndexTTS-2, parrots, index-tts, ComfyUI-kaola-IndexTTS2

Kwai Keye-VL Technical Report

arXiv5 repos

arXiv:2507.01949

Keye, Keye-VL-2.0-30B-A3B, Keye-VL-8B-Preview, Keye-VL-1.5-8B, keye-39e1f0b5

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

arXiv5 repos

arXiv:2507.16116

PusaV1_training, PusaV0.5_Training, Pusa-Wan2.2-V1, PusaV1, Pusa-V0.5

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

arXiv5 repos

arXiv:2507.20939

ShortVid-Bench, ARC-Hunyuan-Video-7B, ARC-Hunyuan-Video-7B, ARC-Qwen-Video-7B, ARC-Qwen-Video-7B-Narrator

PixNerd: Pixel Neural Field Diffusion

arXiv5 repos

arXiv:2507.23268

PixNerd-diffusers, PixNerd-diffusers, PixNerd, PixNerd-XXL-P16-T2I, PixNerd-XL-P16-C2I

Qwen-Image Technical Report

arXiv5 repos

arXiv:2508.02324

Qwen-Image-2512, Qwen-Image, Qwen-Image, Macro-Qwen-Image-Edit, Qwen-Image-Flash

DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrieval

arXiv5 repos

arXiv:2508.07995

Diver-Retriever-4B, Diver, Diver-Retriever-0.6B, Diver-Retriever-4B-1020, Diver-Retriever-1.7B

DINOv3

arXiv5 repos

arXiv:2508.10104

geti-instant-learn, dinov3-vitl16-pretrain-lvd1689m, dinov3, compositio_nn, projet-vision

Thyme: Think Beyond Images

arXiv5 repos

arXiv:2508.11630

Thyme, Thyme-SFT, Thyme-RL, Thyme-SFT, Thyme-RL

Kwai Keye-VL 1.5 Technical Report

arXiv5 repos

arXiv:2509.01563

Keye, Keye-VL-2.0-30B-A3B, Keye-VL-8B-Preview, Keye-VL-1.5-8B, keye-39e1f0b5

WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents

arXiv5 repos

arXiv:2509.06501

MiniMax-M2-BF16, MiniMax-M2.1, MiniMax-M2.1, MiniMax-M2, MiniMax-M2

EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis

arXiv5 repos

arXiv:2510.25628

EHR-R1, EHR-Bench, EHR-Ins-Reasoning, EHR-R1-1.7B, EHR-R1-8B

arXiv:2511.16175

arXiv5 repos

arXiv:2511.16175

Mantis, mantis_libero_lerobot, Mantis-Base, Mantis, wla

Qwen3-VL Technical Report

arXiv5 repos

arXiv:2511.21631

Qwen2.5-VL, Qwen3-VL, Qwen2-VL, nemotron-colembed-vl-4b-v2, RoboSpatial-Eval

Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation

arXiv5 repos

arXiv:2512.23703

Robo-Dopamine-GRM-3B, Robo-Dopamine-GRM-2.0-8B-Preview, Robo-Dopamine-GRM-8B, Robo-Dopamine-GRM-2.0-4B-Preview, Robo-Dopamine-Bench

NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos

arXiv5 repos

arXiv:2601.00393

NeoVerse, NeoVerse, NeoVerse, NeoVerse-archive, neoverse_new

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

arXiv5 repos

arXiv:2601.14251

LightOnOCR-2-1B, LightOnOCR-2-1B-old-church-slavonic-line, LightOnOCR-2-1B-base, LightOnOCR-2-1B-Pinokio, LightOnOCR-mix-0126

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

arXiv5 repos

arXiv:2602.06965

MedMO-8B-Next, MedMO-8B, MedMO-4B, MedMO-4B-Next, MedMO

Revisiting Text Ranking in Deep Research

arXiv5 repos

arXiv:2602.21456

text-ranking-in-deep-research, browsecomp-plus-passage-corpus, browsecomp-plus-passage-corpus-pyserini, browsecomp-plus-indexes, browsecomp-plus-runs

MediX-R1: Open Ended Medical Reinforcement Learning

arXiv5 repos

arXiv:2602.23363

medix-rl-data, MediX-R1-30B, MediX-R1-8B, MediX-R1-2B, MediX-R1

Can Vision-Language Models Solve the Shell Game?

arXiv5 repos

arXiv:2603.08436

shellgame, Molmo2-SGCoT-Demo, Molmo2-SGCoT, vetbench, Molmo2-SGCoT

Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously

arXiv5 repos

arXiv:2603.12262

VST, VST-32B, VST-3B, VST-Training-Data, VST-7B

Small Vision-Language Models are Smart Compressors for Long Video Understanding

arXiv5 repos

arXiv:2604.08120

Tempo, Tempo-6B, Tempo-6B-Stage1, Tempo-6B-Stage2, Tempo-6B-Stage0

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech

arXiv5 repos

arXiv:2605.20830

Raon-OpenTTS-1B, Raon-OpenTTS, Raon-OpenTTS-0.3B, Raon-OpenTTS-Eval, Raon-OpenTTS-Pool

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

arXiv5 repos

arXiv:2606.16533

kairos, Kairos3.1-4B-robot-480P, kairos-4B-robot-RoboTwin2.0, kairos-4B-robot-LIBERO-plus, kairos-sensenova

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

arXiv5 repos

arXiv:2607.24743

ClinFusion-32B, ClinFusion-8B, ClinFusion, clinfusion-medical-vlm, ClinFusion-Eval-Data

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

arXiv5 repos

arXiv:2609.08183

NeoHorse-1-4B, NeoHorse, NeoHorse-1-9B-GGUF, NeoHorse-1-4B-GGUF, NeoHorse-1-9B

A pathology foundation model for cancer diagnosis and prognosis prediction

Nature5 repos

Nature:s41586-024-07894-z

AtlasPatch, KEEP, KEEP, CPathPatchFeature, TITAN

Aligning Large Language Model with Direct Multi-Preference Optimization for Recommendation

ACM4 repos

ACM:3627673.3679611

LlamaFactory, LLaMA-Factory, LLaMA-Factory-personal, LLaMA-Factory-LFS

arXiv:0000.00000

arXiv4 repos

arXiv:0000.00000

granite-4.1-3b, granite-3.3-8b-instruct, granite-4.1-8b, granite-3.1-1b-a400m-instruct

Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs

arXiv4 repos

arXiv:1603.09320

pgvector, pecos, sweet-search, usearch

SearchQA: A New Q&A Dataset Augmented with Context from a Search Engine

arXiv4 repos

arXiv:1704.05179

all-MiniLM-L6-v2, all-MiniLM-L12-v2, CoConflictQA, KoPrivateGPT

TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

arXiv4 repos

arXiv:1705.03551

CoConflictQA, KoPrivateGPT, GraphKV, GraphKV

Proximal Policy Optimization Algorithms

arXiv4 repos

arXiv:1707.06347

minimind, tunix, Reinforcement-Learning-Full-Pipeline, cleanrl

Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

arXiv4 repos

arXiv:1903.12261

MediMeta-C, RobustMedCLIP, RobustMedCLIP, medmnistc-api

Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer

arXiv4 repos

arXiv:1907.01341

Text2Video-Zero, svraster, MiDaS, prisma

Leveraging Pre-trained Checkpoints for Sequence Generation Tasks

arXiv4 repos

arXiv:1907.12461

tf-transformers, wiki_split, bert2bert_L-24_wmt_de_en, Reddit-Sports-Sentiment-Analysis

Release Strategies and the Social Impacts of Language Models

arXiv4 repos

arXiv:1908.09203

DetectLLMSegmentation, L2D, AdaDetectGPT, CodeXGLUE

NEZHA: Neural Contextualized Representation for Chinese Language Understanding

arXiv4 repos

arXiv:1909.00204

nezha-cn-large, nezha-large-wwm, nezha-base-wwm, nezha-cn-base

Libri-Light: A Benchmark for ASR with Limited or No Supervision

arXiv4 repos

arXiv:1912.07875

libriheavy, unispeech-sat-large, libri-light, wavlm-large

On the limits of cross-domain generalization in automated X-ray prediction

arXiv4 repos

arXiv:2002.02497

covid-chestxray-dataset, torchxrayvision, densenet121-res224-chex, torchxrayvision

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

arXiv4 repos

arXiv:2005.11401

gpt-researcher, ROOT-RAG, flan-ul2-dolly, flan-ul2-dolly-lora

ConvBERT: Improving BERT with Span-based Dynamic Convolution

arXiv4 repos

arXiv:2008.02496

berts, turkish-bert, convbert-base-turkish-cased, europeana-bert

What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams

arXiv4 repos

arXiv:2009.13081

MedQA-USMLE-4-options, PodGPT, MedQA, Baichuan2

Prefix-Tuning: Optimizing Continuous Prompts for Generation

arXiv4 repos

arXiv:2101.00190

FasterTransformer, Continual-NExT, SwissArmyTransformer, LoRA

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

arXiv4 repos

arXiv:2101.03961

marker, fiddler, awsome-llm-papers, smolMoELM-custom

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

arXiv4 repos

arXiv:2102.05918

coyo-align-b7-base, coyo-dataset, coyo-700m, align-base

Improved Denoising Diffusion Probabilistic Models

arXiv4 repos

arXiv:2102.09672

PixArt-alpha, pixeart, PixArt-alpha, smalldiffusion

GPT Understands, Too

arXiv4 repos

arXiv:2103.10385

GLM, EasyNLP, Continual-NExT, SwissArmyTransformer

BookSum: A Collection of Datasets for Long-form Narrative Summarization

arXiv4 repos

arXiv:2105.08209

airoboros-summarization, vid2cleantxt, bigbird-pegasus-large-K-booksum, long-t5-tglobal-xl-16384-book-summary

SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval

arXiv4 repos

arXiv:2109.10086

splade, splade_v2_max, splade_v2_distil, pinecone-text

Training Verifiers to Solve Math Word Problems

arXiv4 repos

arXiv:2110.14168

gsm8k, llm-jepa, FTTT, RLPR-Evaluation

Swin Transformer V2: Scaling Up Capacity and Resolution

arXiv4 repos

arXiv:2111.09883

CLIP-ViT-L-14-laion2B-s32B-b82K, MiDaS, swinv2-large-patch4-window12to16-192to256-22kto1k-ft, Swin-Transformer

Locating and Editing Factual Associations in GPT

arXiv4 repos

arXiv:2202.05262

OBLITERATUS, obliteratus, causal_unlearn_llm, KEditVis-LLM-Editing

Competition-Level Code Generation with AlphaCode

arXiv4 repos

arXiv:2203.07814

apps, code_contests, code_contests, code_contests

PLAID: An Efficient Engine for Late Interaction Retrieval

arXiv4 repos

arXiv:2205.09707

fast-plaid, plaidrepro, colbertv2.0, ColBERT

Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset

arXiv4 repos

arXiv:2207.00220

pile-of-law, legalbert-large-1.7M-1, distilbert-base-uncased-finetuned-eoir_privacy, legalbert-large-1.7M-2

Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

arXiv4 repos

arXiv:2209.03003

Matcha-TTS, GR00T-N1.5-3B, nanoMFM, Rectified-Diffusion

Flow Matching for Generative Modeling

arXiv4 repos

arXiv:2210.02747

Matcha-TTS, nanoMFM, com-304-FM-project-2026, Rectified-Diffusion

Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

arXiv4 repos

arXiv:2210.09984

miracl, miracl, miracl-corpus, gte-multilingual-base

High Fidelity Neural Audio Compression

arXiv4 repos

arXiv:2210.13438

encodec, encodec_24khz, encodec_48khz, whisperspeech

Lila: A Unified Benchmark for Mathematical Reasoning

arXiv4 repos

arXiv:2210.17517

Arithmo2-Mistral-7B, Arithmo-Mistral-7B, Arithmo2-Mistral-7B-adapter, Arithmo-Data

OneFormer: One Transformer to Rule Universal Image Segmentation

arXiv4 repos

arXiv:2211.06220

oneformer_ade20k_dinat_large, OneFormer, oneformer_cityscapes_swin_large, Semantic-Segment-Anything

A Time Series is Worth 64 Words: Long-term Forecasting with Transformers

arXiv4 repos

arXiv:2211.14730

patchtst-fm-r1, granite-timeseries-patchtst-fm-r1, PatchTST, FM4Motor

Large Language Models Encode Clinical Knowledge

arXiv4 repos

arXiv:2212.13138

Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, OLAPH, healthsearchqa

Zero-1-to-3: Zero-shot One Image to 3D Object

arXiv4 repos

arXiv:2303.11328

zero123, zero123-live, zero123-weights, zero123-xl-diffusers

FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization

arXiv4 repos

arXiv:2303.14189

coreml-FastViT-T8, coreml-FastViT-MA36, ml-fastvit, fast_donut_KIE

The Vault: A Comprehensive Multilingual Dataset for Advancing Code Understanding and Generation

arXiv4 repos

arXiv:2305.06156

TheVault, the-vault-inline, the-vault-class, the-vault-function

TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

arXiv4 repos

arXiv:2305.07759

esp32s3-distributed-ai, TinyStories, Tiny-Stories-Regional, llama2.ts

arXiv:2305.10703

arXiv4 repos

arXiv:2305.10703

ReGen, news_contrastive_pretrain, wiki_contrastive_pretrain, review_contrastive_pretrain

LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

arXiv4 repos

arXiv:2305.18802

libritts-r-filtered-speaker-descriptions, libritts_r, kanade-tokenizer, libritts_r

MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training

arXiv4 repos

arXiv:2306.00107

MERT-v1-95M, MERT-v1-330M, MERT-v0, YuE

Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

arXiv4 repos

arXiv:2306.00814

vocos-mel-24khz, vocos, whisperspeech, vocos-encodec-24khz

TIES-Merging: Resolving Interference When Merging Models

arXiv4 repos

arXiv:2306.01708

Twin-Merging, Mario, MergeLM, ComfyUI-LoRA-Optimizer

MIMIC-IT: Multi-Modal In-Context Instruction Tuning

arXiv4 repos

arXiv:2306.05425

idefics-80b-instruct, Otter, Otter, idefics-9b-instruct

High-Fidelity Audio Compression with Improved RVQGAN

arXiv4 repos

arXiv:2306.06546

SpeechTokenizer, descript-audio-codec, BigVGAN, nemo-nano-codec-22khz-1.89kbps-21.5fps

Quilt-1M: One Million Image-Text Pairs for Histopathology

arXiv4 repos

arXiv:2306.11207

quilt1m, QuiltNet-B-32, QuiltNet-B-16-PMB, QuiltNet-B-16

Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

arXiv4 repos

arXiv:2306.14289

sd-webui-inpaint-anything, MobileSAM, sd-webui-inpaint-anything, MobileSAM

Large Multimodal Models: Notes on CVPR 2023 Tutorial

arXiv4 repos

arXiv:2306.14895

LLaVA, TriPlaneLLaVA, LLaVA-toy, TinyLLava

ReLoRA: High-Rank Training Through Low-Rank Updates

arXiv4 repos

arXiv:2307.05695

BigDL, ipex-llm, ipex-llm, llm_test

3D-LLM: Injecting the 3D World into Large Language Models

arXiv4 repos

arXiv:2307.12981

PointLLM, ShapeLLM, MiniGPT-3D, 3D-LLM

AltDiffusion: A Multilingual Text-to-Image Diffusion Model

arXiv4 repos

arXiv:2308.09991

AltDiffusion, AltDiffusion-m9, AltDiffusion-m18, AltDiffusion

SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

arXiv4 repos

arXiv:2308.16692

SpeechTokenizer, SpeechGPT, USLM, SpeechTokenizer

PointLLM: Empowering Large Language Models to Understand Point Clouds

arXiv4 repos

arXiv:2308.16911

PointLLM, ShapeLLM, MiniGPT-3D, PointLLM

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

arXiv4 repos

arXiv:2309.00615

PointLLM, ShapeLLM, MiniGPT-3D, Point-Bind_Point-LLM

XGen-7B Technical Report

arXiv4 repos

arXiv:2309.03450

xgen-7b-8k-base, xgen, xgen-7b-4k-base, xgen-7b-8k-inst

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

arXiv4 repos

arXiv:2309.05653

Arithmo2-Mistral-7B, Arithmo-Mistral-7B, Arithmo2-Mistral-7B-adapter, Arithmo-Data

An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

arXiv4 repos

arXiv:2309.09958

LLaVA, TriPlaneLLaVA, LLaVA-toy, TinyLLava

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

arXiv4 repos

arXiv:2309.10020

LLaVA, TriPlaneLLaVA, LLaVA-toy, TinyLLava

OpenChat: Advancing Open-source Language Models with Mixed-Quality Data

arXiv4 repos

arXiv:2309.11235

openchat, openchat_3.5, openchat-3.6-8b-20240522, openchat-3.5-0106

LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

arXiv4 repos

arXiv:2309.11998

FastChat, Nanoflow, multilingual_mt_bench, FastChat

QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models

arXiv4 repos

arXiv:2309.14717

BigDL, ipex-llm, ipex-llm, llm_test

Finite Scalar Quantization: VQ-VAE Made Simple

arXiv4 repos

arXiv:2309.15505

nemo-nano-codec-22khz-1.89kbps-21.5fps, SkinTokens, SkinTokens, whisper_pinyin

OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text

arXiv4 repos

arXiv:2310.06786

proof-pile-2, DeepSeek-Math, OpenWebMath, open-web-math

LCM-LoRA: A Universal Stable-Diffusion Acceleration Module

arXiv4 repos

arXiv:2311.05556

TCD-SDXL-LoRA, lcm-lora-sdxl, lcm-lora-sdv1-5, Arc2Face

Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models

arXiv4 repos

arXiv:2311.06607

Monkey, Monkey, Detailed_Caption, Monkey-Chat

Mustango: Toward Controllable Text-to-Music Generation

arXiv4 repos

arXiv:2311.08355

mustango, mustango, MusicBench, mustango-pretrained

VideoCon: Robust Video-Language Alignment via Contrast Captions

arXiv4 repos

arXiv:2311.10111

owl-con, videocon, videocon, videocon-model

Diffusion Model Alignment Using Direct Preference Optimization

arXiv4 repos

arXiv:2311.12908

dpo-sdxl-text2image-v1, DiffusionDPO, dpo-sd1.5-text2image-v1, GRPO

LM-Cocktail: Resilient Tuning of Language Models via Model Merging

arXiv4 repos

arXiv:2311.13534

bge-large-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5, bge-small-en

EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models

arXiv4 repos

arXiv:2311.15596

EgoThink, EgoThink, embodied-eval, behaviour_subtask

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

arXiv4 repos

arXiv:2311.16502

Yi-34B-Chat, MMMU, MMMU, MMMU

RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance

arXiv4 repos

arXiv:2311.18681

RaDialog_v2_modified, RaDialog_v2, RaDialog, RaDialog-interactive-radiology-report-generation

A Task is Worth One Word: Learning with Task Prompts for High-Quality Versatile Image Inpainting

arXiv4 repos

arXiv:2312.03594

PowerPaint, PowerPaint-v2-1, PowerPaint-v1, PowerPaint_v2

Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos

arXiv4 repos

arXiv:2312.04746

quilt-llava.github.io, quilt1m, quilt-llava, Quilt-Llava-v1.5-7b

DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines

arXiv4 repos

arXiv:2312.13382

dspy, PromptingTools.jl, awesome-dspy, dsp

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

arXiv4 repos

arXiv:2312.15185

emotion2vec_base, emotion2vec, emotion2vec_plus_seed, emotion2vec_plus_base

LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

arXiv4 repos

arXiv:2312.17240

LISA-Llama-3, DINO_LISA, LISA, LISA

Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

arXiv4 repos

arXiv:2401.09417

Vim-small-midclstok, Vim, Vim-tiny-midclstok, Vim-base-midclstok

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

arXiv4 repos

arXiv:2401.10891

Depth-Estimation, coreml-depth-anything-v2-small, coreml-depth-anything-small, Depth-Anything-V2-Small-hf

Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text

arXiv4 repos

arXiv:2401.12070

DetectLLMSegmentation, AdaDetectGPT, LAPD, L2D

AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data

arXiv4 repos

arXiv:2402.00769

AnimateLCM, AnimateLCM-I2V, AnimateLCM, AnimateLCM-SVD-xt

Timer: Generative Pre-trained Transformers Are Large Time Series Models

arXiv4 repos

arXiv:2402.02368

UTSD, timer-base-84m, Large-Time-Series-Model, moment

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

arXiv4 repos

arXiv:2402.03766

MobileVLM_V2-1.7B, MobileVLM, MobileVLM_V2-7B, MobileVLM_V2-3B

MEMORYLLM: Towards Self-Updatable Large Language Models

arXiv4 repos

arXiv:2402.04624

MemoryLLM, memoryllm-8b-chat, memoryllm-8b, ExtendingMemoryLLM

Multilingual E5 Text Embeddings: A Technical Report

arXiv4 repos

arXiv:2402.05672

multilingual-e5-base, multilingual-e5-large, OKEAN, multilingual-e5-large-instruct

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

arXiv4 repos

arXiv:2402.12226

AnyInstruct, AnyGPT-chat, AnyGPT-base, SpeechGPT

BiMediX: Bilingual Medical Mixture of Experts LLM

arXiv4 repos

arXiv:2402.13253

BiMediX, BiMediX-Bi, BiMediX-Eng, BiMediX-Ara

Repetition Improves Language Model Embeddings

arXiv4 repos

arXiv:2402.15449

llm2vec, DermL2V-training, Anchor-Embedding, DermL2V-tmp

Trajectory Consistency Distillation: Improved Latent Consistency Distillation by Semi-Linear Consistency Function with Trajectory Mapping

arXiv4 repos

arXiv:2402.19159

TCD, TCD-SD21-base-LoRA, TCD-SDXL-LoRA, TCD-SD15-LoRA

StarCoder 2 and The Stack v2: The Next Generation

arXiv4 repos

arXiv:2402.19173

evalplus, starchat2-15b-v0.1, speechless-starcoder2-15b, AMALIA-9B-0626-DPO

DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models

arXiv4 repos

arXiv:2402.19481

nunchaku, xDiT, nunchaku, distrifuser-controlnet

Improving Diffusion Models for Authentic Virtual Try-on in the Wild

arXiv4 repos

arXiv:2403.05139

IDM-VTON, IDM-VTON, IDM-VTON-train, modal_pipeline

Fast High-Resolution Image Synthesis with Latent Adversarial Diffusion Distillation

arXiv4 repos

arXiv:2403.12015

Nitro-1, Nitro-1-PixArt, Nitro-1-SD, AMD-Diffusion-Distillation

RULER: What's the Real Context Size of Your Long-Context Language Models?

arXiv4 repos

arXiv:2404.06654

Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, Qwen3-Next-80B-A3B-Instruct

ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving

arXiv4 repos

arXiv:2404.16771

ConsistentID, ConsistentID, FGID, ConsistentID

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

arXiv4 repos

arXiv:2404.16994

PLLaVA, pllava-13b, pllava-34b, pllava-7b

Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

arXiv4 repos

arXiv:2405.05945

Lumina-T2X, Lumina-Next-T2I, Lumina-Next-SFT, Lumina-Next-SFT-diffusers

Improved Distribution Matching Distillation for Fast Image Synthesis

arXiv4 repos

arXiv:2405.14867

FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree, DMD2, DMD2, Qwen-Image-Flash

Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability

arXiv4 repos

arXiv:2405.17398

Vista, Vista, hf-example-vista, Vista

MidiCaps: A large-scale MIDI dataset with text captions

arXiv4 repos

arXiv:2406.02255

Text2midi, text2midi, MidiCaps, MidiCaps

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

arXiv4 repos

arXiv:2406.02430

seed-tts-eval, custom-seed-vc, seed-vc-test, maestro-seedvc

QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead

arXiv4 repos

arXiv:2406.03482

QJL, turboquant_plus, rotorquant, quant.cpp

Improving Alignment and Robustness with Circuit Breakers

arXiv4 repos

arXiv:2406.04313

abliterix, Mistral-7B-Instruct-RR-Abliterated, Llama-3-8B-Instruct-RR-Abliterated, circuit-breakers

MoreHopQA: More Than Multi-hop Reasoning

arXiv4 repos

arXiv:2406.13397

morehopqa, morehopqa, GraphKV, GraphKV

MedMNIST-C: Comprehensive benchmark and improved classifier robustness by simulating realistic image corruptions

arXiv4 repos

arXiv:2406.17536

MediMeta-C, RobustMedCLIP, RobustMedCLIP, medmnistc-api

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

arXiv4 repos

arXiv:2406.20076

evf-sam2, evf-sam, EVF-SAM, evf-sam2-multitask

LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

arXiv4 repos

arXiv:2407.03168

FacePoke_CLONE-THIS-REPO-TO-USE-IT, LivePortrait, FacePoke, FLUXSynID

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

arXiv4 repos

arXiv:2407.08136

echomimic, EchoMimic, EchoMimic, EchoMimic

Qwen2-Audio Technical Report

arXiv4 repos

arXiv:2407.10759

Qwen2-Audio-7B, Qwen2-Audio, Speech-IFEval, Qwen2-Audio-7B-Instruct

Multi-label Cluster Discrimination for Visual Representation Learning

arXiv4 repos

arXiv:2407.17331

mlcd-vit-base-patch32-224, unicom, mlcd-vit-large-patch14-336, MLCD-Embodied-7B

LLaVA-OneVision: Easy Visual Task Transfer

arXiv4 repos

arXiv:2408.03326

llava-onevision-qwen2-7b-si, LLaVA-OneVision-Data, llava-onevision-qwen2-7b-ov-hf, llava-onevision-qwen2-7b-ov

Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

arXiv4 repos

arXiv:2408.04594

Img-Diff, data-juicer, SciDataOS, data-juicer

Scalable Autoregressive Image Generation with Mamba

arXiv4 repos

arXiv:2408.12245

AiM, aim-xlarge, aim-base, aim-large

Towards Evaluating and Building Versatile Large Language Models for Medicine

arXiv4 repos

arXiv:2408.12547

MedS-Ins, MMedS-Llama-3-8B, MedS-Ins, MedS-Bench

Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

arXiv4 repos

arXiv:2408.14774

Instruct-SkillMix, Llama-3-8B-Instruct-SkillMix, Instruct-SkillMix-SDD, Instruct-SkillMix-SDA

CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions

arXiv4 repos

arXiv:2408.16589

ASR-Transcription-Router, CrisperWhisper, faster_CrisperWhisper, CrisperWhisper

Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems

arXiv4 repos

arXiv:2409.06656

FluidAudio, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, diar_sortformer_4spk-v1

OneEncoder: A Lightweight Framework for Progressive Alignment of Modalities

arXiv4 repos

arXiv:2409.11059

OneEncoder-text-image-xray, OneEncoder-text-image, OneEncoder-text-image-audio, OneEncoder-text-image-video

Qwen2.5-Coder Technical Report

arXiv4 repos

arXiv:2409.12186

Qwen2.5-Coder-32B-Instruct, Qwen2.5-Coder-14B-Instruct, Qwen2.5-Coder-7B-Instruct, Qwen2.5-Coder-0.5B

Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts

arXiv4 repos

arXiv:2409.16040

TimeMoE-50M, time-moe, TimeMoE-200M, Time-300B

Emu3: Next-Token Prediction is All You Need

arXiv4 repos

arXiv:2409.18869

Emu3, Emu3-Stage1, Emu3-Chat, Emu3-VisionTokenizer

LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding

arXiv4 repos

arXiv:2410.03355

LANTERN, llamagen_drafter, llamagen2_drafter, anole_drafter

Pyramidal Flow Matching for Efficient Video Generative Modeling

arXiv4 repos

arXiv:2410.05954

pyramid-flow-sd3, Pyramid-Flow, pyramid-flow-miniflux, pyramid-flow

Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

arXiv4 repos

arXiv:2410.08261

Meissonic, Monetico, meissonic, test-time-scaling

Teach Multimodal LLMs to Comprehend Electrocardiographic Images

arXiv4 repos

arXiv:2410.19008

ECGBench, ECGInstruct, PULSE, PULSE-7B

Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation

arXiv4 repos

arXiv:2411.02293

Hunyuan3D-2.1, Hunyuan3D-1, HY3D-Bench, Hunyuan3D-Omni

SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models

arXiv4 repos

arXiv:2411.05007

nunchaku, deepcompressor, nunchaku-qwen-image, nunchaku

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

arXiv4 repos

arXiv:2411.11927

FLAME, FLAME-ReCap-CC3M-MiniCPM-Llama3-V-2_5, FLAME-Mistral-Nemo-ViT-B-16-CC3M, FLAME-ReCap-YFCC15M-MiniCPM-Llama3-V-2_5

AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

arXiv4 repos

arXiv:2411.15738

AnyEdit, AnyEdit, AnySD, AnySD

Arctic-Embed 2.0: Multilingual Retrieval Without Compromise

arXiv4 repos

arXiv:2412.04506

cholesky_encoder, snowflake-arctic-embed-l-v2.0, snowflake-arctic-embed-m-v2.0, arctic-embed

MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

arXiv4 repos

arXiv:2412.05237

MAmmoTH-VL-Instruct-12M, MAmmoTH-VL, MAmmoTH-VL-12M, MAmmoTH-VL-8B

ACT-Bench: Towards Action Controllable World Models for Autonomous Driving

arXiv4 repos

arXiv:2412.05337

ACT-Bench, ACT-Estimator, Terra, ACT-Bench

Concept Bottleneck Large Language Models

arXiv4 repos

arXiv:2412.07992

CB-LLMs, Concept-Bottleneck-LLM, Concept-Bottleneck-LLM, CBLLM-PubMed

Offline Reinforcement Learning for LLM Multi-Step Reasoning

arXiv4 repos

arXiv:2412.16145

OREO, OREO, Qwen2.5-Math-1.5B-OREO, Qwen2.5-Math-1.5B-OREO-Value

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

arXiv4 repos

arXiv:2412.18525

Understanding-Vision-Tasks, UVT-Terminological-based-Vision-Tasks, UVT-Explanatory-based-Vision-Tasks, UVT-7B-448

DeepSeek-V3 Technical Report

arXiv4 repos

arXiv:2412.19437

DeepSeek-V3, DeepSeek-V3-0324, maxtext, DualPipe

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

arXiv4 repos

arXiv:2501.01428

GPT4Scene-qwen2vl_full_sft_mark_32_3D_img512, GPT4Scene, GPT4Scene-All, GPT4Scene-and-VLN-R1

BEN: Using Confidence-Guided Matting for Dichotomous Image Segmentation

arXiv4 repos

arXiv:2501.06230

BEN, BEN, BEN2, BEN2

FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

arXiv4 repos

arXiv:2501.14350

FireRedASR2S, FireRedASR, FireRedASR-LLM-L, FireRedASR-AED-L

Sundial: A Family of Highly Capable Time Series Foundation Models

arXiv4 repos

arXiv:2502.00816

sundial-base-128m, timer-base-84m, Large-Time-Series-Model, Sundial

NitiBench: A Comprehensive Study of LLM Framework Capabilities for Thai Legal Question Answering

arXiv4 repos

arXiv:2502.10868

nitibench, nitibench, nitibench-ccl-human-finetuned-bge-m3, nitibench-ccl-auto-finetuned-bge-m3

AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO

arXiv4 repos

arXiv:2502.14669

Maze-Reasoning-Reset-v0.1, AlphaMaze-v0.2-1.5B, Maze-Reasoning-v0.1, Maze-Reasoning-GRPO-v0.1

FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling

arXiv4 repos

arXiv:2502.14856

Qwen2-7B-Instruct-FR-Spec, FR-Spec, LLaMA3.2-Instruct-1B-FR-Spec, LLaMA3-Instruct-8B-FR-Spec

Mantis: Lightweight Foundation Model for Time Series Classification

arXiv4 repos

arXiv:2502.15637

FM4Motor, mantis, MantisV2Experiments, TiViT

Muon is Scalable for LLM Training

arXiv4 repos

arXiv:2502.16982

Moonlight-16B-A3B, Muon, Moonlight-16B-A3B-Instruct, Emerging-Optimizers

RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete

arXiv4 repos

arXiv:2502.21257

RoboBrain, RoboBrain-LoRA-Affordance, RoboBrain-LoRA-Trajectory, RoboBrain2.5

Scaling Rich Style-Prompted Text-to-Speech Datasets

arXiv4 repos

arXiv:2503.04713

paraspeechcaps, paraspeechcaps, parler-tts-mini-v1-paraspeechcaps, parler-tts-mini-v1-paraspeechcaps-only-base

TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

arXiv4 repos

arXiv:2503.11509

AutomaTikZ, DeTikZify, detikzify-v2.5-8b, detikzify-v2-8b

Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

arXiv4 repos

arXiv:2503.14911

MAKE, DermLIP_PanDerm-base-w-PubMed-256, Derm1M, DermLIP_ViT-B-16

RoboEngine: Plug-and-Play Robot Data Augmentation with Semantic Robot Segmentation and Background Generation

arXiv4 repos

arXiv:2503.18738

roboengine, roboengine-bg-diffusion, roboengine-sam, roboseg

I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders

arXiv4 repos

arXiv:2503.18878

SAE-Reasoning, OpenThoughts-10k-DeepSeek-R1, DeepSeek-R1-Distill-Llama-8B-lmsys-openthoughts-tokenized, deepseek-r1-distill-llama-8b-lmsys-openthoughts

AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset

arXiv4 repos

arXiv:2503.19462

AccVideo, AccVideo, AccVideo-WanX-I2V-480P-14B, AccVideo-WanX-T2V-14B

Qwen2.5-Omni Technical Report

arXiv4 repos

arXiv:2503.20215

Qwen2.5-Omni-7B-AWQ, Qwen2.5-Omni-7B-GPTQ-Int4, Qwen2.5-Omni-3B, Qwen2.5-Omni-7B

Unified Multimodal Discrete Diffusion

arXiv4 repos

arXiv:2503.20853

unidisc, unidisc_hq, unidisc_non_interleaved, unidisc_interleaved

Video-R1: Reinforcing Video Reasoning in MLLMs

arXiv4 repos

arXiv:2503.21776

Video-R1, Video-R1-7B, Qwen2.5-VL-7B-COT-SFT, Video-R1-eval

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

arXiv4 repos

arXiv:2503.23377

JavisBench, JavisDiT, JavisDiT-v1.0-jav, JavisGPT-v0.1-7B-Instruct

ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

arXiv4 repos

arXiv:2504.01934

ILLUME_plus, illume_plus-qwen2_5-3b-hf, illume_plus-qwen2_5-7b-hf, dualvitok

LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation

arXiv4 repos

arXiv:2504.07448

LoRI, LoRI-S_safety_llama3_rank_32, LoRI-D_safety_llama3_rank_32, LoRI-D_code_llama3_rank_32

VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

arXiv4 repos

arXiv:2504.08837

ViRL39K, VL-Rethinker, VL-Rethinker-72B, VL-Rethinker-7B

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

arXiv4 repos

arXiv:2504.15279

VisuLogic, VisuLogic-Train, VisuLogic-Eval, VisuLogic-Train

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

arXiv4 repos

arXiv:2504.16030

LiveSports-3K, LiveCC-7B-Instruct, Live-CC-5M, Live-WhisperX-526K

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

arXiv4 repos

arXiv:2504.17343

TimeChat-Online, TimeChat-Online-139K, TimeChatOnline-7B, TimeChat

Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

arXiv4 repos

arXiv:2504.21356

Nexus-GenV2, Nexus-Gen, Nexus-GenV2-nf4-fp8, diffSynth-studio-notes

EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning

arXiv4 repos

arXiv:2505.04623

AVQA-R1-6K, EchoInk, EchoInk-R1-7B, OmniInstruct_V1_AVQA_R1

MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

arXiv4 repos

arXiv:2505.06152

SkinVL-PubMM, MM-Skin, SkinVL-MM, SkinVL-Pub

MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision

arXiv4 repos

arXiv:2505.13427

MM-EUREKA, MM-PRM, MM-PRM, MM-K12

On the Robustness of Medical Vision-Language Models: Are they Truly Generalizable?

arXiv4 repos

arXiv:2505.15425

MediMeta-C, RobustMedCLIP, RobustMedCLIP, medmnistc-api

ViDoRe Benchmark V2: Raising the Bar for Visual Retrieval

arXiv4 repos

arXiv:2505.17166

esg_reports_v2, biomedical_lectures_v2, economics_reports_v2, esg_reports_human_labeled_v2

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

arXiv4 repos

arXiv:2505.17589

CosyVoice, CosyVoice, CV3-Eval, FastCosyVoice

Distilling LLM Agent into Small Models with Retrieval and Code Tools

arXiv4 repos

arXiv:2505.17612

agent-distillation, Qwen2.5-32B-Instruct_agent_trajectories_2k, agent_distilled_Qwen2.5-1.5B-Instruct, Qwen2.5-32B-Instruct_agent_trajectories_2k_prefix

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start

arXiv4 repos

arXiv:2505.22334

Multimodal-Cold-Start, Multimodal-RL-Data, Qwen2.5VL-7b-RLCS, Qwen2.5VL-3b-RLCS

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

arXiv4 repos

arXiv:2505.22618

d3LLM, DIFFA, dllm, Fast-dLLM

Zero-Shot Vision Encoder Grafting via LLM Surrogates

arXiv4 repos

arXiv:2505.22664

zero, zero-model-checkpoints, llava-1.5-665k-instructions, llava-1.5-665k-genqa-500k-instructions

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

arXiv4 repos

arXiv:2506.15442

Hunyuan3D-2.1, Hunyuan3D-2.1, HY3D-Bench, Hunyuan3D-Omni

RLPR: Extrapolating RLVR to General Domains without Verifiers

arXiv4 repos

arXiv:2506.18254

RLPR, RLPR-Evaluation, RLPR-Qwen2.5-7B-Base, RLPR-Train-Dataset

MindCube: Spatial Mental Modeling from Limited Views

arXiv4 repos

arXiv:2506.21458

SpaceQwen2.5-VL-3B-Instruct, SpaceOm, MindCube, MindCube

OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions

arXiv4 repos

arXiv:2506.23361

Open-OmniVCus, OmniVCus, OmniVCus-Test, OmniVCus-Train

VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents

arXiv4 repos

arXiv:2507.04590

VLM2Vec-V2.0, MMEB-V2, MMEB-V3, TARA

Skywork-R1V3 Technical Report

arXiv4 repos

arXiv:2507.06167

Skywork-R1V-38B-AWQ, Skywork-R1V3-38B, Skywork-R1V3-38B-AWQ, Skywork-R1V3-38B-GGUF

Group Sequence Policy Optimization

arXiv4 repos

arXiv:2507.18071

tunix, MMPR-Tiny, DeepVision-103K, DeepVision-103K

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

arXiv4 repos

arXiv:2507.19457

adk-python, dspy, sweet-search, dsp

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

arXiv4 repos

arXiv:2507.21033

GPT-Image-Edit, GPT-Image-Edit-1.5M, gpt-image-edit-training, gpt-image-edit-benchmark-results

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

arXiv4 repos

arXiv:2507.21802

MixGRPO, flow_grpo, MixGRPO, DanceGRPO

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

arXiv4 repos

arXiv:2508.10433

We-Math2.0-Standard, We-Math, We-Math2.0, We-Math2.0-Pro

ToonOut: Fine-tuned Background-Removal for Anime Characters

arXiv4 repos

arXiv:2509.06839

BiRefNet, toonout, BiRefNet, toonout

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning

arXiv4 repos

arXiv:2509.06888

mmBERT-small, mmBERT-base, mmBERT, mmbert-pretrain-p1-fineweb2-langs

Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models

arXiv4 repos

arXiv:2509.06949

TraDo-8B-Instruct, TraDo-8B-Thinking, TraDo-4B-Instruct, dLLM-RL

Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction

arXiv4 repos

arXiv:2509.15202

abliterix, Llama-3-8B-Instruct-DeepRefusal-Broken, DeepRefusal, Meta-Llama-3-8B-Instruct-DeepRefusal

Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning

arXiv4 repos

arXiv:2509.15279

Fleming-VL-8B, Fleming-R1-7B, Fleming-R1-32B, Fleming-VL-38B

UIPro: Unleashing Superior Interaction Capability For GUI Agents

arXiv4 repos

arXiv:2509.17328

UIPro-7B_Stage2_Web, UIPro_1stage, UIPro-7B_Stage2_Mobile, UIPro-7B_Stage1

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

arXiv4 repos

arXiv:2509.21144

nlp_team, nlp_personal, UniSS, UniST

LLaDA-MoE: A Sparse MoE Diffusion Language Model

arXiv4 repos

arXiv:2509.24389

LLaDA-MoE-7B-A1B-Base, LLaDA-MoE-7B-A1B-Instruct, dllm, dLLM_Cache

Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models

arXiv4 repos

arXiv:2509.25826

Kairos_50m, Kairos_23m, Kairos_10m, Kairos

Mem-α: Learning Memory Construction via Reinforcement Learning

arXiv4 repos

arXiv:2509.25911

Mem-alpha, memalphaljx, tmp-mem-alpha, agentic-memory

ModernVBERT: Towards Smaller Visual Document Retrievers

arXiv4 repos

arXiv:2510.01149

ColModernVBERT-CoreAI, modernvbert, modernvbert, modernvbert

SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision

arXiv4 repos

arXiv:2510.02797

SongFormer, SongFormer, SongFormDB, SheetSage2

dInfer: An Efficient Inference Framework for Diffusion Language Models

arXiv4 repos

arXiv:2510.08666

dInfer, dInfer_adaptive, dinfer_power_sampling, TIDE_DATA_COLLECTION

pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

arXiv4 repos

arXiv:2510.14974

LakonLab, pi-Qwen, pi-FLUX.2, pi-FLUX.1

Chronos-2: From Univariate to Universal Forecasting

arXiv4 repos

arXiv:2510.15821

chronos-forecasting, chronos-2, chronos-2-small, chronos-2-synth

DeepSeek-OCR: Contexts Optical Compression

arXiv4 repos

arXiv:2510.18234

DeepSeek-OCR, devanagari-ocr-benchmark, deepseek-ocr-encoder, DeepSeek-OCR-2

Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models

arXiv4 repos

arXiv:2510.21204

mitra-classifier, mitra-classifier-pipeline, mitra-regressor, mitra-regressor-pipeline

FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use

arXiv4 repos

arXiv:2510.24645

AWorld, FunReason-MT, FunReason-MT, AWorld-RL

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

arXiv4 repos

arXiv:2510.25897

miro, miro-ablations, miro, miro

MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

arXiv4 repos

arXiv:2511.09611

MMaDA, MMaDA-Parallel-A, MMaDA-Parallel-M, MMaDA-Parallel

ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering

arXiv4 repos

arXiv:2511.22715

ReAG, ReAG-Critic, ReAG-3B, ReAG-7B

TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

arXiv4 repos

arXiv:2512.14698

TimeLens-100K, TimeLens-8B, TimeLens-Bench, TimeLens-7B

Recursive Language Models

arXiv4 repos

arXiv:2512.24601

LegalRAG, rlm, rlm-claude, memcp

ObjectForesight: Predicting Future 3D Object Trajectories from Human Videos

arXiv4 repos

arXiv:2601.05237

ObjectForesight-Data, ObjectForesight-EPIC-DiT, ObjectForesight, ObjectForesight-HOT3D-DiT

Orient Anything V2: Unifying Orientation and Rotation Understanding

arXiv4 repos

arXiv:2601.05573

OriAnyV2_Train_Render, OriAnyV2_ckpt, Hunyuan3D-FLUX-Gen, OriAnyV2_Inference

Unifying Speech Recognition, Synthesis and Conversion with Autoregressive Transformers

arXiv4 repos

arXiv:2601.10770

GPA, GPA-v1.5, GPA-v1.5-onnx-runtime, GPA

C-RADIOv4 (Tech Report)

arXiv4 repos

arXiv:2601.17237

RADIO, C-RADIOv4-SO400M, C-RADIOv4-H, CRADIOv4

VIBEVOICE-ASR Technical Report

arXiv4 repos

arXiv:2601.18184

VibeVoice, VibeVoice-ASR, VibeVoice-ASR-HF, VibeVoice

Qwen3-ASR Technical Report

arXiv4 repos

arXiv:2601.21337

Qwen3-ASR-1.7B-hf, Qwen3-ForcedAligner-0.6B-hf, Qwen3-ASR-0.6B-hf, Qwen3-ForcedAligner-0.6B

Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling

arXiv4 repos

arXiv:2602.00594

kanade-tokenizer, kanade-12.5hz, kanade-25hz-clean, kanade-25hz

ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation

arXiv4 repos

arXiv:2602.04279

ECG-R1, ECG-R1-8B-RL, ECG-Protocol-Guided-Grounding-CoT, GEM

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

arXiv4 repos

arXiv:2602.08711

TimeChat-Captioner, Timechat-OmniCaptioner-42K, TimeChat-Captioner-GRPO-7B, Timechat-OmniCaptioner-40K

MOVA: Towards Scalable and Synchronized Video-Audio Generation

arXiv4 repos

arXiv:2602.08794

MOVA, MOVA-360p, MOVA-720p, MOVA_benchmark_for_arena

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

arXiv4 repos

arXiv:2602.19163

JavisDiT, AV-DPO, JavisDiT-v1.0-jav, JavisGPT-v1.0-7B-Instruct

A Very Big Video Reasoning Suite

arXiv4 repos

arXiv:2602.20159

VBVR-Wan2.2-diffsynth, VBVR-Wan2.1-diffsynth, VBVR-Dataset, VBVR-LTX2.3-diffsynth

DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

arXiv4 repos

arXiv:2602.24096

nurec-skills, harmonizer, DiffusionHarmonizer, Harmonizer

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

arXiv4 repos

arXiv:2603.02138

OmniLottie, OmniLottie, MMLottieBench, MMLottie-2M

IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse

arXiv4 repos

arXiv:2603.12201

Hy4-preview, Hy4-preview, GLM-5.2, GLM-5.2-FP8

Fast-WAM: Do World Action Models Need Test-time Future Imagination?

arXiv4 repos

arXiv:2603.16666

FastWAM, fastwam, LIBERO-fastwam, robotwin2.0-fastwam

PolarQuant: Optimal Gaussian Weight Quantization via Hadamard Rotation for LLM Compression

arXiv4 repos

arXiv:2603.29078

Qwen3.5-9B-PolarQuant-MLX-4bit, Qwen3.5-9B-PolarEngine-v4, Qwen3.5-9B-PolarQuant-Q5, polarengine-vllm

arXiv:2603.74245

arXiv4 repos

arXiv:2603.74245

Qwen3.5-9B-EOQ-v3, eoq-quantization, Qwen3.5-9B-PolarQuant-MLX-4bit, Qwen3.5-9B-PolarEngine-v4

REAM: Merging Improves Pruning of Experts in LLMs

arXiv4 repos

arXiv:2604.04356

ream, Qwen3-30B-A3B-Instruct-2507-REAM, Qwen3-Next-80B-A3B-Instruct-REAM, GLM-4.5-Air-REAM

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence

arXiv4 repos

arXiv:2604.07296

OpenSpatial-InternVL3-8B, OpenSpatial-InternVL2.5-8B, OpenSpatial-Qwen3-VL-8B, OpenSpatial-Qwen2.5-VL-7B

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

arXiv4 repos

arXiv:2604.08516

MolmoWeb-8B, MolmoWeb-4B, MolmoWeb-8B-Native, MolmoWeb-4B-Native

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion

arXiv4 repos

arXiv:2604.09450

ECHO_Base_block8, ECHO_block8, ECHO_Base_block4, ECHO_block4

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

arXiv4 repos

arXiv:2604.13416

DF3DV, DI2FIX_HF, DF3DV-1K, DF3DV-1K-Fixer

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

arXiv4 repos

arXiv:2604.18468

nurec-skills, asset-harvester, asset-harvester, asset-harvester

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

arXiv4 repos

arXiv:2604.22245

LAT-Bench, LAT-Audio, LAT-Audio-Base, LAT-Chronicle

DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Models

arXiv4 repos

arXiv:2605.07210

diffretriever-dream-7b-single, diffretriever-llada-8b-single, diffretriever-dream-7b-multi-q4-p16, diffretriever-llada-8b-multi-q4-p4

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

arXiv4 repos

arXiv:2605.09266

SwanLab, SeePhy-Pro, PhysRL, SeePhysPro

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

arXiv4 repos

arXiv:2605.11061

HiDream-O1-Image, HiDream-O1-Image, HiDream-O1-Image-Dev, HiDream-O1-Image-Dev-2604

Post-Trained MoE Can Skip Half Experts via Self-Distillation

arXiv4 repos

arXiv:2605.18643

ZEDA, ZEDA, ZEDA-Qwen3-30B-A3B-Dynamic, ZEDA-GLM-4.7-Flash-Dynamic

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

arXiv4 repos

arXiv:2605.20342

ParaVT, ParaVT-Parquet, ParaVT-Source, ParaVT-8B

ETCHR: Editing To Clarify and Harness Reasoning

arXiv4 repos

arXiv:2605.23897

ETCHR-GRPO-10K, ETCHR-FLUX.2-klein-9B, ETCHR-SFT-400K, DL3DV-2k

CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards

arXiv4 repos

arXiv:2606.00020

ChineseErrorCorrector3-4B, ChineseErrorCorrector4-4B, ChineseErrorDetectorElectra, ChineseErrorCorrector4-4B

Unlimited OCR Works

arXiv4 repos

arXiv:2606.23050

Unlimited-OCR, Unlimited-OCR, devanagari-ocr-benchmark, Unlimited-OCR-ROCm

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

arXiv4 repos

arXiv:2607.24904

Mage, mage-7255fc67, Mage-VL-RTSP, Mage-VL

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

arXiv4 repos

arXiv:2609.03796

LLaDA-Image, LLaDA-Image-FP8, LLaDA-Image-Turbo, LLaDA-Image-Turbo-FP8

A foundation model for clinical-grade computational pathology and rare cancers detection

Nature4 repos

Nature:s41591-024-03141-0

PIANO, TRIDENT, TridentEdited, aegis

DeepSpeed

ACM3 repos

ACM:3394486.3406703

DeepSpeed, DeepSpeed, LLMSurvey

Compositional Semantic Parsing on Semi-Structured Tables

arXiv3 repos

arXiv:1508.00305

tapas-base-finetuned-wtq, WikiTableQuestions, BIPIA

Relation Classification via Recurrent Neural Network

arXiv3 repos

arXiv:1508.01006

IEPile, iepie, iepile

TinyLFU: A Highly Efficient Cache Admission Policy

arXiv3 repos

arXiv:1512.00727

ristretto, ristretto, BitFaster.Caching

Improved Techniques for Training GANs

arXiv3 repos

arXiv:1606.03498

materialgan, VideoGPT, stylegan2-ada-pytorch

STAIR Captions: Constructing a Large-Scale Japanese Image Caption Dataset

arXiv3 repos

arXiv:1705.00823

STAIR-Captions, huggingface-datasets_STAIR-Captions, STAIR-captions

GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

arXiv3 repos

arXiv:1706.08500

materialgan, pytorch-fid, stylegan2-ada-pytorch

Progressive Growing of GANs for Improved Quality, Stability, and Variation

arXiv3 repos

arXiv:1710.10196

materialgan, progressive_growing_of_gans, SkinDeep

ArcFace: Additive Angular Margin Loss for Deep Face Recognition

arXiv3 repos

arXiv:1801.07698

Snap-Safe-Python, FLUXSynID, AuraFace-v1

Improving Distantly Supervised Relation Extraction using Word and Entity Based Attention

arXiv3 repos

arXiv:1804.06987

IEPile, iepie, iepile

GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

arXiv3 repos

arXiv:1804.07461

glue, roberta-large-mnli, lares

A Simple Method for Commonsense Reasoning

arXiv3 repos

arXiv:1806.02847

roberta-large-mnli, roberta-base, roberta-large

XNLI: Evaluating Cross-lingual Sentence Representations

arXiv3 repos

arXiv:1809.05053

roberta-large-mnli, Multilingual-MiniLM-L12-H384, XLM

A Style-Based Generator Architecture for Generative Adversarial Networks

arXiv3 repos

arXiv:1812.04948

materialgan, ffhq-dataset, stylegan2-ada-pytorch

IPRE: a Dataset for Inter-Personal Relationship Extraction

arXiv3 repos

arXiv:1907.12801

IEPile, iepie, iepile

FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age

arXiv3 repos

arXiv:1908.04913

clip-vit-base-patch32, clip-vit-base-patch16, vit_large_patch14_clip_224.openai

Text Summarization with Pretrained Encoders

arXiv3 repos

arXiv:1908.08345

HiWestSum, ATS-islamic-organization-news, BertSum

UER: An Open-Source Toolkit for Pre-training Models

arXiv3 repos

arXiv:1909.05658

t5-base-chinese-cluecorpussmall, t5-small-chinese-cluecorpussmall, gpt2-chinese-cluecorpussmall

PubMedQA: A Dataset for Biomedical Research Question Answering

arXiv3 repos

arXiv:1909.06146

PodGPT, RARE, PubMedQA

PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization

arXiv3 repos

arXiv:1912.08777

GPT_Ranker, long-ke-t5, YoYAK

Lung and Colon Cancer Histopathological Image Dataset (LC25000)

arXiv3 repos

arXiv:1912.12142

LC25000-clean, LC25000, lung_colon_image_set

CLUENER2020: Fine-grained Named Entity Recognition Dataset and Benchmark for Chinese

arXiv3 repos

arXiv:2001.04351

IEPile, iepie, iepile

ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators

arXiv3 repos

arXiv:2003.10555

tf-transformers, europeana-bert, YoYAK

Dense Passage Retrieval for Open-Domain Question Answering

arXiv3 repos

arXiv:2004.04906

bge-m3, DPR, odqa_baseline_code

Longformer: The Long-Document Transformer

arXiv3 repos

arXiv:2004.05150

speechless-starcoder2-15b, final-project-level3-nlp-02, YoYAK

End-to-End Object Detection with Transformers

arXiv3 repos

arXiv:2005.12872

detr-resnet-50, coreml-detr-semantic-segmentation, DINO

DocVQA: A Dataset for VQA on Document Images

arXiv3 repos

arXiv:2007.00398

DocVQA, CoExVQA, DocVQA

Relevance-guided Supervision for OpenQA with ColBERT

arXiv3 repos

arXiv:2007.00814

plaidrepro, colbertv2.0, ColBERT

A Large-Scale Chinese Short-Text Conversation Dataset

arXiv3 repos

arXiv:2008.03946

CDial-GPT_LCCC-base, CDial-GPT_LCCC-large, mmchat

RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

arXiv3 repos

arXiv:2009.11462

real-toxicity-prompts, real-toxicity-prompts, causal_unlearn_llm

Validating UTF-8 In Less Than One Instruction Per Byte

arXiv3 repos

arXiv:2010.03090

simdjson, simdutf, fastvalidate-utf-8

arXiv:2010.05171

arXiv3 repos

arXiv:2010.05171

TIL-2023, wav2vec2-conformer-rel-pos-large-960h-ft, s2t-small-librispeech-asr

mT5: A massively multilingual pre-trained text-to-text transformer

arXiv3 repos

arXiv:2010.11934

mt5-base, mt5-xxl, tf-transformers

Scaled-YOLOv4: Scaling Cross Stage Partial Network

arXiv3 repos

arXiv:2011.08036

darknet, darknet, nanodet

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

arXiv3 repos

arXiv:2101.00204

xnli_bn, squad_bn, banglishbert

Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval

arXiv3 repos

arXiv:2101.00436

plaidrepro, colbertv2.0, ColBERT

Number Parsing at a Gigabyte per Second

arXiv3 repos

arXiv:2101.11408

ffc.h, fast_float, csFastFloat

ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision

arXiv3 repos

arXiv:2102.03334

vilt-b32-finetuned-vqa, ViLT, visual-spatial-reasoning

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

arXiv3 repos

arXiv:2102.08981

dalle-mini, cc12m-wds, conceptual-12m

WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

arXiv3 repos

arXiv:2103.01913

clip-italian, clip-italian, wit

Vision Transformers for Dense Prediction

arXiv3 repos

arXiv:2103.13413

Depth-Estimation, MiDaS, ldm3d-4c

Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

arXiv3 repos

arXiv:2103.14030

FoodSeg103-Benchmark-v1, MiDaS, Swin-Transformer

Towards Measuring Fairness in AI: the Casual Conversations Dataset

arXiv3 repos

arXiv:2104.02821

parakeet-tdt_ctc-1.1b, parakeet-tdt_ctc-110m, parakeet-tdt-1.1b

AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

arXiv3 repos

arXiv:2104.03603

pyannote-audio, 3D-Speaker, 3d-speaker

A Reinforcement Learning Environment For Job-Shop Scheduling

arXiv3 repos

arXiv:2104.03760

JSSEnv, RL-Job-Shop-Scheduling, JSS

Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling

arXiv3 repos

arXiv:2104.06967

splade, pyterrier_dr, pyterrier_dr_jpq

Sample and Computation Redistribution for Efficient Face Detection

arXiv3 repos

arXiv:2105.04714

insightface, face-alignment, Snap-Safe-Python

CogView: Mastering Text-to-Image Generation via Transformers

arXiv3 repos

arXiv:2105.13290

visualglm-6b, VisualGLM-6B, VisualGLM-6B

Structured Denoising Diffusion Models in Discrete State-Spaces

arXiv3 repos

arXiv:2107.03006

dlms-sinks, mdlm, minimal-dlm

Per-Pixel Classification is Not All You Need for Semantic Segmentation

arXiv3 repos

arXiv:2107.06278

mask2former-swin-large-coco-panoptic, maskformer-swin-small-coco, MaskFormer

Contrastive Language-Image Pre-training for the Italian Language

arXiv3 repos

arXiv:2108.08688

clip-italian, clip-italian, clip-italian-demo

Mr. TyDi: A Multi-lingual Benchmark for Dense Retrieval

arXiv3 repos

arXiv:2108.08787

multilingual-e5-base, multilingual-e5-large, KoPrivateGPT

CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation

arXiv3 repos

arXiv:2109.00859

CodeRL, codet5-large-ntp-py, codet5-large

Finetuned Language Models Are Zero-Shot Learners

arXiv3 repos

arXiv:2109.01652

FLAN, promptsource, flan

TruthfulQA: Measuring How Models Mimic Human Falsehoods

arXiv3 repos

arXiv:2109.07958

truthful_qa, RAIN, truthful_qa_de

TorchXRayVision: A library of chest X-ray datasets and models

arXiv3 repos

arXiv:2111.00595

torchxrayvision, densenet121-res224-chex, torchxrayvision

GMFlow: Learning Optical Flow via Global Matching

arXiv3 repos

arXiv:2111.13680

unimatch, gmflow, prisma

Masked-attention Mask Transformer for Universal Image Segmentation

arXiv3 repos

arXiv:2112.01527

mask2former-swin-large-coco-panoptic, Mask2Former, Food_Calories

Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

arXiv3 repos

arXiv:2201.02177

grokfast, grokking, grok

A ConvNet for the 2020s

arXiv3 repos

arXiv:2201.03545

convnext_perceptual_loss, ConvNeXt, CLIP-convnext_large_d_320.laion2B-s29B-b131K-ft-soup

SGPT: GPT Sentence Embeddings for Semantic Search

arXiv3 repos

arXiv:2202.08904

sgpt-bloom-7b1-msmarco, sgpt, sgpt-bloom-1b7-nli

iSTFTNet: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform

arXiv3 repos

arXiv:2203.02395

Kokoro-82M, kokoro-82M-onnx-opt, fish-diffusion

Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

arXiv3 repos

arXiv:2203.05482

Twin-Merging, Mario, MergeLM

BERTopic: Neural topic modeling with a class-based TF-IDF procedure

arXiv3 repos

arXiv:2203.05794

turftopic, turkish-complaint-topic-clustering, BERTopic

ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

arXiv3 repos

arXiv:2203.09509

toxigen-data, toxigen, toxigen-data

Training Compute-Optimal Large Language Models

arXiv3 repos

arXiv:2203.15556

awesome-totally-open-chatgpt, llama2.c, falcon-refinedweb

PaLM: Scaling Language Modeling with Pathways

arXiv3 repos

arXiv:2204.02311

idefics-80b-instruct, transformer-tricks, idefics-9b-instruct

KOBEST: Korean Balanced Evaluation of Significant Tasks

arXiv3 repos

arXiv:2204.04541

polyglot-ko-1.3b, polyglot-ko-3.8b, polyglot-ko-5.8b

The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink

arXiv3 repos

arXiv:2204.05149

Llama-3.1-AlternateTokenizer, Meta-Llama-3.1-8B-bnb-4bit, Meta-Llama-3.1-8B-llamafile

Hierarchical Text-Conditional Image Generation with CLIP Latents

arXiv3 repos

arXiv:2204.06125

coyo-dataset, coyo-700m, BDM1.0

mGPT: Few-Shot Learners Go Multilingual

arXiv3 repos

arXiv:2204.07580

mGPT, mgpt, mGPT

Visual Spatial Reasoning

arXiv3 repos

arXiv:2205.00363

visual-spatial-reasoning, vsr_random, vsr_zeroshot

UL2: Unifying Language Learning Paradigms

arXiv3 repos

arXiv:2205.05131

turkish-bert, bert5urk, flan-ul2

Vectorized and performance-portable Quicksort

arXiv3 repos

arXiv:2205.05982

node, highway, gecko-dev

Pretraining is All You Need for Image-to-Image Translation

arXiv3 repos

arXiv:2205.12952

Stable-Diffusion-FineTuned-zh-v1, Stable-Diffusion-FineTuned-zh-v2, Stable-Diffusion-FineTuned-zh-v0

DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps

arXiv3 repos

arXiv:2206.00927

PixArt-alpha, pixeart, PixArt-alpha

Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation

arXiv3 repos

arXiv:2206.02777

D-FINE-seg, D-FINE-seg, DINO

CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning

arXiv3 repos

arXiv:2207.01780

CodeRL, codet5-large-ntp-py, codet5-large

YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

arXiv3 repos

arXiv:2207.02696

darknet, darknetcv, darknet

arXiv:2207.05987

arXiv3 repos

arXiv:2207.05987

docprompting, tldr, docprompting-conala

Efficient Training of Language Models to Fill in the Middle

arXiv3 repos

arXiv:2207.14255

Qwen2.5-Coder, Qwen3-Coder, speechless-starcoder2-15b

Twitter Topic Classification

arXiv3 repos

arXiv:2209.09824

tweetnlp, tweet_topic_single, tweet_topic_multi

Human Motion Diffusion Model

arXiv3 repos

arXiv:2209.14916

a-mdm-text-to-motion, a-motion-diffusion, motion-diffusion-model

Named Entity Recognition in Twitter: A Dataset and Analysis on Short-Term Temporal Shifts

arXiv3 repos

arXiv:2210.03797

tweetnlp, tweetner7, tner

MedCLIP: Contrastive Learning from Unpaired Medical Images and Text

arXiv3 repos

arXiv:2210.10163

MediMeta-C, RobustMedCLIP, RobustMedCLIP

ESB: A Benchmark For Multi-Domain End-to-End Speech Recognition

arXiv3 repos

arXiv:2210.13352

distil-large-v2, distil-medium.en, distil-small.en

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

arXiv3 repos

arXiv:2211.05100

api-for-open-llm, bloom, GlorIA

Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation

arXiv3 repos

arXiv:2211.12572

diff-mining, plug-and-play, PnP-diffusion-features

InternVideo: General Video Foundation Models via Generative and Discriminative Learning

arXiv3 repos

arXiv:2212.03191

InternVideo, InternVid-Full, internvideo-d2a11ea9

Editing Models with Task Arithmetic

arXiv3 repos

arXiv:2212.04089

Twin-Merging, Mario, MergeLM

Reproducible scaling laws for contrastive language-image learning

arXiv3 repos

arXiv:2212.07143

CLIP_benchmark, CLIP-ViT-bigG-14-laion2B-39B-b160k, TiViT

Objaverse: A Universe of Annotated 3D Objects

arXiv3 repos

arXiv:2212.08051

Cap3D, objaverse, objaverse-xl

Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

arXiv3 repos

arXiv:2212.09689

COIG, COIG, unnatural-instructions

SODA: Million-scale Dialogue Distillation with Social Commonsense Contextualization

arXiv3 repos

arXiv:2212.10465

soda, sodaverse, cosmo-xl

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

arXiv3 repos

arXiv:2301.12503

AudioLDM-S-Full, MMDisCo, audioldm_eval

Accelerating Large Language Model Decoding with Speculative Sampling

arXiv3 repos

arXiv:2302.01318

LLM-Sampling, LLMSpeculativeSampling, RemoteSpeculativeDecoding

MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

arXiv3 repos

arXiv:2302.08113

MultiDiffusion, MultiDiffusion, multidiffusion-region-based

UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers

arXiv3 repos

arXiv:2303.00807

ColBERT, RAG, RAGatouille

Consistency Models

arXiv3 repos

arXiv:2303.01469

TCD, TCD-SDXL-LoRA, LakonLab

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

arXiv3 repos

arXiv:2303.04671

TaskMatrix, visual-chatgpt-zh, TaskMatrix

Tag2Text: Guiding Vision-Language Model via Image Tagging

arXiv3 repos

arXiv:2303.05657

recognize-anything, recognize-anything-plus-model, recognize_anything_model

Erasing Concepts from Diffusion Models

arXiv3 repos

arXiv:2303.07345

erasing, Erasing-Concepts-In-Diffusion, leco

TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

arXiv3 repos

arXiv:2303.11897

tifa, llama2_tifa_question_generation, banana100-additional-iqa-models

Capabilities of GPT-4 on Medical Challenge Problems

arXiv3 repos

arXiv:2303.13375

Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, med42

Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

arXiv3 repos

arXiv:2303.13439

Text2Video-Zero, Text2Video-Zero, Text2Video-Zero

G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

arXiv3 repos

arXiv:2303.16634

RLHF-Korean-Friendly-LLM, KULLM-RLHF, level3_nlp_finalproject-nlp-12

RRHF: Rank Responses to Align Language Models with Human Feedback without tears

arXiv3 repos

arXiv:2304.05302

wombat-7b-gpt4-delta, RRHF, wombat-7b-delta

ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation

arXiv3 repos

arXiv:2304.05977

ImageReward, ImageRewardDB, banana100-additional-iqa-models

Efficient Sequence Transduction by Jointly Predicting Tokens and Durations

arXiv3 repos

arXiv:2304.06795

parakeet-tdt_ctc-1.1b, parakeet-tdt_ctc-110m, parakeet-tdt-1.1b

InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction

arXiv3 repos

arXiv:2304.08085

IEPile, iepie, iepile

FindVehicle and VehicleFinder: A NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system

arXiv3 repos

arXiv:2304.10893

IEPile, iepie, iepile

C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

arXiv3 repos

arXiv:2305.08322

Qwen-14B-Chat, Qwen-7B-Chat, Qwen-1_8B

Common Diffusion Noise Schedules and Sample Steps are Flawed

arXiv3 repos

arXiv:2305.08891

OneTrainer, smalldiffusion, YetAnotherStableDiffusion

Towards Expert-Level Medical Question Answering with Large Language Models

arXiv3 repos

arXiv:2305.09617

Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B, med42

SoundStorm: Efficient Parallel Audio Generation

arXiv3 repos

arXiv:2305.09636

Dia-1.6B-0626, dia, soundstorm-speechtokenizer

Drag Your GAN: Interactive Point-based Manipulation on the Generative Image Manifold

arXiv3 repos

arXiv:2305.10973

DragonDiffusion, InternGPT, InternChat

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

arXiv3 repos

arXiv:2305.12182

glot500-base, Glot500, Glot500

GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

arXiv3 repos

arXiv:2305.13245

tiny-vllm, speechless-starcoder2-15b, llama2.zig

FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation

arXiv3 repos

arXiv:2305.14251

KoHalluLens, HalluLens, llm_factuality_tuning

arXiv:2305.14387

arXiv3 repos

arXiv:2305.14387

alpaca_eval, alpaca_farm, JudgeBench

Calc-X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems

arXiv3 repos

arXiv:2305.15017

zephyr-7b-sft-full124, zephyr-7b-sft-full124_d270, calc-x

HuatuoGPT, towards Taming Language Model to Be a Doctor

arXiv3 repos

arXiv:2305.15075

HuatuoGPT2-SFT-GPT4-140K, HuatuoGPT2-Pretraining-Instruction, HuatuoGPT

Inference-Time Intervention: Eliciting Truthful Answers from a Language Model

arXiv3 repos

arXiv:2306.03341

honest_llama, honest_llama2_chat_7B, 2023FallNLP

Recognize Anything: A Strong Image Tagging Model

arXiv3 repos

arXiv:2306.03514

recognize-anything, recognize-anything-plus-model, recognize_anything_model

PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

arXiv3 repos

arXiv:2306.05087

JudgeBench, Alpaca-7B-v1, PandaLM

Fast Segment Anything

arXiv3 repos

arXiv:2306.12156

sd-webui-inpaint-anything, FastSAM, sd-webui-inpaint-anything

arXiv:2306.14824

arXiv3 repos

arXiv:2306.14824

vqazero, Zero-and-Few-Shot-Visual-Question-Answering, GRIT

LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

arXiv3 repos

arXiv:2306.17107

LLaVAR, LLaVAR_delta, LLaVAR

Provable Robust Watermarking for AI-Generated Text

arXiv3 repos

arXiv:2306.17439

Adversarial-Paraphrasing, impossibility-watermark, lm-watermarking

JourneyDB: A Benchmark for Generative Image Understanding

arXiv3 repos

arXiv:2307.00716

LaVi-Bridge, SEED-Data-Edit-Part2-3, SEED-Data-Edit

Flacuna: Unleashing the Problem Solving Power of Vicuna using FLAN Fine-Tuning

arXiv3 repos

arXiv:2307.02053

flacuna-13b-v1.0, flacuna, flan-mini

SVIT: Scaling up Visual Instruction Tuning

arXiv3 repos

arXiv:2307.04087

Bunny-v1_1-data, Bunny, Bunny-v1_0-data

AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

arXiv3 repos

arXiv:2307.04725

LatentSync-1.5, mcm, MMDisCo

Objaverse-XL: A Universe of 10M+ 3D Objects

arXiv3 repos

arXiv:2307.05663

Cap3D, objaverse-xl, objaverse-xl

MMBench: Is Your Multi-modal Model an All-around Player?

arXiv3 repos

arXiv:2307.06281

idefics-80b-instruct, MMBench, idefics-9b-instruct

Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution

arXiv3 repos

arXiv:2307.06304

Open-Sora-Plan-v1.2.0, Baichuan-Omni-1.5, idefics2-8b

mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs

arXiv3 repos

arXiv:2307.06930

mBLIP, mblip-mt0-xl, mblip-bloomz-7b

QuIP: 2-Bit Quantization of Large Language Models With Guarantees

arXiv3 repos

arXiv:2307.13304

smash, llmtools, llmtools

Med-Flamingo: a Multimodal Medical Few-shot Learner

arXiv3 repos

arXiv:2307.15189

med-flamingo, med-flamingo, med-flamingo

GEMRec: Towards Generative Model Recommendation

arXiv3 repos

arXiv:2308.02205

GEMRec, GEMRec-Roster, GEMRec-PromptBook

UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition

arXiv3 repos

arXiv:2308.03279

universal-ner, UniNER-7B-type-sup, UniNER-7B-all

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

arXiv3 repos

arXiv:2308.03825

SecLists, jailbreak_llms, JailbreakRadar

OctoPack: Instruction Tuning Code Large Language Models

arXiv3 repos

arXiv:2308.07124

reward-bench, bigcode-evaluation-harness, humanevalpack

BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine

arXiv3 repos

arXiv:2308.09442

BioMedGPT-LM-7B, OpenBioMed, OpenBioMed_new

ChatHaruhi: Reviving Anime Character in Reality via Large Language Model

arXiv3 repos

arXiv:2308.09597

Chat-Haruhi-Suzumiya, ChatHaruhi-Expand-118K, ChatHaruhi-54K-Role-Playing-Dialogue

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

arXiv3 repos

arXiv:2308.11596

dissertation-project, seamless-m4t-medium, seamless-m4t-large

MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

arXiv3 repos

arXiv:2309.07915

MIC_full, MIC, MMICL-Instructblip-T5-xxl

LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

arXiv3 repos

arXiv:2309.15103

LaVie, LaVie, videophy

arXiv:2310.01218

arXiv3 repos

arXiv:2310.01218

CoBSAT, SEED, SEED

OceanGPT: A Large Language Model for Ocean Science Tasks

arXiv3 repos

arXiv:2310.02031

OceanGPT-7b, OceanBench, OceanGPT

NEFTune: Noisy Embeddings Improve Instruction Finetuning

arXiv3 repos

arXiv:2310.05914

shisa-7b-v1, Arithmo, Arithmo-Mistral-7B

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

arXiv3 repos

arXiv:2310.06770

opensre, oh-my-claudecode, SWEBench-verified-mini

Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting

arXiv3 repos

arXiv:2310.08278

moment, Lag-Llama, lag-llama

NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails

arXiv3 repos

arXiv:2310.10501

NeMo-Guardrails, Guardrails, Agent-Guard

Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

arXiv3 repos

arXiv:2310.11441

roborazzi, Magma-8B, hacktech24_app_testing

Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation

arXiv3 repos

arXiv:2310.12371

diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, diar_sortformer_4spk-v1

AgentTuning: Enabling Generalized Agent Abilities for LLMs

arXiv3 repos

arXiv:2310.12823

agentlm-7b, agentlm-13b, agentlm-70b

DPM-Solver-v3: Improved Diffusion ODE Solver with Empirical Model Statistics

arXiv3 repos

arXiv:2310.13268

PixArt-alpha, pixeart, PixArt-alpha

Zephyr: Direct Distillation of LM Alignment

arXiv3 repos

arXiv:2310.16944

zephyr-7b-alpha, zephyr-7b-beta, zephyr-7b-gemma-v0.1

Jina Embeddings 2: 8192-Token General-Purpose Text Embeddings for Long Documents

arXiv3 repos

arXiv:2310.19923

late-chunking, mteb-1.34.14, jina-colbert-v1-en

Instruction-Following Evaluation for Large Language Models

arXiv3 repos

arXiv:2311.07911

starchat2-15b-v0.1, llm-action, IFEval

MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

arXiv3 repos

arXiv:2311.16079

Master-Thesis, guidelines, meditron

RETSim: Resilient and Efficient Text Similarity

arXiv3 repos

arXiv:2311.17264

text-dedup, usearch, unisim

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

arXiv3 repos

arXiv:2312.00752

Vim, mamba2-minimal, mdlm

Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

arXiv3 repos

arXiv:2312.11456

Online-RLHF, FsfairX-LLaMA3-RM-v0.1, LLaMA3-iterative-DPO-final

Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions

arXiv3 repos

arXiv:2312.12450

EditPackFT, CanItEdit, CanItEdit

YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction

arXiv3 repos

arXiv:2312.15548

IEPile, iepie, iepile

DB-GPT: Empowering Database Interactions with Private Large Language Models

arXiv3 repos

arXiv:2312.17449

DB-GPT, DB-GPT, pathtraversal_mutation_352

Long Context Compression with Activation Beacon

arXiv3 repos

arXiv:2401.03462

bge-large-zh-v1.5, bge-small-en-v1.5, bge-large-en-v1.5

Mixtral of Experts

arXiv3 repos

arXiv:2401.04088

Awesome-AITools, minimind, Swallow-MX-8x7b-NVE-v0.1

PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models

arXiv3 repos

arXiv:2401.05252

PixArt-alpha, pixeart, PixArt-alpha

VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models

arXiv3 repos

arXiv:2401.09047

VideoCrafter, videophy, MMDisCo

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

arXiv3 repos

arXiv:2401.10935

ScreenSpot, SeeClick, transformer-final-proj

Mercury: A Code Efficiency Benchmark for Code Large Language Models

arXiv3 repos

arXiv:2402.07844

Mercury, Mercury, Venus

CoLLaVO: Crayon Large Language and Vision mOdel

arXiv3 repos

arXiv:2402.11248

CoLLaVO, CoLLaVO-7B, MoAI

Browse and Concentrate: Comprehending Multimodal Content via prior-LLM Context Fusion

arXiv3 repos

arXiv:2402.12195

Brote-pretrain, Brote, Brote-IM-XXL

arXiv:2402.12354

arXiv3 repos

arXiv:2402.12354

llama3-chinese, llama3-chinese, Llama3-Chinese-Lora

FinBen: A Holistic Financial Benchmark for Large Language Models

arXiv3 repos

arXiv:2402.12659

PIXIU, flare-fomc, flare-finer-ord

Training-Free Long-Context Scaling of Large Language Models

arXiv3 repos

arXiv:2402.17463

Qwen3-235B-A22B-Thinking-2507, Qwen3-235B-A22B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507

ShapeLLM: Universal 3D Object Understanding for Embodied Interaction

arXiv3 repos

arXiv:2402.17766

ShapeLLM, MiniGPT-3D, ReConV2

TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space

arXiv3 repos

arXiv:2402.17811

TruthX, TruthX, Llama-2-7b-chat-TruthX

OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

arXiv3 repos

arXiv:2403.01779

MagicClothing, OOTDiffusion, OOTDiffusion

TripoSR: Fast 3D Object Reconstruction from a Single Image

arXiv3 repos

arXiv:2403.02151

TripoSR, TripoSR, TriPlaneCLIP

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

arXiv3 repos

arXiv:2403.03206

F5-TTS, maxdiffusion, Rectified-Diffusion

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

arXiv3 repos

arXiv:2403.04132

FastChat, multilingual_mt_bench, lmsys-arena-human-preference-55k

TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

arXiv3 repos

arXiv:2403.04473

Monkey, MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash

Common 7B Language Models Already Possess Strong Math Capabilities

arXiv3 repos

arXiv:2403.04706

Xwin-LM, Xwin-Math-7B-V1.1, Xwin-Math-70B-V1.1

Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head

arXiv3 repos

arXiv:2403.06892

OmDet, omdet-turbo-swin-tiny-hf, OmDet-Turbo_tiny_SWIN_T

MoAI: Mixture of All Intelligence for Large Language and Vision Models

arXiv3 repos

arXiv:2403.07508

CoLLaVO, MoAI-7B, MoAI

Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation

arXiv3 repos

arXiv:2403.07860

ELLA, LaVi-Bridge, ComfyUI-ELLA-wrapper

Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking

arXiv3 repos

arXiv:2403.09629

quiet-star, quietstar-8-ahead, Adaptive_QuietSTaR

Arc2Face: A Foundation Model for ID-Consistent Human Faces

arXiv3 repos

arXiv:2403.11641

FLUXSynID, Arc2Face, Arc2Face

Generic 3D Diffusion Adapter Using Controlled Multi-View Editing

arXiv3 repos

arXiv:2403.12032

MVEdit, MVEdit, 3D-Adapter

arXiv:2403.13787

arXiv3 repos

arXiv:2403.13787

MM-Eval, prometheus-eval, reward-bench

InstantSplat: Sparse-view Gaussian Splatting in Seconds

arXiv3 repos

arXiv:2403.20309

InstantSplatPP, InstantSplat, InstantSplat

Evaluating Text-to-Visual Generation with Image-to-Text Generation

arXiv3 repos

arXiv:2404.01291

t2v_metrics, clip-flant5-xxl, banana100-additional-iqa-models

CosmicMan: A Text-to-Image Foundation Model for Humans

arXiv3 repos

arXiv:2404.01294

CosmicMan, CosmicMan-SD, CosmicMan-SDXL

Long-context LLMs Struggle with Long In-context Learning

arXiv3 repos

arXiv:2404.02060

CAG, LongICLBench, LongICLBench

JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

arXiv3 repos

arXiv:2404.03027

JailBreakV_28K, JailBreakV_28K, JailBreakV_28K

Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

arXiv3 repos

arXiv:2404.05892

VisualRWKV, rwkv-6-world, rwkv

ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback

arXiv3 repos

arXiv:2404.07987

ControlNet_Plus_Plus, MultiGen-20M_train, flownetpp

Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

arXiv3 repos

arXiv:2404.09956

tango, tango2, tango2-full

MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents

arXiv3 repos

arXiv:2404.10774

MiniCheck, C2D-and-D2C-MiniCheck, Bespoke-MiniCheck-7B

Stepwise Alignment for Constrained Language Model Policy Optimization

arXiv3 repos

arXiv:2404.11049

sacpo, sacpo, p-sacpo

BLINK: Multimodal Large Language Models Can See but Not Perceive

arXiv3 repos

arXiv:2404.12390

BLINK, BLINK_Benchmark, BLINK-ja

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

arXiv3 repos

arXiv:2404.14219

fineweb-edu, Phi-3.5-MoE-instruct, Phi-3.5-vision-instruct

PuLID: Pure and Lightning ID Customization via Contrastive Alignment

arXiv3 repos

arXiv:2404.16022

PuLID, PuLID, FLUXSynID

OpenStreetView-5M: The Many Roads to Global Visual Geolocation

arXiv3 repos

arXiv:2404.18873

osv5m, plonk, baseline

TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

arXiv3 repos

arXiv:2404.19205

granite-4.0-3b-vision, tablevqabench, sa2va_eval

QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving

arXiv3 repos

arXiv:2405.04532

nunchaku, nunchaku, deepcompressor

BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

arXiv3 repos

arXiv:2405.14908

data-juicer, SciDataOS, data-juicer

DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ

arXiv3 repos

arXiv:2405.15306

AutomaTikZ, DeTikZify, detikzify-v2.5-8b

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

arXiv3 repos

arXiv:2405.17220

vision-feedback-mix-binarized, vision-feedback-mix-binarized, OmniLMM

arXiv:2406.00770

arXiv3 repos

arXiv:2406.00770

EvolKit, aurora-m2, slm-innovator-lab

Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

arXiv3 repos

arXiv:2406.00977

Llama-3.1-8B-Dragonfly-Med-v2, Dragonfly, Llama-3.1-8B-Dragonfly-v2

Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration

arXiv3 repos

arXiv:2406.01014

ms-agent, MobileAgent, modelscope-agent

LoFiT: Localized Fine-tuning on LLM Representations

arXiv3 repos

arXiv:2406.01563

lo-fit, llama2_7B_base_lofit_mquake, llama2_7B_base_lofit_truthfulqa

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

arXiv3 repos

arXiv:2406.01574

MMLU-Pro, MMLU-Pro, RLPR-Evaluation

UltraMedical: Building Specialized Generalists in Biomedicine

arXiv3 repos

arXiv:2406.03949

Llama-3-70B-UltraMedical, Llama-3.1-8B-UltraMedical, UltraMedical-Preference

arXiv:2406.04127

arXiv3 repos

arXiv:2406.04127

ZeroEval, mmlu-redux, mmlu-redux

MLVU: Benchmarking Multi-task Long Video Understanding

arXiv3 repos

arXiv:2406.04264

MLVU, UniG2U, lmms-eval-mmllm

The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

arXiv3 repos

arXiv:2406.05761

prometheus-eval, BiGGen-Bench-Results, BiGGen-Bench

Scaling up masked audio encoder learning for general audio classification

arXiv3 repos

arXiv:2406.06992

dasheng, Dasheng, hashing-baseline

MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models

arXiv3 repos

arXiv:2406.07594

MLLMGuard, MLLMGuard, MLLMGuard

LVBench: An Extreme Long Video Understanding Benchmark

arXiv3 repos

arXiv:2406.08035

LVBench, LVBench, LVBench

GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

arXiv3 repos

arXiv:2406.08451

GUI-Odyssey, GUI-Odyssey, GUIOdyssey

Interpreting the Weight Space of Customized Diffusion Models

arXiv3 repos

arXiv:2406.09413

weights2weights, weights2weights, weights2weights

RobustSAM: Segment Anything Robustly on Degraded Images

arXiv3 repos

arXiv:2406.09627

robustsam-vit-large, robustsam-vit-huge, robustsam-vit-base

On the Impacts of Contexts on Repository-Level Code Generation

arXiv3 repos

arXiv:2406.11927

RepoExec, RepoExec, RepoExec-Instruct

arXiv:2406.11939

arXiv3 repos

arXiv:2406.11939

arena-hard-auto, PPE, JudgeBench

CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets

arXiv3 repos

arXiv:2406.13897

Step1X-3D, Hunyuan3D-Omni, UniTEX

ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning

arXiv3 repos

arXiv:2406.14130

ExVideo-CogVideoX-LoRA-129f-v1, ExVideo-SVD-128f-v1, diffSynth-studio-notes

OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer

arXiv3 repos

arXiv:2406.16620

omchat-v2.0-13B-single-beta_hf, omchat, OmAgent

MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?

arXiv3 repos

arXiv:2406.17806

MOSSBench, MOSSBench, MOSSBench

Nomic Embed Vision: Expanding the Latent Space

arXiv3 repos

arXiv:2406.18587

contrastors, nomic-embed-vision-v1.5, contrastors

ProgressGym: Alignment with a Millennium of Moral Progress

arXiv3 repos

arXiv:2406.20087

ProgressGym, ProgressGym-TimelessQA, ProgressGym-HistText

Scaling Synthetic Data Creation with 1,000,000,000 Personas

arXiv3 repos

arXiv:2406.20094

camel, PersonaHub, aurora-m2

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

arXiv3 repos

arXiv:2407.01284

We-Math, We-Math, We-Math2.0

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

arXiv3 repos

arXiv:2407.02371

OpenVid-1M, OpenVid-1M, OpenVid-1M-mapping

Crafting Large Language Models for Enhanced Interpretability

arXiv3 repos

arXiv:2407.04307

CB-LLMs, ConceptBottleneck-GUI-Experiment, Concept-Bottleneck-LLM

The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective

arXiv3 repos

arXiv:2407.08583

data-juicer, SciDataOS, data-juicer

Robotic Control via Embodied Chain-of-Thought Reasoning

arXiv3 repos

arXiv:2407.08693

embodied-CoT, Adaptive-CoT-in-VLA, Fast-ECoT

Panacea: A foundation model for clinical trial search, summarization, design, and recruitment

arXiv3 repos

arXiv:2407.11007

Panacea, Panacea-7B-Chat, TrialAlign

Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development

arXiv3 repos

arXiv:2407.11784

data-juicer, SciDataOS, data-juicer

Sentiment Reasoning for Healthcare

arXiv3 repos

arXiv:2407.21054

Sentiment-Reasoning, Sentiment-Reasoning, Sentiment-Reasoning

OmniParser for Pure Vision Based GUI Agent

arXiv3 repos

arXiv:2408.00203

OmniParser, OmniParser-v2.0, OmniParser

Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

arXiv3 repos

arXiv:2408.02034

MiniMonkey, Monkey, MiniMokney

VidGen-1M: A Large-Scale Dataset for Text-to-video Generation

arXiv3 repos

arXiv:2408.02629

VIDGEN-1M, VidGen, VIDGEN-v1.0

Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

arXiv3 repos

arXiv:2408.05147

gemma-scope-9b-pt-res, gemma-scope, gemma-scope-2b-pt-transcoders

LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs

arXiv3 repos

arXiv:2408.07055

LongWriter-llama3.1-8b, LongWriter-6k, LongWriter-glm4-9b

Docling Technical Report

arXiv3 repos

arXiv:2408.09869

docling, docling, my-docling

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

arXiv3 repos

arXiv:2408.12528

Show-o, show-o-w-clip-vit, show-o

NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks

arXiv3 repos

arXiv:2408.13106

diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1, diar_sortformer_4spk-v1

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models

arXiv3 repos

arXiv:2408.15556

HR-Bench, HR-Bench, sa2va_eval

Target-Driven Distillation: Consistency Distillation with Target Timestep Selection and Decoupled Guidance

arXiv3 repos

arXiv:2409.01347

Target-Driven-Distillation, TDD, Target-Driven-Distillation

LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA

arXiv3 repos

arXiv:2409.02897

LongCite-glm4-9b, LongCite-llama3.1-8b, LongCite-45k

Synthetic continued pretraining

arXiv3 repos

arXiv:2409.07431

Synthetic_Continued_Pretraining, entigraph-quality-corpus, llama-3-8b-entigraph-quality

Phikon-v2, A large and public feature extractor for biomarker prediction

arXiv3 repos

arXiv:2409.09173

PIANO, phikon-v2, SEAL

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

arXiv3 repos

arXiv:2409.12106

ValueLlama-3-8B, gpv, gpv

DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

arXiv3 repos

arXiv:2409.20007

DeSTA2, DeSTA2-8B-beta, Speech-IFEval

Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers

arXiv3 repos

arXiv:2409.20537

HPT, hpt-base, hpt_locoman

Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

arXiv3 repos

arXiv:2410.02073

DepthPro, ml-depth-pro, Depth-Estimation

Distilling an End-to-End Voice Assistant Without Instruction Training Data

arXiv3 repos

arXiv:2410.02678

DiVA-llama-3-v0-8b, Speech-IFEval, canto-audio-llm

SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?

arXiv3 repos

arXiv:2410.03859

SWE-bench, SWE-bench, SWE-bench-fork

IterComp: Iterative Composition-Aware Feedback Learning from Model Gallery for Text-to-Image Generation

arXiv3 repos

arXiv:2410.07171

IterComp, IterComp, RPG-DiffusionMaster

Large-Scale 3D Medical Image Pre-training with Geometric Context Priors

arXiv3 repos

arXiv:2410.09890

PreCT-160K, VoCo, VoComni

Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

arXiv3 repos

arXiv:2410.10733

efficientvit, dc-ae-f32c32-sana-1.0-diffusers, Jazz

Large Continual Instruction Assistant

arXiv3 repos

arXiv:2410.10868

Continual-NExT, CoIN, CoIN_Refined

DepthSplat: Connecting Gaussian Splatting and Depth

arXiv3 repos

arXiv:2410.13862

unimatch, depthsplat, depthsplat

arXiv:2410.15522

arXiv3 repos

arXiv:2410.15522

MM-Eval, prometheus-eval, m-rewardbench

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

arXiv3 repos

arXiv:2410.16153

Pangea, Pangea-7B, PangeaInstruct

Scaling Diffusion Language Models via Adaptation from Autoregressive Models

arXiv3 repos

arXiv:2410.17891

SDAR, dLLM-RL, FreeDave

Scaling up Masked Diffusion Models on Text

arXiv3 repos

arXiv:2410.18514

SMDM, SMDMtry, minimal-dlm

arXiv:2410.21035

arXiv3 repos

arXiv:2410.21035

sdtt, sdtt, SDTT-LaViDa

OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

arXiv3 repos

arXiv:2410.23218

UI-TARS, SeeClick, transformer-final-proj

Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations

arXiv3 repos

arXiv:2411.01661

LLambada, Llambada, LLambada

Cut Your Losses in Large-Vocabulary Language Models

arXiv3 repos

arXiv:2411.09009

ml-cross-entropy, unsloth-zoo, unsloth_zoo

Golden Noise for Diffusion Models: A Learning Framework

arXiv3 repos

arXiv:2411.09502

Golden-Noise-for-Diffusion-Models, GoldenNoiseModel, ComfyUI_Golden-Noise

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

arXiv3 repos

arXiv:2411.10440

LLaVA-CoT-100k, LLaVA-CoT, Llama-3.2V-11B-cot

Adversarial Diffusion Compression for Real-World Image Super-Resolution

arXiv3 repos

arXiv:2411.13383

OSEDiff, AdcSR, AdcSR

Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study

arXiv3 repos

arXiv:2411.13588

xDiT, mochi-xdit, DiTCacheAnalysis

BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models

arXiv3 repos

arXiv:2411.15232

BiomedCoOp, BiomedCoOp, BiomedCoOp

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

arXiv3 repos

arXiv:2411.16537

RoboSpatial-Eval, RoboSpatial-Home, RoboSpatial

GRAPE: Generalizing Robot Policy via Preference Alignment

arXiv3 repos

arXiv:2411.19309

OpenVLA-7B-GRAPE-Simpler, GRAPE, OpenVLA-7B-SFT-Simpler

Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis

arXiv3 repos

arXiv:2411.19509

ditto-talkinghead, ditto-talkinghead, LivePortrait

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

arXiv3 repos

arXiv:2412.01169

OmniFlows, OmniFlow-v0.9, OmniFlow-v0.5

Structured 3D Latents for Scalable and Versatile 3D Generation

arXiv3 repos

arXiv:2412.01506

TRELLIS-image-large, TRELLIS-image-large, c3-trellis-gradio

Switti: Designing Scale-Wise Transformers for Text-to-Image Synthesis

arXiv3 repos

arXiv:2412.01819

Switti, VQVAE-Switti, Switti-AR

Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

arXiv3 repos

arXiv:2412.04431

Infinity, infinity, infinity

Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation

arXiv3 repos

arXiv:2412.04954

Med-CXRGen-F, Med-CXRGen-I, RRG-BioNLP-ACL2024

BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities

arXiv3 repos

arXiv:2412.07769

BiMediX2-8B, BiMediX2-70B, BiMediX2

Learning Flow Fields in Attention for Controllable Person Image Generation

arXiv3 repos

arXiv:2412.08486

Leffa, Leffa, Leffa

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

arXiv3 repos

arXiv:2412.10117

CosyVoice, CosyVoice, FastCosyVoice

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

arXiv3 repos

arXiv:2412.10302

deepseek-vl2, deepseek-vl2-tiny, deepseek-vl2-small

Empowering LLMs to Understand and Generate Complex Vector Graphics

arXiv3 repos

arXiv:2412.11102

OmniSVG, OmniSVG-train, omnisvg-train

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

arXiv3 repos

arXiv:2412.12661

medmax, medmax_eval_data, medmax_data

Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

arXiv3 repos

arXiv:2412.14171

VSI-Bench, embodied-eval, behaviour_subtask

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

arXiv3 repos

arXiv:2412.15838

Align-DS-V, DollyTails-12K, align-anything

Text2midi: Generating Symbolic Music from Captions

arXiv3 repos

arXiv:2412.16526

Text2midi, text2midi, text2midi

HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks

arXiv3 repos

arXiv:2412.17574

data-juicer, SciDataOS, data-juicer

Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching

arXiv3 repos

arXiv:2412.18911

TaylorSeer, ToCa, DuCa

2 OLMo 2 Furious

arXiv3 repos

arXiv:2501.00656

OLMo-2-0425-1B, olmes, qwen3-8b-base

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

arXiv3 repos

arXiv:2501.01108

MuQ-MuLan-large, MuQ, MuQ-large-msd-iter

Sa2VA: Marrying SAM2 with MLLM for Dense Grounded Understanding of Images and Videos

arXiv3 repos

arXiv:2501.04001

Sa2VA, Sa2VA, Sa2VA-Training

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

arXiv3 repos

arXiv:2501.12202

Hunyuan3D-2.1, HY3D-Bench, Hunyuan3D-Omni

UI-TARS: Pioneering Automated GUI Interaction with Native Agents

arXiv3 repos

arXiv:2501.12326

UI-TARS, ScreenSpot-Pro-GUI-Grounding, UI-TARS-7B-DPO

InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

arXiv3 repos

arXiv:2501.12386

InternVideo, InternVL_2_5_HiCo_R16, internvideo-d2a11ea9

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

arXiv3 repos

arXiv:2501.13106

VideoLLaMA2, VideoLLaMA3-7B-Image, VideoLLaMA3-2B-Image

Distilling foundation models for robust and efficient models in digital pathology

arXiv3 repos

arXiv:2501.16239

plism-benchmark, SEAL, plism-dataset

M+: Extending MemoryLLM with Scalable Long-Term Memory

arXiv3 repos

arXiv:2502.00592

MemoryLLM, mplus-8b, ExtendingMemoryLLM

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

arXiv3 repos

arXiv:2502.01051

LPO, LRM, LPO

Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity

arXiv3 repos

arXiv:2502.01776

nunchaku, nunchaku, Sparse-VideoGen

Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined Data

arXiv3 repos

arXiv:2502.04380

data-juicer, SciDataOS, data-juicer

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

arXiv3 repos

arXiv:2502.10248

stepvideo-t2v, Step-Video-T2V, stepvideo-t2v-turbo

Baichuan-M1: Pushing the Medical Capability of Large Language Models

arXiv3 repos

arXiv:2502.12671

Baichuan-M1-14B-Instruct, Baichuan-M1-14B-Base, Med-R1

Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

arXiv3 repos

arXiv:2502.14846

pixmo-docs, CoSyn-point, CoSyn-400K

FeatSharp: Your Vision Model Features, Sharper

arXiv3 repos

arXiv:2502.16025

RADIO, C-RADIOv4-SO400M, C-RADIOv4-H

LettuceDetect: A Hallucination Detection Framework for RAG Applications

arXiv3 repos

arXiv:2502.17125

LettuceDetect, lettucedect-base-modernbert-en-v1, lettucedect-large-modernbert-en-v1

RaceVLA: VLA-based Racing Drone Navigation with Human-like Behaviour

arXiv3 repos

arXiv:2503.02572

RaceVLA, RaceVLA_models, RaceVLA_dataset

Unified Reward Model for Multimodal Understanding and Generation

arXiv3 repos

arXiv:2503.05236

UnifiedReward-qwen-3b, VideoDPO, ShareGPTVideo-DPO

VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control

arXiv3 repos

arXiv:2503.05639

VPBench, VideoPainter, VPData

GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images

arXiv3 repos

arXiv:2503.06073

ECG-Grounding, GEM, GEM

Unleashing the Potential of Large Language Models for Text-to-Image Generation through Autoregressive Representation Alignment

arXiv3 repos

arXiv:2503.07334

ARRA, ARRA, ARRA-Adapt-MIMIC-7B

MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?

arXiv3 repos

arXiv:2503.09499

data-juicer, SciDataOS, data-juicer

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

arXiv3 repos

arXiv:2503.11576

docling, docling, my-docling

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

arXiv3 repos

arXiv:2503.12165

VTON360, VTON360-THuman2.0, VTON360-MVHumanNet

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

arXiv3 repos

arXiv:2503.13377

xiaomi-mimo-vl-miloco, Xiaomi-MiMo-VL-Miloco-7B, Xiaomi-MiMo-VL-Miloco-7B-GGUF

TerraTorch: The Geospatial Foundation Models Toolkit

arXiv3 repos

arXiv:2503.20563

terratorch, terratorch, granite-geospatial-biomass

Hi3DGen: High-fidelity 3D Geometry Generation from Images via Normal Bridging

arXiv3 repos

arXiv:2503.22236

ComfyUI-Hi3DGen, Hi3DGen, Hi3DGen

ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations

arXiv3 repos

arXiv:2504.00824

ScholarCopilot, ScholarCopilot-Data-v1, ScholarCopilot-v1

WikiVideo: Article Generation from Multiple Videos

arXiv3 repos

arXiv:2504.00939

CRAFT, wikivideo, wikivideo

MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

arXiv3 repos

arXiv:2504.00993

MedReason, MedReason-8B, Med-R1

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

arXiv3 repos

arXiv:2504.01805

SpaceR, SpaceR, SpaceR-151k

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

arXiv3 repos

arXiv:2504.02160

UNO, UNO-1M, UNO

Advancing Semantic Caching for LLMs with Domain-Specific Embeddings and Synthetic Data

arXiv3 repos

arXiv:2504.02268

MeanCache, langcache-embed-v1, langcache-embed-v2

R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents

arXiv3 repos

arXiv:2504.07164

MiniMax-M2-BF16, MiniMax-M2, MiniMax-M2

Kimi-VL Technical Report

arXiv3 repos

arXiv:2504.07491

LocateAnything-3B, MoonViT-SO-400M, MMLongBench-Doc

$π_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

arXiv3 repos

arXiv:2504.16054

physical-ai-studio, pi05_base, open-value

ReasonIR: Training Retrievers for Reasoning Tasks

arXiv3 repos

arXiv:2504.20595

ReasonIR, ReasonIR-8B, reasonir-data

In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

arXiv3 repos

arXiv:2504.20690

ICEdit, normal-lora, ICEdit-MoE-LoRA

Llama-Nemotron: Efficient Reasoning Models

arXiv3 repos

arXiv:2505.00949

Llama-Nemotron-Post-Training-Dataset, Llama-3_3-Nemotron-Super-49B-v1_5, Llama-3_3-Nemotron-Super-49B-GenRM

Benchmarking LLMs' Swarm intelligence

arXiv3 repos

arXiv:2505.04364

YuLan-SwarmIntell, swarmbench, swarmbench

Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets

arXiv3 repos

arXiv:2505.07747

Step1X-3D, Step1X-3D, Step1X-3D-obj-data

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

arXiv3 repos

arXiv:2505.08699

granite-speech-3.3-8b, granite-speech-models, granite-speech-studio

SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

arXiv3 repos

arXiv:2505.10793

TuneJury, SongEval, SongEval

Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining

arXiv3 repos

arXiv:2505.11293

B3, B3_Qwen2_7B, B3_Qwen2_2B

MMaDA: Multimodal Large Diffusion Language Models

arXiv3 repos

arXiv:2505.15809

MMaDA-8B-Base, MMaDA-8B-MixCoT, MMaDA

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

arXiv3 repos

arXiv:2505.16915

data-juicer, SciDataOS, data-juicer

Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation

arXiv3 repos

arXiv:2505.18875

nunchaku, nunchaku, Sparse-VideoGen

arXiv:2505.19706

arXiv3 repos

arXiv:2505.19706

PathFinder-PRM, PathFinder-600K, PathFinder-PRM-7B

FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement

arXiv3 repos

arXiv:2505.20192

FunReason, AWorld, AWorld-RL

ImgEdit: A Unified Image Editing Dataset and Benchmark

arXiv3 repos

arXiv:2505.20275

ImgEdit, ImgEdit, ImgEdit_recap_mask

Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence

arXiv3 repos

arXiv:2505.20325

Guided-by-Gut, DS-Qwen-7b-GG-CalibratedConfRL, DS-Qwen-1.5b-GG-CalibratedConfRL

arXiv:2505.20979

arXiv3 repos

arXiv:2505.20979

MelodySim, melodySim, MelodySim

One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models

arXiv3 repos

arXiv:2505.21960

Loopfree, loopfree-sd1.5, loopfree-sd2.1-base

AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views

arXiv3 repos

arXiv:2505.23716

anysplat, AnySplat, AnySplat

TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context Learning

arXiv3 repos

arXiv:2505.23719

TiRex, tirex, TiRex-1.1-gifteval

Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization

arXiv3 repos

arXiv:2505.24111

diarizen-wavlm-large-s80-md, diarizen-wavlm-base-s80-md, diarizen-wavlm-large-s80-md-v2

QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training

arXiv3 repos

arXiv:2506.00711

QoQ-Med-VL-32B, QoQ-Med-VL-7B, QoQ_Med

arXiv:2506.00830

arXiv3 repos

arXiv:2506.00830

SkyReels-V2, SkyReels-A1, SkyReels-A2

Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology

arXiv3 repos

arXiv:2506.02408

CPathPatchFeature, E2E-WSI-ABMILX, CPath_Image

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

arXiv3 repos

arXiv:2506.03150

IllumiCraft, Illumicraft-checkpoints, IllumiCraft

FailureSensorIQ: A Multi-Choice QA Dataset for Understanding Sensor Relationships and Failure Modes

arXiv3 repos

arXiv:2506.03278

FailureSensorIQ, AssetOpsBench, FailureSensorIQ

PixCell: A generative foundation model for digital histopathology images

arXiv3 repos

arXiv:2506.05127

PixCell-1024, PixCell-sample-data, PixCell-256

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

arXiv3 repos

arXiv:2506.07044

Lingshu-32B, Lingshu-7B, ReasonMed

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence

arXiv3 repos

arXiv:2506.07966

SpaceQwen2.5-VL-3B-Instruct, SpaceOm, SpaceThinker-Qwen2.5VL-3B

EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection

arXiv3 repos

arXiv:2506.09827

Voice-Acting-Pipeline, kani-tts-2-en, kani-tts-2-pt

Efficient Part-level 3D Object Generation via Dual Volume Packing

arXiv3 repos

arXiv:2506.09980

PartPacker, PartPacker, PartPacker

3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks

arXiv3 repos

arXiv:2506.11147

3D-RAD, 3D-RAD, M3D-RAD

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning

arXiv3 repos

arXiv:2506.15154

SonicVerse, SonicVerse, SonicVerse

VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models

arXiv3 repos

arXiv:2506.17561

VLA-OS, VLA-OS, VLA-OS-Dataset

Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models

arXiv3 repos

arXiv:2506.18623

diarizen-wavlm-large-s80-md, diarizen-wavlm-base-s80-md, diarizen-wavlm-large-s80-md-v2

Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

arXiv3 repos

arXiv:2506.18898

Tar-7B-v0.1, TA-Tok, Tar-1.5B

WorldVLA: Towards Autoregressive Action World Model

arXiv3 repos

arXiv:2506.21539

WorldVLA, RynnVLA-001, RynnEC

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

arXiv3 repos

arXiv:2507.01006

GLM-V, GLM-4.1V-9B-Thinking, MMLongBench-Doc

DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment

arXiv3 repos

arXiv:2507.02768

DeSTA2, DeSTA2.5-Audio, DeSTA-AQA5M-FROM-Llama3.1-8B-Instruct

Agentic-R1: Distilled Dual-Strategy Reasoning

arXiv3 repos

arXiv:2507.05707

DualDistill, Agentic-R1-SD, Agentic-R1

Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling

arXiv3 repos

arXiv:2507.17801

Lumina-mGPT-2.0, Lumina-mGPT-2.0-Omni, Lumina-mGPT-2.0

Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering

arXiv3 repos

arXiv:2507.18446

WhisperLiveKit, diar_streaming_sortformer_4spk-v2, multitalker-parakeet-streaming-0.6b-v1

DIFFA: Large Language Diffusion Models Can Listen and Understand

arXiv3 repos

arXiv:2507.18452

DIFFA, DIFFA, DIFFA

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

arXiv3 repos

arXiv:2507.20880

JAME, jamify, JAM-0.5

Music Arena: Live Evaluation for Text-to-Music

arXiv3 repos

arXiv:2507.20900

music-arena, music-arena-dataset, TuneJury

Persona Vectors: Monitoring and Controlling Character Traits in Language Models

arXiv3 repos

arXiv:2507.21509

cli, FabricationGuard-linearprobe-qwen36-27b, ReasoningGuard-linearprobe-qwen36-27b

Trade-offs in Image Generation: How Do Different Dimensions Interact?

arXiv3 repos

arXiv:2507.22100

TRIG, TRIG, TRIG

SDMatte: Grafting Diffusion Models for Interactive Matting

arXiv3 repos

arXiv:2508.00443

SDMatte, SDMatte, LiteSDMatte

Marco-Voice Technical Report

arXiv3 repos

arXiv:2508.02038

Marco-Voice, CSEMOTIONS, Marco-Voice

Kronos: A Foundation Model for the Language of Financial Markets

arXiv3 repos

arXiv:2508.02739

Kronos, Kronos-V12, stock_forecasting

Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences

arXiv3 repos

arXiv:2508.03542

speech2latex, Speech2Latex, GemmaApollo

Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation

arXiv3 repos

arXiv:2508.07981

Omni-Effects, Omni-Effects, Omni-VFX

arXiv:2508.08098

arXiv3 repos

arXiv:2508.08098

TBAC-UniImage, TBAC-UniImage-3B, TBAC-UniImage

TopXGen: Topic-Diverse Parallel Data Generation for Low-Resource Machine Translation

arXiv3 repos

arXiv:2508.08680

topxgen-llama-4-scout-and-llama-4-scout, topxgen-gemma-3-27b-and-nllb-3.3b, topxgen

V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center Peaking

arXiv3 repos

arXiv:2508.13634

AWorld, AWorld-RL, V2P-7B

LexSemBridge: Fine-Grained Dense Representation Enhancement through Token-Aware Embedding Augmentation

arXiv3 repos

arXiv:2508.17858

LexSemBridge, LexSemBridge_eval, LexSemBridge_CLR_snowflake

Baichuan-M2: Scaling Medical Capability with Large Verifier System

arXiv3 repos

arXiv:2509.02208

Baichuan-M2-32B, Baichuan-M2-32B, Baichuan-M2-32B-GPTQ-Int4

Sample-efficient Integration of New Modalities into Large Language Models

arXiv3 repos

arXiv:2509.04606

sample-efficient-multimodality, capdels, sample-efficient-multimodality-ckpts

P3-SAM: Native 3D Part Segmentation

arXiv3 repos

arXiv:2509.06784

HY3D-Bench, Hunyuan3D-Part, Hunyuan3D-Part

Continuous Audio Language Models

arXiv3 repos

arXiv:2509.06926

pocket-tts, pocket-tts-korean-300m, pocket-tts-ungated

X-Part: high fidelity and structure coherent shape decomposition

arXiv3 repos

arXiv:2509.08643

HY3D-Bench, Hunyuan3D-Part, Hunyuan3D-Part

Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling

arXiv3 repos

arXiv:2509.08753

tts-1.6b-en_fr, tts_longeval, delayed-streams-modeling

Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models

arXiv3 repos

arXiv:2509.13031

PeBR-R1, PeBR_R1, PeBR_R1_dataset

FinSearchComp: Towards a Realistic, Expert-Level Evaluation of Financial Search and Reasoning

arXiv3 repos

arXiv:2509.13160

MiniMax-M2-BF16, MiniMax-M2, MiniMax-M2

SPATIALGEN: Layout-guided 3D Indoor Scene Generation

arXiv3 repos

arXiv:2509.14981

FLUX.1-Wireframe-dev-lora, SpatialGen-Testset, SpatialGen-1.0

RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation

arXiv3 repos

arXiv:2509.15212

WorldVLA, RynnVLA-002, RynnVLA-002

SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models

arXiv3 repos

arXiv:2509.17664

SD-VLM-7B, SD-VLM, MSMU

EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling

arXiv3 repos

arXiv:2509.23909

EditScore-Reward-Data, EditReward-Bench, EditScore-RL-Data

VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning

arXiv3 repos

arXiv:2509.24650

VoxCPM, VoxCPM2, VoxCPM1.5

SDPose: Exploiting Diffusion Priors for Out-of-Domain and Robust Pose Estimation

arXiv3 repos

arXiv:2509.24980

SDPose-Body, SDPose-Wholebody, SDPose-OOD

arXiv:2509.26346

arXiv3 repos

arXiv:2509.26346

EditReward, EditReward-Data, EditReward-Bench

MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation

arXiv3 repos

arXiv:2509.26642

MLA, MLA_RLBench_post, MLA_pretrain

AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees

arXiv3 repos

arXiv:2510.01268

DetectLLMSegmentation, L2D, AdaDetectGPT

Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation

arXiv3 repos

arXiv:2510.06961

esb-datasets-test-only-sorted, open-asr-leaderboard, asr-leaderboard-longform

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models

arXiv3 repos

arXiv:2510.08531

SpatialLadder-3B, SpatialLadder-26k, SPBench

ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation

arXiv3 repos

arXiv:2510.11000

IMIG-100K, ContextGen, IMIG-Source

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

arXiv3 repos

arXiv:2510.12784

SRUM_BAGEL_7B_MoT, SRUM, SRUM_6k_CompBench_Train

UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE

arXiv3 repos

arXiv:2510.13344

Uni-MoE, UniMoE-Audio-preview, UMOE-Scaling-Unified-Multimodal-LLMs

BADAS: Context Aware Collision Prediction Using Real-World Dashcam Data

arXiv3 repos

arXiv:2510.14876

Cosmos-Sentinel, Cosmos_Sentinel, Nvidia-Cosmos-Cookoff

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

arXiv3 repos

arXiv:2510.15742

Ditto, Ditto-1M, Ditto_models

Skyfall-GS: Synthesizing Immersive 3D Urban Scenes from Satellite Imagery

arXiv3 repos

arXiv:2510.15869

Skyfall-GS-datasets, Skyfall-GS-eval, Skyfall-GS-ply

DeepAnalyze: Agentic Large Language Models for Autonomous Data Science

arXiv3 repos

arXiv:2510.16872

DeepAnalyze, DataScience-Instruct-500K, DeepAnalyze-8B

Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain

arXiv3 repos

arXiv:2510.17801

RoboBench, RoboBench-Results, RoboBench

UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation

arXiv3 repos

arXiv:2510.18701

UniGenBench-EvalModel-qwen3vl-32b-v1, UniGenBench-EvalModel-qwen-72b-v1, UniGenBench-Eval-Images

Video-As-Prompt: Unified Semantic Control for Video Generation

arXiv3 repos

arXiv:2510.20888

Video-As-Prompt, Video-As-Prompt-Wan2.1-14B, Video-As-Prompt-CogVideoX-5B

DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching

arXiv3 repos

arXiv:2510.22950

diffrhythm2, DiffRhythm2, DiffRhythm2

Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization

arXiv3 repos

arXiv:2510.25616

BlindVLA, openvla-7b-warmup-checkpoint_lora_002000, openvla_1k-dataset

Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization

arXiv3 repos

arXiv:2511.01588

PDF-VLM2Vec, PDF-VLM2Vec-Qwen2VL-2B, PDF-VLM2Vec-Qwen2VL-7B

SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia

arXiv3 repos

arXiv:2511.01670

SeaLLMs-Audio, SeaBench-Audio, SeaLLMs-Audio-7B

Step-Audio-EditX Technical Report

arXiv3 repos

arXiv:2511.03601

Step-Audio-EditX, Step-Audio-EditX, Step-Audio-EditX-AWQ-4bit

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

arXiv3 repos

arXiv:2511.11434

weave, Weave, Bagel-weave

GroupRank: A Groupwise Paradigm for Effective and Efficient Passage Reranking with LLMs

arXiv3 repos

arXiv:2511.11653

Diver, Diver-GroupRank-7B, Diver-GroupRank-32B

Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data

arXiv3 repos

arXiv:2511.12609

zen5, Uni-MoE, UMOE-Scaling-Unified-Multimodal-LLMs

RynnVLA-002: A Unified Vision-Language-Action and World Model

arXiv3 repos

arXiv:2511.17502

WorldVLA, RynnVLA-001, RynnVLA-002

Fara-7B: An Efficient Agentic Model for Computer Use

arXiv3 repos

arXiv:2511.19663

WebTailBench, Fara-7b, Fara_Test

LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

arXiv3 repos

arXiv:2511.20785

LongVT, LongVT-Parquet, LongVT-Source

G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning

arXiv3 repos

arXiv:2511.21688

G2VLM-2B-MoT, G2VLM, g2vlm

ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration

arXiv3 repos

arXiv:2511.21689

ToolOrchestra, Orchestrator-8B, ToolScale

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

arXiv3 repos

arXiv:2512.02556

Hy4-preview, Hy4-preview, maxtext

Light-X: Generative 4D Video Rendering with Camera and Illumination Control

arXiv3 repos

arXiv:2512.05115

Light-X-Uni, Light-X, Light-Syn

Singing Timbre Popularity Assessment Based on Multimodal Large Foundation Model

arXiv3 repos

arXiv:2512.06999

QwenFeat-Vocal-Score, QwenFeat-Vocal-Score, Singing-Aesthetic-Assessment

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

arXiv3 repos

arXiv:2512.09874

pdf-parse-bench, wikipedia-latex-formulas-319k, pdf-parse-bench

EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing

arXiv3 repos

arXiv:2512.11715

EditMGT, EditMGT, CrispEdit-2M

arXiv:2512.11831

arXiv3 repos

arXiv:2512.11831

ExplicitShortCut, ESC-XL2, ESC-B2

Image Diffusion Preview with Consistency Solver

arXiv3 repos

arXiv:2512.13592

consolver, consolver, EditReward

Bolmo: Byteifying the Next Generation of Language Models

arXiv3 repos

arXiv:2512.15586

Bolmo-7B, Bolmo-1B, bolmo_mix

Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition

arXiv3 repos

arXiv:2512.15603

Qwen-Image, Qwen-Image-Layered, Qwen-Image-Layered

JustRL: Scaling a 1.5B LLM with a Simple RL Recipe

arXiv3 repos

arXiv:2512.16649

UltraData-RL-2609, MiniCPM5-1B-GGUF, MiniCPM5-1B

SAM Audio: Segment Anything in Audio

arXiv3 repos

arXiv:2512.18099

sam-audio, sam-audio, Sam-Audio

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation

arXiv3 repos

arXiv:2512.24551

Open-PhyGDPO, PhyGDPO, PhyGDPO

From Failure to Mastery: Generating Hard Samples for Tool-use Agents

arXiv3 repos

arXiv:2601.01498

AWorld, FunReason-MT, AWorld-RL

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

arXiv3 repos

arXiv:2601.02456

InternVLA-A1-3B-RoboTwin, InternVLA-A1-3B, internvla

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning

arXiv3 repos

arXiv:2601.09536

Omni-Bench, Omni-R1-Zero, Omni-R1

The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models

arXiv3 repos

arXiv:2601.10387

karma-electric-project, karma-electric-llama31-8b, drowse

F-Actor: Controllable Conversational Behaviour in Full-Duplex Models

arXiv3 repos

arXiv:2601.11329

f-actor-behavior-sd-nanocodec, f-actor, f-actor-behavior-sd-mimi

TeleStyle: Content-Preserving Style Transfer in Images and Videos

arXiv3 repos

arXiv:2601.20175

TeleStyle, TeleStyleV2, TeleStyleV2-Qwen-Image-Edit-2511-bf16

RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models

arXiv3 repos

arXiv:2602.00443

RVCBench, RVCBench, RVCBench

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio

arXiv3 repos

arXiv:2602.03328

GuardReasoner-Omni, GuardReasoner-Omni-7B, GuardReasoner-Omni-3B

LLaDA2.1: Speeding Up Text Diffusion via Token Editing

arXiv3 repos

arXiv:2602.08676

LLaDA2.1-mini, LLaDA2.1-flash, dllm

StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors

arXiv3 repos

arXiv:2602.08934

StealthRL, StealthRL-Benchmark, StealthRL

RoboInter: A Holistic Intermediate Representation Suite Towards Robotic Manipulation

arXiv3 repos

arXiv:2602.09973

RoboInter-VLM_llavaov_7B, RoboInter-VLM, RoboInter-VLM_qwenvl25_3b

ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation

arXiv3 repos

arXiv:2602.10113

ConsID-Gen, ConsID-Gen, ConsIDVid

Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

arXiv3 repos

arXiv:2602.20161

Mobile-O-SFT, Mobile-O-Pre-Train, Mobile-O-Post-Train

DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation

arXiv3 repos

arXiv:2602.22839

PPTAgent, DeepPresenter-9B-GGUF, DeepPresenter-9B

TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment

arXiv3 repos

arXiv:2602.23068

tada-1b, tada-codec, tada-3b-ml

Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design

arXiv3 repos

arXiv:2603.00152

Dr-Seg, Dr_Seg, coconut

Fish Audio S2 Technical Report

arXiv3 repos

arXiv:2603.08823

s2-pro, fish-speech, RVCBench

$Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation

arXiv3 repos

arXiv:2603.12263

Psi0, psi0, Psi0

Overcoming the Modality Gap in Context-Aided Forecasting

arXiv3 repos

arXiv:2603.12451

DoubleCast, CAF_7M, DoubleCast

Multimodal OCR: Parse Anything from Documents

arXiv3 repos

arXiv:2603.13032

dots.ocr, dots.mocr, dots.mocr-svg

Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models

arXiv3 repos

arXiv:2603.21426

beta-kd, Cosine-Beta-KD-Instance, Cosine-Beta-KD-Task

AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese

arXiv3 repos

arXiv:2603.26511

AMALIA-9B-0626-DPO, amalia-lm-eval, AMALIA

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding

arXiv3 repos

arXiv:2603.27064

ChartNet, granite-vision-4.1-4b, granite-4.0-3b-vision

An Empirical Recipe for Universal Phone Recognition

arXiv3 repos

arXiv:2603.29042

PhoneticXeus, PhoneticXeus, PhoneticXeus

Memory Intelligence Agent

arXiv3 repos

arXiv:2604.04503

MIA, MIA, MIA

FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

arXiv3 repos

arXiv:2604.06757

FlowInOne, VisPrompt5M, VPBench

Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

arXiv3 repos

arXiv:2604.10905

audio-flamingo-next-hf, audio-flamingo-next-captioner-hf, audio-flamingo-next-think-hf

Introspective Diffusion Language Models

arXiv3 repos

arXiv:2604.11035

I-DLM-8B, I-DLM-32B, I-DLM-8B-lora-r128

MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

arXiv3 repos

arXiv:2604.12928

moshi-rag, moshika-rag-pytorch-bf16, moshika-rag-candle-bf16

UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

arXiv3 repos

arXiv:2604.18518

UDM-GRPO, URSA-1.7B-IBQ512-UDMGRPO-GenEval, URSA-1.7B-IBQ512-UDMGRPO-PickScore

Vista4D: Video Reshooting with 4D Point Clouds

arXiv3 repos

arXiv:2604.21915

Vista4D, Vista4D, Vista4D-Eval-Data

GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction

arXiv3 repos

arXiv:2604.23941

GoClick, GoClick-Large, GoClick-Base

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

arXiv3 repos

arXiv:2604.24954

Eagle, Nemotron-3-Nano-Omni-30B-A3B-Reasoning-FP8, Nemotron-3-Nano-Omni-30B-A3B-Reasoning-NVFP4

Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

arXiv3 repos

arXiv:2604.28123

PRISM, gemini_distill, gemini_public_mmr1

GLiGuard: Schema-Conditioned Classification for LLM Safeguard

arXiv3 repos

arXiv:2605.07982

gliguard-LLMGuardrails-300M, GLiNER2-Guardrails-PII-Multi, GLiNER2-Guardrails-PII-Multi-onnx

GLiNER2-PII: A Multilingual Model for Personally Identifiable Information Extraction

arXiv3 repos

arXiv:2605.09973

gliner2-privacy-filter-PII-multi, GLiNER2-Guardrails-PII-Multi, GLiNER2-Guardrails-PII-Multi-onnx

Allegory of the Cave: Measurement-Grounded Vision-Language Learning

arXiv3 repos

arXiv:2605.11727

PRISM-VL, PRSIMVL-LoRA-V1, MeasL-150K-V1

Asymmetric Flow Models

arXiv3 repos

arXiv:2605.12964

AsymFLUX.2-klein-9B, LakonLab, AsymFLUX.2-klein

Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image

arXiv3 repos

arXiv:2605.14984

Sat3DGen, Sat3DGen, VIGOR_SAT3DGEN_add_skymask_DSM_satdepth

ReactiveGWM: Steering NPC in Reactive Game World Models

arXiv3 repos

arXiv:2605.15256

ReactiveGWM-Models, ReactiveGWM-Datasets, stable-retro

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

arXiv3 repos

arXiv:2605.19075

CRAFT, CRAFT-MAGMaR, CRAFT-WikiVideo

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

arXiv3 repos

arXiv:2605.20176

ClinSeekAgent, ClinSeek-35B-A3B, ClinSeek-Bench

GEM: Generative Supervision Helps Embodied Intelligence

arXiv3 repos

arXiv:2605.28548

GEM, GEM-250K, GEM-2B

dots.tts Technical Report

arXiv3 repos

arXiv:2606.07080

dots.tts-mf-2steps, dots.tts-mf-1step, dots.tts-mf-2steps-stts

Kwai Keye-VL-2.0 Technical Report

arXiv3 repos

arXiv:2606.10651

Keye, Keye-VL-2.0-30B-A3B, keye-39e1f0b5

i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

arXiv3 repos

arXiv:2606.11289

i1-3B, i1, i1-captions

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

arXiv3 repos

arXiv:2606.12195

InternVideo, InternVideo3-8B-Instruct, InternVideo3_Dataset

Modality Forcing for Scalable Spatial Generation

arXiv3 repos

arXiv:2606.13676

modality-forcing, modality_forcing, modality_forcing

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

arXiv3 repos

arXiv:2606.17006

TuneJury, tunejury, release-scores

Fara-1.5: Scalable Learning Environments for Computer Use Agents

arXiv3 repos

arXiv:2606.20785

fara, Fara1.5-4B, Fara1.5-9B

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

arXiv3 repos

arXiv:2606.21138

SEED, SEED, GenText-Forensics-3rd-Place

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

arXiv3 repos

arXiv:2607.00895

lettucedect-v2-mmbert-base, lettucedetect-code-hallucination, lettucedect-v2-qwen-2b

Representation Distribution Matching for One-Step Visual Generation

arXiv3 repos

arXiv:2607.02375

RDM, flux2-klein-1step-demo, flux2-klein-1step-rdm

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation

arXiv3 repos

arXiv:2607.02642

giga-world-1, Giga-World-1, Giga-World-1-Toydata

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

arXiv3 repos

arXiv:2607.05147

SpecForge, kimi-k3-dspark, DeepSpec

FreyaTTS: A Compact Tokenizer-Free Flow-Matching Transformer for Turkish-First Speech Synthesis

arXiv3 repos

arXiv:2607.09530

FreyaTTS, freya-tr-eval, freya-tts

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

arXiv3 repos

arXiv:2607.14952

Megatron-Bridge, Megatron-Bridge, Megatron-Bridge

SWE-Pruner Pro: The Coder LLM Already Knows What to Prune

arXiv3 repos

arXiv:2607.18213

swe-pruner-pro-mimo-v2-flash-head, swe-pruner-pro-training-corpus, swe-pruner-pro-qwen3-coder-next-head

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

arXiv3 repos

arXiv:2607.19064

Mage, mage-7255fc67, Mage-VL-RTSP

arXiv:2607.29679

arXiv3 repos

arXiv:2607.29679

GenAI-Caption-Pipeline, SP-PE-Qwen3.5-35B-A3B, Qwen-Image-SP

TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity

arXiv3 repos

arXiv:2608.15767

tinycast, tinycast, tinycast-forecaster

GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture

arXiv3 repos

arXiv:2608.15875

giga-brain-0, GigaBrain-0.7-SampleData, GigaBrain-0.7-3.5B-Base

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

arXiv3 repos

arXiv:2608.20958

TLive-Omni, TLive-Omni-9B, TLive-Omni-4B

WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

arXiv3 repos

arXiv:2608.24053

WeMM-Embedding-9B, WeMM-Embedding-4B, WeMM-Embedding-2B

Detecting hallucinations in large language models using semantic entropy

Nature3 repos

Nature:s41586-024-07421-0

VASE, uqlm, Spnda

A vision–language foundation model for precision oncology

Nature3 repos

Nature:s41586-024-08378-w

PIANO, AtlasPatch, PathPT

Tanks and temples

ACM2 repos

ACM:3072959.3073599

awesome-mvs, InstantSplat

Microsoft recommenders

ACM2 repos

ACM:3298689.3346967

recommenders, Recommenders

Raha

ACM2 repos

ACM:3299869.3324956

Jellyfish-13B, raha

Microsoft Recommenders: Best Practices for Production-Ready Recommendation Systems

ACM2 repos

ACM:3366424.3382692

recommenders, Recommenders

WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning

ACM2 repos

ACM:3404835.3463257

test-big-dataset, wit

Automating QUIC Interoperability Testing

ACM2 repos

ACM:3405796.3405826

quic-interop-runner, quic-interop-runner

SyRust: automatic testing of Rust libraries with semantic-aware program synthesis

ACM2 repos

ACM:3453483.3454084

miri, rust

ZeRO-infinity

ACM2 repos

ACM:3458817.3476205

DeepSpeed, DeepSpeed

Hammer

ACM2 repos

ACM:3489517.3530672

chipyard, hammer

Retrieval-Based Gradient Boosting Decision Trees for Disease Risk Assessment

ACM2 repos

ACM:3534678.3539052

awesome-decision-tree-papers, awesome-gradient-boosting-papers

A Stochastic Shortest Path Algorithm for Optimizing Spaced Repetition Scheduling

ACM2 repos

ACM:3534678.3539081

awesome-fsrs, fsrs4anki

WebUI: A Dataset for Enhancing Visual UI Understanding with Web Semantics

ACM2 repos

ACM:3544548.3581158

webui-all, webui

A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training

ACM2 repos

ACM:3577193.3593704

DeepSpeed, DeepSpeed

Demo: ARFlow: A Framework for Simplifying AR Experimentation Workflow

ACM2 repos

ACM:3638550.3643617

rerun, Dalaran

Verified Extraction from Coq to OCaml

ACM2 repos

ACM:3656379

MetaCoq, metarocq

System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

ACM2 repos

ACM:3662158.3662806

DeepSpeed, DeepSpeed

Crabtree: Rust API Test Synthesis Guided by Coverage and Type

ACM2 repos

ACM:3689733

rust, miri

Rustlantis: Randomized Differential Testing of the Rust Compiler

ACM2 repos

ACM:3689780

rust, miri

Correct and Complete Type Checking and Certified Erasure for Coq , in Coq

ACM2 repos

ACM:3706056

MetaCoq, metarocq

SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips

ACM2 repos

ACM:3760250.3762217

DeepSpeed, DeepSpeed

Tactics for Reasoning modulo AC in Coq

arXiv2 repos

arXiv:1106.4448

aac-tactics, aac-tactics

BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data

arXiv2 repos

arXiv:1203.5485

awesome-bigdata, awesome-bigdata

arXiv:1209.2137

arXiv2 repos

arXiv:1209.2137

JavaFastPFOR, FastPFor

Efficient Estimation of Word Representations in Vector Space

arXiv2 repos

arXiv:1301.3781

Word-Embeddings-Repository-for-Turkish, ReAGent

arXiv:1401.6399

arXiv2 repos

arXiv:1401.6399

JavaFastPFOR, FastPFor

A Fast, Minimal Memory, Consistent Hash Algorithm

arXiv2 repos

arXiv:1406.2294

hash4j, grenier

Generative Adversarial Networks

arXiv2 repos

arXiv:1406.2661

ocaml-torch, Data-Science

Neural Machine Translation by Jointly Learning to Align and Translate

arXiv2 repos

arXiv:1409.0473

nl2bash, nmt

Leveraging Cloud Data to Mitigate User Experience from "Breaking Bad"

arXiv2 repos

arXiv:1411.7955

breakout, breakout-ruby

arXiv:1502.01916

arXiv2 repos

arXiv:1502.01916

FastPFor, JavaFastPFOR

Distilling the Knowledge in a Neural Network

arXiv2 repos

arXiv:1503.02531

hasktorch, distilgpt2

arXiv:1503.07387

arXiv2 repos

arXiv:1503.07387

JavaFastPFOR, FastPFor

DeepFont: Identify Your Font from An Image

arXiv2 repos

arXiv:1507.03196

YuzuMarker.FontDetection, YuzuMarker.FontDetection

An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition

arXiv2 repos

arXiv:1507.05717

yas, EasyOCR

Unsupervised Deep Embedding for Clustering Analysis

arXiv2 repos

arXiv:1511.06335

awesome-datascience, awesome-datascience

Perceptual Losses for Real-Time Style Transfer and Super-Resolution

arXiv2 repos

arXiv:1603.08155

convnext_perceptual_loss, SkinDeep

Neural Language Correction with Character-Based Attention

arXiv2 repos

arXiv:1603.09727

pycorrector, repo-5146-pycorrector

arXiv:1606.06031

arXiv2 repos

arXiv:1606.06031

sdtt, SDTT-LaViDa

Bag of Tricks for Efficient Text Classification

arXiv2 repos

arXiv:1607.01759

xlsum, fasttext-language-identification

Enriching Word Vectors with Subword Information

arXiv2 repos

arXiv:1607.04606

Word-Embeddings-Repository-for-Turkish, fasttext-language-identification

Pointer Sentinel Mixture Models

arXiv2 repos

arXiv:1609.07843

esp32s3-distributed-ai, wikitext

The HoTT Library: A formalization of homotopy type theory in Coq

arXiv2 repos

arXiv:1610.04591

HoTT, Coq-HoTT

Axiomatic Attribution for Deep Networks

arXiv2 repos

arXiv:1703.01365

shap, shap

Tacotron: Towards End-to-End Speech Synthesis

arXiv2 repos

arXiv:1703.10135

Real-Time-Voice-Cloning, MockingBird

Learning Important Features Through Propagating Activation Differences

arXiv2 repos

arXiv:1704.02685

shap, shap

A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

arXiv2 repos

arXiv:1704.05426

roberta-large-mnli, t5-large-encoder-only-bf16

Efficient Natural Language Response Suggestion for Smart Reply

arXiv2 repos

arXiv:1705.00652

splade-ecommerce-esci, FinBot

ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases

arXiv2 repos

arXiv:1705.02315

srt-cxr14-pooled-probe, diff-mining

SmoothGrad: removing noise by adding noise

arXiv2 repos

arXiv:1706.03825

shap, shap

Rethinking Atrous Convolution for Semantic Image Segmentation

arXiv2 repos

arXiv:1706.05587

EdgeSeg, spectra

GPU-acceleration for Large-scale Tree Boosting

arXiv2 repos

arXiv:1706.08359

LightGBM, LightGBM

Crowdsourcing Multiple Choice Science Questions

arXiv2 repos

arXiv:1707.06209

sciq, COMP4222-Course-Project

A Distributional Perspective on Reinforcement Learning

arXiv2 repos

arXiv:1707.06887

open-value, cleanrl

S$^3$FD: Single Shot Scale-invariant Face Detector

arXiv2 repos

arXiv:1708.05237

S3FD.pytorch, face-alignment

Squeeze-and-Excitation Networks

arXiv2 repos

arXiv:1709.01507

Multi-fake-detective, Waifu2x

arXiv:1709.08990

arXiv2 repos

arXiv:1709.08990

JavaFastPFOR, FastPFor

Searching for Activation Functions

arXiv2 repos

arXiv:1710.05941

pegasus-x-base-synthsumm_open-16k, llama2.zig

Generalized End-to-End Loss for Speaker Verification

arXiv2 repos

arXiv:1710.10467

Real-Time-Voice-Cloning, MockingBird

Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

arXiv2 repos

arXiv:1712.05884

tacotron2, seq2seq_accent_conversion_model

The NarrativeQA Reading Comprehension Challenge

arXiv2 repos

arXiv:1712.07040

GraphKV, GraphKV

DVQA: Understanding Data Visualizations via Question Answering

arXiv2 repos

arXiv:1801.08163

DVQA_dataset, PlotQA

Efficient Neural Audio Synthesis

arXiv2 repos

arXiv:1802.08435

Real-Time-Voice-Cloning, MockingBird

Shampoo: Preconditioned Stochastic Tensor Optimization

arXiv2 repos

arXiv:1802.09568

modded-nanogpt, Emerging-Optimizers

xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems

arXiv2 repos

arXiv:1803.05170

recommenders, Recommenders

YOLOv3: An Incremental Improvement

arXiv2 repos

arXiv:1804.02767

darknet, darknetcv

MMDenseLSTM: An efficient combination of convolutional and recurrent neural networks for audio source separation

arXiv2 repos

arXiv:1805.02410

demucs, demucs

arXiv:1805.08949

arXiv2 repos

arXiv:1805.08949

docprompting-conala, conala

Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis

arXiv2 repos

arXiv:1806.04558

Real-Time-Voice-Cloning, MockingBird

The latest gossip on BFT consensus

arXiv2 repos

arXiv:1807.04938

cometbft, tendermint

Learning to Describe Differences Between Pairs of Similar Images

arXiv2 repos

arXiv:1808.10584

idefics-80b-instruct, idefics-9b-instruct

Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

arXiv2 repos

arXiv:1809.08887

llm-jepa, sql-create-context

HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

arXiv2 repos

arXiv:1809.09600

GraphKV, GraphKV

How Powerful are Graph Neural Networks?

arXiv2 repos

arXiv:1810.00826

MCBG, Malware_survey_my_experiments

Model Cards for Model Reporting

arXiv2 repos

arXiv:1810.03993

moonshine, galactica-120b

WikiHow: A Large Scale Text Summarization Dataset

arXiv2 repos

arXiv:1810.09305

all-MiniLM-L6-v2, all-MiniLM-L12-v2

arXiv:1810.12368

arXiv2 repos

arXiv:1810.12368

geolm-base-toponym-recognition, Pragmatic-Guide-to-Geoparsing-Evaluation

The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

arXiv2 repos

arXiv:1811.00982

SEED-Data-Edit-Part2-3, SEED-Data-Edit

SAFE: Self-Attentive Function Embeddings for Binary Similarity

arXiv2 repos

arXiv:1811.05296

MCBG, Malware_survey_my_experiments

Neural Abstractive Text Summarization with Sequence-to-Sequence Models

arXiv2 repos

arXiv:1812.02303

pycorrector, repo-5146-pycorrector

TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents

arXiv2 repos

arXiv:1901.08149

CDial-GPT_LCCC-large, CDial-GPT_LCCC-base

Parameter-Efficient Transfer Learning for NLP

arXiv2 repos

arXiv:1902.00751

Continual-NExT, LoRA

Computing Extremely Accurate Quantiles Using t-Digests

arXiv2 repos

arXiv:1902.04023

t-digest-c, t-digest

arXiv:1902.04043

arXiv2 repos

arXiv:1902.04043

smac, smacv2

LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

arXiv2 repos

arXiv:1904.02882

F5-TTS, libritts

Publicly Available Clinical BERT Embeddings

arXiv2 repos

arXiv:1904.03323

ehr_deidentification, deid_bert_i2b2

A Repository of Conversational Datasets

arXiv2 repos

arXiv:1904.06472

all-MiniLM-L6-v2, all-MiniLM-L12-v2

Improved Precision and Recall Metric for Assessing Generative Models

arXiv2 repos

arXiv:1904.06991

materialgan, stylegan2-ada-pytorch

Learning to Prove Theorems via Interacting with Proof Assistants

arXiv2 repos

arXiv:1905.09381

coq-serapi, coq-serapi

GLTR: Statistical Detection and Visualization of Generated Text

arXiv2 repos

arXiv:1906.04043

L2D, AdaDetectGPT

Pre-Training with Whole Word Masking for Chinese BERT

arXiv2 repos

arXiv:1906.08101

chinese-roberta-wwm-ext, chinese-bert-wwm-ext

MediaPipe: A Framework for Building Perception Pipelines

arXiv2 repos

arXiv:1906.08172

mediapipe, mediapipe

Generating Correctness Proofs with Neural Networks

arXiv2 repos

arXiv:1907.07794

coq-serapi, coq-serapi

LVIS: A Dataset for Large Vocabulary Instance Segmentation

arXiv2 repos

arXiv:1908.03195

LocateAnything-3B, lvis-api

TabNet: Attentive Interpretable Tabular Learning

arXiv2 repos

arXiv:1908.07442

pytorch-frame, Trompt

PlotQA: Reasoning over Scientific Plots

arXiv2 repos

arXiv:1909.00997

VisRAG-Ret-Train-In-domain-data, PlotQA

Fine-Tuning Language Models from Human Preferences

arXiv2 repos

arXiv:1909.08593

Starling-LM-7B-beta, trlx

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

arXiv2 repos

arXiv:1909.11942

tf-transformers, sahajBERT

Machine Learning in Python: Main developments and technology trends in data science, machine learning, and artificial intelligence

arXiv2 repos

arXiv:2002.04803

cuml, cuml

A Simple Framework for Contrastive Learning of Visual Representations

arXiv2 repos

arXiv:2002.05709

vilmedic, BiomedVLP-CXR-BERT-specialized

UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

arXiv2 repos

arXiv:2002.06353

UniVL-video-captioning, video-captioning

MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers

arXiv2 repos

arXiv:2002.10957

torchserve-all-minilm-l6-v2, Multilingual-MiniLM-L12-H384

TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages

arXiv2 repos

arXiv:2003.05002

squad_bn, tydiqa-primary-task-xlm-roberta-large

PathVQA: 30000+ Questions for Medical Visual Question Answering

arXiv2 repos

arXiv:2003.10286

path-vqa, LLaDA-MedV

KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding

arXiv2 repos

arXiv:2004.03289

kor_nli, kor_nlu

Deep Learning Models for Multilingual Hate Speech Detection

arXiv2 repos

arXiv:2004.06465

dehatebert-mono-english, DE-LIMIT

Deep Generation of Coq Lemma Names Using Elaborated Terms

arXiv2 repos

arXiv:2004.07761

coq-serapi, coq-serapi

Multi-Dimensional Gender Bias Classification

arXiv2 repos

arXiv:2005.00614

STAIR-Captions, rebel-dataset

Conformer: Convolution-augmented Transformer for Speech Recognition

arXiv2 repos

arXiv:2005.08100

GigaAM, stt_eu_conformer_ctc_large

Visual Transformers: Token-based Image Representation and Processing for Computer Vision

arXiv2 repos

arXiv:2006.03677

vit-base-patch16-224-in21k, vit-large-patch16-224-in21k

Unsupervised Cross-lingual Representation Learning for Speech Recognition

arXiv2 repos

arXiv:2006.13979

fairseq, wav2vec2-large-xlsr-53

Learning to Format Coq Code Using Language Models

arXiv2 repos

arXiv:2006.16743

coq-serapi, coq-serapi

Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval

arXiv2 repos

arXiv:2007.00808

pyterrier_dr, pyterrier_dr_jpq

Big Bird: Transformers for Longer Sequences

arXiv2 repos

arXiv:2007.14062

final-project-level3-nlp-02, Block-Sparse-Attention

Multilingual Translation with Extensible Multilingual Pretraining and Finetuning

arXiv2 repos

arXiv:2008.00401

EasyNMT, mbart-large-50-many-to-many-mmt

TransNet V2: An effective deep network architecture for fast shot transition detection

arXiv2 repos

arXiv:2008.04838

shotplan, TransNetV2

KILT: a Benchmark for Knowledge Intensive Language Tasks

arXiv2 repos

arXiv:2009.02252

instructor-embedding, instructor-embedding

Unconstrained Text Detection in Manga: a New Dataset and Baseline

arXiv2 repos

arXiv:2009.04042

YuzuMarker.FontDetection, YuzuMarker.FontDetection

Large-Scale Intelligent Microservices

arXiv2 repos

arXiv:2009.08044

awesome-spark, SynapseML

Autoregressive Entity Retrieval

arXiv2 repos

arXiv:2010.00904

OKEAN, trusted_ke

D3Net: Densely connected multidilated DenseNet for music source separation

arXiv2 repos

arXiv:2010.01733

demucs, demucs

Denoising Diffusion Implicit Models

arXiv2 repos

arXiv:2010.02502

DDPM_vs_DDIM, smalldiffusion

Deformable DETR: Deformable Transformers for End-to-End Object Detection

arXiv2 repos

arXiv:2010.04159

rf-detr, DINO

Distilling Dense Representations for Ranking using Tightly-Coupled Teachers

arXiv2 repos

arXiv:2010.11386

pyterrier_dr, pyterrier_dr_jpq

GPUTreeShap: Massively Parallel Exact Calculation of SHAP Scores for Tree Ensembles

arXiv2 repos

arXiv:2010.13972

shap, shap

AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

arXiv2 repos

arXiv:2010.15980

Magic_Words, imodelsX

Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

arXiv2 repos

arXiv:2011.01060

GraphKV, GraphKV

Multimodal Pretraining for Dense Video Captioning

arXiv2 repos

arXiv:2011.11760

TimeChat-Online-139K, Video-Timeline-Tags-ViTT

Score-Based Generative Modeling through Stochastic Differential Equations

arXiv2 repos

arXiv:2011.13456

DragonDiffusion, DDPM_vs_DDIM

The Third DIHARD Diarization Challenge

arXiv2 repos

arXiv:2012.01477

pyannote-audio, hf-speaker-diarization-3.1

TabTransformer: Tabular Data Modeling Using Contextual Embeddings

arXiv2 repos

arXiv:2012.06678

pytorch-frame, Trompt

Extracting Smart Contracts Tested and Verified in Coq

arXiv2 repos

arXiv:2012.09138

ConCert, coq-rust-extraction

DeepHateExplainer: Explainable Hate Speech Detection in Under-resourced Bengali Language

arXiv2 repos

arXiv:2012.14353

bangla-bert-base, bangla-bert

Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup

arXiv2 repos

arXiv:2101.06983

pylate, langcache-embed-v2

arXiv:2102.01454

arXiv2 repos

arXiv:2102.01454

sdtt, SDTT-LaViDa

PAQ: 65 Million Probably-Asked Questions and What You Can Do With Them

arXiv2 repos

arXiv:2102.07033

all-MiniLM-L6-v2, all-MiniLM-L12-v2

Zero-Shot Text-to-Image Generation

arXiv2 repos

arXiv:2102.12092

DALL-E, vilmedic

Roosterize: Suggesting Lemma Names for Coq Verification Projects Using Deep Learning

arXiv2 repos

arXiv:2103.01346

coq-serapi, coq-serapi

Perceiver: General Perception with Iterative Attention

arXiv2 repos

arXiv:2103.03206

idefics2-8b, PathBench-MIL

Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training

arXiv2 repos

arXiv:2104.01027

fairseq, wav2vec2-large-robust

Aggregated Contextual Transformations for High-Resolution Image Inpainting

arXiv2 repos

arXiv:2104.01431

comic-translate, UnComicTranslate

AST: Audio Spectrogram Transformer

arXiv2 repos

arXiv:2104.01778

ast-finetuned-audioset-10-10-0.4593, ast

FUDGE: Controlled Text Generation With Future Discriminators

arXiv2 repos

arXiv:2104.05218

constrDecoding, naacl-2021-fudge-controlled-generation

LocalViT: Analyzing Locality in Vision Transformers

arXiv2 repos

arXiv:2104.05707

TexTok-DiT, transformer_latent_diffusion

Emotion Classification in a Resource Constrained Language Using Transformer-based Approach

arXiv2 repos

arXiv:2104.08613

bangla-bert-base, bangla-bert

GooAQ: Open Question Answering with Diverse Answer Types

arXiv2 repos

arXiv:2104.08727

all-MiniLM-L6-v2, all-MiniLM-L12-v2

GermanQuAD and GermanDPR: Improving Non-English Question Answering and Passage Retrieval

arXiv2 repos

arXiv:2104.12741

germandpr, germanquad

MLP-Mixer: An all-MLP Architecture for Vision

arXiv2 repos

arXiv:2105.01601

vision_transformer, Swin-Transformer

A Large-Scale Benchmark for Food Image Segmentation

arXiv2 repos

arXiv:2105.05409

FoodSeg103, FoodSeg103-Benchmark-v1

Measuring Coding Challenge Competence With APPS

arXiv2 repos

arXiv:2105.09938

apps, apps

CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks

arXiv2 repos

arXiv:2105.12655

Project_CodeNet, code_contests

CTSpine1K: A Large-Scale Dataset for Spinal Vertebrae Segmentation in Computed Tomography

arXiv2 repos

arXiv:2105.14711

CTSpine1K, CTSpine1K

The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

arXiv2 repos

arXiv:2106.03193

fleurs, flores_101

GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

arXiv2 repos

arXiv:2106.06909

unispeech-sat-large, wavlm-large

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

arXiv2 repos

arXiv:2106.07447

tmh, hubert-large-ls960-ft

Revisiting Deep Learning Models for Tabular Data

arXiv2 repos

arXiv:2106.11959

pytorch-frame, Trompt

Variational Diffusion Models

arXiv2 repos

arXiv:2107.00630

TexTok-DiT, transformer_latent_diffusion

ERNIE 3.0: Large-scale Knowledge Enhanced Pre-training for Language Understanding and Generation

arXiv2 repos

arXiv:2107.02137

ernie-3.0-base-zh, ernie-3.0-nano-zh

A Review of Bangla Natural Language Processing Tasks and the Utility of Transformer Models

arXiv2 repos

arXiv:2107.03844

bangla-bert-base, bangla-bert

SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking

arXiv2 repos

arXiv:2107.05720

bge-m3, splade

Deduplicating Training Data Makes Language Models Better

arXiv2 repos

arXiv:2107.06499

falcon-refinedweb, KoCommercial-Dataset

QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries

arXiv2 repos

arXiv:2107.09609

TimeChat-Online-139K, moment_detr

Extracting functional programs from Coq, in Coq

arXiv2 repos

arXiv:2108.02995

ConCert, coq-rust-extraction

MMChat: Multi-Modal Chat Dataset on Social Media

arXiv2 repos

arXiv:2108.07154

MMChat, mmchat

Program Synthesis with Large Language Models

arXiv2 repos

arXiv:2108.07732

FTTT, mbpp

Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models

arXiv2 repos

arXiv:2108.08877

sentence-t5-base, sentence-t5-xxl

An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA

arXiv2 repos

arXiv:2109.05014

idefics-80b-instruct, idefics-9b-instruct

Decoupling Magnitude and Phase Estimation with Deep ResUNet for Music Source Separation

arXiv2 repos

arXiv:2109.05418

demucs, demucs

Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

arXiv2 repos

arXiv:2109.10686

turkish-bert, bert5urk

Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

arXiv2 repos

arXiv:2109.11680

fairseq, wav2vec2-xlsr-53-espeak-cv-ft

Swiss-Judgment-Prediction: A Multilingual Legal Judgment Prediction Benchmark

arXiv2 repos

arXiv:2110.00806

SwissJudgementPrediction, swiss_judgment_prediction

UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training

arXiv2 repos

arXiv:2110.05752

UniSpeech, unispeech-sat-large

Mengzi: Towards Lightweight yet Ingenious Pre-trained Models for Chinese

arXiv2 repos

arXiv:2110.06696

EasyNLP, t5-chinese-couplet

ByteTrack: Multi-Object Tracking by Associating Every Detection Box

arXiv2 repos

arXiv:2110.06864

D-FINE-seg, CoreML-Models

Ego4D: Around the World in 3,000 Hours of Egocentric Video

arXiv2 repos

arXiv:2110.07058

pyannote-audio, Ego4d

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

arXiv2 repos

arXiv:2110.07205

speecht5_tts, SpeechT5

Pre-training Molecular Graph Representation with 3D Geometry

arXiv2 repos

arXiv:2110.07728

OpenBioMed, OpenBioMed_new

Hybrid Spectrogram and Waveform Source Separation

arXiv2 repos

arXiv:2111.03600

demucs, demucs

Prune Once for All: Sparse Pre-Trained Language Models

arXiv2 repos

arXiv:2111.05754

inteLearn_ML, intel-extension-for-transformers

Merging Models with Fisher-Weighted Averaging

arXiv2 repos

arXiv:2111.09832

Mario, MergeLM

AVA-AVD: Audio-Visual Speaker Diarization in the Wild

arXiv2 repos

arXiv:2111.14448

pyannote-audio, hf-speaker-diarization-3.1

A General Language Assistant as a Laboratory for Alignment

arXiv2 repos

arXiv:2112.00861

stack-exchange-preferences, MERA

Bidimensional Leaderboards: Generate and Evaluate Language Hand in Hand

arXiv2 repos

arXiv:2112.04139

instructor-embedding, instructor-embedding

FaceFormer: Speech-Driven 3D Facial Animation with Transformers

arXiv2 repos

arXiv:2112.05329

FaceFormer, faceformer-emo

LongT5: Efficient Text-To-Text Transformer for Long Sequences

arXiv2 repos

arXiv:2112.07916

long-t5-tglobal-xl-16384-book-summary, long-ke-t5

QuALITY: Question Answering with Long Input Texts, Yes!

arXiv2 repos

arXiv:2112.08608

Synthetic_Continued_Pretraining, SoE

Unsupervised Dense Information Retrieval with Contrastive Learning

arXiv2 repos

arXiv:2112.09118

contriever, contriever-msmarco

Image Segmentation Using Text and Image Prompts

arXiv2 repos

arXiv:2112.10003

Semantic-Segment-Anything, FastSAM

Efficient Large Scale Language Modeling with Mixtures of Experts

arXiv2 repos

arXiv:2112.10684

fairseq-dense-13B, fairseq-dense-2.7B

Scaling Language Models: Methods, Analysis & Insights from Training Gopher

arXiv2 repos

arXiv:2112.11446

falcon-refinedweb, GlorIA

Collapse by Conditioning: Training Class-conditional GANs with Limited Data

arXiv2 repos

arXiv:2201.06578

wound-stylegan, transitional-cGAN

Synchromesh: Reliable code generation from pre-trained language models

arXiv2 repos

arXiv:2201.11227

syncode, syncode-cypher-lark

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

arXiv2 repos

arXiv:2201.11903

prompt-engineering, tree-of-thought-prompting

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

arXiv2 repos

arXiv:2202.03052

Levels_image_captioning_NICE, OFA-Chinese

DiT: Self-supervised Pre-training for Document Image Transformer

arXiv2 repos

arXiv:2203.02378

TexTok-DiT, transformer_latent_diffusion

Dawn of the transformer era in speech emotion recognition: closing the valence gap

arXiv2 repos

arXiv:2203.07378

wav2vec2-large-robust-12-ft-emotion-msp-dim, w2v2-how-to

RELIC: Retrieving Evidence for Literary Claims

arXiv2 repos

arXiv:2203.10053

RAG, RAGatouille

ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

arXiv2 repos

arXiv:2203.10244

VisRAG-Ret-Train-In-domain-data, ChartQA

MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

arXiv2 repos

arXiv:2203.14371

PodGPT, medmcqa

Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$

arXiv2 repos

arXiv:2203.17189

xmtf, t5x

AdaFace: Quality Adaptive Margin for Face Recognition

arXiv2 repos

arXiv:2204.00964

AdaFace, FLUXSynID

Making the Most of Text Semantics to Improve Biomedical Vision--Language Processing

arXiv2 repos

arXiv:2204.09817

BiomedVLP-BioViL-T, BiomedVLP-CXR-BERT-specialized

Evaluating Interpolation and Extrapolation Performance of Neural Retrieval Models

arXiv2 repos

arXiv:2204.11447

RAG, RAGatouille

StyleGAN-Human: A Data-Centric Odyssey of Human Generation

arXiv2 repos

arXiv:2204.11823

CosmicMan, DeepFashion-MultiModal

Flamingo: a Visual Language Model for Few-Shot Learning

arXiv2 repos

arXiv:2204.14198

idefics-80b-instruct, idefics-9b-instruct

CoCa: Contrastive Captioners are Image-Text Foundation Models

arXiv2 repos

arXiv:2205.01917

VLSA, VL-KE-T5

MS-Shift: An Analysis of MS MARCO Distribution Shifts on Neural Retrieval

arXiv2 repos

arXiv:2205.02870

RAG, RAGatouille

arXiv:2205.09911

arXiv2 repos

arXiv:2205.09911

Jellyfish-13B, fm_data_tasks

hmBERT: Historical Multilingual Language Models for Named Entity Recognition

arXiv2 repos

arXiv:2205.15575

bert-base-historic-multilingual-cased, clef-hipe

Elucidating the Design Space of Diffusion-Based Generative Models

arXiv2 repos

arXiv:2206.00364

playground-v2.5-1024px-aesthetic, YetAnotherStableDiffusion

No Parameter Left Behind: How Distillation and Model Size Affect Zero-Shot Retrieval

arXiv2 repos

arXiv:2206.02873

monot5-3b-msmarco-10k, scaling-zero-shot-retrieval

CLAP: Learning Audio Concepts From Natural Language Supervision

arXiv2 repos

arXiv:2206.04769

vocalsound, shira_audio

SkexGen: Autoregressive Generation of CAD Construction Sequences with Disentangled Codebooks

arXiv2 repos

arXiv:2207.04632

FlexCAD, SkexGen

WISE: Whitebox Image Stylization by Example-based Learning

arXiv2 repos

arXiv:2207.14606

Whitebox-Style-Transfer-Editing, wise

Evaluating Table Structure Recognition: A New Perspective

arXiv2 repos

arXiv:2208.00385

granite-vision-4.1-4b, granite-4.0-3b-vision

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

arXiv2 repos

arXiv:2208.07339

vigogne, textsum

Selective Annotation Makes Language Models Better Few-Shot Learners

arXiv2 repos

arXiv:2209.01975

instructor-embedding, instructor-embedding

AudioLM: a Language Modeling Approach to Audio Generation

arXiv2 repos

arXiv:2209.03143

bark, ultravox

An Empirical Study on Cross-X Transfer for Legal Judgment Prediction

arXiv2 repos

arXiv:2209.12325

SwissJudgementPrediction, swiss_judgment_prediction

Music Source Separation with Band-split RNN

arXiv2 repos

arXiv:2209.15174

demucs, demucs

SMaLL-100: Introducing Shallow Multilingual Machine Translation Model for Low-Resource Languages

arXiv2 repos

arXiv:2210.11621

small100, ContraDecode

Contrastive Search Is What You Need For Neural Text Generation

arXiv2 repos

arXiv:2210.14140

Adaptive-Contrastive-Search, 4th-Bookathon-The-Unbearable-Heaviness-of-GPT

Autoregressive Structured Prediction with Language Models

arXiv2 repos

arXiv:2210.14698

trusted_ke, knowledge-graph-on-research-paper

QuaLA-MiniLM: a Quantized Length Adaptive MiniLM

arXiv2 repos

arXiv:2210.17114

inteLearn_ML, intel-extension-for-transformers

Large Language Models Are Human-Level Prompt Engineers

arXiv2 repos

arXiv:2211.01910

YiVal, PRL-Prompts-from-Reinforcement-Learning

Efficient Spatially Sparse Inference for Conditional GANs and Diffusion Models

arXiv2 repos

arXiv:2211.02048

nunchaku, nunchaku

Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

arXiv2 repos

arXiv:2211.05105

i2p, i2p

Fast DistilBERT on CPUs

arXiv2 repos

arXiv:2211.07715

inteLearn_ML, intel-extension-for-transformers

Hybrid Transformers for Music Source Separation

arXiv2 repos

arXiv:2211.08553

demucs, demucs

InstructPix2Pix: Learning to Follow Image Editing Instructions

arXiv2 repos

arXiv:2211.09800

instruct-pix2pix, SEED

Solving math word problems with process- and outcome-based feedback

arXiv2 repos

arXiv:2211.14275

Qwen2.5-Math-7B-Instruct-PRM-0.2, Qwen2.5-Math-1.5B-Instruct-PRM-0.2

Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models

arXiv2 repos

arXiv:2212.03860

cool-japan-diffusion-for-learning-2-0, picasso-diffusion-1-1

Fast Number Parsing Without Fallback

arXiv2 repos

arXiv:2212.06644

ffc.h, fast_float

RTMDet: An Empirical Study of Designing Real-Time Object Detectors

arXiv2 repos

arXiv:2212.07784

mmyolo, mmdetection

Discovering Language Model Behaviors with Model-Written Evaluations

arXiv2 repos

arXiv:2212.09251

unintentional-unalignment, model-written-evals

MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation

arXiv2 repos

arXiv:2212.09478

MMDisCo, MM-Diffusion

The case for 4-bit precision: k-bit Inference Scaling Laws

arXiv2 repos

arXiv:2212.09720

airllm, GPTQ-for-LLaMa

One Embedder, Any Task: Instruction-Finetuned Text Embeddings

arXiv2 repos

arXiv:2212.09741

instructor-embedding, instructor-embedding

Dataless Knowledge Fusion by Merging Weights of Language Models

arXiv2 repos

arXiv:2212.09849

Mario, MergeLM

DDColor: Towards Photo-Realistic Image Colorization via Dual Decoders

arXiv2 repos

arXiv:2212.11613

DDColor, DDColor-models

Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP

arXiv2 repos

arXiv:2212.14024

dspy, dsp

Muse: Text-To-Image Generation via Masked Generative Transformers

arXiv2 repos

arXiv:2301.00704

open-muse, NTU_ADL_Team11_Final

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

arXiv2 repos

arXiv:2301.02111

ChatTTS, bark

ExcelFormer: A neural network surpassing GBDTs on tabular data

arXiv2 repos

arXiv:2301.02819

pytorch-frame, Trompt

FullStop:Punctuation and Segmentation Prediction for Dutch with Transformers

arXiv2 repos

arXiv:2301.03319

fullstop-punctuation-multilingual-base, fullstop-dutch-punctuation-prediction

Mastering Diverse Domains through World Models

arXiv2 repos

arXiv:2301.04104

d-dreamerv3-world-model, dreamerv3

SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images

arXiv2 repos

arXiv:2301.04883

sa2va_eval, VisRAG-Ret-Train-In-domain-data

MusicLM: Generating Music From Text

arXiv2 repos

arXiv:2301.11325

MusicCaps, open-musiclm

Multimodal Chain-of-Thought Reasoning in Language Models

arXiv2 repos

arXiv:2302.00923

Awesome-Multimodal-Prompts, MUStReason

Black Box Adversarial Prompting for Foundation Models

arXiv2 repos

arXiv:2302.04237

adversarial_prompting, JailbreakLab

Q-Diffusion: Quantizing Diffusion Models

arXiv2 repos

arXiv:2302.04304

nunchaku, nunchaku

A Text-guided Protein Design Framework

arXiv2 repos

arXiv:2302.04611

ChatDrug, ProteinCLAP_pretrain_EBM_NCE_downstream_property_prediction

Large Language Models for Code: Security Hardening and Adversarial Testing

arXiv2 repos

arXiv:2302.05319

sven_modified, sven

Scaling Vision Transformers to 22 Billion Parameters

arXiv2 repos

arXiv:2302.05442

idefics-80b-instruct, idefics-9b-instruct

Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

arXiv2 repos

arXiv:2302.09664

moralchoice, Spnda

On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective

arXiv2 repos

arXiv:2302.12095

Julia_bench, robustlearn

ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth

arXiv2 repos

arXiv:2302.12288

MiDaS, prisma

WhisperX: Time-Accurate Speech Transcription of Long-Form Audio

arXiv2 repos

arXiv:2303.00747

whisperX, dissertation-project

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

arXiv2 repos

arXiv:2303.04137

diffusion_pusht, stable-worldmodel

Eliciting Latent Predictions from Transformers with the Tuned Lens

arXiv2 repos

arXiv:2303.08112

TransformerLens, notebooks

GPT-4 Technical Report

arXiv2 repos

arXiv:2303.08774

openchat_3.5, openchat-3.5-0106

EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face Animation

arXiv2 repos

arXiv:2303.11089

EmoTalk_release, 3DFaceAnimation

CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion

arXiv2 repos

arXiv:2303.11916

CompoDiff, CompoDiff-Charlie

VideoXum: Cross-modal Visual and Textural Summarization of Videos

arXiv2 repos

arXiv:2303.12060

videoxum, videoxum

On the De-duplication of LAION-2B

arXiv2 repos

arXiv:2303.12733

idefics-80b-instruct, idefics-9b-instruct

EVA-CLIP: Improved Training Techniques for CLIP at Scale

arXiv2 repos

arXiv:2303.15389

eva02_large_patch14_448.mim_m38m_ft_in1k, EVA-CLIP

The Stable Signature: Rooting Watermarks in Latent Diffusion Models

arXiv2 repos

arXiv:2303.15435

impossibility-watermark, WMCopier

HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

arXiv2 repos

arXiv:2303.17580

HuggingGPT, JARVIS

Generative Agents: Interactive Simulacra of Human Behavior

arXiv2 repos

arXiv:2304.03442

khms-memory, generative_agents

RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

arXiv2 repos

arXiv:2304.06767

LMFlow, FsfairX-LLaMA3-RM-v0.1

Chinese Open Instruction Generalist: A Preliminary Release

arXiv2 repos

arXiv:2304.07987

COIG-CQIA, COIG

Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca

arXiv2 repos

arXiv:2304.08177

Chinese-LLaMA-Alpaca-2, Chinese-LLaMA-Alpaca-3

UPGPT: Universal Diffusion Model for Person Image Generation, Editing and Pose Transfer

arXiv2 repos

arXiv:2304.08870

upgpt, upgpt

Measuring Massive Multitask Chinese Understanding

arXiv2 repos

arXiv:2304.12986

ChatGLM-6B, ChatGLM-6B

Towards Automated Circuit Discovery for Mechanistic Interpretability

arXiv2 repos

arXiv:2304.14997

TransformerLens, Automatic-Circuit-Discovery

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

arXiv2 repos

arXiv:2304.15010

LLaMA-Adapter, LLaMA2-Accessory

LMEye: An Interactive Perception Network for Large Language Models

arXiv2 repos

arXiv:2305.03701

LingCloud, Multimodal_Instruction_data_v1

Otter: A Multi-Modal Model with In-Context Instruction Tuning

arXiv2 repos

arXiv:2305.03726

Otter, Otter

VCSUM: A Versatile Chinese Meeting Summarization Dataset

arXiv2 repos

arXiv:2305.05280

Qwen-7B-Chat, Qwen-14B-Chat

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

arXiv2 repos

arXiv:2305.05662

InternGPT, InternChat

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

arXiv2 repos

arXiv:2305.06500

LAVIS, instructblip-vicuna-7b

AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages

arXiv2 repos

arXiv:2305.06897

afriqa, afriqa

Evaluating Object Hallucination in Large Vision-Language Models

arXiv2 repos

arXiv:2305.10355

POPE, POPE

PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering

arXiv2 repos

arXiv:2305.10415

MedAI-project, ClinicalNLP_PMCVQA

arXiv:2305.10427

arXiv2 repos

arXiv:2305.10427

LookaheadDecoding, LookAhead

Tree of Thoughts: Deliberate Problem Solving with Large Language Models

arXiv2 repos

arXiv:2305.10601

tree-of-thought-prompting, minihf

LDM3D: Latent Diffusion Model for 3D

arXiv2 repos

arXiv:2305.10853

MiDaS, ldm3d-4c

UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

arXiv2 repos

arXiv:2305.11147

MultiGen-20M_train, UniControl

AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation

arXiv2 repos

arXiv:2305.11408

WhisperLiveKit, naist-simulst

TheoremQA: A Theorem-driven Question Answering dataset

arXiv2 repos

arXiv:2305.12524

RLPR-Evaluation, KOpen-platypus

Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning

arXiv2 repos

arXiv:2305.13971

syncode, syncode-cypher-lark

WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia

arXiv2 repos

arXiv:2305.14292

WikiChat, wikipedia

This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models

arXiv2 repos

arXiv:2305.14610

borderlines, borderlines

NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario

arXiv2 repos

arXiv:2305.14836

nuscenes-qa-mini, NuScenes-QA

Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language Models

arXiv2 repos

arXiv:2305.15023

LaVIN, A-Lightweight-Unified-Autoregressive-MLLM

ViTMatte: Boosting Image Matting with Pretrained Plain Vision Transformers

arXiv2 repos

arXiv:2305.15272

vitmatte-small-composition-1k, ViTMatte

Expand, Rerank, and Retrieve: Query Reranking for Open-Domain Question Answering

arXiv2 repos

arXiv:2305.17080

EAR, query_decomposer

Trompt: Towards a Better Deep Neural Network for Tabular Data

arXiv2 repos

arXiv:2305.18446

pytorch-frame, Trompt

Grammar Prompting for Domain-Specific Language Generation with Large Language Models

arXiv2 repos

arXiv:2305.19234

grammar-prompting, blendsql

VideoComposer: Compositional Video Synthesis with Motion Controllability

arXiv2 repos

arXiv:2306.02018

videocomposer, i2vgen-xl

Large-Scale Cell Representation Learning via Divide-and-Conquer Contrastive Learning

arXiv2 repos

arXiv:2306.04371

OpenBioMed, OpenBioMed_new

DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text

arXiv2 repos

arXiv:2306.05540

L2D, AdaDetectGPT

StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

arXiv2 repos

arXiv:2306.07691

Kokoro-82M, kokoro-82M-onnx-opt

CMMLU: Measuring massive multitask language understanding in Chinese

arXiv2 repos

arXiv:2306.09212

cmmlu, PodGPT

Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

arXiv2 repos

arXiv:2306.09341

HPDv2, banana100-additional-iqa-models

Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

arXiv2 repos

arXiv:2306.14565

LRV-Instruction, prismatic-vlms

FunQA: Towards Surprising Video Comprehension

arXiv2 repos

arXiv:2306.14899

FunQA, FunQA

3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

arXiv2 repos

arXiv:2306.15354

3D-Speaker, 3d-speaker

One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization

arXiv2 repos

arXiv:2306.16928

torchsparse, One-2-3-45

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

arXiv2 repos

arXiv:2306.17203

Diff-Foley, Diff-Foley

BatGPT: A Bidirectional Autoregessive Talker from Generative Pre-trained Transformer

arXiv2 repos

arXiv:2307.00360

CMMLU, cmmlu-debug

DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models

arXiv2 repos

arXiv:2307.02421

DragonDiffusion, DragonDiffusion

Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers

arXiv2 repos

arXiv:2307.03183

whisper-at, whisper-at

RADAR: Robust AI-Text Detection via Adversarial Learning

arXiv2 repos

arXiv:2307.03838

L2D, AdaDetectGPT

Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++

arXiv2 repos

arXiv:2307.07686

HPC_Fortran_CPP, OpenMP-Fortran-CPP-Translation

Planting a SEED of Vision in Large Language Model

arXiv2 repos

arXiv:2307.08041

SEED, SEED

MolFM: A Multimodal Molecular Foundation Model

arXiv2 repos

arXiv:2307.09484

OpenBioMed, OpenBioMed_new

Evaluating the Moral Beliefs Encoded in LLMs

arXiv2 repos

arXiv:2307.14324

moralchoice, contextual_moralchoice

Prot2Text: Multimodal Protein's Function Generation with GNNs and Transformers

arXiv2 repos

arXiv:2307.14367

Prot2Text-Data, Prot2Text

MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

arXiv2 repos

arXiv:2307.14460

Depth-Estimation, MiDaS

Robust Distortion-free Watermarks for Language Models

arXiv2 repos

arXiv:2307.15593

impossibility-watermark, watermark

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

arXiv2 repos

arXiv:2307.16125

SEED-Bench, SEED-Bench

arXiv:2308.00264

arXiv2 repos

arXiv:2308.00264

MMML, eval-moshi

XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models

arXiv2 repos

arXiv:2308.01263

jailbreakbench, reward-bench

Tweet Insights: A Visualization Platform to Extract Temporal Insights from Twitter

arXiv2 repos

arXiv:2308.02142

twitter-roberta-base-2022-154m, twitter-roberta-large-2022-154m

PIPPA: A Partially Synthetic Conversational Dataset

arXiv2 repos

arXiv:2308.05884

PIPPA-shareGPT, PIPPA

BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

arXiv2 repos

arXiv:2308.05960

BOLAA, AgentLite

IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

arXiv2 repos

arXiv:2308.06721

IP-Adapter, IP-Adapter-FaceID

MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

arXiv2 repos

arXiv:2308.08544

MeViSv2, MeViS

Chinese Spelling Correction as Rephrasing Language Model

arXiv2 repos

arXiv:2308.08796

lemon, ReLM

CMB: A Comprehensive Medical Benchmark in Chinese

arXiv2 repos

arXiv:2308.08833

CMB, CMB

Steering Language Models With Activation Engineering

arXiv2 repos

arXiv:2308.10248

OBLITERATUS, obliteratus

arXiv:2308.11276

arXiv2 repos

arXiv:2308.11276

MU-LLaMA, MU-LLaMA

Large Language Models Vote: Prompting for Rare Disease Identification

arXiv2 repos

arXiv:2308.12890

llms-vote, llms-vote

SAM-Med2D

arXiv2 repos

arXiv:2308.16184

SA-Med2D-20M, SAM-Med2D

Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis

arXiv2 repos

arXiv:2308.16705

CREHate, CREHate

Baseline Defenses for Adversarial Attacks Against Aligned Language Models

arXiv2 repos

arXiv:2309.00614

jailbreakbench, llm-jailbreaking-defense

Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities

arXiv2 repos

arXiv:2309.00952

Bridge_Diffusion_Model, BDM1.0

Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models

arXiv2 repos

arXiv:2309.02218

DeepFakeFace, DeepFakeFace

Textbooks Are All You Need II: phi-1.5 technical report

arXiv2 repos

arXiv:2309.05463

cosmopedia, phi-1_5

Natural Language Supervision for General-Purpose Audio Representations

arXiv2 repos

arXiv:2309.05767

msclap, CLAP

Annotating Data for Fine-Tuning a Neural Ranker? Current Active Learning Strategies are not Better than Random Selection

arXiv2 repos

arXiv:2309.06131

RAG, RAGatouille

Mitigating Hallucinations and Off-target Machine Translation with Source-Contrastive and Language-Contrastive Decoding

arXiv2 repos

arXiv:2309.07098

ContraDecode, ContraDecode

Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context

arXiv2 repos

arXiv:2309.08105

libriheavy, libriheavy

LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

arXiv2 repos

arXiv:2309.12307

LongLoRA, LongLoRA

ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs

arXiv2 repos

arXiv:2309.13007

opencrabs, rightmind

Aligning Large Multimodal Models with Factually Augmented RLHF

arXiv2 repos

arXiv:2309.14525

vision-feedback-mix-binarized, vision-feedback-mix-binarized

ModuLoRA: Finetuning 2-Bit LLMs on Consumer GPUs by Integrating with Modular Quantizers

arXiv2 repos

arXiv:2309.16119

llmtools, llmtools

Data Filtering Networks

arXiv2 repos

arXiv:2309.17425

ml-mobileclip, hashing-baseline

AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ

arXiv2 repos

arXiv:2310.00367

AutomaTikZ, AutomaTikZ

TimeGPT-1

arXiv2 repos

arXiv:2310.03589

moment, nixtla

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

arXiv2 repos

arXiv:2310.03684

jailbreakbench, llm-jailbreaking-defense

DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

arXiv2 repos

arXiv:2310.03714

dspy, dsp

Aligning Text-to-Image Diffusion Models with Reward Backpropagation

arXiv2 repos

arXiv:2310.03739

CogVideoX-Fun-V1.1-Reward-LoRAs, EasyAnimateV5-Reward-LoRAs

Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

arXiv2 repos

arXiv:2310.04378

TCD-SDXL-LoRA, PixArt-LCM-XL-2-1024-MS

RECOMP: Improving Retrieval-Augmented LMs with Compression and Selective Augmentation

arXiv2 repos

arXiv:2310.04408

recomp, Prompt-Compression

UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

arXiv2 repos

arXiv:2310.05126

LLaVA-OneVision-Data, mPLUG-DocOwl

Scaling Laws of RoPE-based Extrapolation

arXiv2 repos

arXiv:2310.05209

Llama-3-70B-Instruct-Gradient-262k, Llama-3-70B-Instruct-Gradient-1048k

Compressing Context to Enhance Inference Efficiency of Large Language Models

arXiv2 repos

arXiv:2310.06201

Selective_Context, trimwise

DKEC: Domain Knowledge Enhanced Multi-Label Classification for Diagnosis Prediction

arXiv2 repos

arXiv:2310.07059

EMS-Pipeline, DKEC

CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving

arXiv2 repos

arXiv:2310.07240

CacheGen, CacheGen-CV

BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language Associations

arXiv2 repos

arXiv:2310.07276

OpenBioMed, OpenBioMed_new

arXiv:2310.07641

arXiv2 repos

arXiv:2310.07641

LLMBar, reward-bench

Jailbreaking Black Box Large Language Models in Twenty Queries

arXiv2 repos

arXiv:2310.08419

jailbreakbench, llm-jailbreaking-defense

BitNet: Scaling 1-bit Transformers for Large Language Models

arXiv2 repos

arXiv:2310.11453

BitNet, zeta

GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

arXiv2 repos

arXiv:2310.11513

geneval, banana100-additional-iqa-models

SALMONN: Towards Generic Hearing Abilities for Large Language Models

arXiv2 repos

arXiv:2310.13289

Speech-IFEval, SALMONN

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

arXiv2 repos

arXiv:2310.15008

Wonder3D, wonder3d_archive

DISC-FinLLM: A Chinese Financial Large Language Model based on Multiple Experts Fine-tuning

arXiv2 repos

arXiv:2310.15205

DISC-FinLLM, DISC-FinLLMa

DALE: Generative Data Augmentation for Low-Resource Legal NLP

arXiv2 repos

arXiv:2310.15799

DALE, DALE

ControlLLM: Augment Language Models with Tools by Searching on Graphs

arXiv2 repos

arXiv:2310.17796

ControlLLM, ControlLLM

Punica: Multi-Tenant LoRA Serving

arXiv2 repos

arXiv:2310.18547

lorax, llm-lora-hotswap

Foundation Models for Generalist Geospatial Artificial Intelligence

arXiv2 repos

arXiv:2310.18660

Prithvi-100M, granite-geospatial-biomass

Efficient LLM Inference on CPUs

arXiv2 repos

arXiv:2311.00502

inteLearn_ML, intel-extension-for-transformers

S-LoRA: Serving Thousands of Concurrent LoRA Adapters

arXiv2 repos

arXiv:2311.03285

S-LoRA, llm-lora-hotswap

Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

arXiv2 repos

arXiv:2311.03348

JBB-Behaviors, jailbreakbench

OtterHD: A High-Resolution Multi-modality Model

arXiv2 repos

arXiv:2311.04219

Otter, Otter

Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models

arXiv2 repos

arXiv:2311.04378

watermarks-remover, impossibility-watermark

LRM: Large Reconstruction Model for Single Image to 3D

arXiv2 repos

arXiv:2311.04400

TripoSR, Real3D

SA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks

arXiv2 repos

arXiv:2311.11969

SA-Med2D-20M, SAM-Med2D

Text-Guided Texturing by Synchronized Multi-View Diffusion

arXiv2 repos

arXiv:2311.12891

FlexiSyncMVD, SyncMVD

GeoChat: Grounded Large Vision-Language Model for Remote Sensing

arXiv2 repos

arXiv:2311.15826

GeoChat, GeoChat_Instruct

Self-correcting LLM-controlled Diffusion Models

arXiv2 repos

arXiv:2311.16090

Omost, SLD

Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine

arXiv2 repos

arXiv:2311.16452

UltraMedical, promptbase

Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding

arXiv2 repos

arXiv:2311.16922

VideoLLaMA2, AVProunRLForVideoLLaMa2

Visual Anagrams: Generating Multi-View Optical Illusions with Diffusion Models

arXiv2 repos

arXiv:2311.17919

visual_anagrams, LatentGenerativeAnamorphoses

Sequential Modeling Enables Scalable Learning for Large Vision Models

arXiv2 repos

arXiv:2312.00785

LVM, LVM_ckpts

OpenVoice: Versatile Instant Voice Cloning

arXiv2 repos

arXiv:2312.01479

OpenVoiceV2_Webui_resemble_enhance, OpenVoice

With Great Humor Comes Great Developer Engagement

arXiv2 repos

arXiv:2312.01680

faker, faker

TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

arXiv2 repos

arXiv:2312.02051

TimeChat-7b, TimeIT

PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth Estimation

arXiv2 repos

arXiv:2312.02284

PatchFusion, PatchFusion

Analyzing and Improving the Training Dynamics of Diffusion Models

arXiv2 repos

arXiv:2312.02696

EDM2-diffusers, edm2

MotionCtrl: A Unified and Flexible Motion Controller for Video Generation

arXiv2 repos

arXiv:2312.03641

MotionCtrl, MotionCtrl

Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

arXiv2 repos

arXiv:2312.03818

AlphaCLIP, MaskImageNet

PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

arXiv2 repos

arXiv:2312.04461

PhotoMaker, PhotoMaker-V2

Grounded Question-Answering in Long Egocentric Videos

arXiv2 repos

arXiv:2312.06505

GroundVQA, GroundVQA

Steering Llama 2 via Contrastive Activation Addition

arXiv2 repos

arXiv:2312.06681

OBLITERATUS, obliteratus

EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

arXiv2 repos

arXiv:2312.06722

embodied-eval, behaviour_subtask

SGLang: Efficient Execution of Structured Language Model Programs

arXiv2 repos

arXiv:2312.07104

sgl-learning-materials, sglang-vla

Modality Plug-and-Play: Elastic Modality Adaptation in Multimodal LLMs for Embodied AI

arXiv2 repos

arXiv:2312.07886

mPnP-LLM, nuscenes-qa-mini

Distributed Inference and Fine-tuning of Large Language Models Over The Internet

arXiv2 repos

arXiv:2312.08361

petals, bloombee_add_models

Triplane Meets Gaussian Splatting: Fast and Generalizable Single-View 3D Reconstruction with Transformers

arXiv2 repos

arXiv:2312.09147

TriplaneGaussian, TriplaneGaussian

IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages

arXiv2 repos

arXiv:2312.09508

RAG, RAGatouille

VidToMe: Video Token Merging for Zero-Shot Video Editing

arXiv2 repos

arXiv:2312.10656

VidToMe, VidToMe

Silkie: Preference Distillation for Large Visual Language Models

arXiv2 repos

arXiv:2312.10665

vision-feedback-mix-binarized, vision-feedback-mix-binarized

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

arXiv2 repos

arXiv:2312.12456

prosparse-llama-2-13b, prosparse-llama-2-7b

Generative Multimodal Models are In-Context Learners

arXiv2 repos

arXiv:2312.13286

Emu2, CoBSAT

HD-Painter: High-Resolution and Prompt-Faithful Text-Guided Image Inpainting with Diffusion Models

arXiv2 repos

arXiv:2312.14091

HD-Painter, HD-Painter

LingoQA: Visual Question Answering for Autonomous Driving

arXiv2 repos

arXiv:2312.14115

LingoQA, Cosmos-Reason2-32B

LangSplat: 3D Language Gaussian Splatting

arXiv2 repos

arXiv:2312.16084

LangSplat, LangSurf

Towards Better Monolingual Japanese Retrievers with Multi-Vector Models

arXiv2 repos

arXiv:2312.16144

RAG, RAGatouille

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

arXiv2 repos

arXiv:2401.00849

cosmo, Howto-Interlink7M

A Comprehensive Study of Knowledge Editing for Large Language Models

arXiv2 repos

arXiv:2401.01286

KnowLM, KnowEdit

AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI

arXiv2 repos

arXiv:2401.01651

AdaptiveDiffusion, Sampled_AIGCBench_text2image_ar_0.625

TinyLlama: An Open-Source Small Language Model

arXiv2 repos

arXiv:2401.02385

TinyLlama_v1.1, tinyllama-embed

Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks

arXiv2 repos

arXiv:2401.02731

speechless, speechless-sparsetral-16x7b-MoE

InstantID: Zero-shot Identity-Preserving Generation in Seconds

arXiv2 repos

arXiv:2401.07519

InstantID, InstantID

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

arXiv2 repos

arXiv:2401.10774

LLM-Sampling, llm_project

In-Context Learning for Extreme Multi-Label Classification

arXiv2 repos

arXiv:2401.12178

dspy, dsp

A Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray Interpretation

arXiv2 repos

arXiv:2401.12208

RadPhi-2, CheXagent-2-3b

Raidar: geneRative AI Detection viA Rewriting

arXiv2 repos

arXiv:2401.12970

RAFT, L2D

WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

arXiv2 repos

arXiv:2401.13919

UI-TARS, WebVoyager

TURNA: A Turkish Encoder-Decoder Language Model for Enhanced Understanding and Generation

arXiv2 repos

arXiv:2401.14373

bert5urk, turkish-lm-tuner

MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

arXiv2 repos

arXiv:2401.15391

GraphKV, GraphKV

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

arXiv2 repos

arXiv:2401.16420

internlm-xcomposer2-vl-7b, internlm-xcomposer2-4khd-7b

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval

arXiv2 repos

arXiv:2401.18059

drbrain, LARS

Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

arXiv2 repos

arXiv:2402.00159

dolma, llm_project

DiffEditor: Boosting Accuracy and Flexibility on Diffusion-based Image Editing

arXiv2 repos

arXiv:2402.02583

DragonDiffusion, DragonDiffusion

KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache

arXiv2 repos

arXiv:2402.02750

glq, KIVI

EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models

arXiv2 repos

arXiv:2402.03049

EasyInstruct, KnowLM

ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

arXiv2 repos

arXiv:2402.03804

prosparse-llama-2-13b, prosparse-llama-2-7b

EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

arXiv2 repos

arXiv:2402.04252

EVA-CLIP-8B, EVA-CLIP-18B

QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks

arXiv2 repos

arXiv:2402.04396

quip-sharp, llmtools

EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss

arXiv2 repos

arXiv:2402.05008

efficientvit, efficientvit-sam

LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation

arXiv2 repos

arXiv:2402.05054

LGM, LGM

JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

arXiv2 repos

arXiv:2402.05668

JailbreakRadar_Backup, JailbreakRadar

OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models

arXiv2 repos

arXiv:2402.06044

OpenToM, OpenToM

Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations

arXiv2 repos

arXiv:2402.07023

Llama3-OpenBioLLM-8B, Llama3-OpenBioLLM-70B

An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

arXiv2 repos

arXiv:2402.08846

slam_asr_pytorch, SMIT

Generative Representational Instruction Tuning

arXiv2 repos

arXiv:2402.09906

sgpt, MEDI2

Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

arXiv2 repos

arXiv:2402.11411

vision-feedback-mix-binarized, vision-feedback-mix-binarized

FiT: Flexible Vision Transformer for Diffusion Model

arXiv2 repos

arXiv:2402.12376

Open-Sora-Plan-v1.2.0, FiT-diffusers

The Revolution of Multimodal Large Language Models: A Survey

arXiv2 repos

arXiv:2402.12451

LLaVA-MORE, LLaVA_MORE-gemma_2_9b-finetuning

Visual Style Prompting with Swapping Self-Attention

arXiv2 repos

arXiv:2402.12974

StyleKeeper, visual-style-prompting

Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

arXiv2 repos

arXiv:2402.13064

safe-guard-prompt-injection, slm-innovator-lab

Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

arXiv2 repos

arXiv:2402.14804

MathVision, MATH-V

PALO: A Polyglot Large Multimodal Model for 5B People

arXiv2 repos

arXiv:2402.14818

palo_multilingual_dataset, PALO

arXiv:2402.15391

arXiv2 repos

arXiv:2402.15391

Jazz, 1xgpt

Nemotron-4 15B Technical Report

arXiv2 repos

arXiv:2402.16819

Curator, NeMo-Curator

Transparent Image Layer Diffusion using Latent Transparency

arXiv2 repos

arXiv:2402.17113

Diffuser-layerdiffuse, layer_diffusers

Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

arXiv2 repos

arXiv:2402.17245

playground-v2.5-1024px-aesthetic, MJHQ-30K

Evaluating Very Long-Term Conversational Memory of LLM Agents

arXiv2 repos

arXiv:2402.17753

REALTALK, memex

ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models

arXiv2 repos

arXiv:2403.02084

res-adapter, res-adapter

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

arXiv2 repos

arXiv:2403.05135

ELLA, LaVi-Bridge

DeepSeek-VL: Towards Real-World Vision-Language Understanding

arXiv2 repos

arXiv:2403.05525

deepseek-vl-7b-chat, deepseek-vl-1.3b-chat

SPLADE-v3: New baselines for SPLADE

arXiv2 repos

arXiv:2403.06789

splade-v3-lexical-mlx, splade-v3-distilbert

ORPO: Monolithic Preference Optimization without Reference Model

arXiv2 repos

arXiv:2403.07691

tunix, MedicalGPT

AutoDev: Automated AI-Driven Development

arXiv2 repos

arXiv:2403.08299

pentagi, codel

CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model

arXiv2 repos

arXiv:2403.08350

Continual-NExT, CoIN_Refined

Can We Talk Models Into Seeing the World Differently?

arXiv2 repos

arXiv:2403.09193

vlm_shapebias, frequency-cue-conflict

DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers

arXiv2 repos

arXiv:2403.10266

MindSpeed-MM, OpenDiT

BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics

arXiv2 repos

arXiv:2403.10380

BirdSet, BirdSet

arXiv:2403.13164

arXiv2 repos

arXiv:2403.13164

VL-ICL, VL-ICL

Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity

arXiv2 repos

arXiv:2403.14403

Adaptive-RAG, raglite

MyVLM: Personalizing VLMs for User-Specific Queries

arXiv2 repos

arXiv:2403.14599

MyVLM, MyVLM

InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

arXiv2 repos

arXiv:2403.15377

InternVideo, internvideo-d2a11ea9

Protecting Copyrighted Material with Unique Identifiers in Large Language Model Training

arXiv2 repos

arXiv:2403.15740

RefAlign, Llama-2-7b-hf-conf-refalign

Explore until Confident: Efficient Exploration for Embodied Question Answering

arXiv2 repos

arXiv:2403.15941

embodied-eval, behaviour_subtask

COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

arXiv2 repos

arXiv:2403.18058

COIG-CQIA, ruozhiba_gpt4

SDSAT: Accelerating LLM Inference through Speculative Decoding with Semantic Adaptive Tokens

arXiv2 repos

arXiv:2403.18647

SDSAT, CodeLlama-SDSAT_L7_13B

Are We on the Right Way for Evaluating Large Vision-Language Models?

arXiv2 repos

arXiv:2403.20330

MMStar, MMStar

Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models

arXiv2 repos

arXiv:2403.20331

UPD, MM-UPD

PyTorch Frame: A Modular Framework for Multi-Modal Tabular Learning

arXiv2 repos

arXiv:2404.00776

pytorch-frame, Trompt

Release of Pre-Trained Models for the Japanese Language

arXiv2 repos

arXiv:2404.01657

bilingual-gpt-neox-4b-minigpt4, youri-7b-chat-gptq

Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack

arXiv2 repos

arXiv:2404.01833

GA, TUP-detection

Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

arXiv2 repos

arXiv:2404.02905

VAR, var

ReFT: Representation Finetuning for Language Models

arXiv2 repos

arXiv:2404.03592

reft_ethos, reft_chat7b_1k

Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese

arXiv2 repos

arXiv:2404.07824

heron-chat-git-ja-stablelm-base-7b-v1, Japanese-Heron-Bench

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

arXiv2 repos

arXiv:2404.07972

UI-TARS, OSWorld

View Selection for 3D Captioning via Diffusion Ranking

arXiv2 repos

arXiv:2404.07984

DiffSplat, Cap3D

Toward a Theory of Tokenization in LLMs

arXiv2 repos

arXiv:2404.08335

translation-api, writing-assistance-apis

Magic Clothing: Controllable Garment-Driven Image Synthesis

arXiv2 repos

arXiv:2404.09512

MagicClothing, MagicClothing

Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

arXiv2 repos

arXiv:2404.09967

Ctrl-Adapter, Ctrl-Adapter

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

arXiv2 repos

arXiv:2404.09990

HQ-Edit, HQ-Edit

Masked Autoencoders for Microscopy are Scalable Learners of Cellular Biology

arXiv2 repos

arXiv:2404.10242

maes_microscopy, rxrx3-core

MathWriting: A Dataset For Handwritten Mathematical Expression Recognition

arXiv2 repos

arXiv:2404.10690

MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash

A Study of Undefined Behavior Across Foreign Function Boundaries in Rust Libraries

arXiv2 repos

arXiv:2404.11671

miri, rust

STaRK: Benchmarking LLM Retrieval on Textual and Relational Knowledge Bases

arXiv2 repos

arXiv:2404.13207

stark, stark

From Local to Global: A Graph RAG Approach to Query-Focused Summarization

arXiv2 repos

arXiv:2404.16130

graphrag, agentic-graphrag

"Ask Me Anything": How Comcast Uses LLMs to Assist Agents in Real Time

arXiv2 repos

arXiv:2405.00801

haystack, haystack-ai

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

arXiv2 repos

arXiv:2405.01535

prometheus-7b-v2.0, JudgeBench

What matters when building vision-language models?

arXiv2 repos

arXiv:2405.02246

idefics2-8b, Idefics3-8B-Llama3

SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

arXiv2 repos

arXiv:2405.04007

SEED-X, SEED-Data-Edit

Granite Code Models: A Family of Open Foundation Models for Code Intelligence

arXiv2 repos

arXiv:2405.04324

granite-20b-code-instruct-8k, granite-20b-code-base-8k

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

arXiv2 repos

arXiv:2405.04434

minimind, Tiny-R2

SketchDream: Sketch-based Text-to-3D Generation and Editing

arXiv2 repos

arXiv:2405.06461

SketchDream, sketch-to-multiview

Exploring the Capabilities of Large Multimodal Models on Dense Text

arXiv2 repos

arXiv:2405.06706

MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash

LangCell: Language-Cell Pre-training for Cell Identity Understanding

arXiv2 repos

arXiv:2405.06708

OpenBioMed, OpenBioMed_new

A Comprehensive Analysis of Static Word Embeddings for Turkish

arXiv2 repos

arXiv:2405.07778

Word-Embeddings-Repository-for-Turkish, bert-turkish-x2static

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

arXiv2 repos

arXiv:2405.07940

sloptotal, raid

Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

arXiv2 repos

arXiv:2405.10300

efficientvit, Rex-Omni

Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

arXiv2 repos

arXiv:2405.11273

UMOE-Scaling-Unified-Multimodal-LLMs, Uni-MoE

RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search

arXiv2 repos

arXiv:2405.12497

turbovec, RaBitQ-Library

OLAPH: Improving Factuality in Biomedical Long-form Question Answering

arXiv2 repos

arXiv:2405.12701

OLAPH, MedLFQA

Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models

arXiv2 repos

arXiv:2405.16645

Diffusion4D, Diffusion4D

Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs

arXiv2 repos

arXiv:2405.16700

ima-lmms, IMA-DePALM

GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping

arXiv2 repos

arXiv:2405.17251

genwarp, genwarp

MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation

arXiv2 repos

arXiv:2405.17842

MMDisCo, MMDisCo

SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation

arXiv2 repos

arXiv:2405.18503

soundctm, soundctm

Xwin-LM: Strong and Scalable Alignment Practice for LLMs

arXiv2 repos

arXiv:2405.20335

Xwin-LM, Xwin-Math-70B-V1.0

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

arXiv2 repos

arXiv:2405.21060

mamba2-minimal, nugie-jax-nemotron-3-nano

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

arXiv2 repos

arXiv:2405.21075

video-tt, Video-MME

AudioLCM: Text-to-Audio Generation with Latent Consistency Models

arXiv2 repos

arXiv:2406.00356

AudioLCM, AudioLCM

YODAS: Youtube-Oriented Dataset for Audio and Speech

arXiv2 repos

arXiv:2406.00899

yodas, yodas2

SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning

arXiv2 repos

arXiv:2406.01006

semcoder_s_1030, semcoder_1030

V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

arXiv2 repos

arXiv:2406.02511

V-Express, V-Express

Parrot: Multilingual Visual Instruction Tuning

arXiv2 repos

arXiv:2406.02539

Ovis, Ovis

Wings: Learning Multimodal LLMs without Text-only Forgetting

arXiv2 repos

arXiv:2406.03496

Ovis, Ovis

VideoPhy: Evaluating Physical Commonsense for Video Generation

arXiv2 repos

arXiv:2406.03520

videophy, videophy_autoeval_scores

SilentCipher: Deep Audio Watermarking

arXiv2 repos

arXiv:2406.03822

silentcipher, SilentCipher

VideoTetris: Towards Compositional Text-to-Video Generation

arXiv2 repos

arXiv:2406.04277

VideoTetris, VideoTetris-long

VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

arXiv2 repos

arXiv:2406.04292

bge-visualized, VISTA_Evaluation_FineTuning

DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs

arXiv2 repos

arXiv:2406.04334

granite-vision-4.1-4b, granite-4.0-3b-vision

MAIRA-2: Grounded Radiology Report Generation

arXiv2 repos

arXiv:2406.04449

libra-maira-2, llava-rad

GenAI Arena: An Open Evaluation Platform for Generative Models

arXiv2 repos

arXiv:2406.04485

ImagenHub, VideoGenHub

F-LMM: Grounding Frozen Large Multimodal Models

arXiv2 repos

arXiv:2406.05821

F-LMM, F-LMM

MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

arXiv2 repos

arXiv:2406.06565

MixEval-Archon, MixEval

Spectrum: Targeted Training on Signal to Noise Ratio

arXiv2 repos

arXiv:2406.06623

spectrum, donutloop-genesis

PatchRefiner: Leveraging Synthetic Data for Real-Domain High-Resolution Monocular Metric Depth Estimation

arXiv2 repos

arXiv:2406.06679

PatchRefiner, PatchFusion

Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation

arXiv2 repos

arXiv:2406.06890

mcm, mcm

MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance

arXiv2 repos

arXiv:2406.07209

MS-Bench, MS-Diffusion

Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

arXiv2 repos

arXiv:2406.07546

Commonsense-T2I, CommonsensenT2I

An Image is Worth 32 Tokens for Reconstruction and Generation

arXiv2 repos

arXiv:2406.07550

tokenizer_titok_l32_imagenet, clustermark_1d-tokenizer

What If We Recaption Billions of Web Images with LLaMA-3?

arXiv2 repos

arXiv:2406.08478

Recap-DataComp-1B, ViT-L-16-HTxt-Recap-CLIP

DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage

arXiv2 repos

arXiv:2406.08820

disfluency_speech_german, disfluency_speech_english

Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs

arXiv2 repos

arXiv:2406.10216

GRM-Llama3.2-3B-rewardmodel-ft, GRM-Gemma-2B-rewardmodel-ft

SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations

arXiv2 repos

arXiv:2406.11171

scpp, SugarCrepe_pp

Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models

arXiv2 repos

arXiv:2406.11230

MMNeedle, multimodal-needle-in-a-haystack

Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs

arXiv2 repos

arXiv:2406.11695

dspy, dsp

Prompts as Auto-Optimized Training Hyperparameters: Training Best-in-Class IR Models from Scratch with 10 Gold Labels

arXiv2 repos

arXiv:2406.11706

dspy, dsp

DataComp-LM: In search of the next generation of training sets for language models

arXiv2 repos

arXiv:2406.11794

dclm-baseline-1.0, dclm-baseline-1.0-parquet

WebCanvas: Benchmarking Web Agents in Online Environments

arXiv2 repos

arXiv:2406.12373

Mind2Web_Live_SeeAct_V, mind2web-live-seeact-v

Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

arXiv2 repos

arXiv:2406.12742

MIRB, MIRB

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

arXiv2 repos

arXiv:2406.13352

felonybench, moltshield

xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics

arXiv2 repos

arXiv:2406.14553

xCOMET-lite, XCOMET-lite

AEM: Attention Entropy Maximization for Multiple Instance Learning based Whole Slide Image Classification

arXiv2 repos

arXiv:2406.15303

AEM, AEM-dataset

Improving Text-To-Audio Models with Synthetic Captions

arXiv2 repos

arXiv:2406.15487

tango-af-ac-ft-ac, tango-music-af-ft-mc

HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis

arXiv2 repos

arXiv:2406.16192

HEST, hest-tissue-seg

arXiv:2406.18521

arXiv2 repos

arXiv:2406.18521

CharXiv, CharXiv

arXiv:2406.18665

arXiv2 repos

arXiv:2406.18665

RouteLLM, llm-router

EmPO: Emotion Grounding for Empathetic Response Generation through Preference Optimization

arXiv2 repos

arXiv:2406.19071

zephyr-7b-sft-full124, zephyr-7b-sft-full124_d270

RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs

arXiv2 repos

arXiv:2406.19232

RuBLiMP, rublimp

ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos

arXiv2 repos

arXiv:2406.19392

ReXTime, ReXTime

OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding

arXiv2 repos

arXiv:2407.04923

omchat-v2.0-13B-single-beta_hf, omchat

Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?

arXiv2 repos

arXiv:2407.05134

Formulate_and_Solve, BeyondX

MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions

arXiv2 repos

arXiv:2407.06358

MiraData, MiraData

Vision language models are blind: Failing to translate detailed visual features into words

arXiv2 repos

arXiv:2407.06581

vision-llms-are-blind, vlmsareblind

Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions

arXiv2 repos

arXiv:2407.06723

ml-gbc, GBC10M-PromptGen-200M

Cue Point Estimation using Object Detection

arXiv2 repos

arXiv:2407.06823

cue-detr, edm-cue

WildGaussians: 3D Gaussian Splatting in the Wild

arXiv2 repos

arXiv:2407.08447

wild-gaussians, nerfonthego-undistorted

Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together

arXiv2 repos

arXiv:2407.10930

dspy, dsp

Goldfish: Vision-Language Understanding of Arbitrarily Long Videos

arXiv2 repos

arXiv:2407.12679

MiniGPT4-Video, TVQA-Long

IMAGDressing-v1: Customizable Virtual Dressing

arXiv2 repos

arXiv:2407.12705

IMAGDressing, IMAGDressing

LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

arXiv2 repos

arXiv:2407.12772

UniG2U, lmms-eval-mmllm

Phi-3 Safety Post-Training: Aligning Language Models with a "Break-Fix" Cycle

arXiv2 repos

arXiv:2407.13833

Phi-3.5-MoE-instruct, Phi-4-multimodal-instruct

arXiv:2407.14435

arXiv2 repos

arXiv:2407.14435

dictionary_learning, CLT-Forge

NV-Retriever: Improving text embedding models with effective hard-negative mining

arXiv2 repos

arXiv:2407.15831

LateOn-Code, LateOn-Code-edge

UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models

arXiv2 repos

arXiv:2407.18391

UOUO, UOUO-Bench

Towards A Generalizable Pathology Foundation Model via Unified Knowledge Distillation

arXiv2 repos

arXiv:2407.18449

GPFM, GPFM

mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval

arXiv2 repos

arXiv:2407.19669

gte-large-en-v1.5, gte-multilingual-base

Towards Localized Fine-Grained Control for Facial Expression Generation

arXiv2 repos

arXiv:2407.20175

fineface, fineface

Learning Feature-Preserving Portrait Editing from Generated Pairs

arXiv2 repos

arXiv:2407.20455

feature-preserve-portrait-editing, feature-preserve-portrait-editing

Black-Box Adversarial Attacks on LLM-Based Code Completion

arXiv2 repos

arXiv:2408.02509

insec, insec-vulnerability

MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

arXiv2 repos

arXiv:2408.02718

MMIU, MMIU-Benchmark

EXAONE 3.0 7.8B Instruction Tuned Language Model

arXiv2 repos

arXiv:2408.03541

KoMT-Bench, KoMT-Bench

Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling

arXiv2 repos

arXiv:2408.03695

openstorypp, OpenstoryPlusPlus

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

arXiv2 repos

arXiv:2408.03910

ms-agent, modelscope-agent

Med42-v2: A Suite of Clinical LLMs

arXiv2 repos

arXiv:2408.06142

Llama3-Med42-70B, Llama3-Med42-8B

ChemVLM: Exploring the Power of Multimodal Large Language Models in Chemistry Area

arXiv2 repos

arXiv:2408.07246

ChemVLM_test_data, ChemVLM-26B-1-2

Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning

arXiv2 repos

arXiv:2408.09600

Vaccine, Lisa

RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data

arXiv2 repos

arXiv:2408.12109

Vision-LLM-Alignment, robust_visual_reward_model

Sapiens: Foundation for Human Vision Models

arXiv2 repos

arXiv:2408.12569

sapiens-pose-1b-torchscript, sapiens

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

arXiv2 repos

arXiv:2408.13257

Video-MME, sa2va_eval

Benchmarking foundation models as feature extractors for weakly-supervised computational pathology

arXiv2 repos

arXiv:2408.15823

STAMP, STAMP_attention_ui

CSGO: Content-Style Composition in Text-to-Image Generation

arXiv2 repos

arXiv:2408.16766

CSGO, CSGO

VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters

arXiv2 repos

arXiv:2408.17253

uni2ts, uni2ts

OnlySportsLM: Optimizing Sports-Domain Language Models with SOTA Performance under Billion Parameters

arXiv2 repos

arXiv:2409.00286

OnlySportsLM, OnlySports_Dataset

MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

arXiv2 repos

arXiv:2409.00750

MaskGCT-Windows, maskgct

ToolACE: Winning the Points of LLM Function Calling

arXiv2 repos

arXiv:2409.00920

ToolACE, ToolACE-8B

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

arXiv2 repos

arXiv:2409.02813

MMMU, MMMU_Pro

NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls

arXiv2 repos

arXiv:2409.03797

NESTFUL, nestful

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

arXiv2 repos

arXiv:2409.06666

LLaMA-Omni, Llama-3.1-8B-Omni

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

arXiv2 repos

arXiv:2409.08425

SoloAudio, SoloAudio

Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models

arXiv2 repos

arXiv:2409.09788

SpaceThinker-Qwen2.5VL-3B, Q-Spatial-Bench-code

Eureka: Evaluating and Understanding Large Foundation Models

arXiv2 repos

arXiv:2409.10566

eureka-ml-insights, Eureka-Bench-Logs

EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer

arXiv2 repos

arXiv:2409.10819

EzAudio, EzAudio

OmniGen: Unified Image Generation

arXiv2 repos

arXiv:2409.11340

OmniGen, OmniGen-v1

Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion

arXiv2 repos

arXiv:2409.11406

Phidias-Diffusion, Phidias-Diffusion

Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

arXiv2 repos

arXiv:2409.12122

Qwen2.5-Math-7B-Instruct, Qwen2.5-Math-RM-72B

StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation

arXiv2 repos

arXiv:2409.12576

StoryMaker, StoryMaker

Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm

arXiv2 repos

arXiv:2409.12951

SimSIMD, numkong

MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding

arXiv2 repos

arXiv:2409.14818

mobilevlm_test, Mobile3M

DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D Diffusion

arXiv2 repos

arXiv:2409.17145

DreamWaltz-G, DreamWaltz-G

LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

arXiv2 repos

arXiv:2409.18125

LLaVA-3D-7B, LLaVA-3D

arXiv:2409.19425

arXiv2 repos

arXiv:2409.19425

freeze-align, concept_coverage_laion_6m

MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation

arXiv2 repos

arXiv:2409.19483

MedCLIP-SAMv2, MedCLIP-SAMv2

arXiv:2409.19603

arXiv2 repos

arXiv:2409.19603

VideoLISA, VideoLISA-3.8B

Preserving Generalization of Language models in Few-shot Continual Relation Extraction

arXiv2 repos

arXiv:2410.00334

sirus_v2, Minion

Do Music Generation Models Encode Music Theory?

arXiv2 repos

arXiv:2410.00872

syntheory, syntheory

nGPT: Normalized Transformer with Representation Learning on the Hypersphere

arXiv2 repos

arXiv:2410.01131

SimSIMD, numkong

Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks

arXiv2 repos

arXiv:2410.01744

Leopard-Idefics2, Leopard-LLaVA

SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration

arXiv2 repos

arXiv:2410.02367

SageAttention, FastVideo

Recent Advances in Speech Language Models: A Survey

arXiv2 repos

arXiv:2410.03751

VoxEval, VoxEval

Timer-XL: Long-Context Transformers for Unified Time Series Forecasting

arXiv2 repos

arXiv:2410.04803

timer-base-84m, Large-Time-Series-Model

Accelerating Diffusion Transformers with Token-wise Feature Caching

arXiv2 repos

arXiv:2410.05317

ToCa, DuCa

T2V-Turbo-v2: Enhancing Video Generation Model Post-Training through Data, Reward, and Conditional Guidance Design

arXiv2 repos

arXiv:2410.05677

t2v-turbo, T2V-Turbo-v2-no-MG

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

arXiv2 repos

arXiv:2410.06940

budget-flow-matching, REPA

Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow

arXiv2 repos

arXiv:2410.07303

Rectified-Diffusion, Rectified-Diffusion

MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts

arXiv2 repos

arXiv:2410.07348

zen5, Chat-UniVi

Lost in Time: A New Temporal Benchmark for VideoLLMs

arXiv2 repos

arXiv:2410.07752

TVBench, tvbench

SPA: 3D Spatial-Awareness Enables Effective Embodied Representation

arXiv2 repos

arXiv:2410.08208

SPA, SPA

Reconstructive Visual Instruction Tuning

arXiv2 repos

arXiv:2410.09575

ross-qwen2-7b, Visual-Region

Training-Free Adaptive Diffusion with Bounded Difference Approximation Strategy

arXiv2 repos

arXiv:2410.09873

AdaptiveDiffusion, Sampled_AIGCBench_text2image_ar_0.625

GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation

arXiv2 repos

arXiv:2410.10393

gift-eval, GiftEval

Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

arXiv2 repos

arXiv:2410.10469

uni2ts, uni2ts

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

arXiv2 repos

arXiv:2410.11623

embodied-eval, behaviour_subtask

CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos

arXiv2 repos

arXiv:2410.11831

co-tracker, CoTracker3_Kubric

JudgeBench: A Benchmark for Evaluating LLM-based Judges

arXiv2 repos

arXiv:2410.12784

JudgeBench, JudgeBench

The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

arXiv2 repos

arXiv:2410.12787

VideoLLaMA2, AVProunRLForVideoLLaMa2

D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement

arXiv2 repos

arXiv:2410.13842

D-FINE-seg, pydfine

BrainTransformers: SNN-LLM

arXiv2 repos

arXiv:2410.14687

BrainTransformers-SNN-LLM, BrainTransformers-3B-Chat

Allegro: Open the Black Box of Commercial-Level Video Generation Model

arXiv2 repos

arXiv:2410.15458

Allegro, Allegro-TI2V

Frontiers in Intelligent Colonoscopy

arXiv2 repos

arXiv:2410.17241

Project-Imaging-X, VPS

MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models

arXiv2 repos

arXiv:2410.17578

MMQA, prometheus-eval

3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation

arXiv2 repos

arXiv:2410.18974

MVEdit, 3D-Adapter

MarDini: Masked Autoregressive Diffusion for Video Generation at Scale

arXiv2 repos

arXiv:2410.20280

OpenVid-1M, OpenVid

ARLON: Boosting Diffusion Transformers with Autoregressive Models for Long Video Generation

arXiv2 repos

arXiv:2410.20502

OpenVid-1M, OpenVid

InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models

arXiv2 repos

arXiv:2410.22770

moltshield, NotInject

MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering

arXiv2 repos

arXiv:2410.22949

OpenBioMed, OpenBioMed_new

In-Context LoRA for Diffusion Transformers

arXiv2 repos

arXiv:2410.23775

catvton-flux, catvton-unstudio-flux

$π_0$: A Vision-Language-Action Flow Model for General Robot Control

arXiv2 repos

arXiv:2410.24164

GR00T-N1.5-3B, GR00T-N1.6-3B

DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

arXiv2 repos

arXiv:2411.00836

DynaMath_Sample, DynaMath

Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

arXiv2 repos

arXiv:2411.01156

fish-speech, fish-speech-1.5

xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism

arXiv2 repos

arXiv:2411.01738

xDiT, mochi-xdit

GenXD: Generating Any 3D and 4D Scenes

arXiv2 repos

arXiv:2411.02319

OpenVid-1M, OpenVid

DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion

arXiv2 repos

arXiv:2411.04928

OpenVid-1M, OpenVid

LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation

arXiv2 repos

arXiv:2411.04997

LLM2CLIP, LLM2CLIP-Llama-3-8B-Instruct-CC-Finetuned

FlexCAD: Unified and Versatile Controllable CAD Generation with Fine-tuned Large Language Models

arXiv2 repos

arXiv:2411.05823

FlexCAD, FlexCAD

Improved Video VAE for Latent Video Diffusion Model

arXiv2 repos

arXiv:2411.06449

OpenVid-1M, OpenVid

OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

arXiv2 repos

arXiv:2411.07199

OmniEdit-Filtered-1.2M, OmniEdit

Zero-shot Voice Conversion with Diffusion Transformers

arXiv2 repos

arXiv:2411.09943

maestro-seedvc, seed-vc-test

EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation

arXiv2 repos

arXiv:2411.10061

echomimic_v2, EchoMimicV2

CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic Segmentation

arXiv2 repos

arXiv:2411.10086

CorrCLIPv2, CorrCLIP

OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models

arXiv2 repos

arXiv:2411.10501

OnlyFlow, onlyflow

Multimodal Autoregressive Pre-training of Large Vision Encoders

arXiv2 repos

arXiv:2411.14402

ml-aim, aimv2-large-patch14-native

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI

arXiv2 repos

arXiv:2411.14522

GMAI-VL-5.5M, GMAI-VL

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

arXiv2 repos

arXiv:2411.16157

MVGenMaster, MVGenMaster

BayLing 2: A Multilingual Large Language Model with Efficient Language Alignment

arXiv2 repos

arXiv:2411.16300

BayLing, BayLing

One Diffusion to Generate Them All

arXiv2 repos

arXiv:2411.16318

OneDiffusion, OneDiffusion

AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning

arXiv2 repos

arXiv:2411.16495

AtomR, BlendQA

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

arXiv2 repos

arXiv:2411.16508

ALM-Bench, ALM-Bench

DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive Characters

arXiv2 repos

arXiv:2411.17423

DRiVE, DRiVE

Pushing the Limits of Large Language Model Quantization via the Linearity Theorem

arXiv2 repos

arXiv:2411.17525

quant.cpp, turboquant-vllm

MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

arXiv2 repos

arXiv:2411.17945

MARVEL-FX3D, MARVEL-40M

Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

arXiv2 repos

arXiv:2411.19108

TeaCache, FastVideo

Video Depth without Video Models

arXiv2 repos

arXiv:2411.19189

RollingDepth, rollingdepth-v1-0

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

arXiv2 repos

arXiv:2411.19325

GEO-Bench-VLM, GEOBench-VLM

Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

arXiv2 repos

arXiv:2412.00733

hallo3, hallo3_training_data

CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking

arXiv2 repos

arXiv:2412.01007

LateOn-Code, LateOn-Code-edge

OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation

arXiv2 repos

arXiv:2412.02592

OHR-Bench, OHR-Bench

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

arXiv2 repos

arXiv:2412.03304

Global-MMLU-Lite, Global-MMLU-Lite

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

arXiv2 repos

arXiv:2412.03558

MIDI-3D, MIDI-3D

MV-Adapter: Multi-view Consistent Image Generation Made Easy

arXiv2 repos

arXiv:2412.03632

mv-adapter, Objaverse-Rand6View

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios

arXiv2 repos

arXiv:2412.04447

embodied-eval, behaviour_subtask

The BrowserGym Ecosystem for Web Agent Research

arXiv2 repos

arXiv:2412.05467

AL_for_hallucination, AgentLab

Flex Attention: A Programming Model for Generating Optimized Attention Kernels

arXiv2 repos

arXiv:2412.05496

SparseD, szl-khipu

LLaVA-SpaceSGG: Visual Instruct Tuning for Open-vocabulary Scene Graph Generation with Enhanced Spatial Relations

arXiv2 repos

arXiv:2412.06322

LLaVA-SpaceSGG, LLaVA-SpaceSGG

arXiv:2412.06410

arXiv2 repos

arXiv:2412.06410

dictionary_learning, notebooks

Around the World in 80 Timesteps: A Generative Approach to Global Visual Geolocation

arXiv2 repos

arXiv:2412.06781

plonk, PLONK

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

arXiv2 repos

arXiv:2412.07626

OmniDocBench, granite-4.0-3b-vision

3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation

arXiv2 repos

arXiv:2412.07759

3DTrajMaster, 3DTrajMaster

From Slow Bidirectional to Fast Autoregressive Video Diffusion Models

arXiv2 repos

arXiv:2412.07772

CausVid, CausVid

TryOffAnyone: Tiled Cloth Generation from a Dressed Person

arXiv2 repos

arXiv:2412.08573

try-off-anyone, tryOffAnyone

FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models

arXiv2 repos

arXiv:2412.08629

FlowEdit, Wan2.1

Arbitrary-steps Image Super-resolution via Diffusion Inversion

arXiv2 repos

arXiv:2412.09013

InvSR, InvSR

SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos

arXiv2 repos

arXiv:2412.09401

slam3r_i2p, slam3r_l2w

Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation

arXiv2 repos

arXiv:2412.09585

VisPer-LM, OLA-VLM

V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding

arXiv2 repos

arXiv:2412.09616

InternVL3-2B, InternVL3-38B

WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model

arXiv2 repos

arXiv:2412.09951

WiseAD, WiseAD_training_data

AniSora: Exploring the Frontiers of Animation Video Generation in the Sora Era

arXiv2 repos

arXiv:2412.10255

Index-anisora, Index-anisora

Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection

arXiv2 repos

arXiv:2412.10432

L2D, AdaDetectGPT

SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile Device

arXiv2 repos

arXiv:2412.10494

OpenVid-1M, OpenVid

SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation

arXiv2 repos

arXiv:2412.12693

SPHERE-VLM, SPHERE-VLM

ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction

arXiv2 repos

arXiv:2412.12888

ArtAug-lora-FLUX.1dev-v1, diffSynth-studio-notes

Knowledge-enhanced Pretraining for Vision-language Pathology Foundation Model on Cancer Diagnosis

arXiv2 repos

arXiv:2412.13126

KEEP, PathPT

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models

arXiv2 repos

arXiv:2412.13702

llama3.1-typhoon2-audio-8b-instruct, typhoon2-audio

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

arXiv2 repos

arXiv:2412.15190

EarthDial, Land-Change-Detection

Don't Do RAG: When Cache-Augmented Generation is All You Need for Knowledge Tasks

arXiv2 repos

arXiv:2412.15605

CAG, Cache-Augmented-Generation-Granite

Personalized Representation from Personalized Generation

arXiv2 repos

arXiv:2412.16156

personalized-rep, PODS

DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

arXiv2 repos

arXiv:2412.17644

DreamFit, Dreamfit

YuLan-Mini: An Open Data-efficient Language Model

arXiv2 repos

arXiv:2412.17743

ProX, program-every-example

Towards Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST)

arXiv2 repos

arXiv:2412.18367

MultilingualAITerminology, multilingual-terminology

Open-Sora: Democratizing Efficient Video Production for All

arXiv2 repos

arXiv:2412.20404

Open-Sora, Open-Sora-v2

LTX-Video: Realtime Video Latent Diffusion

arXiv2 repos

arXiv:2501.00103

LTX-Video, ltx2-vidgen-skill

LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models

arXiv2 repos

arXiv:2501.00874

lusifer, LUSIFER

EliGen: Entity-Level Controlled Image Generation with Regional Attention

arXiv2 repos

arXiv:2501.01097

EliGen, diffSynth-studio-notes

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

arXiv2 repos

arXiv:2501.01957

VITA, VITA-1.5

Cosmos World Foundation Model Platform for Physical AI

arXiv2 repos

arXiv:2501.03575

Cosmos1GP, Cosmos-Tokenizer

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

arXiv2 repos

arXiv:2501.04670

colva_internvl2_4b, CoLVA

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

arXiv2 repos

arXiv:2501.04962

VoxEval, VoxEval

Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model

arXiv2 repos

arXiv:2501.05122

Centurio, Synthdog-Multilingual-100

FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors

arXiv2 repos

arXiv:2501.08225

FramePainter, FramePainter

Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models

arXiv2 repos

arXiv:2501.08453

Vchitect-2.0, Vchitect_T2V_DataVerse

VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

arXiv2 repos

arXiv:2501.12327

VARGPT, VARGPT_datasets

Kimi k1.5: Scaling Reinforcement Learning with LLMs

arXiv2 repos

arXiv:2501.12599

MathVision, MATH-V

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

arXiv2 repos

arXiv:2501.13826

lmms-eval, UniG2U

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

arXiv2 repos

arXiv:2501.13956

graphiti, post-graph-rag

Visual Generation Without Guidance

arXiv2 repos

arXiv:2501.15420

GFT, GFT

Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation

arXiv2 repos

arXiv:2501.15907

kani-tts-2-en, kani-tts-2-pt

Molecular-driven Foundation Model for Oncologic Pathology

arXiv2 repos

arXiv:2501.16652

TridentEdited, aegis

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation

arXiv2 repos

arXiv:2501.16764

DiffSplat, DiffSplat

BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights

arXiv2 repos

arXiv:2501.17790

BreezyVoice, BreezyVoice

Diverse Preference Optimization

arXiv2 repos

arXiv:2501.18101

DiversityTuning, fairseq2

MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding

arXiv2 repos

arXiv:2501.18362

MedXpertQA, MedXpertQA

GuardReasoner: Towards Reasoning-based LLM Safeguards

arXiv2 repos

arXiv:2501.18492

GuardReasoner-Omni-7B, GuardReasoner-Omni-3B

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

arXiv2 repos

arXiv:2501.18954

LLMDet, LLMDet

Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models

arXiv2 repos

arXiv:2501.19054

CADFusion, CADFusion

Process Reinforcement through Implicit Rewards

arXiv2 repos

arXiv:2502.01456

OpenRLHF, PRIME

Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation

arXiv2 repos

arXiv:2502.02464

reranking-datasets, reranking-datasets-light

LIMO: Less is More for Reasoning

arXiv2 repos

arXiv:2502.03387

open-korean-instructions, ko-limo

Detecting Strategic Deception Using Linear Probes

arXiv2 repos

arXiv:2502.03407

FabricationGuard-linearprobe-qwen36-27b, ReasoningGuard-linearprobe-qwen36-27b

DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation

arXiv2 repos

arXiv:2502.03930

VoxCPM, dots.tts

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

arXiv2 repos

arXiv:2502.05171

ProX, program-every-example

LANTERN++: Enhancing Relaxed Speculative Decoding with Static Tree Drafting for Visual Auto-regressive Models

arXiv2 repos

arXiv:2502.06352

LANTERN, anole_drafter

Accelerating Data Processing and Benchmarking of AI Models for Pathology

arXiv2 repos

arXiv:2502.06750

TRIDENT, TridentEdited

History-Guided Video Diffusion

arXiv2 repos

arXiv:2502.06764

DFoT, diffusion-forcing-transformer

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

arXiv2 repos

arXiv:2502.06782

Lumina-Video, Lumina-Video-f24R960

Self-Supervised Prompt Optimization

arXiv2 repos

arXiv:2502.06855

MetaGPT, MetaGPT

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

arXiv2 repos

arXiv:2502.08047

WorldGUI, WorldGUI-Bench

Universal Model Routing for Efficient LLM Inference

arXiv2 repos

arXiv:2502.08773

Aurora-AI-local-router, router

EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling

arXiv2 repos

arXiv:2502.09509

EQ-SDXL-VAE, HakuLatent

EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

arXiv2 repos

arXiv:2502.09560

behaviour_subtask, embodied-eval

HealthGPT: A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation

arXiv2 repos

arXiv:2502.09838

VL-Health, HealthGPT

KernelBench: Can LLMs Write Efficient GPU Kernels?

arXiv2 repos

arXiv:2502.10517

autokernel, KernelBench

arXiv:2502.10841

arXiv2 repos

arXiv:2502.10841

SkyReels-A1, SkyReels-A1

Phantom: Subject-consistent video generation via cross-modal alignment

arXiv2 repos

arXiv:2502.11079

Phantom, Phantom

How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training

arXiv2 repos

arXiv:2502.11196

KnowledgeCircuits, DynamicKnowledgeCircuits

BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model

arXiv2 repos

arXiv:2502.11798

BackdoorDM, BackdoorDM

Atom of Thoughts for Markov LLM Test-Time Scaling

arXiv2 repos

arXiv:2502.12018

MetaGPT, MetaGPT

FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views

arXiv2 repos

arXiv:2502.12138

FLARE_NVS, FLARE

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

arXiv2 repos

arXiv:2502.12148

HermesFlow, HermesFlow

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

arXiv2 repos

arXiv:2502.12558

MomentSeeker, MomentSeeker

MMTEB: Massive Multilingual Text Embedding Benchmark

arXiv2 repos

arXiv:2502.13595

mteb, ru_sci_bench_mteb

Language Model Re-rankers are Fooled by Lexical Similarities

arXiv2 repos

arXiv:2502.17036

rerankers-and-lexical-similarities, rerankers-and-lexical-similarities

The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence

arXiv2 repos

arXiv:2502.17420

OBLITERATUS, obliteratus

MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

arXiv2 repos

arXiv:2502.17422

mllms_know, TextVQA_GT_bbox

Chain of Draft: Thinking Faster by Writing Less

arXiv2 repos

arXiv:2502.18600

fabric, CoDE-Stop

UniTok: A Unified Tokenizer for Visual Generation and Understanding

arXiv2 repos

arXiv:2502.20321

unitok_mllm, unitok_tokenizer

HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning

arXiv2 repos

arXiv:2503.00912

HiBench, HiBench

Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

arXiv2 repos

arXiv:2503.01710

Spark-TTS-0.5B, Spark-TTS

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

arXiv2 repos

arXiv:2503.01743

Phi-4-multimodal-instruct, Phi-4-mini-instruct

IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval

arXiv2 repos

arXiv:2503.04644

IFIR, IFIR

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

arXiv2 repos

arXiv:2503.04721

Full-Duplex-Bench, personaplex

Dynamic Knowledge Integration for Evidence-Driven Counter-Argument Generation with Large Language Models

arXiv2 repos

arXiv:2503.05328

counter-argument, counter-argument-generation

UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces

arXiv2 repos

arXiv:2503.06157

UrbanVideo-Bench, UrbanVideo-Bench.code

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

arXiv2 repos

arXiv:2503.06800

videophy, Cosmos-Reason2-32B

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

arXiv2 repos

arXiv:2503.07265

UniWorld-V1, UniWorld-V1-NF4

MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

arXiv2 repos

arXiv:2503.07365

MMK12, MM-EUREKA

NullFace: Training-Free Localized Face Anonymization

arXiv2 repos

arXiv:2503.08478

nullface, nullface-test-set

OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models

arXiv2 repos

arXiv:2503.08686

OmniMamba, OmniMamba

Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

arXiv2 repos

arXiv:2503.09642

Open-Sora, Open-Sora-v2

VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

arXiv2 repos

arXiv:2503.10291

VisualPRM-8B, VisualPRM400K

Distilling Diversity and Control in Diffusion Models

arXiv2 repos

arXiv:2503.10637

stable-confusion, distillation

Taming Knowledge Conflicts in Language Models

arXiv2 repos

arXiv:2503.10996

JUICE, ParaConfilct

Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models

arXiv2 repos

arXiv:2503.11073

PURE, PURE

Beyond the Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering

arXiv2 repos

arXiv:2503.11117

behaviour_subtask, embodied-eval

Context-Aware Rule Mining Using a Dynamic Transformer-Based Framework

arXiv2 repos

arXiv:2503.11125

awesome-datascience, awesome-datascience

FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis

arXiv2 repos

arXiv:2503.13265

FlexWorld, FlexWorld

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

arXiv2 repos

arXiv:2503.13444

VideoMind, VideoMind-Dataset

Where do Large Vision-Language Models Look at when Answering Questions?

arXiv2 repos

arXiv:2503.13891

LVLM_Interpretation, LVLM_Interpretation

AIGVE-Tool: AI-Generated Video Evaluation Toolkit with Multifaceted Benchmark

arXiv2 repos

arXiv:2503.14064

AIGVE_Tool, AIGVE-Bench

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

arXiv2 repos

arXiv:2503.14935

FAVOR, FAVOR-Bench

UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction

arXiv2 repos

arXiv:2503.15661

UI-Vision, ui-vision

InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

arXiv2 repos

arXiv:2503.16418

InfiniteYou, InfiniteYou

AMD-Hummingbird: Towards an Efficient Text-to-Video Model

arXiv2 repos

arXiv:2503.18559

Hummingbird, AMD-Hummingbird-T2V

ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

arXiv2 repos

arXiv:2503.19470

ReSearch, ReCall

Gemma 3 Technical Report

arXiv2 repos

arXiv:2503.19786

SLU_pipeline, open-value

Understanding R1-Zero-Like Training: A Critical Perspective

arXiv2 repos

arXiv:2503.20783

tunix, understand-r1-zero

Empowering Retrieval-based Conversational Recommendation with Contrasting User Preferences

arXiv2 repos

arXiv:2503.22005

CORAL, CORAL

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

arXiv2 repos

arXiv:2503.22020

VLAC, AI539_NLP

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

arXiv2 repos

arXiv:2503.22152

behaviour_subtask, embodied-eval

VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models

arXiv2 repos

arXiv:2503.23064

VGRP-Bench, VGRP-Bench

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

arXiv2 repos

arXiv:2504.00072

chapter-llama, chapter-llama

Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead

arXiv2 repos

arXiv:2504.00294

eureka-ml-insights, Eureka-Bench-Logs

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection

arXiv2 repos

arXiv:2504.00470

LIMA, SMDL-Attribution

BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing

arXiv2 repos

arXiv:2504.01786

BlenderGym-Open, BG_bench_data

arXiv:2504.02436

arXiv2 repos

arXiv:2504.02436

SkyReels-A2, SkyReels-A2

Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models

arXiv2 repos

arXiv:2504.02821

sae-for-vlm, sae-for-vlm

VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning

arXiv2 repos

arXiv:2504.02949

VARGPT, VARGPT_datasets

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

arXiv2 repos

arXiv:2504.03624

NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-3-Nano-4B-GGUF

Training state-of-the-art pathology foundation models with orders of magnitude less data

arXiv2 repos

arXiv:2504.05186

midnight, Midnight

POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction

arXiv2 repos

arXiv:2504.05692

POMATO, POMATO

DDT: Decoupled Diffusion Transformer

arXiv2 repos

arXiv:2504.05741

DDT-XL-22en6de-R512, DDT-XL-22en6de-R256

GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography

arXiv2 repos

arXiv:2504.07083

GenDoP, DataDoP

ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

arXiv2 repos

arXiv:2504.07981

UI-TARS, LocateAnything-3B

ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model

arXiv2 repos

arXiv:2504.09421

MedFound-176B, ClinicalGPT-R1-Qwen-7B-EN-preview

SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

arXiv2 repos

arXiv:2504.09644

EarthReason, SegEarth-R1

OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape Generation

arXiv2 repos

arXiv:2504.09975

octgpt, OctGPT

ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness

arXiv2 repos

arXiv:2504.10514

ColorBench, ColorBench

DataDecide: How to Predict Best Pretraining Data with Small Experiments

arXiv2 repos

arXiv:2504.11393

ProX, program-every-example

ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

arXiv2 repos

arXiv:2504.11536

Open-AgentRL, ReTool

arXiv:2504.13074

arXiv2 repos

arXiv:2504.13074

SkyCaptioner-V1, SkyReels-V2-I2V-14B-720P

LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard

arXiv2 repos

arXiv:2504.13125

Awesome-Chinese-LLM, Awesome-Chinese-LLM

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

arXiv2 repos

arXiv:2504.13143

Complex-Edit, Complex-Edit

St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World

arXiv2 repos

arXiv:2504.13152

St4RTrack, St4rTrack

Sleep-time Compute: Beyond Inference Scaling at Test-time

arXiv2 repos

arXiv:2504.13171

letta-code, khms-memory

Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models

arXiv2 repos

arXiv:2504.15271

Eagle, GR00T-N1.6-3B

Towards Understanding Camera Motions in Any Video

arXiv2 repos

arXiv:2504.15376

t2v_metrics, CameraBench

TTRL: Test-Time Reinforcement Learning

arXiv2 repos

arXiv:2504.16084

TTRL, EMPO

DreamO: A Unified Framework for Image Customization

arXiv2 repos

arXiv:2504.16915

DreamO, dreamo

Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs

arXiv2 repos

arXiv:2504.17432

ms-agent, modelscope-agent

HalluLens: LLM Hallucination Benchmark

arXiv2 repos

arXiv:2504.17550

KoHalluLens, HalluLens

Kimi-Audio Technical Report

arXiv2 repos

arXiv:2504.18425

Kimi-Audio, dissertation-project

PixelHacker: Image Inpainting with Structural and Semantic Consistency

arXiv2 repos

arXiv:2504.20438

PixelHacker, PixelHacker

YoChameleon: Personalized Vision and Language Generation

arXiv2 repos

arXiv:2504.20998

YoChameleon, Mini-YoChameleon-Data

Phi-4-reasoning Technical Report

arXiv2 repos

arXiv:2504.21318

Phi-4-reasoning, Eureka-Bench-Logs

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation

arXiv2 repos

arXiv:2505.01456

UnLOK-VQA, UnLOK-VQA

A survey of agent interoperability protocols: Model Context Protocol (MCP), Agent Communication Protocol (ACP), Agent-to-Agent Protocol (A2A), and Agent Network Protocol (ANP)

arXiv2 repos

arXiv:2505.02279

guaca, awesome-agentic-payments

arXiv:2505.02387

arXiv2 repos

arXiv:2505.02387

RM-R1, RM-R1-Qwen2.5-Instruct-7B

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning

arXiv2 repos

arXiv:2505.03318

VideoDPO, ShareGPTVideo-DPO

FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios

arXiv2 repos

arXiv:2505.03730

FlexiAct, FlexiAct

Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data

arXiv2 repos

arXiv:2505.05427

Ultra-FineWeb-L1, Ultra-FineWeb

Flow-GRPO: Training Flow Matching Models via Online RL

arXiv2 repos

arXiv:2505.05470

HY-SOAR, GRPO

The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization

arXiv2 repos

arXiv:2505.06371

zeus, benchmark

DanceGRPO: Unleashing GRPO on Visual Generation

arXiv2 repos

arXiv:2505.07818

DanceGRPO, GRPO

HealthBench: Evaluating Large Language Models Towards Improved Human Health

arXiv2 repos

arXiv:2505.08775

AntAngelMed, AntAngelMed

StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation

arXiv2 repos

arXiv:2505.10292

QwenStoryteller, StoryReasoning

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning

arXiv2 repos

arXiv:2505.11049

GuardReasoner-Omni-7B, GuardReasoner-Omni-3B

BLEUBERI: BLEU is a surprisingly effective reward for instruction following

arXiv2 repos

arXiv:2505.11080

BLEUBERI, BLEUBERI-Tulu3-50k

Video-GPT via Next Clip Diffusion

arXiv2 repos

arXiv:2505.12489

Video-GPT, Video-GPT

CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models

arXiv2 repos

arXiv:2505.12504

MM-EUREKA, CPGD-7B

Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents

arXiv2 repos

arXiv:2505.12632

monday, MONDAY

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

arXiv2 repos

arXiv:2505.13439

VTBench, VTBench

This Time is Different: An Observability Perspective on Time Series Foundation Models

arXiv2 repos

arXiv:2505.14766

toto, toto

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

arXiv2 repos

arXiv:2505.15404

LRM-Safety-Study, LRM-Safety-Study

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets

arXiv2 repos

arXiv:2505.15517

robo2VLM, Robo2VLM-1

Efficient PRM Training Data Synthesis via Formal Verification

arXiv2 repos

arXiv:2505.15960

Qwen-2.5-7B-FoVer-PRM-2026, Llama-3.1-8B-FoVer-PRM-2026

FreshRetailNet-50K: A Stockout-Annotated Censored Demand Dataset for Latent Demand Recovery and Forecasting in Fresh Retail

arXiv2 repos

arXiv:2505.16319

FreshRetailNet-50K, frn-50k-baseline

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

arXiv2 repos

arXiv:2505.16933

LLaDA-V, LLaDA-V

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

arXiv2 repos

arXiv:2505.17426

DistilCodec, DistilCodec-v1.0

MMMG: a Comprehensive and Reliable Evaluation Suite for Multitask Multimodal Generation

arXiv2 repos

arXiv:2505.17613

MMMG, MMMG

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

arXiv2 repos

arXiv:2505.17952

AlphaMed-7B-instruct-rl, AlphaMed-8B-instruct-rl

Enhancing Training Data Attribution with Representational Optimization

arXiv2 repos

arXiv:2505.18513

AirRep, AirRep-Flan-Small

LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

arXiv2 repos

arXiv:2505.19223

LLaDA-1.5, LLaDA-1.5

Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models

arXiv2 repos

arXiv:2505.19743

MARA_AGENTS, MARA

What Can RL Bring to VLA Generalization? An Empirical Study

arXiv2 repos

arXiv:2505.19789

RL4VLA, TTT_VLARL

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration

arXiv2 repos

arXiv:2505.20256

Omni-R1, Omni-R1

SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution

arXiv2 repos

arXiv:2505.20732

SPA-RL-Agent, mobile-agent-rl

Rendering-Aware Reinforcement Learning for Vector Graphics Generation

arXiv2 repos

arXiv:2505.20793

detikzify-v2.5-8b, star-vector

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

arXiv2 repos

arXiv:2505.22232

JQL-Annotation-Pipeline, JQL-Edu-Heads

ATI: Any Trajectory Instruction for Controllable Video Generation

arXiv2 repos

arXiv:2505.22944

ATI, ATI

Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

arXiv2 repos

arXiv:2505.22954

codegraff, darwinian_evolver

EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge

arXiv2 repos

arXiv:2505.23009

higgs-tts-2-3b-base, higgs-audio-v2-generation-3B-base

FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing

arXiv2 repos

arXiv:2505.23145

FlowEdit, Wan2.1

TrackVLA: Embodied Visual Tracking in the Wild

arXiv2 repos

arXiv:2505.23189

OmTrackVLA, TrackVLA

Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking

arXiv2 repos

arXiv:2505.24857

diffusion-gemma-lab, ParallelBench

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL

arXiv2 repos

arXiv:2505.24875

ReasonGen-R1, ReasonGen-R1-SFT

COSMIC: Generalized Refusal Direction Identification in LLM Activations

arXiv2 repos

arXiv:2506.00085

OBLITERATUS, obliteratus

VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos

arXiv2 repos

arXiv:2506.02448

ShotPlan, shotplan

ORV: 4D Occupancy-centric Robot Video Generation

arXiv2 repos

arXiv:2506.03079

orv-gen-model, ORV

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

arXiv2 repos

arXiv:2506.03135

SpaceQwen2.5-VL-3B-Instruct, SpaceThinker-Qwen2.5VL-3B

A Foundation Model for Spatial Proteomics

arXiv2 repos

arXiv:2506.03373

CARTA, KRONOS

HMAR: Efficient Hierarchical Masked Auto-Regressive Image Generation

arXiv2 repos

arXiv:2506.04421

HMAR, HMAR

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

arXiv2 repos

arXiv:2506.04779

Qwen2.5-Omni, Qwen-2.5-7b

MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm

arXiv2 repos

arXiv:2506.05218

Monkey, MonkeyOCR

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts

arXiv2 repos

arXiv:2506.06211

PuzzleWorld, PuzzleWorld

STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis

arXiv2 repos

arXiv:2506.06276

ml-starflow, starflow

CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning

arXiv2 repos

arXiv:2506.06290

CellCLIP, CellCLIP

dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching

arXiv2 repos

arXiv:2506.06295

dLLM-cache, dLLM_Cache

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

arXiv2 repos

arXiv:2506.06962

AR-RAG, arrag_faid

Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models

arXiv2 repos

arXiv:2506.07334

GraphKV, GraphKV

Real-Time Execution of Action Chunking Flow Policies

arXiv2 repos

arXiv:2506.07339

RLDX-1, RLDX-FineAct

Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse

arXiv2 repos

arXiv:2506.07639

Adaptive-CoT-in-VLA, Fast-ECoT

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

arXiv2 repos

arXiv:2506.07977

OneIG-Benchmark, OneIG-Bench

PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

arXiv2 repos

arXiv:2506.10741

PosterCraft, Poster100K

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?

arXiv2 repos

arXiv:2506.10912

ToxiMol, ToxiMol-benchmark

Towards Building General Purpose Embedding Models for Industry 4.0 Agents

arXiv2 repos

arXiv:2506.12607

FailureSensorIQ, AssetOpsBench

LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D

arXiv2 repos

arXiv:2506.13766

LHM-plusplus, LHM-plusplus

Comprehensive Verilog Design Problems: A Next-Generation Benchmark Dataset for Evaluating Large Language Models and Agents on RTL Design and Verification

arXiv2 repos

arXiv:2506.14074

cvdp_client, cvdp_benchmark

Sekai: A Video Dataset towards World Exploration

arXiv2 repos

arXiv:2506.15675

Sekai, sekai-codebase

Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details

arXiv2 repos

arXiv:2506.16504

Hunyuan3D-2, qsdsd

Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search

arXiv2 repos

arXiv:2506.16962

Chiron-o1-8B, Chiron-o1-2B

VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning

arXiv2 repos

arXiv:2506.17221

GPT4Scene, GPT4Scene-and-VLN-R1

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

arXiv2 repos

arXiv:2506.18088

robotwin2.0-fastwam, RoboTwin

ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

arXiv2 repos

arXiv:2506.18095

ShareGPT-4o-Image, Janus-4o-7B

LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning

arXiv2 repos

arXiv:2506.18841

LongWriter, LongWriter-Zero-32B

TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer

arXiv2 repos

arXiv:2506.18904

TC-Light, TC-Light

SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning

arXiv2 repos

arXiv:2506.21355

smmile, SMMILE

Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR

arXiv2 repos

arXiv:2506.22646

multitalker-parakeet-streaming-0.6b-v1, multitalker-parakeet-streaming-0.6b-v1-onnx-int8

Token Activation Map to Visually Explain Multimodal LLMs

arXiv2 repos

arXiv:2506.23270

TAM, TAM

arXiv:2506.23869

arXiv2 repos

arXiv:2506.23869

aria, aria-medium-base

Understanding and Improving Length Generalization in Recurrent Models

arXiv2 repos

arXiv:2507.02782

AI21-Jamba-Reasoning-3B-GGUF, AI21-Jamba-Reasoning-3B

LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

arXiv2 repos

arXiv:2507.02813

LangScene-X, LangScene-X

RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism

arXiv2 repos

arXiv:2507.02962

AWorld, AWorld-RL

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

arXiv2 repos

arXiv:2507.03905

EchoMimicV3, echomimic_v3

Neural-Driven Image Editing

arXiv2 repos

arXiv:2507.05397

loongx, L-Mind

Scaling RL to Long Videos

arXiv2 repos

arXiv:2507.07966

Long-RL, LongVILA-R1-7B

Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective

arXiv2 repos

arXiv:2507.08801

Lumos, Lumos-1

Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation

arXiv2 repos

arXiv:2507.10524

mixture_of_recursions, R3-ViT

Text-ADBench: Text Anomaly Detection Benchmark Based on LLM Embeddings

arXiv2 repos

arXiv:2507.12295

Text-Anomaly-Detection-Benchmark, Text-ADBench

Voxtral

arXiv2 repos

arXiv:2507.13264

dissertation-project, Voxtral-Small-24B-2507

$π^3$: Permutation-Equivariant Visual Geometry Learning

arXiv2 repos

arXiv:2507.13347

Pi3, Pi3X

VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

arXiv2 repos

arXiv:2507.13353

VideoITG, VideoITG-8B

Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning

arXiv2 repos

arXiv:2507.16746

Zebra-CoT, Reasoning-Visual-World

Towards Robust Foundation Models for Digital Pathology

arXiv2 repos

arXiv:2507.17845

PathoROB, croma

Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding

arXiv2 repos

arXiv:2507.19427

Step-3.5-Flash, Step-3.5-Flash

Agentic Reinforced Policy Optimization

arXiv2 repos

arXiv:2507.19849

Tool-Star, Tool-Star

arXiv:2507.20534

arXiv2 repos

arXiv:2507.20534

checkpoint-engine, Emerging-Optimizers

Latent Inter-User Difference Modeling for LLM Personalization

arXiv2 repos

arXiv:2507.20849

DEP, DEP-model

SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment

arXiv2 repos

arXiv:2507.20984

SmallThinker-4BA0.6B-Instruct, SmallThinker-21BA3B-Instruct

TTS-1 Technical Report

arXiv2 repos

arXiv:2507.21138

tts, Anime-XCodec2-44.1kHz-v2

Benchmarking LLMs for Unit Test Generation from Real-World Functions

arXiv2 repos

arXiv:2508.00408

UnLeakedTestBench, unleakedtestbench

LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding

arXiv2 repos

arXiv:2508.01617

LLaDA-MedV, LLaDA-MedV

Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images

arXiv2 repos

arXiv:2508.03643

Uni3R, Uni3R

FlowState: Sampling-Rate-Equivariant Time-Series Forecasting

arXiv2 repos

arXiv:2508.05287

flowstate, granite-timeseries-flowstate-r1

Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation

arXiv2 repos

arXiv:2508.07901

Stand-In, Stand-In

Ovis2.5 Technical Report

arXiv2 repos

arXiv:2508.11737

Ovis2.5-9B, Ovis2.5-2B

Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models

arXiv2 repos

arXiv:2508.12945

Lumen, Lumen

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

arXiv2 repos

arXiv:2508.14033

InfiniteTalk, modal-infinitetalk

Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration

arXiv2 repos

arXiv:2508.14483

Vivid-VR, Vivid-VR

Dream 7B: Diffusion Large Language Models

arXiv2 repos

arXiv:2508.15487

d3LLM, dllm

LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind

arXiv2 repos

arXiv:2508.15601

SimSIMD, numkong

AWorld: Orchestrating the Training Recipe for Agentic AI

arXiv2 repos

arXiv:2508.20404

Qwen3-32B-AWorld, AWorld

MSRS: Evaluating Multi-Source Retrieval-Augmented Generation

arXiv2 repos

arXiv:2508.20867

MSRS, MSRS

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models

arXiv2 repos

arXiv:2508.20869

OLMoASR, OLMoASR-Pool

RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching

arXiv2 repos

arXiv:2508.21258

open-r-lens, RelP

Is this chart lying to me? Automating the detection of misleading visualizations

arXiv2 repos

arXiv:2508.21675

acl2026-misleading-visualizations, acl2026-misviz

CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of Sparse Autoencoders

arXiv2 repos

arXiv:2509.00691

CE-Bench, contrastive-stories-v4

REFRAG: Rethinking RAG based Decoding

arXiv2 repos

arXiv:2509.01092

ruvector, AXRU

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

arXiv2 repos

arXiv:2509.02544

UI-TARS, VeOmni

Real-Time Detection of Hallucinated Entities in Long-Form Generation

arXiv2 repos

arXiv:2509.03531

hallucination_probes, hallucination-probes

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

arXiv2 repos

arXiv:2509.03740

BiomedCoOp, CLIP-SVD

WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

arXiv2 repos

arXiv:2509.03959

WenetSpeech-Yue, WSYue-TTS

Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching

arXiv2 repos

arXiv:2509.05952

MixGRPO, flow_grpo

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

arXiv2 repos

arXiv:2509.06818

UMO, UMO

Multiple Instance Learning Framework with Masked Hard Instance Mining for Gigapixel Histopathology Image Analysis

arXiv2 repos

arXiv:2509.11526

MHIM-MIL, CPathPatchFeature

Embodied Navigation Foundation Model

arXiv2 repos

arXiv:2509.12129

OmTrackVLA, OmTrackVLA-0.6B

Aegis: Automated Error Generation and Attribution for Multi-Agent Systems

arXiv2 repos

arXiv:2509.14295

AEGIS, AEGIS

Lynx: Towards High-Fidelity Personalized Video Generation

arXiv2 repos

arXiv:2509.15496

lynx, lynx

Discovering Top-k Periodic and High-Utility Patterns

arXiv2 repos

arXiv:2509.15732

awesome-datascience, awesome-datascience

Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing

arXiv2 repos

arXiv:2509.17052

Sidon, sidon_raw_weight

FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction

arXiv2 repos

arXiv:2509.18362

FastMTP, speculators

Hunyuan3D-Omni: A Unified Framework for Controllable Generation of 3D Assets

arXiv2 repos

arXiv:2509.21245

HY3D-Bench, Hunyuan3D-Omni

LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer

arXiv2 repos

arXiv:2509.22414

LucidFlux, LucidFlux

Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning

arXiv2 repos

arXiv:2509.25052

deepscaler, rllm

Fine-Grained GRPO for Precise Preference Alignment in Flow Models

arXiv2 repos

arXiv:2510.01982

Granular-GRPO, G2RPO

xLSTM Scaling Laws: Competitive Performance with Linear Time-Complexity

arXiv2 repos

arXiv:2510.02228

xlstm_scaling_laws, xlstm_scaling_laws

VideoNSA: Native Sparse Attention Scales Video Understanding

arXiv2 repos

arXiv:2510.02295

VideoNSA, VideoNSA

EditLens: Quantifying the Extent of AI Editing in Text

arXiv2 repos

arXiv:2510.03154

sloptotal, EditLens

ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion

arXiv2 repos

arXiv:2510.04706

Arc2Face, Arc2Face

Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning

arXiv2 repos

arXiv:2510.04786

ttc, verifiable-corpus

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

arXiv2 repos

arXiv:2510.06308

Lumina-DiMOO, Lumina-DiMOO

PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs

arXiv2 repos

arXiv:2510.09507

PhysToolBench, PhysToolBench

Don't Just Fine-tune the Agent, Tune the Environment

arXiv2 repos

arXiv:2510.10197

AWorld, AWorld-RL

Chart-RVR: Reinforcement Learning with Verifiable Rewards for Explainable Chart Reasoning

arXiv2 repos

arXiv:2510.10973

chart-rvr-3b, chart-rvr-hard-3b

arXiv:2510.12747

arXiv2 repos

arXiv:2510.12747

FlashVSR, FlashVSR-v1.1

UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy

arXiv2 repos

arXiv:2510.13745

UniCalli, UniCalli_Dev

Agentic Entropy-Balanced Policy Optimization

arXiv2 repos

arXiv:2510.14545

Tool-Star, Tool-Star

WithAnyone: Towards Controllable and ID Consistent Image Generation

arXiv2 repos

arXiv:2510.14975

WithAnyone, MultiID-Bench

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

arXiv2 repos

arXiv:2510.15710

UniMedVL, UniMedVL

Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

arXiv2 repos

arXiv:2510.16888

UniWorld, UniWorld-V1

OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction

arXiv2 repos

arXiv:2510.17532

Clinical-Reasoning-LLMs, cancer-reasoning-traces

From Charts to Code: A Hierarchical Benchmark for Multimodal Models

arXiv2 repos

arXiv:2510.17932

Chart2Code, Chart2Code

Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs

arXiv2 repos

arXiv:2510.20064

hedgespec, hedgespec_eagle_drafters

Generative Reasoning Recommendation via LLMs

arXiv2 repos

arXiv:2510.20815

GRRM, GREAM_data

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation

arXiv2 repos

arXiv:2510.22115

Ling-1T, ArtifactsBenchmark

LongCat-Video Technical Report

arXiv2 repos

arXiv:2510.22200

LongCat-Video, LongCat-Video

GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping

arXiv2 repos

arXiv:2510.22319

flow_grpo, GRPO

LimRank: Less is More for Reasoning-Intensive Information Reranking

arXiv2 repos

arXiv:2510.23544

limrank, limrank

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

arXiv2 repos

arXiv:2510.23607

Concerto, Concerto

FullPart: Generating each 3D Part at Full Resolution

arXiv2 repos

arXiv:2510.26140

fullpart, partversexl

Pre-trained Forecasting Models: Strong Zero-Shot Feature Extractors for Time Series Classification

arXiv2 repos

arXiv:2510.26777

TiRex, tirex

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

arXiv2 repos

arXiv:2511.00916

Fleming-VL-8B, Fleming-VL-38B

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

arXiv2 repos

arXiv:2511.04727

IndicVisionBench, IndicVisionBench

TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning

arXiv2 repos

arXiv:2511.05489

TimeSearch-R, TimeSearch-R

Expert-Guided Prompting and Retrieval-Augmented Generation for Emergency Medical Service Question Answering

arXiv2 repos

arXiv:2511.10900

EMS-MCQA, EMS-Knowledge

Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition

arXiv2 repos

arXiv:2511.11139

SAP2-ASR, SAP2-ASR

Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection

arXiv2 repos

arXiv:2511.13027

Skills, NeMo-Skills

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

arXiv2 repos

arXiv:2511.14439

AntAngelMed, AntAngelMed

$π^{*}_{0.6}$: a VLA That Learns From Experience

arXiv2 repos

arXiv:2511.14759

open-value, lerobot

arXiv:2511.15186

arXiv2 repos

arXiv:2511.15186

ROSALIA, ROSALIA-7B-v1

arXiv:2511.15684

arXiv2 repos

arXiv:2511.15684

walrus, walrus

SAM 3: Segment Anything with Concepts

arXiv2 repos

arXiv:2511.16719

geti-instant-learn, rf-detr

ChineseErrorCorrector3-4B: State-of-the-Art Chinese Spelling and Grammar Corrector

arXiv2 repos

arXiv:2511.17562

ChineseErrorCorrector3-4B, ChineseErrorCorrector

UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

arXiv2 repos

arXiv:2511.18050

UltraFlux-v1, UltraFlux-v1-1-Transformer

MedVision: Benchmarking Quantitative Medical Image Analysis

arXiv2 repos

arXiv:2511.18676

MedVision, MedVision

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on

arXiv2 repos

arXiv:2511.18957

Eevee, Eevee

STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows

arXiv2 repos

arXiv:2511.20462

ml-starflow, starflow

MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices

arXiv2 repos

arXiv:2511.21475

MobileI2V, MobileI2V

ReasonEdit: Towards Reasoning-Enhanced Image Editing Models

arXiv2 repos

arXiv:2511.22625

Step1X-Edit, Step1X-Edit-v1p2

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

arXiv2 repos

arXiv:2511.22663

AIA, AIA

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

arXiv2 repos

arXiv:2511.22699

Z-Image-Turbo, Z-Image

PanFlow: Decoupled Motion Control for Panoramic Video Generation

arXiv2 repos

arXiv:2512.00832

PanFlow, PanFlow

Accelerating Streaming Video Large Language Models via Hierarchical Token Compression

arXiv2 repos

arXiv:2512.00891

VidCom2, STC

InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision

arXiv2 repos

arXiv:2512.01342

InternVideo, internvideo-d2a11ea9

Improved Mean Flows: On the Challenges of Fastforward Generative Models

arXiv2 repos

arXiv:2512.02012

iMF-diffusers, imeanflow

M3DR: Towards Universal Multilingual Multimodal Document Retrieval

arXiv2 repos

arXiv:2512.03514

ColNetraEmbed, NetraEmbed

Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning

arXiv2 repos

arXiv:2512.03667

Project-Imaging-X, VPS

SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs

arXiv2 repos

arXiv:2512.04746

auto-round, Qwen3.8-27B-bpw2.8-AutoRound

Aligned but Stereotypical? How System Prompts Shape Demographic Bias in LLM-Based Text-to-Image Models

arXiv2 repos

arXiv:2512.04981

fairpro, fairpro

Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark

arXiv2 repos

arXiv:2512.05091

Sa2VA, Sa2VA

Training-Time Action Conditioning for Efficient Real-Time Chunking

arXiv2 repos

arXiv:2512.05964

RLDX-1, RLDX-FineAct

PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory

arXiv2 repos

arXiv:2512.06688

PersonaMem-v2, PersonaMem-v2

Unified Camera Positional Encoding for Controlled Video Generation

arXiv2 repos

arXiv:2512.07237

PanShot, prope

LongCat-Image Technical Report

arXiv2 repos

arXiv:2512.07584

LongCat-Image-Edit-Turbo, GenAI-Caption-Pipeline

InfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain Generation

arXiv2 repos

arXiv:2512.08309

terrain-diffusion-30m, terrain-diffusion

Deterministic and Exact Fully-dynamic Minimum Cut of Superpolylogarithmic Size in Subpolynomial Time

arXiv2 repos

arXiv:2512.13105

AXRU, ruvector

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

arXiv2 repos

arXiv:2512.13604

LongVie, LongVie2

RePo: Language Models with Context Re-Positioning

arXiv2 repos

arXiv:2512.14391

repo, RePo-OLMo2-1B-stage2-L5

Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs

arXiv2 repos

arXiv:2512.16378

hearing2translate, hearing2translate-humeval

Gabliteration: Adaptive Multi-Directional Neural Weight Modification for Selective Behavioral Alteration in Large Language Models

arXiv2 repos

arXiv:2512.18901

OBLITERATUS, obliteratus

Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

arXiv2 repos

arXiv:2512.20848

nugie-jax-nemotron-3-nano, NVIDIA-Nemotron-3-Nano-4B-GGUF

dUltra: Ultra-Fast Diffusion Language Models via Reinforcement Learning

arXiv2 repos

arXiv:2512.21446

dUltra-os, dUltra-math-b128

DiRL: An Efficient Post-Training Framework for Diffusion Language Models

arXiv2 repos

arXiv:2512.22234

DiRL, DiRL-8B-Instruct

TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish

arXiv2 repos

arXiv:2512.23065

TabiBERT, Tabibert

SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation

arXiv2 repos

arXiv:2512.23379

SoulX-FlashTalk, SoulX-FlashTalk-14B

Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation

arXiv2 repos

arXiv:2512.23705

DKT, TransPhy3D

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

arXiv2 repos

arXiv:2601.01528

DrivingGen, DrivingGen

Yuan3.0 Flash: An Open Multimodal Large Language Model for Enterprise Applications

arXiv2 repos

arXiv:2601.01718

Yuan3.0-Flash, Yuan3.0-Flash-4bit

K-EXAONE Technical Report

arXiv2 repos

arXiv:2601.01739

K-EXAONE-236B-A23B, K-EXAONE

Pearmut: Human Evaluation of Translation Made Trivial

arXiv2 repos

arXiv:2601.02933

hearing2translate-humeval, pearmut

LTX-2: Efficient Joint Audio-Visual Foundation Model

arXiv2 repos

arXiv:2601.03233

LTX-2, LTX-2

MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents

arXiv2 repos

arXiv:2601.03236

rlm-claude, memcp

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

arXiv2 repos

arXiv:2601.05138

VerseCrafter, VerseCrafter

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

arXiv2 repos

arXiv:2601.05593

Step-3.5-Flash, Step-3.5-Flash

Affostruction: 3D Affordance Grounding with Generative Reconstruction

arXiv2 repos

arXiv:2601.09211

Affostruction, Affostruction

HeartMuLa: A Family of Open Sourced Music Foundation Models

arXiv2 repos

arXiv:2601.10547

heartlib, HeartMuLa_ComfyUI

ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation

arXiv2 repos

arXiv:2601.12983

acl2026-misleading-visualizations, acl2026-misviz

CURE-Med: Curriculum-Informed Reinforcement Learning for Multilingual Medical Reasoning

arXiv2 repos

arXiv:2601.13262

cure-med, CUREMED-BENCH

Numina-Lean-Agent: An Open and General Agentic Reasoning System for Formal Mathematics

arXiv2 repos

arXiv:2601.14027

open-atp, lean-lsp-mcp

SAMTok: Representing Any Mask with Two Words

arXiv2 repos

arXiv:2601.16093

Sa2VA, Sa2VA

ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion

arXiv2 repos

arXiv:2601.16148

ActionMesh, actionbench

Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders

arXiv2 repos

arXiv:2601.16208

scale-rae-data, Scale-RAE

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

arXiv2 repos

arXiv:2601.19194

TS-ASR-Whisper, DiCoW_v3_3_large

Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration

arXiv2 repos

arXiv:2601.19506

Pref_Restore, Pref-Restore-PhaseA-Fidelity

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

arXiv2 repos

arXiv:2601.19834

VisWorld-Eval, Reasoning-Visual-World

One-step Latent-free Image Generation with Pixel Mean Flows

arXiv2 repos

arXiv:2601.22158

pMF-diffusers, imeanflow

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

arXiv2 repos

arXiv:2601.22599

FlowSep-hive, AudioSep-hive

arXiv:2601.22710

arXiv2 repos

arXiv:2601.22710

AlienLM, llama3-8b-instruct-alienlm-full

DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding

arXiv2 repos

arXiv:2601.23161

DIFFA, DIFFA-2

Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models

arXiv2 repos

arXiv:2602.01025

UltraBreak, UltraBreak-Repro

P-EAGLE: Parallel-Drafting EAGLE with Scalable Training

arXiv2 repos

arXiv:2602.01469

SpecForge, Qwen3-8B-speculator.peagle

WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

arXiv2 repos

arXiv:2602.02537

WorldVQA, WorldVQA

InfMem: Learning System-2 Memory Control for Long-Context Agent

arXiv2 repos

arXiv:2602.02704

InfMem, infmem_superlong

SWE-World: Building Software Engineering Agents in Docker-Free Environments

arXiv2 repos

arXiv:2602.03419

SWE-Master, SWE-World

See-through: Single-image Layer Decomposition for Anime Characters

arXiv2 repos

arXiv:2602.03749

see-through-demo, ComfyUI-See-through

Nemotron ColEmbed V2: Top-Performing Late Interaction Embedding Models for Visual Document Retrieval

arXiv2 repos

arXiv:2602.03992

vllm-factory, nemotron-colembed-vl-4b-v2

Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems

arXiv2 repos

arXiv:2602.05176

model_collaboration, AmongUs

Revisiting the Generic Transformer: Deconstructing a Strong Baseline for Time Series Foundation Models

arXiv2 repos

arXiv:2602.06909

patchtst-fm-r1, granite-timeseries-patchtst-fm-r1

VideoVeritas: AI-Generated Video Detection via Perception Pretext Reinforcement Learning

arXiv2 repos

arXiv:2602.08828

Veritas, VideoVeritas

Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters

arXiv2 repos

arXiv:2602.10604

Step-3.5-Flash, Step-3.5-Flash

Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning

arXiv2 repos

arXiv:2602.11149

data-repetition, olmo3-7b_data-repetition

arXiv:2602.11910

arXiv2 repos

arXiv:2602.11910

steer-audio, patching-music-musiccaps-prompts

GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristics

arXiv2 repos

arXiv:2602.12617

GeoAgent, GeoAgent

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges

arXiv2 repos

arXiv:2602.13576

Rubrics-as-an-Attack-Surface, ripd-dataset

Experiential Reinforcement Learning

arXiv2 repos

arXiv:2602.13949

deepscaler, rllm

VLANeXt: Recipes for Building Strong VLA Models

arXiv2 repos

arXiv:2602.18532

VLANeXt, VLANeXt

WildOS: Open-Vocabulary Object Search in the Wild

arXiv2 repos

arXiv:2602.19308

nebula2-wildos, wildos

arXiv:2602.20113

arXiv2 repos

arXiv:2602.20113

StyleStream, StyleStream

PreScience: A Dataset and Benchmark for Scientific Forecasting

arXiv2 repos

arXiv:2602.20459

prescience, prescience

D-FINE-seg: Object Detection and Instance Segmentation Framework with multi-backend deployment

arXiv2 repos

arXiv:2602.23043

D-FINE-seg, D-FINE-seg

SenCache: Accelerating Diffusion Model Inference via Sensitivity-Aware Caching

arXiv2 repos

arXiv:2602.24208

maxdiffusion, ltx2-vidgen-skill

LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model

arXiv2 repos

arXiv:2603.01068

LLaDA-o, LLaDA-V

According to Me: Long-Term Personalized Referential Memory QA

arXiv2 repos

arXiv:2603.01990

ATM-Bench, ATM-Bench

Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training

arXiv2 repos

arXiv:2603.02208

reasoning-core, reasoning_core

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

arXiv2 repos

arXiv:2603.03269

LoGeR, LoGeR

Utonia: Toward One Encoder for All Point Clouds

arXiv2 repos

arXiv:2603.03283

Utonia, Utonia

$V_1$: Unifying Generation and Self-Verification for Parallel Reasoners

arXiv2 repos

arXiv:2603.04304

deepscaler, rllm

MSpoofTTS: Multi-Resolution Spoof-Guided Inference for Discrete Speech Synthesis

arXiv2 repos

arXiv:2603.05373

MSpoofTTS, MSpoofTTS

Scalable Training of Mixture-of-Experts Models with Megatron Core

arXiv2 repos

arXiv:2603.07685

Megatron-LM, geodesic-megatron

DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining

arXiv2 repos

arXiv:2603.08216

dualturn, dualturn-qwen2.5-mimi-0.5B

LLM2Vec-Gen: Generative Embeddings from Large Language Models

arXiv2 repos

arXiv:2603.10913

llm2vec, LLM2Vec-Gen-Qwen3-8B

arXiv:2603.11661

arXiv2 repos

arXiv:2603.11661

Resonate, Resonate

Real-World Point Tracking with Verifier-Guided Pseudo-Labeling

arXiv2 repos

arXiv:2603.12217

track_on, track_on_r

InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization

arXiv2 repos

arXiv:2603.13375

InfiniteDance, InfiniteDance

GeoNVS: Geometry Grounded Video Diffusion for Novel View Synthesis

arXiv2 repos

arXiv:2603.14965

GeoNVS, GeoNVS

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

arXiv2 repos

arXiv:2603.15118

varex-bench, VAREX

Gym-V: A Unified Vision Environment System for Agentic Vision Research

arXiv2 repos

arXiv:2603.15432

Game-RL, GameQA-140K

SegviGen: Repurposing 3D Generative Model for Part Segmentation

arXiv2 repos

arXiv:2603.16869

SegviGen, SegviGen

Beyond String Matching: Semantic Evaluation of PDF Table Extraction

arXiv2 repos

arXiv:2603.18652

pdf-parse-bench, pdf-parse-bench

Breeze Taigi: Benchmarks and Models for Taiwanese Hokkien Speech Recognition and Synthesis

arXiv2 repos

arXiv:2603.19259

Breeze-ASR-26, faster-whisper-Breeze-ASR-26

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

arXiv2 repos

arXiv:2603.19312

mlx-tune, lewm-pusht

The Universal Normal Embedding

arXiv2 repos

arXiv:2603.21786

UNE, NoiseZoo

ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention

arXiv2 repos

arXiv:2603.22016

ROM, ROM

MemDLM: Memory-Enhanced DLM Training

arXiv2 repos

arXiv:2603.22241

LLaDA-MoE-7B-A1B-Base-MemDLM, LLaDA2.1-mini-MemDLM

MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens

arXiv2 repos

arXiv:2603.23516

MSA, MSA-4B

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks

arXiv2 repos

arXiv:2603.24755

unlazy, scb-check

FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants

arXiv2 repos

arXiv:2603.26008

FairLLaVA, FairLLaVA

SonoWorld: From One Image to a 3D Audio-Visual Scene

arXiv2 repos

arXiv:2603.28757

sonoworld, SonoScene360

Generative World Renderer

arXiv2 repos

arXiv:2604.02329

AlayaRenderer, AlayaRenderer

Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?

arXiv2 repos

arXiv:2604.03016

Agentic-MME, Agentic-MME

TORA: Topological Representation Alignment for 3D Shape Assembly

arXiv2 repos

arXiv:2604.04050

tora, tora

Synthetic Sandbox for Training Machine Learning Engineering Agents

arXiv2 repos

arXiv:2604.04872

deepscaler, rllm

Squeez: Task-Conditioned Tool-Output Pruning for Coding Agents

arXiv2 repos

arXiv:2604.04979

tool-output-extraction-swebench, squeez-2b

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space

arXiv2 repos

arXiv:2604.05030

npcpy, qllm2

The Art of Building Verifiers for Computer Use Agents

arXiv2 repos

arXiv:2604.06240

CUAVerifierBench, WebTailBench

Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM

arXiv2 repos

arXiv:2604.06832

Fast-dLLM, Fast_dVLM_3B

PhysInOne: Visual Physics Learning and Reasoning in One Suite

arXiv2 repos

arXiv:2604.09415

PhysInOne, PhysInOne

Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing

arXiv2 repos

arXiv:2604.10708

Audio-Omni, Audio-Omni

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

arXiv2 repos

arXiv:2604.11804

OmniShow, HOIVG-Bench

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

arXiv2 repos

arXiv:2604.13016

MiniCPM5-1B-GGUF, MiniCPM5-1B

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation

arXiv2 repos

arXiv:2604.18486

OneVL_training, onevl

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

arXiv2 repos

arXiv:2604.19587

SmartPhotoCrafter, SmartPhotoCrafter

Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers

arXiv2 repos

arXiv:2604.21592

Sculpt4D, Sculpt4D

Building a Precise Video Language with Human-AI Oversight

arXiv2 repos

arXiv:2604.21718

t2v_metrics, CHAI_testset

CADFit: Precise Mesh-to-CAD Program Generation with Hybrid Optimization

arXiv2 repos

arXiv:2605.01171

CADFit, CADFit

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

arXiv2 repos

arXiv:2605.04128

JoyAI-Image, JoyAI-Image-Edit

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?

arXiv2 repos

arXiv:2605.06068

vibesys, vibe-train-sample

Relit-LiVE: Relight Video by Jointly Learning Environment Video

arXiv2 repos

arXiv:2605.06658

Relit-LiVE, Relit-LiVE

Pixal3D: Pixel-Aligned 3D Generation from Images

arXiv2 repos

arXiv:2605.10922

ComfyUI-Pixal3D, Pixal3D

Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning

arXiv2 repos

arXiv:2605.14386

Darwin-TTS-1.7B-Cross-Qwen3Tokenizer, Darwin-TTS-1.7B-Cross

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization

arXiv2 repos

arXiv:2605.17757

OSCAR, OSCAR-RotationZoo

PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis

arXiv2 repos

arXiv:2605.17916

PanoWorld, PanoWorld

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload

arXiv2 repos

arXiv:2605.20179

TIDE, TIDE_DATA_COLLECTION

AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild

arXiv2 repos

arXiv:2605.22715

AnyMo, AnyMo-Bench

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving

arXiv2 repos

arXiv:2605.23163

Fast-dDrive, Fast-dLLM

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

arXiv2 repos

arXiv:2605.23904

darwin-skill, SkillOpt

Raon-Speech Technical Report

arXiv2 repos

arXiv:2605.23912

Raon-Speech-9B, Raon-SpeechChat-9B

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding

arXiv2 repos

arXiv:2605.27365

LocateAnything-3B, LocateAnything

Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation

arXiv2 repos

arXiv:2605.28642

ESRT-4B, esrt

DriveMA: Driving Vision-Language-Action Models with verifiable Meta-Actions

arXiv2 repos

arXiv:2605.31271

DriveMA-2B, DriveMA_Datasets

MOSS-Audio Technical Report

arXiv2 repos

arXiv:2606.01802

MOSS-Audio, MOSS-audio-swift

Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses

arXiv2 repos

arXiv:2606.02373

harness-1, harness-1

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

arXiv2 repos

arXiv:2606.02624

TadA-Bench, TadA-Bench

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

arXiv2 repos

arXiv:2606.09079

FlashMemory-Deepseek-V4, FlashMemory-Deepseek-V4

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

arXiv2 repos

arXiv:2606.09186

DuplexOmni, DuplexOmni-Data

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

arXiv2 repos

arXiv:2606.12886

Game-RL, GameQA-140K

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

arXiv2 repos

arXiv:2606.14516

every_eval_ever, EEE_datastore

ReportQA: QA-Based Radiology Report Evaluation

arXiv2 repos

arXiv:2606.15037

ReportQA, ReportQA

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

arXiv2 repos

arXiv:2606.19047

AWorld-RL, Qwen3-4B-RODS

PorTEXTO: A European Portuguese Benchmark for Visual Text Extraction

arXiv2 repos

arXiv:2606.19096

amalia-vl-eval, PorTEXTO

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

arXiv2 repos

arXiv:2606.19195

Moebius, Moebius

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

arXiv2 repos

arXiv:2606.19348

DeepSeek-V4-Flash-0731, DeepSeek-V4-Pro

Improved Large Language Diffusion Models

arXiv2 repos

arXiv:2606.25331

iLLaDA-8B-Base, iLLaDA-8B-Instruct

Frequency-Aware Self-Supervised Music Representation Learning

arXiv2 repos

arXiv:2606.25713

MERT-v2-30s, MERT-v2-FullSong

Orca: The World is in Your Mind

arXiv2 repos

arXiv:2606.30534

Orca, Orca-4B

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

arXiv2 repos

arXiv:2606.31435

data-juicer-hub, CDR-Bench

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts

arXiv2 repos

arXiv:2606.31986

CoLT, CoLT-8B

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

arXiv2 repos

arXiv:2607.04064

speaker_disentangled_hubert, SylReg-LM-7B

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

arXiv2 repos

arXiv:2607.04884

HunyuanOCR, HunyuanOCR

ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching

arXiv2 repos

arXiv:2607.07119

ColorFM, ColorFM

Persona Cartography: Charting Language Model Personality Traits in Weight Space

arXiv2 repos

arXiv:2607.07916

persona-cartography, monorepo

MuScriptor: An Open Model for Multi-Instrument Music Transcription

arXiv2 repos

arXiv:2607.08168

qinglong-captions, HOT-Step-CPP

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

arXiv2 repos

arXiv:2607.08716

deepscaler, rllm

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

arXiv2 repos

arXiv:2607.11562

MonkeyOCRv2, MonkeyOCRv2-B-Parsing-DFlash

Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate

arXiv2 repos

arXiv:2607.12266

Qwen3.8-27B-K4, Qwen3.8-27B-EXL3-K5K6

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

arXiv2 repos

arXiv:2607.14935

VideoChat-Flash, Ask-Anything

ShotPlan: Cinematic Video Generation with Learnable Planning Token

arXiv2 repos

arXiv:2607.17675

ShotPlan-Wan2.2-T2V-A14B-HighNoise, ShotPlan-Wan2.1-T2V-14B

A Controlled Study of Attention-Only Transformers

arXiv2 repos

arXiv:2607.18363

mimimodel, needle

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

arXiv2 repos

arXiv:2607.18934

CrisperWhisper, CrisperWhisper2.0_large

AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters

arXiv2 repos

arXiv:2607.19223

AdaFlash, Qwen3-8B-AdaFlash

VibeVoice-ASR-BitNet Technical Report

arXiv2 repos

arXiv:2607.21075

VibeVoice, VibeVoice-ASR-BitNet

ID-V2V: Identity-Preserving Video Restylization

arXiv2 repos

arXiv:2607.22830

ID-V2V, ID-V2V

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

arXiv2 repos

arXiv:2607.23855

OmniVAE, OmniVAE

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding

arXiv2 repos

arXiv:2607.25852

TorchSpec, TorchSpec

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

arXiv2 repos

arXiv:2607.27113

Veritas, VideoVeritas

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

arXiv2 repos

arXiv:2607.28595

Game-RL, GameQA-140K

PhiZero: A World Model Built Around Physical Language

arXiv2 repos

arXiv:2607.28624

PhiZero, PhiZero

Bridging the English-Arabic Medical Knowledge Gap: Targeted Low-Rank Adaptation via Causal Layer Selection

arXiv2 repos

arXiv:2608.00207

Adaptation, TLoRA-Adaptation

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

arXiv2 repos

arXiv:2608.02673

dots.tts, dots.tts.edit

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

arXiv2 repos

arXiv:2608.04205

MatrAIx-Persona-8B, MatrAIx_Persona_1M

DarwinX: Evolving Agent Harnesses Through Natural Selection

arXiv2 repos

arXiv:2608.07545

Beagle, darwinx

Instruction-Based Video Editing by Repurposing an Image Editing Model

arXiv2 repos

arXiv:2608.14790

Qwen-Video-Edit, Qwen-Video-Edit

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

arXiv2 repos

arXiv:2608.16157

FreeToken, sparklab

FixAnything: 3D-Consistent Rendering Refinement via Video Generative Priors

arXiv2 repos

arXiv:2608.23549

fix-anything, fix-anything

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

arXiv2 repos

arXiv:2609.03047

shelf-benchmark, SHELF

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

arXiv2 repos

arXiv:2609.05405

WearableQA, WearableQA

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

arXiv2 repos

arXiv:2609.08936

AuK, AuK-Flash

Essential Incompleteness of Arithmetic Verified by Coq

arXiv2 repos

arXiv:cs/0505034

hydra-battles, hydra-battles

MAPS: pathologist-level cell type annotation from tissue images through machine learning

Nature2 repos

Nature:s41467-023-44188-w

CORAL, KRONOS2

GrandQC: A comprehensive solution to quality control problem in digital pathology

Nature2 repos

Nature:s41467-024-54769-y

TRIDENT, TridentEdited

Explainable machine-learning predictions for the prevention of hypoxaemia during surgery

Nature2 repos

Nature:s41551-018-0304-0

shap, shap

Data-efficient and weakly supervised computational pathology on whole-slide images

Nature2 repos

Nature:s41551-020-00682-w

CLAM, Mussel

Large language models encode clinical knowledge

Nature2 repos

Nature:s41586-023-06291-2

Awesome-Medical-Dataset, healthsearchqa

World and Human Action Models towards gameplay ideation

Nature2 repos

Nature:s41586-025-08600-3

Awesome-From-Video-Generation-to-World-Model, wham

The Virtual Tissues foundation model resolves spatial proteomics across scales

Nature2 repos

Nature:s41586-026-10884-y

virtues, virtues

A multimodal whole-slide foundation model for pathology

Nature2 repos

Nature:s41591-025-03982-3

AtlasPatch, TITAN

NeuroMechFly v2: simulating embodied sensorimotor control in adult Drosophila

Nature2 repos

Nature:s41592-024-02497-y

flygym, fly

From whole-slide image to biomarker prediction: end-to-end weakly supervised deep learning in computational pathology

Nature2 repos

Nature:s41596-024-01047-2

STAMP, STAMP_attention_ui

HyperKvasir, a comprehensive multi-class image and video dataset for gastrointestinal endoscopy

Nature2 repos

Nature:s41597-020-00622-y

Endo-FM, hyper-kvasir

A high-resolution temporal transcriptomic and imaging dataset of porcine wound healing

Nature2 repos

Nature:s41597-025-05921-w

wound-forecasting, porcine-wound-forecasting-processed

From local explanations to global understanding with explainable AI for trees

Nature2 repos

Nature:s42256-019-0138-9

shap, shap

A Conversation with Mary Lou Jepsen

ACM1 repo

ACM:1331287.1331291

Low-power-E-Paper-OS

Reciprocal rank fusion outperforms condorcet and individual rank learning methods

ACM1 repo

ACM:1571941.1572114

ogham-mcp

OXenstored

ACM1 repo

ACM:1631687.1596581

irmin

Printing floating-point numbers quickly and accurately with integers

ACM1 repo

ACM:1806596.1806623

godot

Manifold bootstrapping for SVBRDF capture

ACM1 repo

ACM:1833349.1778835

Awesome-InverseRendering

Learning to efficiently rank

ACM1 repo

ACM:1835449.1835475

RLT4Reranking

A standardised polarisation visualisation for images

ACM1 repo

ACM:1925059.1925070

Awesome-Polarization

A mixed reality system for virtual glasses try-on

ACM1 repo

ACM:2087756.2087816

awesome-virtual-try-on

Virtual try-on of eyeglasses using 3D model of the head

ACM1 repo

ACM:2087756.2087838

awesome-virtual-try-on

Minimizing row displacement dispatch tables

ACM1 repo

ACM:217839.217851

sdk

Printing floating-point numbers quickly and accurately

ACM1 repo

ACM:231379.231397

godot

Beyond random walk and metropolis-hastings samplers

ACM1 repo

ACM:2318857.2254795

littleballoffur

A volumetric method for building complex models from range images

ACM1 repo

ACM:237170.237269

awesome-mvs

Practical SVBRDF capture in the frequency domain

ACM1 repo

ACM:2461912.2461978

Awesome-InverseRendering

Metric convergence in social network sampling

ACM1 repo

ACM:2491159.2491168

littleballoffur

Network Sampling

ACM1 repo

ACM:2601438

littleballoffur

Garment Replacement in Monocular Video Sequences

ACM1 repo

ACM:2634212

awesome-virtual-try-on

ESC

ACM1 repo

ACM:2733373.2806390

BEANS-Zero

Two-shot SVBRDF capture for stationary materials

ACM1 repo

ACM:2766967

Awesome-InverseRendering

Learning to Represent Knowledge Graphs with Gaussian Embedding

ACM1 repo

ACM:2806416.2806502

pykeen

The use of MMR, diversity-based reranking for reordering documents and producing summaries

ACM1 repo

ACM:290941.291025

ogham-mcp

Asymmetric Transitivity Preserving Graph Embedding

ACM1 repo

ACM:2939672.2939751

karateclub

ACM:296806.296824

ACM1 repo

ACM:296806.296824

szl-mesh

KickStarter

ACM1 repo

ACM:3037697.3037748

awesome-dynamic-graphs

GPU Virtualization and Scheduling Methods

ACM1 repo

ACM:3068281

awesome-gpu-engineering

Anserini

ACM1 repo

ACM:3077136.3080721

anserini

Selfie and the basics

ACM1 repo

ACM:3133850.3133857

selfie

Verifying strong eventual consistency in distributed systems

ACM1 repo

ACM:3133933

awesome-local-first

FreeGuard

ACM1 repo

ACM:3133956.3133957

mimalloc-bench

Random sampling with a reservoir

ACM1 repo

ACM:3147.3165

qsv

TinyLFU

ACM1 repo

ACM:3149371

caffeine

ACM:3164135.3164139

ACM1 repo

ACM:3164135.3164139

awesome-dynamic-graphs

On optimistic methods for concurrency control

ACM1 repo

ACM:319566.319567

miniflare

Graphtides: a framework for evaluating stream-based graph processing platforms

ACM1 repo

ACM:3210259.3210262

awesome-dynamic-graphs

Partisan

ACM1 repo

ACM:3231104.3231106

partisan

Plan3D

ACM1 repo

ACM:3233794

awesome-mvs

Anserini

ACM1 repo

ACM:3239571

anserini

Adaptive Software Cache Management

ACM1 repo

ACM:3274808.3274816

caffeine

Software multiplexing: share your libraries and statically link them too

ACM1 repo

ACM:3276524

linux_distro_tests

GraphBolt

ACM1 repo

ACM:3302424.3303974

awesome-dynamic-graphs

Learning to Represent the Evolution of Dynamic Graphs with Recurrent Models

ACM1 repo

ACM:3308560.3316581

pytorch_geometric_temporal

Gen: a general-purpose probabilistic programming system with programmable inference

ACM1 repo

ACM:3314221.3314642

Gen.jl

PASE: PostgreSQL Ultra-High-Dimensional Approximate Nearest Neighbor Search Extension

ACM1 repo

ACM:3318464.3386131

pgvector

Deeper Text Understanding for IR with Contextual Neural Language Modeling

ACM1 repo

ACM:3331184.3331303

RLT4Reranking

Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations

ACM1 repo

ACM:3336191.3371856

pypantera

Massively Parallel ANS Decoding on GPUs

ACM1 repo

ACM:3337821.3337888

dietgpu

Enhancing Graph Neural Network-based Fraud Detectors against Camouflaged Fraudsters

ACM1 repo

ACM:3340531.3411903

antifraud

ElderReact: A Multimodal Dataset for Recognizing Emotional Response in Aging Adults

ACM1 repo

ACM:3340555.3353747

ElderReact

CATSLU: The 1st Chinese Audio-Textual Spoken Language Understanding Challenge

ACM1 repo

ACM:3340555.3356098

ChineseNLPCorpus

Cubical agda: a dependently typed programming language with univalence and higher inductive types

ACM1 repo

ACM:3341691

cubical

An Assumption-Free Approach to the Dynamic Truncation of Ranked Lists

ACM1 repo

ACM:3341981.3344234

RLT4Reranking

Virtually Trying on New Clothing with Arbitrary Poses

ACM1 repo

ACM:3343031.3350946

awesome-virtual-try-on

A Framework for Effective Known-item Search in Video

ACM1 repo

ACM:3343031.3351046

TransNetV2

FashionOn

ACM1 repo

ACM:3343031.3351075

awesome-virtual-try-on

G raph O ne

ACM1 repo

ACM:3364180

awesome-dynamic-graphs

Extracting Knowledge from Web Text with Monte Carlo Tree Search

ACM1 repo

ACM:3366423.3380010

awesome-monte-carlo-tree-search-papers

Overview of the HASOC track at FIRE 2019

ACM1 repo

ACM:3368567.3368584

Tutorial-Resources

An equational theory for weak bisimulation via generalized parameterized coinduction

ACM1 repo

ACM:3372885.3373813

paco

The Deceptive Potential of Common Design Tactics Used in Data Visualizations

ACM1 repo

ACM:3380851.3416762

acl2026-misleading-visualizations

A robust and flexible operating system compatibility architecture

ACM1 repo

ACM:3381052.3381327

kerla

IMACS - an <u>i</u>nteractive cognitive assistant <u>m</u>odule for <u>c</u>ardiac <u>a</u>rrest cases in emergency medical <u>s</u>ervice

ACM1 repo

ACM:3384419.3430451

EMS-Pipeline

Down to the Last Detail

ACM1 repo

ACM:3394171.3413514

awesome-virtual-try-on

VideoIC: A Video Interactive Comments Dataset and Multimodal Multitask Learning for Comments Generation

ACM1 repo

ACM:3394171.3413890

VideoIC

Geodesic Forests

ACM1 repo

ACM:3394486.3403094

awesome-decision-tree-papers

Predicting Temporal Sets with Deep Neural Networks

ACM1 repo

ACM:3394486.3403152

pytorch_geometric_temporal

Choppy: Cut Transformer for Ranked List Truncation

ACM1 repo

ACM:3397271.3401188

RLT4Reranking

Evidence Weighted Tree Ensembles for Text Classification

ACM1 repo

ACM:3397271.3401229

awesome-decision-tree-papers

Huffman Coding with Gap Arrays for GPU Acceleration

ACM1 repo

ACM:3404397.3404429

dietgpu

Propensity-scored Probabilistic Label Trees

ACM1 repo

ACM:3404835.3463084

napkinXC

Tools, Tricks, and Hacks: Exploring Novel Digital Fabrication Workflows on #PlotterTwitter

ACM1 repo

ACM:3411764.3445653

awesome-plotters

Offsite aerial path planning for efficient urban scene reconstruction

ACM1 repo

ACM:3414685.3417791

awesome-mvs

Single image portrait relighting via explicit multiple reflectance channel modeling

ACM1 repo

ACM:3414685.3417824

Awesome-InverseRendering

Studying Politeness across Cultures using English Twitter and Mandarin Weibo

ACM1 repo

ACM:3415190

funNLP

Probabilistic Gradient Boosting Machines for Large-Scale Probabilistic Regression

ACM1 repo

ACM:3447548.3467278

awesome-decision-tree-papers

BLOCKSET (Block-Aligned Serialized Trees)

ACM1 repo

ACM:3447548.3467368

awesome-decision-tree-papers

ControlBurn

ACM1 repo

ACM:3447548.3467387

awesome-decision-tree-papers

Unikraft

ACM1 repo

ACM:3447786.3456248

awesome-os

Retrofitting effect handlers onto OCaml

ACM1 repo

ACM:3453483.3454039

effects-examples

Efficient large-scale language model training on GPU clusters using megatron-LM

ACM1 repo

ACM:3458817.3476209

awesome-gpu-engineering

Learning to Pack

ACM1 repo

ACM:3459637.3481933

awesome-monte-carlo-tree-search-papers

Geometric Heuristics for Transfer Learning in Decision Trees

ACM1 repo

ACM:3459637.3482259

awesome-decision-tree-papers

Fairness-Aware Training of Decision Trees by Abstract Interpretation

ACM1 repo

ACM:3459637.3482342

awesome-decision-tree-papers

Unsupervised Domain Adaptation for Static Malware Detection based on Gradient Boosting Trees

ACM1 repo

ACM:3459637.3482400

awesome-gradient-boosting-papers

BNN

ACM1 repo

ACM:3459637.3482414

awesome-gradient-boosting-papers

CrossVul: a cross-language vulnerability dataset with commit data

ACM1 repo

ACM:3468264.3473122

FinalYearProject

Marcelle: Composing Interactive Machine Learning Workflows and Interfaces

ACM1 repo

ACM:3472749.3474734

IR-Lens

Screen2Words: Automatic Mobile UI Summarization with Multimodal Learning

ACM1 repo

ACM:3472749.3474765

RefinedVision

Relevance under the Iceberg

ACM1 repo

ACM:3477495.3531767

pecos

From Distillation to Hard Negative Sampling

ACM1 repo

ACM:3477495.3531857

Rankify

InPars: Unsupervised Dataset Generation for Information Retrieval

ACM1 repo

ACM:3477495.3531863

InPars

Multimodal Entity Linking with Gated Hierarchical Fusion and Contrastive Training

ACM1 repo

ACM:3477495.3531867

GEMEL

Incorporating Retrieval Information into the Truncation of Ranking Lists for Better Legal Search

ACM1 repo

ACM:3477495.3531998

RLT4Reranking

Trimming Data Sets: a Verified Algorithm for Robust Mean Estimation

ACM1 repo

ACM:3479394.3479412

infotheo

Enterprise-Scale Search: Accelerating Inference for Sparse Extreme Multi-Label Ranking Trees

ACM1 repo

ACM:3485447.3511973

pecos

GPU Accelerated Boosted Trees and Deep Neural Networks for Better Recommender Systems

ACM1 repo

ACM:3487572.3487605

xgboost

MtCut

ACM1 repo

ACM:3488560.3498466

RLT4Reranking

Transform, Warp, and Dress: A New Transformation-guided Model for Virtual Try-on

ACM1 repo

ACM:3491226

awesome-virtual-try-on

D3

ACM1 repo

ACM:3492321.3519576

erdos

Lightweight Robust Size Aware Cache Management

ACM1 repo

ACM:3507920

caffeine

Cross-category Virtual Try-on Technology Research Based on PF-AFN

ACM1 repo

ACM:3511176.3511201

awesome-virtual-try-on

Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?

ACM1 repo

ACM:3531146.3533192

Q16

Integrity Authentication in Tree Models

ACM1 repo

ACM:3534678.3539428

awesome-decision-tree-papers

Online Clustering

ACM1 repo

ACM:3534678.3542600

river

SciTS: A Benchmark for Time-Series Databases in Scientific Experiments and Industrial Internet of Things

ACM1 repo

ACM:3538712.3538723

ClickBench

Surprise: Result List Truncation via Extreme Value Theory

ACM1 repo

ACM:3539618.3592066

RLT4Reranking

FINGER: Fast Inference for Graph-based Approximate Nearest Neighbor Search

ACM1 repo

ACM:3543507.3583318

pecos

Filtered-DiskANN: Graph Algorithms for Approximate Nearest Neighbor Search with Filters

ACM1 repo

ACM:3543507.3583552

pgvectorscale

EasySpider: A No-Code Visual System for Crawling the Web

ACM1 repo

ACM:3543873.3587345

EasySpider

CALVI: Critical Thinking Assessment for Literacy in Visualizations

ACM1 repo

ACM:3544548.3581406

acl2026-misleading-visualizations

Aeneas: Rust verification by functional translation

ACM1 repo

ACM:3547647

aeneas

Make Your Own Sprites

ACM1 repo

ACM:3550454.3555482

Pixelization

Optimization Techniques for GPU Programming

ACM1 repo

ACM:3570638

awesome-gpu-engineering

Omnisemantics: Smooth Handling of Nondeterminism

ACM1 repo

ACM:3579834

bedrock2

Taming the Domain Shift in Multi-source Learning for Energy Disaggregation

ACM1 repo

ACM:3580305.3599910

awesome-nilm

Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data Poisoning

ACM1 repo

ACM:3581783.3612108

BackdoorDM

MER 2023: Multi-label Learning, Modality Robustness, and Semi-Supervised Learning

ACM1 repo

ACM:3581783.3612836

MERTools

FORGE: Pre-Training Open Foundation Models for Science

ACM1 repo

ACM:3581784.3613215

gpt-neox

EMSAssist: An End-to-End Mobile Voice Assistant at the Edge for Emergency Medical Services

ACM1 repo

ACM:3581791.3596853

EMS-Pipeline

Interactive Latent Diffusion Model

ACM1 repo

ACM:3583131.3590471

sdstudio

Practice on Effectively Extracting NLP Features for Click-Through Rate Prediction

ACM1 repo

ACM:3583780.3614707

BAIU

PyABSA: A Modularized Framework for Reproducible Aspect-based Sentiment Analysis

ACM1 repo

ACM:3583780.3614752

PyABSA

SE-PQA: Personalized Community Question Answering

ACM1 repo

ACM:3589335.3651445

SE-PQA

PG-Schema: Schemas for Property Graphs

ACM1 repo

ACM:3589778

rudof

CryptOpt: Verified Compilation with Randomized Program Search for Cryptographic Primitives

ACM1 repo

ACM:3591272

CryptOpt

Delilah: eBPF-offload on Computational Storage

ACM1 repo

ACM:3592980.3595319

awesome-ebpf

SSProve: A Foundational Framework for Modular Cryptographic Proofs in Coq

ACM1 repo

ACM:3594735

ssprove

Detail-Preserving Video-based Virtual Try-On (DPV-VTON)

ACM1 repo

ACM:3599589.3599599

awesome-virtual-try-on

PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation

ACM1 repo

ACM:3600006.3613139

MInference

A Grounded Conceptual Model for Ownership Types in Rust

ACM1 repo

ACM:3622841

aquascope

Understanding Any Time Series Classifier with a Subsequence-based Explainer

ACM1 repo

ACM:3624480

Awesome-Time-Series-Explainability

CIVSCOPE: Analyzing Potential Memory Corruption Bugs in Compartment Interfaces

ACM1 repo

ACM:3625275.3625399

publications

ACM TechBrief: Generative Artificial Intelligence

ACM1 repo

ACM:3626110

awesome-ai4lam

OpenIVM: a SQL-to-SQL Compiler for Incremental Computations

ACM1 repo

ACM:3626246.3654743

openivm

ALP: Adaptive Lossless floating-Point Compression

ACM1 repo

ACM:3626717

ALP

MACRec: A Multi-Agent Collaboration Framework for Recommendation

ACM1 repo

ACM:3626772.3657669

MACRec

Resources for Brewing BEIR: Reproducible Reference Models and Statistical Analyses

ACM1 repo

ACM:3626772.3657862

beir

Ranked List Truncation for Large Language Model-based Re-Ranking

ACM1 repo

ACM:3626772.3657864

RLT4Reranking

Fine-Tuning LLaMA for Multi-Stage Text Retrieval

ACM1 repo

ACM:3626772.3657951

reranker-as-judge

pyPANTERA: A Python PAckage for Natural language obfuscaTion Enforcing pRivacy & Anonymization

ACM1 repo

ACM:3627673.3679173

pypantera

Distributed Boosting: An Enhancing Method on Dataset Distillation

ACM1 repo

ACM:3627673.3679897

awesome-gradient-boosting-papers

Exploring Performance and Cost Optimization with ASIC-Based CXL Memory

ACM1 repo

ACM:3627703.3650061

lightllm

Guided Equality Saturation

ACM1 repo

ACM:3632900

equational_theories

Endoprocess: Programmable and Extensible Subprocess Isolation

ACM1 repo

ACM:3633500.3633507

publications

PEMBOT: Pareto-Ensembled Multi-task Boosted Trees

ACM1 repo

ACM:3637528.3671619

awesome-gradient-boosting-papers

ImputeFormer: Low Rankness-Induced Transformers for Generalizable Spatiotemporal Imputation

ACM1 repo

ACM:3637528.3671751

frn-50k-baseline

Iterative Weak Learnability and Multiclass AdaBoost

ACM1 repo

ACM:3637528.3671842

awesome-gradient-boosting-papers

Uplift Modelling via Gradient Boosting

ACM1 repo

ACM:3637528.3672019

awesome-gradient-boosting-papers

FaultProfIT: Hierarchical Fault Profiling of Incident Tickets in Large-scale Cloud Systems

ACM1 repo

ACM:3639477.3639754

awesome-LLM-AIOps

Hypermedia Controls: Feral to Formal

ACM1 repo

ACM:3648188.3675127

fixi

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

ACM1 repo

ACM:3658644.3670388

jailbreak_llms

ClassID: Enabling Student Behavior Attribution from Ambient Classroom Sensing Systems

ACM1 repo

ACM:3659586

ClassID

Leveraging Large Language Models for the Auto-remediation of Microservice Applications: An Experimental Study

ACM1 repo

ACM:3663529.3663855

awesome-LLM-AIOps

Which Neurons Matter in IR? Applying Integrated Gradients-based Methods to Understand Cross-Encoders

ACM1 repo

ACM:3664190.3672528

IR-Lens

BrainRAM: Cross-Modality Retrieval-Augmented Image Reconstruction from Human Brain Activity

ACM1 repo

ACM:3664647.3681296

BrainRAM

Sound Borrow-Checking for Rust via Symbolic Semantics

ACM1 repo

ACM:3674640

aeneas

Design and Implementation of a Coverage-Guided Ruby Fuzzer

ACM1 repo

ACM:3675741.3675749

ruzzy

Past-Future Scheduler for LLM Serving under SLA Guarantees

ACM1 repo

ACM:3676641.3716011

lightllm

Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Trustworthy Response Generation in Chinese

ACM1 repo

ACM:3686807

Huatuo-Llama-Med-Chinese

Reproducibility Report for ACM SIGMOD 2024 Paper: 'ALP: Adaptive Lossless Floating-Point Compression'

ACM1 repo

ACM:3687998.3717057

ALP

MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition

ACM1 repo

ACM:3689092.3689959

MERTools

StarMalloc: Verifying a Modern, Hardened Memory Allocator

ACM1 repo

ACM:3689773

mimalloc-bench

CoqPilot, a plugin for LLM-based generation of proofs

ACM1 repo

ACM:3691620.3695357

coqpilot

The Potential of One-Shot Failure Root Cause Analysis: Collaboration of the Large Language Model and Small Classifier

ACM1 repo

ACM:3691620.3695475

awesome-LLM-AIOps

LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism

ACM1 repo

ACM:3694715.3695948

lightllm

Reducing Energy Bloat in Large Model Training

ACM1 repo

ACM:3694715.3695970

zeus

Unearthing Semantic Checks for Cloud Infrastructure-as-Code Programs

ACM1 repo

ACM:3694715.3695974

awesome-LLM-AIOps

Grad: Guided Relation Diffusion Generation for Graph Augmentation in Graph Fraud Detection

ACM1 repo

ACM:3696410.3714520

antifraud

Hybrid, Unified and Iterative: A Novel Framework for Text-based Person Anomaly Retrieval

ACM1 repo

ACM:3701716.3717653

Hybrid-Unified-and-Iterative-A-Novel-Framework-for-Text-based-Person-Anomaly-Retrieval

A Survey of Geometric Optimization for Deep Learning: From Euclidean Space to Riemannian Manifold

ACM1 repo

ACM:3708498

ai-agent-book

SymphonyQG: Towards Symphonious Integration of Quantization and Graph for Approximate Nearest Neighbor Search

ACM1 repo

ACM:3709730

RaBitQ-Library

Helping the Helper : Supporting Peer Counselors via AI-Empowered Practice and Feedback

ACM1 repo

ACM:3710993

CARE

Multi-modal Time Series Analysis: A Tutorial and Survey

ACM1 repo

ACM:3711896.3736567

TS-RAG

Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language Models

ACM1 repo

ACM:3726302.3730059

PAD

TourMLLM: A Retrieval-Augmented Multimodal Large Language Model for Multitask Learning in the Tourism Domain

ACM1 repo

ACM:3731715.3733450

TourMLLM

Semantics of Probabilistic Programs Using s-Finite Kernels in Dependent Type Theory

ACM1 repo

ACM:3732291

analysis

G-ALP: Rethinking Light-weight Encodings for GPUs

ACM1 repo

ACM:3736227.3736242

fastlanes

MER 2025: When Affective Computing Meets Large Language Models

ACM1 repo

ACM:3746027.3762007

MERTools

Federated Gradient Boosting for Financial Fraud Detection: An Empirical Study in the Banking Sector

ACM1 repo

ACM:3746252.3760891

awesome-gradient-boosting-papers

FairRegBoost: An End-to-End Data Processing Framework for Fair and Scalable Regression

ACM1 repo

ACM:3746252.3761277

awesome-gradient-boosting-papers

Cloud Infrastructure Management in the Age of AI Agents

ACM1 repo

ACM:3759441.3759443

awesome-LLM-AIOps

A Joint Classification Method for Traditional Chinese Medicine Diseases and Syndromes Based on BertChinese-RCNNATTN

ACM1 repo

ACM:3759972.3759979

SwanLab

Multi-Agent LLM Reasoning for Clinical Procedure Sequencing from High-Granularity EHR Data

ACM1 repo

ACM:3765612.3767238

Multiagent_Procedure_MIMIC-III

AgentSight: System-Level Observability for AI Agents Using eBPF

ACM1 repo

ACM:3766882.3767169

agentsight

Representation-Aware Root Cause Analysis with Large Language Models (Position Paper)

ACM1 repo

ACM:3777911.3801108

awesome-LLM-AIOps

Cylindrical Algebraic Decomposition in Coq/Rocq

ACM1 repo

ACM:3779031.3779100

analysis

Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization

ACM1 repo

ACM:3786161.3788460

boatse-extractor

A Replicability Study of Joint Product Quantisation for Effective Space-Efficient Dense Retrieval

ACM1 repo

ACM:3805712.3808565

pyterrier_dr

Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior in Generative Retrieval

ACM1 repo

ACM:3805712.3808567

lost-in-decoding

Programming languages for distributed computing systems

ACM1 repo

ACM:72551.72552

nextflow

Format-directed list processing in LISP

ACM1 repo

ACM:800005.807968

regal

Random sampling from hash files

ACM1 repo

ACM:93605.98746

mongo

Light Logics and Optimal Reduction: Completeness and Complexity

arXiv1 repo

arXiv:0704.2448

ESCoC

Power-law distributions in empirical data

arXiv1 repo

arXiv:0706.1062

tokenizer-flores-validation

Steganography of VoIP Streams

arXiv1 repo

arXiv:0805.2938

awesome-rtc-hacking

Estimating and Sampling Graphs with Multidimensional Random Walks

arXiv1 repo

arXiv:1002.1751

littleballoffur

Type Classes for Mathematics in Type Theory

arXiv1 repo

arXiv:1102.1323

math-classes

Deciding Kleene Algebras in Coq

arXiv1 repo

arXiv:1105.4537

atbr

RTED: A Robust Algorithm for the Tree Edit Distance

arXiv1 repo

arXiv:1201.0230

opendataloader-bench

A Synthesis of the Procedural and Declarative Styles of Interactive Theorem Proving

arXiv1 repo

arXiv:1201.3601

hol-light

Harmony Explained: Progress Towards A Scientific Theory of Music

arXiv1 repo

arXiv:1202.4212

awesome-music-production

Programming with Algebraic Effects and Handlers

arXiv1 repo

arXiv:1203.1539

eff

Automatic facial feature extraction and expression recognition based on neural network

arXiv1 repo

arXiv:1204.2073

awesome-affective-computing

BPR: Bayesian Personalized Ranking from Implicit Feedback

arXiv1 repo

arXiv:1205.2618

implicit

Public Key Cryptography Standards: PKCS

arXiv1 repo

arXiv:1207.5446

awesome-standards

PaxosLease: Diskless Paxos for Leases

arXiv1 repo

arXiv:1209.4187

translations

Fast Packed String Matching for Short Patterns

arXiv1 repo

arXiv:1209.6449

gecko-dev

Sequence Transduction with Recurrent Neural Networks

arXiv1 repo

arXiv:1211.3711

GigaAM

A Multilingual Semantic Wiki Based on Attempto Controlled English and Grammatical Framework

arXiv1 repo

arXiv:1303.4293

logicmoo_workspace

Functional Package Management with Guix

arXiv1 repo

arXiv:1305.4584

Functional-Programming

Bounding the Estimation Error of Sampling-based Shapley Value Approximation

arXiv1 repo

arXiv:1306.4265

shapley

Generating Sequences With Recurrent Neural Networks

arXiv1 repo

arXiv:1308.0850

char-rnn

Distributed Representations of Words and Phrases and their Compositionality

arXiv1 repo

arXiv:1310.4546

Word-Embeddings-Repository-for-Turkish

Asymmetric numeral systems: entropy coding combining speed of Huffman coding with compression rate of arithmetic coding

arXiv1 repo

arXiv:1311.2540

FiniteStateEntropy

Computing in Operations Research using Julia

arXiv1 repo

arXiv:1312.1431

JuMP.jl

Experience Implementing a Performant Category-Theory Library in Coq

arXiv1 repo

arXiv:1401.7694

agda-categories

Better bitmap performance with Roaring bitmaps

arXiv1 repo

arXiv:1402.6407

roaring-rs

Principles of Antifragile Software

arXiv1 repo

arXiv:1404.3056

awesome-chaos-engineering

NILMTK: An Open Source Toolkit for Non-intrusive Load Monitoring

arXiv1 repo

arXiv:1404.3878

awesome-nilm

Static Analysis for Regular Expression Exponential Runtime via Substructural Logics (Extended)

arXiv1 repo

arXiv:1405.7058

fancy-regex

Koka: Programming with Row Polymorphic Effect Types

arXiv1 repo

arXiv:1406.2061

koka

Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition

arXiv1 repo

arXiv:1406.2227

Meta-SelfLearning

Convolutional Neural Networks for Sentence Classification

arXiv1 repo

arXiv:1408.5882

simple-ntc

Very Deep Convolutional Networks for Large-Scale Image Recognition

arXiv1 repo

arXiv:1409.1556

materialgan

Going Deeper with Convolutions

arXiv1 repo

arXiv:1409.4842

caffe

Memory Networks

arXiv1 repo

arXiv:1410.3916

nmt

Conditional Generative Adversarial Nets

arXiv1 repo

arXiv:1411.1784

ocaml-torch

Show and Tell: A Neural Image Caption Generator

arXiv1 repo

arXiv:1411.4555

mscoco-it

Learn Physics by Programming in Haskell

arXiv1 repo

arXiv:1412.4880

Functional-Programming

ORB-SLAM: a Versatile and Accurate Monocular SLAM System

arXiv1 repo

arXiv:1502.00956

stella_vslam

Gated Feedback Recurrent Neural Networks

arXiv1 repo

arXiv:1502.02367

parallel-ss-dep

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

arXiv1 repo

arXiv:1502.03044

LaTeX_OCR

arXiv:1503.03465

arXiv1 repo

arXiv:1503.03465

clhash

U-Net: Convolutional Networks for Biomedical Image Segmentation

arXiv1 repo

arXiv:1505.04597

Landslide4Sense-2022

Spatial Transformer Networks

arXiv1 repo

arXiv:1506.02025

stn3d

Teaching Machines to Read and Comprehend

arXiv1 repo

arXiv:1506.03340

rc-data

Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

arXiv1 repo

arXiv:1506.06724

DiagnosisCoding

Skip-Thought Vectors

arXiv1 repo

arXiv:1506.06726

SentEval

Causal Decision Trees

arXiv1 repo

arXiv:1508.03812

Data-Science

Effective Approaches to Attention-based Neural Machine Translation

arXiv1 repo

arXiv:1508.04025

nmt

A large annotated corpus for learning natural language inference

arXiv1 repo

arXiv:1508.05326

roberta-large-mnli

A Neural Algorithm of Artistic Style

arXiv1 repo

arXiv:1508.06576

ocaml-torch

Continuous control with deep reinforcement learning

arXiv1 repo

arXiv:1509.02971

cleanrl

Spatially Encoding Temporal Correlations to Classify Temporal Data Using Convolutional Neural Networks

arXiv1 repo

arXiv:1509.07481

AI-assisted-chemical-sensing

Evasion and Hardening of Tree Ensemble Classifiers

arXiv1 repo

arXiv:1509.07892

RobustTrees

Fast Algorithms for Convolutional Neural Networks

arXiv1 repo

arXiv:1509.09308

mace

Semi-supervised Sequence Learning

arXiv1 repo

arXiv:1511.01432

bert

Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

arXiv1 repo

arXiv:1511.06434

ocaml-torch

Rethinking the Inception Architecture for Computer Vision

arXiv1 repo

arXiv:1512.00567

materialgan

ShapeNet: An Information-Rich 3D Model Repository

arXiv1 repo

arXiv:1512.03012

Cap3D

Programming in logic without logic programming

arXiv1 repo

arXiv:1601.00529

logicmoo_workspace

Convolutional Pose Machines

arXiv1 repo

arXiv:1602.00134

openpose

Are Elephants Bigger than Butterflies? Reasoning about Sizes of Objects

arXiv1 repo

arXiv:1602.00753

Kosmos-X

EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos

arXiv1 repo

arXiv:1602.03012

TF-Cholec80

Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

arXiv1 repo

arXiv:1602.07332

LRV-Instruction

XGBoost: A Scalable Tree Boosting System

arXiv1 repo

arXiv:1603.02754

xgboost

The sphere packing problem in dimension 8

arXiv1 repo

arXiv:1603.04246

glq

Incorporating Copying Mechanism in Sequence-to-Sequence Learning

arXiv1 repo

arXiv:1603.06393

nl2bash

Consistently faster and smaller compressed bitmaps with Roaring

arXiv1 repo

arXiv:1603.06549

RoaringFormatSpec

A Diagram Is Worth A Dozen Images

arXiv1 repo

arXiv:1603.07396

ai2d

S-hull: a fast radial sweep-hull routine for Delaunay triangulation

arXiv1 repo

arXiv:1604.01428

torch_delaunay

GLEU Without Tuning

arXiv1 repo

arXiv:1605.02592

gec-metrics

Neural Network Translation Models for Grammatical Error Correction

arXiv1 repo

arXiv:1606.00189

CTCResources

OpenAI Gym

arXiv1 repo

arXiv:1606.01540

gym

Natural Language Generation enhances human decision-making with uncertain information

arXiv1 repo

arXiv:1606.03254

awesome-nlg

Neural Generation of Regular Expressions from Natural Language with Minimal Domain Knowledge

arXiv1 repo

arXiv:1608.03000

llm-jepa

Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks

arXiv1 repo

arXiv:1608.04207

SentEval

SandBlaster: Reversing the Apple Sandbox

arXiv1 repo

arXiv:1608.04303

sandblaster

Using the Output Embedding to Improve Language Models

arXiv1 repo

arXiv:1608.05859

GPT-2

Polysemous codes

arXiv1 repo

arXiv:1609.01882

faiss

Predicting the future relevance of research institutions - The winning solution of the KDD Cup 2016

arXiv1 repo

arXiv:1609.02728

xgboost

Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network

arXiv1 repo

arXiv:1609.04802

MAX-Image-Resolution-Enhancer

Image-to-Markup Generation with Coarse-to-Fine Attention

arXiv1 repo

arXiv:1609.04938

LaTeX-OCR

Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

arXiv1 repo

arXiv:1609.08144

nmt

Non-Intrusive Load Monitoring: A Review and Outlook

arXiv1 repo

arXiv:1610.01191

awesome-nilm

Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models

arXiv1 repo

arXiv:1610.02424

fairseq

ORB-SLAM2: an Open-Source SLAM System for Monocular, Stereo and RGB-D Cameras

arXiv1 repo

arXiv:1610.06475

stella_vslam

Towards Automatic Resource Bound Analysis for OCaml

arXiv1 repo

arXiv:1611.00692

plutus

Cubical Type Theory: a constructive interpretation of the univalence axiom

arXiv1 repo

arXiv:1611.02108

cubical

Image-to-Image Translation with Conditional Adversarial Networks

arXiv1 repo

arXiv:1611.07004

pachyderm

Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

arXiv1 repo

arXiv:1611.08050

openpose

NewsQA: A Machine Comprehension Dataset

arXiv1 repo

arXiv:1611.09830

CoConflictQA

Self-critical Sequence Training for Image Captioning

arXiv1 repo

arXiv:1612.00563

vilmedic

Overcoming catastrophic forgetting in neural networks

arXiv1 repo

arXiv:1612.00796

Continual-NExT

FMA: A Dataset For Music Analysis

arXiv1 repo

arXiv:1612.01840

fma

FastText.zip: Compressing text classification models

arXiv1 repo

arXiv:1612.03651

fasttext-language-identification

Fast keyed hash/pseudo-random function using SIMD multiply and permute

arXiv1 repo

arXiv:1612.06257

highwayhash

OpenNMT: Open-Source Toolkit for Neural Machine Translation

arXiv1 repo

arXiv:1701.02810

nmt

Fast Exact k-Means, k-Medians and Bregman Divergence Clustering in 1D

arXiv1 repo

arXiv:1701.07204

ml-stable-diffusion

Emotion Recognition From Speech With Recurrent Neural Networks

arXiv1 repo

arXiv:1701.08071

awesome-affective-computing

arXiv:1701.08398

arXiv1 repo

arXiv:1701.08398

TIL-2023

New cardinality estimation algorithms for HyperLogLog sketches

arXiv1 repo

arXiv:1702.01284

hash4j

Software Engineering at Google

arXiv1 repo

arXiv:1702.01715

awesome-cto

Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning

arXiv1 repo

arXiv:1702.03118

tiny-vllm

Chaos Engineering

arXiv1 repo

arXiv:1702.05843

awesome-chaos-engineering

A Platform for Automating Chaos Experiments

arXiv1 repo

arXiv:1702.05849

awesome-chaos-engineering

ERA: A Framework for Economic Resource Allocation for the Cloud

arXiv1 repo

arXiv:1702.07311

papers-notebook

Billion-scale similarity search with GPUs

arXiv1 repo

arXiv:1702.08734

faiss

Neural Machine Translation and Sequence-to-sequence Models: A Tutorial

arXiv1 repo

arXiv:1703.01619

nmt

Massive Exploration of Neural Machine Translation Architectures

arXiv1 repo

arXiv:1703.03906

nmt

Automated Hate Speech Detection and the Problem of Offensive Language

arXiv1 repo

arXiv:1703.04009

toxic-bert

Understanding Black-box Predictions via Influence Functions

arXiv1 repo

arXiv:1703.04730

influence_boosting

Prototypical Networks for Few-shot Learning

arXiv1 repo

arXiv:1703.05175

dialog-mteb

Modeling Long- and Short-Term Temporal Patterns with Deep Neural Networks

arXiv1 repo

arXiv:1703.07015

Timeseries-PILE

Deep Photo Style Transfer

arXiv1 repo

arXiv:1703.07511

Github-Ranking

Transfer learning for music classification and regression tasks

arXiv1 repo

arXiv:1703.09179

fma

Towards Automatic Learning of Procedures from Web Instructional Videos

arXiv1 repo

arXiv:1703.09788

TimeChat-Online-139K

Survey of the State of the Art in Natural Language Generation: Core tasks, applications and evaluation

arXiv1 repo

arXiv:1703.09902

awesome-nlg

Reading Wikipedia to Answer Open-Domain Questions

arXiv1 repo

arXiv:1704.00051

odqa_baseline_code

A Survey of Distributed Message Broker Queues

arXiv1 repo

arXiv:1704.00411

awesome-scalability

Faster Base64 Encoding and Decoding Using AVX2 Instructions

arXiv1 repo

arXiv:1704.00605

simdutf

Get To The Point: Summarization with Pointer-Generator Networks

arXiv1 repo

arXiv:1704.04368

cnn-dailymail

RACE: Large-scale ReAding Comprehension Dataset From Examinations

arXiv1 repo

arXiv:1704.04683

lares

Learning to Reason: End-to-End Module Networks for Visual Question Answering

arXiv1 repo

arXiv:1704.05526

n2nmn

Accelerated Nearest Neighbor Search with Quick ADC

arXiv1 repo

arXiv:1704.07355

RaBitQ-Library

Automatic Anomaly Detection in the Cloud Via Statistical Learning

arXiv1 repo

arXiv:1704.07706

AnomalyDetection.rb

Hand Keypoint Detection in Single Images using Multiview Bootstrapping

arXiv1 repo

arXiv:1704.07809

openpose

A Novel Hybrid Quicksort Algorithm Vectorized using AVX-512 on Intel Skylake

arXiv1 repo

arXiv:1704.08579

x86-simd-sort

Dense-Captioning Events in Videos

arXiv1 repo

arXiv:1705.00754

ActivityNet_Captions

TALL: Temporal Activity Localization via Language Query

arXiv1 repo

arXiv:1705.02101

TALL

Supervised Learning of Universal Sentence Representations from Natural Language Inference Data

arXiv1 repo

arXiv:1705.02364

SentEval

Large-scale, Fast and Accurate Shot Boundary Detection through Spatio-temporal Convolutional Neural Networks

arXiv1 repo

arXiv:1705.03281

TransNetV2

Learning how to explain neural networks: PatternNet and PatternAttribution

arXiv1 repo

arXiv:1705.05598

dianna

Engineering Record And Replay For Deployability: Extended Technical Report

arXiv1 repo

arXiv:1705.05937

rr

Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics

arXiv1 repo

arXiv:1705.07115

Fine-Grained_Features_Alignment_via_Constrastive_Learning

Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

arXiv1 repo

arXiv:1705.07750

JavisBench

pix2code: Generating Code from a Graphical User Interface Screenshot

arXiv1 repo

arXiv:1705.07962

Screenshot-to-code

arXiv:1705.08926

arXiv1 repo

arXiv:1705.08926

smac

Marmara Turkish Coreference Corpus and Coreference Resolution Baseline

arXiv1 repo

arXiv:1706.01863

g4t0r2-nlp

Deep reinforcement learning from human preferences

arXiv1 repo

arXiv:1706.03741

prompt-engineering

On Calibration of Modern Neural Networks

arXiv1 repo

arXiv:1706.04599

ogham-mcp

SuperMinHash - A New Minwise Hashing Algorithm for Jaccard Similarity Estimation

arXiv1 repo

arXiv:1706.05698

hash4j

Developing Bug-Free Machine Learning Systems With Formal Mathematics

arXiv1 repo

arXiv:1706.08605

certigrad

Causal Structure Learning

arXiv1 repo

arXiv:1706.09141

Data-Science

CatBoost: unbiased boosting with categorical features

arXiv1 repo

arXiv:1706.09516

catboost

SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation

arXiv1 repo

arXiv:1708.00055

t5-large-encoder-only-bf16

Localizing Moments in Video with Natural Language

arXiv1 repo

arXiv:1708.01641

LocalizingMoments

StarCraft II: A New Challenge for Reinforcement Learning

arXiv1 repo

arXiv:1708.04782

pysc2

Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

arXiv1 repo

arXiv:1709.00103

MirageTVQA

ProSLAM: Graph SLAM from a Programmer's Perspective

arXiv1 repo

arXiv:1709.04377

stella_vslam

The First Evaluation of Chinese Human-Computer Dialogue Technology

arXiv1 repo

arXiv:1709.10217

ChineseNLPCorpus

DisSent: Sentence Representation Learning from Explicit Discourse Relations

arXiv1 repo

arXiv:1710.04334

SentEval

Representation Learning of Music Using Artist Labels

arXiv1 repo

arXiv:1710.06648

fma

FigureQA: An Annotated Figure Dataset for Visual Reasoning

arXiv1 repo

arXiv:1710.07300

PlotQA

Souper: A Synthesizing Superoptimizer

arXiv1 repo

arXiv:1711.04422

souper

Emotional End-to-End Neural Speech Synthesizer

arXiv1 repo

arXiv:1711.05447

FastSpeech2-Plus

Total Haskell is Reasonable Coq

arXiv1 repo

arXiv:1711.09286

hs-to-coq

Are GANs Created Equal? A Large-Scale Study

arXiv1 repo

arXiv:1711.10337

pytorch-fid

Occam's razor is insufficient to infer the preferences of irrational agents

arXiv1 repo

arXiv:1712.05812

minihf

SuperPoint: Self-Supervised Interest Point Detection and Description

arXiv1 repo

arXiv:1712.07629

gtsfm

Demystifying MMD GANs

arXiv1 repo

arXiv:1801.01401

stylegan2-ada-pytorch

MobileNetV2: Inverted Residuals and Linear Bottlenecks

arXiv1 repo

arXiv:1801.04381

mobilenet_v2_1.4_224

Universal Language Model Fine-tuning for Text Classification

arXiv1 repo

arXiv:1801.06146

indonesian-language-models

Generating Wikipedia by Summarizing Long Sequences

arXiv1 repo

arXiv:1801.10198

transformer-tricks

DensePose: Dense Human Pose Estimation In The Wild

arXiv1 repo

arXiv:1802.00434

DensePose

On Higher Inductive Types in Cubical Type Theory

arXiv1 repo

arXiv:1802.01170

cubical

Classification and Disease Localization in Histopathology Using Only Global Labels: A Weakly-Supervised Approach

arXiv1 repo

arXiv:1802.02212

HistoSSLscaling

Revisiting the Inverted Indices for Billion-Scale Approximate Nearest Neighbors

arXiv1 repo

arXiv:1802.02422

hnswlib

Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation

arXiv1 repo

arXiv:1802.02611

spectra

Partisan: Enabling Cloud-Scale Erlang Applications

arXiv1 repo

arXiv:1802.02652

partisan

Attention-based Deep Multiple Instance Learning

arXiv1 repo

arXiv:1802.04712

HistoSSLscaling

Deep contextualized word representations

arXiv1 repo

arXiv:1802.05365

Word-Embeddings-Repository-for-Turkish

Finding Influential Training Samples for Gradient Boosted Decision Trees

arXiv1 repo

arXiv:1802.06640

influence_boosting

Learning Word Vectors for 157 Languages

arXiv1 repo

arXiv:1802.06893

fasttext-language-identification

Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration

arXiv1 repo

arXiv:1802.08802

android_world_seeact_v

Back to Basics: Benchmarking Canonical Evolution Strategies for Playing Atari

arXiv1 repo

arXiv:1802.08842

Canonical_ES_Atari

Addressing Function Approximation Error in Actor-Critic Methods

arXiv1 repo

arXiv:1802.09477

cleanrl

Chest X-Ray Analysis of Tuberculosis by Deep Learning with Segmentation and Augmentation

arXiv1 repo

arXiv:1803.01199

Project-Imaging-X

GONet: A Semi-Supervised Deep Learning Approach For Traversability Estimation

arXiv1 repo

arXiv:1803.03254

UniWM_Dataset

Deep-FSMN for Large Vocabulary Continuous Speech Recognition

arXiv1 repo

arXiv:1803.05030

fsmn-vad-onnx

OSINT Analysis of the TOR Foundation

arXiv1 repo

arXiv:1803.05201

non-typical-OSINT-guide

Learning to Recognize Musical Genre from Audio

arXiv1 repo

arXiv:1803.05337

fma

SentEval: An Evaluation Toolkit for Universal Sentence Representations

arXiv1 repo

arXiv:1803.05449

SentEval

Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

arXiv1 repo

arXiv:1803.05457

ai2_arc

Complex-YOLO: Real-time 3D Object Detection on Point Clouds

arXiv1 repo

arXiv:1803.06199

Complex-YOLOv4-Pytorch

Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

arXiv1 repo

arXiv:1803.09017

MockingBird

Universal Sentence Encoder

arXiv1 repo

arXiv:1803.11175

SentEval

arXiv:1803.11485

arXiv1 repo

arXiv:1803.11485

smac

ESPnet: End-to-End Speech Processing Toolkit

arXiv1 repo

arXiv:1804.00015

kan-bayashi_ljspeech_vits

Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning

arXiv1 repo

arXiv:1804.00079

SentEval

SBFT: a Scalable and Decentralized Trust Infrastructure

arXiv1 repo

arXiv:1804.01626

concord-bft

Flexible and Scalable Deep Learning with MMLSpark

arXiv1 repo

arXiv:1804.04031

SynapseML

arXiv:1804.05839

arXiv1 repo

arXiv:1804.05839

llm_test

Phrase-Based & Neural Unsupervised Machine Translation

arXiv1 repo

arXiv:1804.07755

XLM

An Aggregated Multicolumn Dilated Convolution Network for Perspective-Free Counting

arXiv1 repo

arXiv:1804.07821

marker

Gender Bias in Coreference Resolution

arXiv1 repo

arXiv:1804.09301

model-written-evals

Link and code: Fast indexing with graphs and compact regression codes

arXiv1 repo

arXiv:1804.09996

faiss

Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates

arXiv1 repo

arXiv:1804.10959

sentencepiece

What you can cram into a single vector: Probing sentence embeddings for linguistic properties

arXiv1 repo

arXiv:1805.01070

SentEval

Online normalizer calculation for softmax

arXiv1 repo

arXiv:1805.02867

GPT-2

TED-LIUM 3: twice as much data and corpus repartition for experiments on speaker adaptation

arXiv1 repo

arXiv:1805.04699

asr_consilium

A Chaos Engineering System for Live Analysis and Falsification of Exception-handling in the JVM

arXiv1 repo

arXiv:1805.05246

awesome-chaos-engineering

arXiv:1805.08318

arXiv1 repo

arXiv:1805.08318

SkinDeep

COCO-CN for Cross-Lingual Image Tagging, Captioning and Retrieval

arXiv1 repo

arXiv:1805.08661

coco-cn

AutoAugment: Learning Augmentation Policies from Data

arXiv1 repo

arXiv:1805.09501

Chinese-CLIP

Neural Network Acceptability Judgments

arXiv1 repo

arXiv:1805.12471

t5-large-encoder-only-bf16

Digging Into Self-Supervised Monocular Depth Estimation

arXiv1 repo

arXiv:1806.01260

Depth-Estimation

Generative Adversarial Networks for Realistic Synthesis of Hyperspectral Samples

arXiv1 repo

arXiv:1806.02583

Data-Science

Cell Detection with Star-convex Polygons

arXiv1 repo

arXiv:1806.03535

stardist

Know What You Don't Know: Unanswerable Questions for SQuAD

arXiv1 repo

arXiv:1806.03822

squad_v2

The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

arXiv1 repo

arXiv:1806.09514

EmoV-DB

The relativistic discriminator: a key element missing from standard GAN

arXiv1 repo

arXiv:1807.00734

ocaml-torch

CAIL2018: A Large-Scale Legal Dataset for Judgment Prediction

arXiv1 repo

arXiv:1807.02478

CAIL

The Double Sphere Camera Model

arXiv1 repo

arXiv:1807.08957

kalibr

HiDDeN: Hiding Data With Deep Networks

arXiv1 repo

arXiv:1807.09937

WMCopier

Unified Perceptual Parsing for Scene Understanding

arXiv1 repo

arXiv:1807.10221

FoodSeg103-Benchmark-v1

Speaker Recognition from Raw Waveform with SincNet

arXiv1 repo

arXiv:1808.00158

NIPS4Bplus

One Billion Apples' Secret Sauce: Recipe for the Apple Wireless Direct Link Ad hoc Protocol

arXiv1 repo

arXiv:1808.03156

stop-stutter

Fast Video Shot Transition Localization with Deep Structured Models

arXiv1 repo

arXiv:1808.04234

TransNetV2

WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations

arXiv1 repo

arXiv:1808.09121

t5-large-encoder-only-bf16

Evaluating Theory of Mind in Question Answering

arXiv1 repo

arXiv:1808.09352

ToMi

Story Ending Generation with Incremental Encoding and Commonsense Knowledge

arXiv1 repo

arXiv:1808.10113

AwesomeSEG

Toward a Standardized and More Accurate Indonesian Part-of-Speech Tagging

arXiv1 repo

arXiv:1809.03391

indonlu

Addressing the Fundamental Tension of PCGML with Discriminative Learning

arXiv1 repo

arXiv:1809.04432

WaveFunctionCollapse

BRAVO -- Biased Locking for Reader-Writer Locks

arXiv1 repo

arXiv:1810.01553

VictoriaMetrics

The UCR Time Series Archive

arXiv1 repo

arXiv:1810.07758

UTSD

Don't Unroll Adjoint: Differentiating SSA-Form Programs

arXiv1 repo

arXiv:1810.07951

Zygote.jl

From Louvain to Leiden: guaranteeing well-connected communities

arXiv1 repo

arXiv:1810.08473

sweet-search

MMLSpark: Unifying Machine Learning Ecosystems at Massive Scales

arXiv1 repo

arXiv:1810.08744

SynapseML

pair2vec: Compositional Word-Pair Embeddings for Cross-Sentence Inference

arXiv1 repo

arXiv:1810.08854

relbert

3D MRI brain tumor segmentation using autoencoder regularization

arXiv1 repo

arXiv:1810.11654

TriALS

Audio inpainting of music by means of neural networks

arXiv1 repo

arXiv:1810.12138

fma

ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension

arXiv1 repo

arXiv:1810.12885

t5-large-encoder-only-bf16

Exploration by Random Network Distillation

arXiv1 repo

arXiv:1810.12894

cleanrl

WaveGlow: A Flow-based Generative Network for Speech Synthesis

arXiv1 repo

arXiv:1811.00002

waveglow

Image Chat: Engaging Grounded Conversations

arXiv1 repo

arXiv:1811.00945

sirius-spring2021-image2chat

Extended Isolation Forest

arXiv1 repo

arXiv:1811.02141

isolation-forest

Subtask Gated Networks for Non-Intrusive Load Monitoring

arXiv1 repo

arXiv:1811.06692

awesome-nilm

GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism

arXiv1 repo

arXiv:1811.06965

awsome-llm-papers

A New Cervical Cytology Dataset for Nucleus Detection and Image Classification (Cervix93) and Methods for Cervical Nucleus Detection

arXiv1 repo

arXiv:1811.09651

cytology_dataset

Visual Entailment Task for Visually-Grounded Language Learning

arXiv1 repo

arXiv:1811.10582

SNLI-VE

CCNet: Criss-Cross Attention for Semantic Segmentation

arXiv1 repo

arXiv:1811.11721

FoodSeg103-Benchmark-v1

Learning from a tiny dataset of manual annotations: a teacher/student approach for surgical phase recognition

arXiv1 repo

arXiv:1812.00033

TF-Cholec80

Scalable Graph Learning for Anti-Money Laundering: A First Look

arXiv1 repo

arXiv:1812.00076

AMLSim

Transferring Knowledge across Learning Processes

arXiv1 repo

arXiv:1812.01054

xfer

A micro Lie theory for state estimation in robotics

arXiv1 repo

arXiv:1812.01537

optik

Towards Accurate Generative Models of Video: A New Metric & Challenges

arXiv1 repo

arXiv:1812.01717

VideoGPT

Soft Actor-Critic Algorithms and Applications

arXiv1 repo

arXiv:1812.05905

cleanrl

Inverse Cooking: Recipe Generation from Food Images

arXiv1 repo

arXiv:1812.06164

inversecooking

The Adverse Effects of Code Duplication in Machine Learning Models of Code

arXiv1 repo

arXiv:1812.06469

Project_CodeNet

OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields

arXiv1 repo

arXiv:1812.08008

openpose

Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

arXiv1 repo

arXiv:1812.08466

fadtk

Pansori: ASR Corpus Generation from Open Online Video Contents

arXiv1 repo

arXiv:1812.09798

kabooks

A Poisson-Gaussian Denoising Dataset with Real Fluorescence Microscopy Images

arXiv1 repo

arXiv:1812.10366

denoising-fluorescence

TripleAgent: Monitoring, Perturbation and Failure-obliviousness for Automated Resilience Improvement in Java Applications

arXiv1 repo

arXiv:1812.10706

awesome-chaos-engineering

A Comprehensive Survey on Graph Neural Networks

arXiv1 repo

arXiv:1901.00596

Data-Science

Learning From Less Data: A Unified Data Subset Selection and Active Learning Framework for Computer Vision

arXiv1 repo

arXiv:1901.01151

cords

Panoptic Feature Pyramid Networks

arXiv1 repo

arXiv:1901.02446

FoodSeg103-Benchmark-v1

Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

arXiv1 repo

arXiv:1901.02860

transfo-xl-wt103

From Plots to Endings: A Reinforced Pointer Generator for Story Ending Generation

arXiv1 repo

arXiv:1901.03459

AwesomeSEG

Passage Re-ranking with BERT

arXiv1 repo

arXiv:1901.04085

pygaggle

TensorFlow.js: Machine Learning for the Web and Beyond

arXiv1 repo

arXiv:1901.05350

awesome-tensorflow-js

Visual Entailment: A Novel Task for Fine-Grained Image Understanding

arXiv1 repo

arXiv:1901.06706

SNLI-VE

MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs

arXiv1 repo

arXiv:1901.07042

baple

Cross-lingual Language Model Pretraining

arXiv1 repo

arXiv:1901.07291

XLM

Personalized Dialogue Generation with Diversified Traits

arXiv1 repo

arXiv:1901.09672

BoB

Information Operations Recognition: from Nonlinear Analysis to Decision-making

arXiv1 repo

arXiv:1901.10876

non-typical-OSINT-guide

UcoSLAM: Simultaneous Localization and Mapping by Fusion of KeyPoints and Squared Planar Markers

arXiv1 repo

arXiv:1902.03729

stella_vslam

BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model

arXiv1 repo

arXiv:1902.04094

temporal-robustness

Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology

arXiv1 repo

arXiv:1902.06543

Midnight

Parsing Gigabytes of JSON per Second

arXiv1 repo

arXiv:1902.08318

simdjson

Wavenilm: A causal neural network for power disaggregation from the complex power signal

arXiv1 repo

arXiv:1902.08736

awesome-nilm

A large annotated medical image dataset for the development and evaluation of segmentation algorithms

arXiv1 repo

arXiv:1902.09063

TotalSegmentator

Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression

arXiv1 repo

arXiv:1902.09630

Complex-YOLOv4-Pytorch

EvolveGCN: Evolving Graph Convolutional Networks for Dynamic Graphs

arXiv1 repo

arXiv:1902.10191

AMLSim

Accelerating Self-Play Learning in Go

arXiv1 repo

arXiv:1902.10565

KataGo

Robust Decision Trees Against Adversarial Examples

arXiv1 repo

arXiv:1902.10660

RobustTrees

COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis

arXiv1 repo

arXiv:1903.02874

TimeChat-Online-139K

Fine-tune BERT for Extractive Summarization

arXiv1 repo

arXiv:1903.10318

BertSum

arXiv:1903.11269

arXiv1 repo

arXiv:1903.11269

css10

Learning Discrete Structures for Graph Neural Networks

arXiv1 repo

arXiv:1903.11960

fma

Habitat: A Platform for Embodied AI Research

arXiv1 repo

arXiv:1904.01201

habitat-lab

Character Region Awareness for Text Detection

arXiv1 repo

arXiv:1904.01941

EasyOCR

arXiv:1904.02285

arXiv1 repo

arXiv:1904.02285

Jellyfish-13B

Speech Model Pre-training for End-to-End Spoken Language Understanding

arXiv1 repo

arXiv:1904.03670

openWakeWord

StegaStamp: Invisible Hyperlinks in Physical Photographs

arXiv1 repo

arXiv:1904.05343

WMCopier

The Android Platform Security Model (2023)

arXiv1 repo

arXiv:1904.05572

apparmor.d

wav2vec: Unsupervised Pre-training for Speech Recognition

arXiv1 repo

arXiv:1904.05862

dissertation-project

DocBERT: BERT for Document Classification

arXiv1 repo

arXiv:1904.08398

NLP-DocBERT-financial-news-trading

SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition

arXiv1 repo

arXiv:1904.08779

s2t-small-librispeech-asr

Cantor-Bernstein implies Excluded Middle

arXiv1 repo

arXiv:1904.09193

topology

ERNIE: Enhanced Representation through Knowledge Integration

arXiv1 repo

arXiv:1904.09223

ernie-1.0-base-zh

SocialIQA: Commonsense Reasoning about Social Interactions

arXiv1 repo

arXiv:1904.09728

DeepEnlighten

arXiv:1904.09751

arXiv1 repo

arXiv:1904.09751

llama2.zig

Generating Long Sequences with Sparse Transformers

arXiv1 repo

arXiv:1904.10509

VideoGPT

Genet: A Quickly Scalable Fat-Tree Overlay for Personal Volunteer Computing using WebRTC

arXiv1 repo

arXiv:1904.11402

simple-peer

Local Relation Networks for Image Recognition

arXiv1 repo

arXiv:1904.11491

Swin-Transformer

Style Transfer by Relaxed Optimal Transport and Self-Similarity

arXiv1 repo

arXiv:1904.12785

wise

Drug-Drug Adverse Effect Prediction with Graph Co-Attention

arXiv1 repo

arXiv:1905.00534

chemicalx

RetinaFace: Single-stage Dense Face Localisation in the Wild

arXiv1 repo

arXiv:1905.00641

face-alignment

Searching for MobileNetV3

arXiv1 repo

arXiv:1905.02244

yas

Taming Pretrained Transformers for Extreme Multi-label Text Classification

arXiv1 repo

arXiv:1905.02331

pecos

Does Environmental Economics lead to patentable research?

arXiv1 repo

arXiv:1905.02875

Kosmos-X

Harvey: A Greybox Fuzzer for Smart Contracts

arXiv1 repo

arXiv:1905.06944

echidna

MR-GNN: Multi-Resolution and Dual Graph Neural Network for Predicting Structured Entity Interactions

arXiv1 repo

arXiv:1905.09558

chemicalx

BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

arXiv1 repo

arXiv:1905.10044

t5-large-encoder-only-bf16

CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition

arXiv1 repo

arXiv:1905.11235

ASR-Knowledge-Transferring

Mixed Precision DNNs: All you need is a good parametrization

arXiv1 repo

arXiv:1905.11452

ai-research-code

EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

arXiv1 repo

arXiv:1905.11946

yolov5

Racial Bias in Hate Speech and Abusive Language Detection Datasets

arXiv1 repo

arXiv:1905.12516

toxic-bert

Latent Retrieval for Weakly Supervised Open Domain Question Answering

arXiv1 repo

arXiv:1906.00300

llm-jepa

Coresets for Data-efficient Training of Machine Learning Models

arXiv1 repo

arXiv:1906.01827

cords

Gradio: Hassle-Free Sharing and Testing of ML Models in the Wild

arXiv1 repo

arXiv:1906.02569

gradio

Visualizing and Measuring the Geometry of BERT

arXiv1 repo

arXiv:1906.02715

adv-glue-plus-plus

Mesh R-CNN

arXiv1 repo

arXiv:1906.02739

pytorch3d

Multi-hop Reading Comprehension through Question Decomposition and Rescoring

arXiv1 repo

arXiv:1906.02916

query_decomposer

TransNet: A deep network for fast detection of common shot transitions

arXiv1 repo

arXiv:1906.03363

TransNetV2

Robustness Verification of Tree-based Models

arXiv1 repo

arXiv:1906.03849

RobustTrees

GluonTS: Probabilistic Time Series Models in Python

arXiv1 repo

arXiv:1906.05264

gluonts

Image Captioning: Transforming Objects into Words

arXiv1 repo

arXiv:1906.05963

object-relation-transformer

Scheduled Sampling for Transformers

arXiv1 repo

arXiv:1906.07651

final-project-level3-nlp-02

The Second DIHARD Diarization Challenge: Dataset, task, and baselines

arXiv1 repo

arXiv:1906.07839

speaker-diarization-benchmark

Multi-Span Acoustic Modelling using Raw Waveform Signals

arXiv1 repo

arXiv:1906.11047

ami

The Indirect Convolution Algorithm

arXiv1 repo

arXiv:1907.02129

XNNPACK

Zero-shot Learning for Audio-based Music Classification and Tagging

arXiv1 repo

arXiv:1907.02670

fma

XGBoostLSS -- An extension of XGBoost to probabilistic forecasting

arXiv1 repo

arXiv:1907.03178

LightGBMLSS

BlazeFace: Sub-millisecond Neural Face Detection on Mobile GPUs

arXiv1 repo

arXiv:1907.05047

face-alignment

Large Memory Layers with Product Keys

arXiv1 repo

arXiv:1907.05242

XLM

Hello, It's GPT-2 -- How Can I Help You? Towards the Use of Pretrained Language Models for Task-Oriented Dialogue Systems

arXiv1 repo

arXiv:1907.05774

KoGPT2-chatbot

Facebook FAIR's WMT19 News Translation Task Submission

arXiv1 repo

arXiv:1907.06616

wmt19-en-ru

Efficient Pipeline for Camera Trap Image Review

arXiv1 repo

arXiv:1907.06772

Depth-Estimation

SpanBERT: Improving Pre-training by Representing and Predicting Spans

arXiv1 repo

arXiv:1907.10529

odqa_baseline_code

WinoGrande: An Adversarial Winograd Schema Challenge at Scale

arXiv1 repo

arXiv:1907.10641

Kosmos-X

ConCert: A Smart Contract Certification Framework in Coq

arXiv1 repo

arXiv:1907.10674

ConCert

MaskGAN: Towards Diverse and Interactive Facial Image Manipulation

arXiv1 repo

arXiv:1907.11922

CelebAMask-HQ

Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

arXiv1 repo

arXiv:1907.11932

DeBERTa_TxtClassifier

Observability and Chaos Engineering on System Calls for Containerized Applications in Docker

arXiv1 repo

arXiv:1907.13039

awesome-chaos-engineering

arXiv:1907.13440

arXiv1 repo

arXiv:1907.13440

minerl

VisualBERT: A Simple and Performant Baseline for Vision and Language

arXiv1 repo

arXiv:1908.03557

visual-spatial-reasoning

Neural Text Generation with Unlikelihood Training

arXiv1 repo

arXiv:1908.04319

calibrating-summaries

Aspect and Opinion Terms Extraction Using Double Embeddings and Attention Mechanism for Indonesian Hotel Reviews

arXiv1 repo

arXiv:1908.04899

indonlu

LXMERT: Learning Cross-Modality Encoder Representations from Transformers

arXiv1 repo

arXiv:1908.07490

visual-spatial-reasoning

An Empirical Study of Incorporating Pseudo Data into Grammatical Error Correction

arXiv1 repo

arXiv:1909.00502

CTCResources

Robust Invisible Video Watermarking with Attention

arXiv1 repo

arXiv:1909.01285

WMCopier

Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop

arXiv1 repo

arXiv:1909.01362

ConvoKit

An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction

arXiv1 repo

arXiv:1909.02027

task-aware-embedding-refinement

Neural Machine Translation with Byte-Level Subwords

arXiv1 repo

arXiv:1909.03341

vocab-coverage

KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension

arXiv1 repo

arXiv:1909.07005

squad_kor_v1

Ludwig: a type-based declarative deep learning toolbox

arXiv1 repo

arXiv:1909.07930

ludwig

Aspect and Opinion Term Extraction for Hotel Reviews using Transfer Learning and Auxiliary Labels

arXiv1 repo

arXiv:1909.11879

indonlu

A Pilot Study for Chinese SQL Semantic Parsing

arXiv1 repo

arXiv:1909.13293

ChineseNLPCorpus

Interpretations are useful: penalizing explanations to align neural networks with prior knowledge

arXiv1 repo

arXiv:1909.13584

imodelsX

Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

arXiv1 repo

arXiv:1910.00177

trax

OpenVSLAM: A Versatile Visual SLAM Framework

arXiv1 repo

arXiv:1910.01122

stella_vslam

MLPerf Training Benchmark

arXiv1 repo

arXiv:1910.01500

training

NGBoost: Natural Gradient Boosting for Probabilistic Prediction

arXiv1 repo

arXiv:1910.03225

ngboost

Base64 encoding and decoding at almost the speed of a memory copy

arXiv1 repo

arXiv:1910.05109

simdutf

vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

arXiv1 repo

arXiv:1910.05453

dissertation-project

JSDoop and TensorFlow.js: Volunteer Distributed Web Browser-Based Neural Network Training

arXiv1 repo

arXiv:1910.07402

awesome-tensorflow-js

MLQA: Evaluating Cross-lingual Extractive Question Answering

arXiv1 repo

arXiv:1910.07475

Multilingual-MiniLM-L12-H384

Can I teach a robot to replicate a line art

arXiv1 repo

arXiv:1910.07860

awesome-plotters

Using Speech Synthesis to Train End-to-End Spoken Language Understanding Models

arXiv1 repo

arXiv:1910.09463

openWakeWord

Hierarchical Transformers for Long Document Classification

arXiv1 repo

arXiv:1910.10781

NLP-DocBERT-financial-news-trading

A Unified MRC Framework for Named Entity Recognition

arXiv1 repo

arXiv:1910.11476

nlp-bazel-tutorial

Confident Learning: Estimating Uncertainty in Dataset Labels

arXiv1 repo

arXiv:1911.00068

cleanlab

CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data

arXiv1 repo

arXiv:1911.00359

falcon-refinedweb

On the Measure of Intelligence

arXiv1 repo

arXiv:1911.01547

logicmoo_workspace

arXiv:1911.01601

arXiv1 repo

arXiv:1911.01601

MTP-Codebase

Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports

arXiv1 repo

arXiv:1911.02541

vilmedic

MLPerf Inference Benchmark

arXiv1 repo

arXiv:1911.02549

inference

Scalable Zero-shot Entity Linking with Dense Entity Retrieval

arXiv1 repo

arXiv:1911.03814

trusted_ke

Knowledge Guided Text Retrieval and Reading for Open Domain Question Answering

arXiv1 repo

arXiv:1911.03868

DPR

Smart Contract Interactions in Coq

arXiv1 repo

arXiv:1911.04732

ConCert

Identification of Rhetorical Roles of Sentences in Indian Legal Judgments

arXiv1 repo

arXiv:1911.05405

InLegalBERT

CASTER: Predicting Drug Interactions with Chemical Substructure Representation

arXiv1 repo

arXiv:1911.06446

chemicalx

Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model

arXiv1 repo

arXiv:1911.08265

xwm

Semantic Hierarchy Emerges in Deep Generative Representations for Scene Synthesis

arXiv1 repo

arXiv:1911.09267

materialgan

Fast Sparse ConvNets

arXiv1 repo

arXiv:1911.09723

XNNPACK

Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack Integration

arXiv1 repo

arXiv:1911.09925

chipyard

Causality for Machine Learning

arXiv1 repo

arXiv:1911.10500

Data-Science

Image-based table recognition: data, model, and evaluation

arXiv1 repo

arXiv:1911.10683

opendataloader-bench

PIQA: Reasoning about Physical Commonsense in Natural Language

arXiv1 repo

arXiv:1911.11641

Kosmos-X

SuperGlue: Learning Feature Matching with Graph Neural Networks

arXiv1 repo

arXiv:1911.11763

gtsfm

CSPNet: A New Backbone that can Enhance Learning Capability of CNN

arXiv1 repo

arXiv:1911.11929

darknet

Pythia: AI-assisted Code Completion System

arXiv1 repo

arXiv:1912.00742

CodeXGLUE

Analyzing and Improving the Image Quality of StyleGAN

arXiv1 repo

arXiv:1912.04958

materialgan

Common Voice: A Massively-Multilingual Speech Corpus

arXiv1 repo

arXiv:1912.06670

covost

SynSin: End-to-end View Synthesis from a Single Image

arXiv1 repo

arXiv:1912.08804

pytorch3d

Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting

arXiv1 repo

arXiv:1912.09363

frn-50k-baseline

PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition

arXiv1 repo

arXiv:1912.10211

audio-embeddings

Big Transfer (BiT): General Visual Representation Learning

arXiv1 repo

arXiv:1912.11370

bit-50

Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro

arXiv1 repo

arXiv:1912.11554

numpyro

A Gentle Introduction to Deep Learning for Graphs

arXiv1 repo

arXiv:1912.12693

Data-Science

LayoutLM: Pre-training of Text and Layout for Document Image Understanding

arXiv1 repo

arXiv:1912.13318

layoutlm-base-cased

The LDBC Social Network Benchmark

arXiv1 repo

arXiv:2001.02299

duckdb

Efficient Memory Management for Deep Neural Net Inference

arXiv1 repo

arXiv:2001.03288

XNNPACK

The Two-Pass Softmax Algorithm

arXiv1 repo

arXiv:2001.04438

XNNPACK

Reformer: The Efficient Transformer

arXiv1 repo

arXiv:2001.04451

trax

SQLFlow: A Bridge between SQL and Machine Learning

arXiv1 repo

arXiv:2001.06846

sqlflow

Fast Sequence-Based Embedding with Diffusion Graphs

arXiv1 repo

arXiv:2001.07463

littleballoffur

Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference

arXiv1 repo

arXiv:2001.07676

pet

Scaling Laws for Neural Language Models

arXiv1 repo

arXiv:2001.08361

parameter-golf

Towards Measuring Supply Chain Attacks on Package Managers for Interpreted Languages

arXiv1 repo

arXiv:2002.01139

guarddog

CoVoST: A Diverse Multilingual Speech-To-Text Translation Corpus

arXiv1 repo

arXiv:2002.01320

covost

Training Keyword Spotters with Limited and Synthesized Speech Data

arXiv1 repo

arXiv:2002.01322

openWakeWord

Using Fractal Neural Networks to Play SimCity 1 and Conway's Game of Life at Variable Scales

arXiv1 repo

arXiv:2002.03896

MicropolisCore

A Simple General Approach to Balance Task Difficulty in Multi-Task Learning

arXiv1 repo

arXiv:2002.04792

Fine-Grained_Features_Alignment_via_Constrastive_Learning

Causality in cognitive neuroscience: concepts, challenges, and distributional robustness

arXiv1 repo

arXiv:2002.06060

Data-Science

CodeBERT: A Pre-Trained Model for Programming and Natural Languages

arXiv1 repo

arXiv:2002.08155

CodeXGLUE

Scalable Second Order Optimization for Deep Learning

arXiv1 repo

arXiv:2002.09018

modded-nanogpt

Unsupervised Question Decomposition for Question Answering

arXiv1 repo

arXiv:2002.09758

query_decomposer

Resources for Turkish Dependency Parsing: Introducing the BOUN Treebank and the BoAT Annotation Tool

arXiv1 repo

arXiv:2002.10416

Word-Embeddings-Repository-for-Turkish

CausalML: Python Package for Causal Machine Learning

arXiv1 repo

arXiv:2002.11631

causalml

Gradient Boosted Normalizing Flows

arXiv1 repo

arXiv:2002.11896

gradient-boosted-normalizing-flows

OpEn: Code Generation for Embedded Nonconvex Optimization

arXiv1 repo

arXiv:2003.00292

optimization-engine

CheXclusion: Fairness gaps in deep chest X-ray classifiers

arXiv1 repo

arXiv:2003.00827

Fine-Grained_Features_Alignment_via_Constrastive_Learning

Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap

arXiv1 repo

arXiv:2003.02405

wordcab-transcribe

Combining GHOST and Casper

arXiv1 repo

arXiv:2003.03052

consensus-specs

Document Ranking with a Pretrained Sequence-to-Sequence Model

arXiv1 repo

arXiv:2003.06713

pygaggle

Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

arXiv1 repo

arXiv:2003.07082

stanza

Dash: Scalable Hashing on Persistent Memory

arXiv1 repo

arXiv:2003.07302

dragonfly

NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis

arXiv1 repo

arXiv:2003.08934

nerf

A Trustful Monad for Axiomatic Reasoning with Probability and Nondeterminism

arXiv1 repo

arXiv:2003.09993

monae

Two-stage Discriminative Re-ranking for Large-scale Landmark Retrieval

arXiv1 repo

arXiv:2003.11211

google-landmark

Similarity of Neural Networks with Gradients

arXiv1 repo

arXiv:2003.11498

xfer

COVID-19 Image Data Collection

arXiv1 repo

arXiv:2003.11597

covid-chestxray-dataset

RAFT: Recurrent All-Pairs Field Transforms for Optical Flow

arXiv1 repo

arXiv:2003.12039

prisma

SaccadeNet: A Fast and Accurate Object Detector

arXiv1 repo

arXiv:2003.12125

foveate

Computer Aided Detection for Pulmonary Embolism Challenge (CAD-PE)

arXiv1 repo

arXiv:2003.13440

Project-Imaging-X

Information Leakage in Embedding Models

arXiv1 repo

arXiv:2004.00053

langtest

Merkle-CRDTs: Merkle-DAGs meet CRDTs

arXiv1 repo

arXiv:2004.00107

defradb

CurricularFace: Adaptive Curriculum Learning Loss for Deep Face Recognition

arXiv1 repo

arXiv:2004.00288

FLUXSynID

arXiv:2004.00584

arXiv1 repo

arXiv:2004.00584

Jellyfish-13B

RisGraph: A Real-Time Streaming System for Evolving Graphs to Support Sub-millisecond Per-update Analysis at Millions Ops/s

arXiv1 repo

arXiv:2004.00803

awesome-dynamic-graphs

Google Landmarks Dataset v2 -- A Large-Scale Benchmark for Instance-Level Recognition and Retrieval

arXiv1 repo

arXiv:2004.01804

google-landmark

TAPAS: Weakly Supervised Table Parsing via Pre-training

arXiv1 repo

arXiv:2004.02349

tapas-base-finetuned-wtq

Deep Learning Based Text Classification: A Comprehensive Review

arXiv1 repo

arXiv:2004.03705

OpenTextClassification

On the Effect of Dropping Layers of Pre-trained Transformer Models

arXiv1 repo

arXiv:2004.03844

final-project-level3-nlp-02

Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence

arXiv1 repo

arXiv:2004.03974

turftopic

Unveiling COVID-19 from Chest X-ray with deep learning: a hurdles race with small data

arXiv1 repo

arXiv:2004.05405

covid-chestxray-dataset

Minimizing FLOPs to Learn Efficient Sparse Representations

arXiv1 repo

arXiv:2004.05665

splade-ecommerce-esci

Benchmarking Unsupervised Outlier Detection with Realistic Synthetic Data

arXiv1 repo

arXiv:2004.06947

Data-Science

Image Quality Assessment: Unifying Structure and Texture Similarity

arXiv1 repo

arXiv:2004.07728

DISTS

Classification Benchmarks for Under-resourced Bengali Language based on Multichannel Convolutional-LSTM Network

arXiv1 repo

arXiv:2004.07807

bangla-electra

SongNet: Rigid Formats Controlled Text Generation

arXiv1 repo

arXiv:2004.08022

textgen

arXiv:2004.08483

arXiv1 repo

arXiv:2004.08483

Block-Sparse-Attention

Data Efficient and Weakly Supervised Computational Pathology on Whole Slide Images

arXiv1 repo

arXiv:2004.09666

CLAM

DIET: Lightweight Language Understanding for Dialogue Systems

arXiv1 repo

arXiv:2004.09936

datacopilot

Yoga-82: A New Dataset for Fine-grained Classification of Human Poses

arXiv1 repo

arXiv:2004.10362

ai-yoga-trainer

ivis Dimensionality Reduction Framework for Biomacromolecular Simulations

arXiv1 repo

arXiv:2004.10718

SciencePlots

DuReader_robust: A Chinese Dataset Towards Evaluating Robustness and Generalization of Machine Reading Comprehension in Real-World Applications

arXiv1 repo

arXiv:2004.11142

DuReader

Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation

arXiv1 repo

arXiv:2004.11867

opus-100

Formal Adventures in Convex and Conical Spaces

arXiv1 repo

arXiv:2004.12713

infotheo

"Call me sexist, but...": Revisiting Sexism Detection Using Psychological Scales and Adversarial Samples

arXiv1 repo

arXiv:2004.12764

DeepLearningProject

A Critic Evaluation of Methods for COVID-19 Automatic Detection from X-Ray Images

arXiv1 repo

arXiv:2004.12823

covid-chestxray-dataset

Reevaluating Adversarial Examples in Natural Language

arXiv1 repo

arXiv:2004.14174

TextAttack

VGGSound: A Large-scale Audio-Visual Dataset

arXiv1 repo

arXiv:2004.14368

VGGSound

Fact or Fiction: Verifying Scientific Claims

arXiv1 repo

arXiv:2004.14974

scifact

UnifiedQA: Crossing Format Boundaries With a Single QA System

arXiv1 repo

arXiv:2005.00700

mmlu

Feature Selection Methods for Uplift Modeling and Heterogeneous Treatment Effect

arXiv1 repo

arXiv:2005.03447

causalml

Beyond Accuracy: Behavioral Testing of NLP models with CheckList

arXiv1 repo

arXiv:2005.04118

langtest

arXiv:2005.05110

arXiv1 repo

arXiv:2005.05110

misp-galaxy

TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP

arXiv1 repo

arXiv:2005.05909

TextAttack

Arabic Dialect Identification in the Wild

arXiv1 repo

arXiv:2005.06557

Glot500

ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

arXiv1 repo

arXiv:2005.07143

Displace2024_baseline_updated

IntelliCode Compose: Code Generation Using Transformer

arXiv1 repo

arXiv:2005.08025

CodeXGLUE

Efficient Wait-k Models for Simultaneous Machine Translation

arXiv1 repo

arXiv:2005.08595

translation-api

U$^2$-Net: Going Deeper with Nested U-Structure for Salient Object Detection

arXiv1 repo

arXiv:2005.09007

U-2-Net

Backstabber's Knife Collection: A Review of Open Source Software Supply Chain Attacks

arXiv1 repo

arXiv:2005.09535

guarddog

Lung Segmentation from Chest X-rays using Variational Data Imputation

arXiv1 repo

arXiv:2005.10052

covid-chestxray-dataset

BERTweet: A pre-trained language model for English Tweets

arXiv1 repo

arXiv:2005.10200

tweeteval

arXiv:2005.10356

arXiv1 repo

arXiv:2005.10356

ReVOS-api

LibriMix: An Open-Source Dataset for Generalizable Speech Separation

arXiv1 repo

arXiv:2005.11262

LibriMix

Predicting COVID-19 Pneumonia Severity on Chest X-ray with Deep Learning

arXiv1 repo

arXiv:2005.11856

covid-chestxray-dataset

NDD20: A large-scale few-shot dolphin dataset for coarse and fine-grained categorisation

arXiv1 repo

arXiv:2005.13359

RobustSAM

NuClick: A Deep Learning Framework for Interactive Segmentation of Microscopy Images

arXiv1 repo

arXiv:2005.14511

DinoBloom

Massive Choice, Ample Tasks (MaChAmp): A Toolkit for Multi-task Learning in NLP

arXiv1 repo

arXiv:2005.14672

machamp

A Scalable and Cloud-Native Hyperparameter Tuning System

arXiv1 repo

arXiv:2006.02085

awesome-kubeflow

Unsupervised Translation of Programming Languages

arXiv1 repo

arXiv:2006.03511

CodeXGLUE

Little Ball of Fur: A Python Library for Graph Sampling

arXiv1 repo

arXiv:2006.04311

littleballoffur

Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection

arXiv1 repo

arXiv:2006.04388

nanodet

Maximizing Error Injection Realism for Chaos Engineering with System Calls

arXiv1 repo

arXiv:2006.04444

awesome-chaos-engineering

BS-Net: learning COVID-19 pneumonia severity on a large Chest X-Ray dataset

arXiv1 repo

arXiv:2006.04603

covid-chestxray-dataset

Active Invariant Causal Prediction: Experiment Selection through Stability

arXiv1 repo

arXiv:2006.05690

Data-Science

TableQA: a Large-Scale Chinese Text-to-SQL Dataset for Table-Aware SQL Generation

arXiv1 repo

arXiv:2006.06434

ChineseNLPCorpus

Training Generative Adversarial Networks with Limited Data

arXiv1 repo

arXiv:2006.06676

stylegan2-ada-pytorch

FastPitch: Parallel Text-to-speech with Pitch Prediction

arXiv1 repo

arXiv:2006.06873

tts_en_fastpitch

Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training

arXiv1 repo

arXiv:2006.09092

openchat

New Vietnamese Corpus for Machine Reading Comprehension of Health News Articles

arXiv1 repo

arXiv:2006.11138

ViQG

COVID-19 Image Data Collection: Prospective Predictions Are the Future

arXiv1 repo

arXiv:2006.11988

covid-chestxray-dataset

Learning to Combine Top-Down and Bottom-Up Signals in Recurrent Neural Networks with Attention over Modules

arXiv1 repo

arXiv:2006.16981

parallel-ss-dep

Playing with Words at the National Library of Sweden -- Making a Swedish BERT

arXiv1 repo

arXiv:2007.01658

bert-base-swedish-cased

Detailed spectroscopy of doubly magic $^{132}$Sn

arXiv1 repo

arXiv:2007.03029

awsome-llm-papers

MosAIc: Finding Artistic Connections across Culture with Conditional Image Retrieval

arXiv1 repo

arXiv:2007.07177

SynapseML

LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

arXiv1 repo

arXiv:2007.08124

logiqa

Accelerating 3D Deep Learning with PyTorch3D

arXiv1 repo

arXiv:2007.08501

pytorch3d

CoVoST 2 and Massively Multilingual Speech-to-Text Translation

arXiv1 repo

arXiv:2007.10310

covost

Biomedical and Clinical English Model Packages in the Stanza Python NLP Library

arXiv1 repo

arXiv:2007.14640

stanza

MKQA: A Linguistically Diverse Benchmark for Multilingual Open Domain Question Answering

arXiv1 repo

arXiv:2007.15207

gte-multilingual-base

Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing

arXiv1 repo

arXiv:2007.15779

BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext

A Survey on Text Classification: From Shallow to Deep Learning

arXiv1 repo

arXiv:2008.00364

OpenTextClassification

Aligning AI With Shared Human Values

arXiv1 repo

arXiv:2008.02275

mmlu

Shonan Rotation Averaging: Global Optimality by Surfing $SO(p)^n$

arXiv1 repo

arXiv:2008.02737

gtsfm

A Parallel Evaluation Data Set of Software Documentation with Document Structure Annotation

arXiv1 repo

arXiv:2008.04550

Glot500

Object Detection for Graphical User Interface: Old Fashioned or Deep Learning or a Combination?

arXiv1 repo

arXiv:2008.05132

transformer-final-proj

An Experimental Study of Deep Neural Network Models for Vietnamese Multiple-Choice Reading Comprehension

arXiv1 repo

arXiv:2008.08810

ViQG

Top2Vec: Distributed Representations of Topics

arXiv1 repo

arXiv:2008.09470

turftopic

A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild

arXiv1 repo

arXiv:2008.10010

Wav2Lip

An Ensemble of Simple Convolutional Neural Network Models for MNIST Digit Recognition

arXiv1 repo

arXiv:2008.10400

MaaAI

Generating (Factual?) Narrative Summaries of RCTs: Experiments with Neural Multi-Document Summarization

arXiv1 repo

arXiv:2008.11293

mslr-shared-task

Automatic Yara Rule Generation Using Biclustering

arXiv1 repo

arXiv:2009.03779

awesome-ai-security-tools

Phasic Policy Gradient

arXiv1 repo

arXiv:2009.04416

cleanrl

An Open-Source Platform for High-Performance Non-Coherent On-Chip Communication

arXiv1 repo

arXiv:2009.05334

axi

Searching for a Search Method: Benchmarking Search Algorithms for Generating NLP Adversarial Examples

arXiv1 repo

arXiv:2009.06368

TextAttack

It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners

arXiv1 repo

arXiv:2009.07118

pet

DLBCL-Morph: Morphological features computed using deep learning for an annotated digital DLBCL image set

arXiv1 repo

arXiv:2009.08123

DLBCL-Morph

FewJoint: A Few-shot Learning Benchmark for Joint Language Understanding

arXiv1 repo

arXiv:2009.08138

ChineseNLPCorpus

Using the Hammer Only on Nails: A Hybrid Method for Evidence Retrieval for Question Answering

arXiv1 repo

arXiv:2009.10791

odqa_baseline_code

Qlib: An AI-oriented Quantitative Investment Platform

arXiv1 repo

arXiv:2009.11189

qlib

Probabilistic Label Trees for Extreme Multi-label Classification

arXiv1 repo

arXiv:2009.11218

napkinXC

Current Time Series Anomaly Detection Benchmarks are Flawed and are Creating the Illusion of Progress

arXiv1 repo

arXiv:2009.13807

Large-Time-Series-Model

PettingZoo: Gym for Multi-Agent Reinforcement Learning

arXiv1 repo

arXiv:2009.14471

PettingZoo

A Vietnamese Dataset for Evaluating Machine Reading Comprehension

arXiv1 repo

arXiv:2009.14725

ViQG

Understanding tables with intermediate pre-training

arXiv1 repo

arXiv:2010.00571

tapas-base-finetuned-wtq

Contrastive Learning of Medical Visual Representations from Paired Images and Text

arXiv1 repo

arXiv:2010.00747

clip-image-search

Sharpness-Aware Minimization for Efficiently Improving Generalization

arXiv1 repo

arXiv:2010.01412

vision_transformer

Learning from Context or Names? An Empirical Study on Neural Relation Extraction

arXiv1 repo

arXiv:2010.01923

RE-Context-or-Names

Constraining Logits by Bounded Function for Adversarial Robustness

arXiv1 repo

arXiv:2010.02558

frn-50k-baseline

Inductive Entity Representations from Text via Link Prediction

arXiv1 repo

arXiv:2010.03496

mkb

ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

arXiv1 repo

arXiv:2010.03768

alfworld

All for One and One for All: Improving Music Separation by Bridging Networks

arXiv1 repo

arXiv:2010.04228

ai-research-code

Load What You Need: Smaller Versions of Multilingual BERT

arXiv1 repo

arXiv:2010.05609

smaller-transformers

PECOS: Prediction for Enormous and Correlated Output Spaces

arXiv1 repo

arXiv:2010.05878

pecos

arXiv:2010.07115

arXiv1 repo

arXiv:2010.07115

WasmEdge

Dimsum @LaySumm 20: BART-based Approach for Scientific Document Summarization

arXiv1 repo

arXiv:2010.09252

Laysumm

DiDiSpeech: A Large Scale Mandarin Speech Corpus

arXiv1 repo

arXiv:2010.09275

seed-tts-eval

PySBD: Pragmatic Sentence Boundary Disambiguation

arXiv1 repo

arXiv:2010.09657

pySBD

Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation

arXiv1 repo

arXiv:2010.10042

vilmedic

German's Next Language Model

arXiv1 repo

arXiv:2010.10906

berts

Uncovering the Hidden Dangers: Finding Unsafe Go Code in the Wild

arXiv1 repo

arXiv:2010.11242

go-recipes

Self-training and Pre-training are Complementary for Speech Recognition

arXiv1 repo

arXiv:2010.11430

fairseq

Generating Plausible Counterfactual Explanations for Deep Transformers in Financial Text Classification

arXiv1 repo

arXiv:2010.12512

PIXIU

Pre-trained Summarization Distillation

arXiv1 repo

arXiv:2010.13002

kotoba-whisper

Out-of-core Training for Extremely Large-Scale Neural Networks With Adaptive Window-Based Scheduling

arXiv1 repo

arXiv:2010.14109

ai-research-code

Transporter Networks: Rearranging the Visual World for Robotic Manipulation

arXiv1 repo

arXiv:2010.14406

saycanpay

Generating Radiology Reports via Memory-driven Transformer

arXiv1 repo

arXiv:2010.16056

vilmedic

Joint Masked CPC and CTC Training for ASR

arXiv1 repo

arXiv:2011.00093

voxpopuli

IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP

arXiv1 repo

arXiv:2011.00677

indobert-base-uncased

Tabular Transformers for Modeling Multivariate Time Series

arXiv1 repo

arXiv:2011.01843

TabFormer

Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding

arXiv1 repo

arXiv:2011.02523

ml-hypersim

A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text Systems

arXiv1 repo

arXiv:2011.03992

awesome-nlg

Dual-stream Multiple Instance Learning Network for Whole Slide Image Classification with Self-supervised Contrastive Learning

arXiv1 repo

arXiv:2011.08939

HistoSSLscaling

QuerYD: A video dataset with high-quality text and audio narrations

arXiv1 repo

arXiv:2011.11071

TimeChat-Online-139K

Densely connected multidilated convolutional networks for dense prediction tasks

arXiv1 repo

arXiv:2011.11844

ai-research-code

CPM: A Large-scale Generative Chinese Pre-trained Language Model

arXiv1 repo

arXiv:2012.00413

CPM-Generate

Exploring the Effect of Image Enhancement Techniques on COVID-19 Detection using Chest X-rays Images

arXiv1 repo

arXiv:2012.02238

baple

FloodNet: A High Resolution Aerial Imagery Dataset for Post Flood Scene Understanding

arXiv1 repo

arXiv:2012.02951

FloodNet-Challenge-EARTHVISION2021

MLS: A Large-Scale Multilingual Dataset for Speech Research

arXiv1 repo

arXiv:2012.03411

multilingual_librispeech

Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting

arXiv1 repo

arXiv:2012.07436

informer-tourism-monthly

Session-Aware Query Auto-completion using Extreme Multi-label Ranking

arXiv1 repo

arXiv:2012.07654

pecos

Taming Transformers for High-Resolution Image Synthesis

arXiv1 repo

arXiv:2012.09841

dalle-mini

SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning

arXiv1 repo

arXiv:2012.09852

deepcompressor

Few-Shot Text Generation with Pattern-Exploiting Training

arXiv1 repo

arXiv:2012.11926

pet

Learning Dense Representations of Phrases at Scale

arXiv1 repo

arXiv:2012.12624

SimCSE

Training data-efficient image transformers & distillation through attention

arXiv1 repo

arXiv:2012.12877

Vim

Towards Fully Automated Manga Translation

arXiv1 repo

arXiv:2012.14271

manga-image-translator

arXiv:2012.14913

arXiv1 repo

arXiv:2012.14913

KEditVis-LLM-Editing

How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

arXiv1 repo

arXiv:2012.15613

hgiyt

Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection

arXiv1 repo

arXiv:2012.15761

roberta-hate-speech-dynabench-r4-target

Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers

arXiv1 repo

arXiv:2012.15840

FoodSeg103-Benchmark-v1

I-BERT: Integer-only BERT Quantization

arXiv1 repo

arXiv:2101.01321

gemmini

Dynamic Hybrid Relation Network for Cross-Domain Context-Dependent Semantic Parsing

arXiv1 repo

arXiv:2101.01686

DAMO-ConvAI

The Shapley Value of Classifiers in Ensemble Games

arXiv1 repo

arXiv:2101.02153

shapley

Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies

arXiv1 repo

arXiv:2101.02235

query_decomposer

Towards Real-World Blind Face Restoration with Generative Facial Prior

arXiv1 repo

arXiv:2101.04061

GFPGAN

MLGO: a Machine Learning Guided Compiler Optimizations Framework

arXiv1 repo

arXiv:2101.04808

ml-compiler-opt

The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models

arXiv1 repo

arXiv:2101.05667

MonoQwen2-VL-v0.1

Persistent Anti-Muslim Bias in Large Language Models

arXiv1 repo

arXiv:2101.05783

clip-italian

UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data

arXiv1 repo

arXiv:2101.07597

unispeech-large-1500h-cv

WangchanBERTa: Pretraining transformer-based Thai Language Models

arXiv1 repo

arXiv:2101.09635

SEA-PILE-v1

VisualMRC: Machine Reading Comprehension on Document Images

arXiv1 repo

arXiv:2101.11272

VisualMRC

MultiRocket: Multiple pooling operators and transformations for fast and effective time series classification

arXiv1 repo

arXiv:2102.00457

FM4Motor

Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details

arXiv1 repo

arXiv:2102.01066

GLIP

PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation

arXiv1 repo

arXiv:2102.01243

ast

CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

arXiv1 repo

arXiv:2102.04664

CodeXGLUE

Proof Artifact Co-training for Theorem Proving with Language Models

arXiv1 repo

arXiv:2102.06203

llmstep-mathlib4-pythia2.8b

Neural Network Libraries: A Deep Learning Framework Designed from Engineers' Perspectives

arXiv1 repo

arXiv:2102.06725

nnabla

Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm

arXiv1 repo

arXiv:2102.07350

awesome-prompt-engineering

Top-$k$ eXtreme Contextual Bandits with Arm Hierarchy

arXiv1 repo

arXiv:2102.07800

pecos

TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models

arXiv1 repo

arXiv:2102.07988

xDiT

SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering

arXiv1 repo

arXiv:2102.09542

LLaDA-MedV

Pyserini: An Easy-to-Use Python Toolkit to Support Replicable IR Research with Sparse and Dense Representations

arXiv1 repo

arXiv:2102.10073

odqa_baseline_code

LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation

arXiv1 repo

arXiv:2102.10815

univnet

LogME: Practical Assessment of Pre-trained Models for Transfer Learning

arXiv1 repo

arXiv:2102.11005

FM4Motor

Design and Analysis of a Logless Dynamic Reconfiguration Protocol

arXiv1 repo

arXiv:2102.11960

mongo

Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

arXiv1 repo

arXiv:2102.12122

FoodSeg103-Benchmark-v1

PyCG: Practical Call Graph Generation in Python

arXiv1 repo

arXiv:2103.00587

PyCG

Fast Adaptation with Linearized Neural Networks

arXiv1 repo

arXiv:2103.01439

xfer

Predicting Video with VQVAE

arXiv1 repo

arXiv:2103.01950

Kosmos-X

Who Can Find My Devices? Security and Privacy of Apple's Crowd-Sourced Bluetooth Location Tracking System

arXiv1 repo

arXiv:2103.02282

openhaystack

Ribbon filter: practically smaller than Bloom and Xor

arXiv1 repo

arXiv:2103.02515

gecko-dev

Catala: A Programming Language for the Law

arXiv1 repo

arXiv:2103.03198

catala

Nutrition5k: Towards Automatic Nutritional Understanding of Generic Food

arXiv1 repo

arXiv:2103.03375

MacroScope-Portion-Scout

Measuring Mathematical Problem Solving With the MATH Dataset

arXiv1 repo

arXiv:2103.03874

FTTT

ModelingToolkit: A Composable Graph Transformation System For Equation-Based Modeling

arXiv1 repo

arXiv:2103.05244

ModelingToolkit.jl

Unknown Object Segmentation from Stereo Images

arXiv1 repo

arXiv:2103.06796

humanoid_grasping

EXSCLAIM! -- An automated pipeline for the construction of labeled materials imaging datasets from literature

arXiv1 repo

arXiv:2103.10631

exsclaim2.0

MuRIL: Multilingual Representations for Indian Languages

arXiv1 repo

arXiv:2103.10730

IndicAbusive

Data Cleansing for Deep Neural Networks with Storage-efficient Approximation of Influence Functions

arXiv1 repo

arXiv:2103.11807

ai-research-code

Open Domain Question Answering over Tables via Dense Retrieval

arXiv1 repo

arXiv:2103.12011

tapas

Contact-GraspNet: Efficient 6-DoF Grasp Generation in Cluttered Scenes

arXiv1 repo

arXiv:2103.14127

humanoid_grasping

Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks

arXiv1 repo

arXiv:2103.14749

cleanlab

VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization

arXiv1 repo

arXiv:2103.16874

VITON-HD

Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval

arXiv1 repo

arXiv:2104.00650

webvid

LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference

arXiv1 repo

arXiv:2104.01136

MiDaS

A flexible and fast PyTorch toolkit for simulating training and inference on analog crossbar arrays

arXiv1 repo

arXiv:2104.02184

aihwkit

Image Composition Assessment with Saliency-augmented Multi-pattern Pooling

arXiv1 repo

arXiv:2104.03133

compositio_nn

Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings

arXiv1 repo

arXiv:2104.03502

tmh

Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

arXiv1 repo

arXiv:2104.04473

MindSpeed-MM

arXiv:2104.04691

arXiv1 repo

arXiv:2104.04691

ReVOS-api

An Efficient 2D Method for Training Super-Large Deep Learning Models

arXiv1 repo

arXiv:2104.05343

ColossalAI

Getting to the Point. Index Sets and Parallelism-Preserving Autodiff for Pointful Array Programming

arXiv1 repo

arXiv:2104.05372

dex-lang

Rapid Exploration for Open-World Navigation with Latent Goal Models

arXiv1 repo

arXiv:2104.05859

UniWM_Dataset

Learning and Planning in Complex Action Spaces

arXiv1 repo

arXiv:2104.06303

xwm

MS2: Multi-Document Summarization of Medical Studies

arXiv1 repo

arXiv:2104.06486

mslr-shared-task

TSDAE: Using Transformer-based Sequential Denoising Auto-Encoder for Unsupervised Sentence Embedding Learning

arXiv1 repo

arXiv:2104.06979

Vorbereitung

EAT: Enhanced ASR-TTS for Self-supervised Speech Recognition

arXiv1 repo

arXiv:2104.07474

parrots

The Power of Scale for Parameter-Efficient Prompt Tuning

arXiv1 repo

arXiv:2104.08691

Continual-NExT

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

arXiv1 repo

arXiv:2104.08758

falcon-refinedweb

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

arXiv1 repo

arXiv:2104.08860

CLIP4Clip

Evaluating Deep Neural Networks Trained on Clinical Images in Dermatology with the Fitzpatrick 17k Dataset

arXiv1 repo

arXiv:2104.09957

fitzpatrick17k

VideoGPT: Video Generation using VQ-VAE and Transformers

arXiv1 repo

arXiv:2104.10157

VideoGPT

ImageNet-21K Pretraining for the Masses

arXiv1 repo

arXiv:2104.10972

convnext

Multiscale Vision Transformers

arXiv1 repo

arXiv:2104.11227

pytorchvideo

Morph Call: Probing Morphosyntactic Content of Multilingual Transformers

arXiv1 repo

arXiv:2104.12847

morph-call

TRECVID 2020: A comprehensive campaign for evaluating video retrieval tasks across multiple application domains

arXiv1 repo

arXiv:2104.13473

ladi-overview

arXiv:2104.13921

arXiv1 repo

arXiv:2104.13921

Echo-ViLD

Emerging Properties in Self-Supervised Vision Transformers

arXiv1 repo

arXiv:2104.14294

dino

Conversational Machine Reading Comprehension for Vietnamese Healthcare Texts

arXiv1 repo

arXiv:2105.01542

ViQG

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism

arXiv1 repo

arXiv:2105.02446

fish-diffusion

What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus

arXiv1 repo

arXiv:2105.02732

whatsinthebox

A Dataset of Information-Seeking Questions and Answers Anchored in Research Papers

arXiv1 repo

arXiv:2105.03011

qasper

ResMLP: Feedforward networks for image classification with data-efficient training

arXiv1 repo

arXiv:2105.03404

Swin-Transformer

High-performance symbolic-numerics via multiple dispatch

arXiv1 repo

arXiv:2105.03949

Symbolics.jl

Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning

arXiv1 repo

arXiv:2105.04165

InterGPS

VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning

arXiv1 repo

arXiv:2105.04906

xwm

Diffusion Models Beat GANs on Image Synthesis

arXiv1 repo

arXiv:2105.05233

guided-diffusion

Snipuzz: Black-box Fuzzing of IoT Firmware via Message Snippet Inference

arXiv1 repo

arXiv:2105.05445

awesome-connected-things-sec

Monash Time Series Forecasting Archive

arXiv1 repo

arXiv:2105.06643

UTSD

A cost-benefit analysis of cross-lingual transfer methods

arXiv1 repo

arXiv:2105.06813

mmarco

Few-NERD: A Few-Shot Named Entity Recognition Dataset

arXiv1 repo

arXiv:2105.07464

Few-NERD

NExT-QA:Next Phase of Question-Answering to Explaining Temporal Actions

arXiv1 repo

arXiv:2105.08276

NExT-QA

Progressively Normalized Self-Attention Network for Video Polyp Segmentation

arXiv1 repo

arXiv:2105.08468

VPS

Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech

arXiv1 repo

arXiv:2105.09040

UniSpeech

Methods for Detoxification of Texts for the Russian Language

arXiv1 repo

arXiv:2105.09052

rudetoxifier

DeepCAD: A Deep Generative Network for Computer-Aided Design Models

arXiv1 repo

arXiv:2105.09492

DeepCAD

FreshDiskANN: A Fast and Accurate Graph-Based ANN Index for Streaming Similarity Search

arXiv1 repo

arXiv:2105.09613

slater

Unsupervised Speech Recognition

arXiv1 repo

arXiv:2105.11084

fairseq

Read, Listen, and See: Leveraging Multimodal Information Helps Chinese Spell Checking

arXiv1 repo

arXiv:2105.12306

CTCResources

Sequence Parallelism: Long Sequence Training from System Perspective

arXiv1 repo

arXiv:2105.13120

ColossalAI

ILDC for CJPE: Indian Legal Documents Corpus for Court Judgment Prediction and Explanation

arXiv1 repo

arXiv:2105.13562

InLegalBERT

An Attention Free Transformer

arXiv1 repo

arXiv:2105.14103

rwkv

Maximizing Parallelism in Distributed Training for Huge Neural Networks

arXiv1 repo

arXiv:2105.14450

ColossalAI

Tesseract: Parallelize the Tensor Parallelism Efficiently

arXiv1 repo

arXiv:2105.14500

ColossalAI

Exploration and Exploitation: Two Ways to Improve Chinese Spelling Correction Models

arXiv1 repo

arXiv:2105.14813

CTCResources

DoT: An efficient Double Transformer for NLP tasks with tables

arXiv1 repo

arXiv:2106.00479

tapas

SpanNER: Named Entity Re-/Recognition as Span Prediction

arXiv1 repo

arXiv:2106.00641

nlp-bazel-tutorial

Enabling Efficiency-Precision Trade-offs for Label Trees in Extreme Classification

arXiv1 repo

arXiv:2106.00730

pecos

TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification

arXiv1 repo

arXiv:2106.00908

HistoSSLscaling

NVC-Net: End-to-End Adversarial Voice Conversion

arXiv1 repo

arXiv:2106.00992

ai-research-code

When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations

arXiv1 repo

arXiv:2106.01548

vision_transformer

Fre-GAN: Adversarial Frequency-consistent Audio Synthesis

arXiv1 repo

arXiv:2106.02297

MockingBird

Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation

arXiv1 repo

arXiv:2106.03153

vad_score_prediction

arXiv:2106.03609

arXiv1 repo

arXiv:2106.03609

HEBO

End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering

arXiv1 repo

arXiv:2106.05346

art

Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

arXiv1 repo

arXiv:2106.06103

vitsgpt-vits

TrafficStream: A Streaming Traffic Flow Forecasting Framework Based on Graph Neural Networks and Continual Learning

arXiv1 repo

arXiv:2106.06273

A2TTA

RETRIEVE: Coreset Selection for Efficient and Robust Semi-Supervised Learning

arXiv1 repo

arXiv:2106.07760

cords

UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation

arXiv1 repo

arXiv:2106.07889

univnet

CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

arXiv1 repo

arXiv:2106.08087

SDAK

BEiT: BERT Pre-Training of Image Transformers

arXiv1 repo

arXiv:2106.08254

MiDaS

End-to-End Semi-Supervised Object Detection with Soft Teacher

arXiv1 repo

arXiv:2106.09018

Swin-Transformer

Large-Scale Chemical Language Representations Capture Molecular Structure and Properties

arXiv1 repo

arXiv:2106.09553

molformer

XCiT: Cross-Covariance Image Transformers

arXiv1 repo

arXiv:2106.09681

dino

Bad Characters: Imperceptible NLP Attacks

arXiv1 repo

arXiv:2106.09898

garak

Distributed Deep Learning in Open Collaborations

arXiv1 repo

arXiv:2106.10207

DeDLOC

How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

arXiv1 repo

arXiv:2106.10270

vision_transformer

Surgical data science for safe cholecystectomy: a protocol for segmentation of hepatocystic anatomy and assessment of the critical view of safety

arXiv1 repo

arXiv:2106.10916

Endoscapes

BARTScore: Evaluating Generated Text as Text Generation

arXiv1 repo

arXiv:2106.11520

BARTScore

It's All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning

arXiv1 repo

arXiv:2106.12066

xwinograd

Alias-Free Generative Adversarial Networks

arXiv1 repo

arXiv:2106.12423

ffhq-dataset

Extreme Multi-label Learning for Semantic Matching in Product Search

arXiv1 repo

arXiv:2106.12657

pecos

Label Disentanglement in Partition-based Extreme Multilabel Classification

arXiv1 repo

arXiv:2106.12751

pecos

Unsupervised Topic Segmentation of Meetings with BERT Embeddings

arXiv1 repo

arXiv:2106.12978

E2E-Video-Processing-system

Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting

arXiv1 repo

arXiv:2106.13008

autoformer-tourism-monthly

AudioCLIP: Extending CLIP to Image, Text and Audio

arXiv1 repo

arXiv:2106.13043

JavisBench

Video Swin Transformer

arXiv1 repo

arXiv:2106.13230

Swin-Transformer

panda-gym: Open-source goal-conditioned environments for robotic learning

arXiv1 repo

arXiv:2106.13687

panda-gym

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages

arXiv1 repo

arXiv:2106.13822

Glot500

Habitat 2.0: Training Home Assistants to Rearrange their Habitat

arXiv1 repo

arXiv:2106.14405

habitat-lab

R-Drop: Regularized Dropout for Neural Networks

arXiv1 repo

arXiv:2106.14448

final-project-level3-nlp-02

A Few Brief Notes on DeepImpact, COIL, and a Conceptual Framework for Information Retrieval Techniques

arXiv1 repo

arXiv:2106.14807

bge-m3

AutoFormer: Searching Transformers for Visual Recognition

arXiv1 repo

arXiv:2107.00651

Cream

Data Centric Domain Adaptation for Historical Text with OCR Errors

arXiv1 repo

arXiv:2107.00927

historic-domain-adaptation-icdar

Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors

arXiv1 repo

arXiv:2107.01545

UniSpeech

DeepDDS: deep graph neural network with attention mechanism to predict synergistic drug combinations

arXiv1 repo

arXiv:2107.02467

chemicalx

Depth-supervised NeRF: Fewer Views and Faster Training for Free

arXiv1 repo

arXiv:2107.02791

svraster

SoundStream: An End-to-End Neural Audio Codec

arXiv1 repo

arXiv:2107.03312

moshi

Trusting RoBERTa over BERT: Insights from CheckListing the Natural Language Inference Task

arXiv1 repo

arXiv:2107.07229

awesome-cybersecurity-agentic-ai

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation

arXiv1 repo

arXiv:2107.07651

LAVIS

Declarative Machine Learning Systems

arXiv1 repo

arXiv:2107.08148

ludwig

YOLOX: Exceeding YOLO Series in 2021

arXiv1 repo

arXiv:2107.08430

YOLOX

Frequency-Domain Data-Driven Controller Synthesis for Unstable LPV Systems

arXiv1 repo

arXiv:2107.09712

speech-emotion-recognition

Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data

arXiv1 repo

arXiv:2107.10833

Real-ESRGAN

SurfaceNet: Adversarial SVBRDF Estimation from a Single Image

arXiv1 repo

arXiv:2107.11298

surfacenet

Open-Ended Learning Leads to Generally Capable Agents

arXiv1 repo

arXiv:2107.12808

WaveFunctionCollapse

Domain-matched Pre-training Tasks for Dense Retrieval

arXiv1 repo

arXiv:2107.13602

dpr-scale

Perceiver IO: A General Architecture for Structured Inputs & Outputs

arXiv1 repo

arXiv:2107.14795

FM4Motor

PyEuroVoc: A Tool for Multilingual Legal Document Classification with EuroVoc Descriptors

arXiv1 repo

arXiv:2108.01139

pyeurovoc

Object Wake-up: 3D Object Rigging from a Single Image

arXiv1 repo

arXiv:2108.02708

object-wakeup

BERT-based distractor generation for Swedish reading comprehension questions using a small-scale dataset

arXiv1 repo

arXiv:2108.03973

rc-answer-generation

Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models

arXiv1 repo

arXiv:2108.04024

CIRR

Util::Lookup: Exploiting key decoding in cryptographic libraries

arXiv1 repo

arXiv:2108.04600

bc-rust

PatrickStar: Parallel Training of Pre-trained Models via Chunk-based Memory Management

arXiv1 repo

arXiv:2108.05818

ColossalAI

Semantic Answer Similarity for Evaluating Question Answering Models

arXiv1 repo

arXiv:2108.06130

Medical-Assistant

Conditional DETR for Fast Training Convergence

arXiv1 repo

arXiv:2108.06152

DINO

Pixel Difference Networks for Efficient Edge Detection

arXiv1 repo

arXiv:2108.07009

pidinet

On the Opportunities and Risks of Foundation Models

arXiv1 repo

arXiv:2108.07258

gpttools

EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning

arXiv1 repo

arXiv:2108.08842

llama.cpp-eden

A Survey on Common Threats in npm and PyPi Registries

arXiv1 repo

arXiv:2108.09576

guarddog

An Empirical Assessment of Endpoint Security Systems Against Advanced Persistent Threats Attack Vectors

arXiv1 repo

arXiv:2108.10422

awesome-edr-bypass

Rewrite Rule Inference Using Equality Saturation

arXiv1 repo

arXiv:2108.10436

equational_theories

One TTS Alignment To Rule Them All

arXiv1 repo

arXiv:2108.10447

tts_en_fastpitch

LLVIP: A Visible-infrared Paired Dataset for Low-light Vision

arXiv1 repo

arXiv:2108.10831

LLVIP

Meta Self-Learning for Multi-Source Domain Adaptation: A Benchmark

arXiv1 repo

arXiv:2108.10840

Meta-SelfLearning

mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset

arXiv1 repo

arXiv:2108.13897

mmarco

FinQA: A Dataset of Numerical Reasoning over Financial Data

arXiv1 repo

arXiv:2109.00122

MirageTVQA

WebQA: Multihop and Multimodal QA

arXiv1 repo

arXiv:2109.00590

ReMuQ

Learning to Prompt for Vision-Language Models

arXiv1 repo

arXiv:2109.01134

BiomedCoOp

Similarity of Sentence Representations in Multilingual LMs: Resolving Conflicting Literature and Case Study of Baltic Languages

arXiv1 repo

arXiv:2109.01207

xsim

Fast Succinct Retrieval and Approximate Membership using Ribbon

arXiv1 repo

arXiv:2109.01892

gecko-dev

MATE: Multi-view Attention for Table Transformer Efficiency

arXiv1 repo

arXiv:2109.04312

tapas

Smoothed Contrastive Learning for Unsupervised Sentence Embedding

arXiv1 repo

arXiv:2109.04321

sentemb

ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding

arXiv1 repo

arXiv:2109.04380

sentemb

Packed Levitated Marker for Entity and Relation Extraction

arXiv1 repo

arXiv:2109.06067

trusted_ke

LM-Critic: Language Models for Unsupervised Grammatical Error Correction

arXiv1 repo

arXiv:2109.06822

CTCResources

Resolution-robust Large Mask Inpainting with Fourier Convolutions

arXiv1 repo

arXiv:2109.07161

lama

Towards Zero-shot Cross-lingual Image Retrieval and Tagging

arXiv1 repo

arXiv:2109.07622

Multilingual-CLIP

ROS-X-Habitat: Bridging the ROS Ecosystem with Embodied AI

arXiv1 repo

arXiv:2109.07703

habitat-lab

Towards Zero and Few-shot Knowledge-seeking Turn Detection in Task-orientated Dialogue Systems

arXiv1 repo

arXiv:2109.08820

KoPrivateGPT

TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models

arXiv1 repo

arXiv:2109.10282

trocr-base-handwritten

Transcoding Billions of Unicode Characters per Second with SIMD Instructions

arXiv1 repo

arXiv:2109.10433

simdutf

Pix2seq: A Language Modeling Framework for Object Detection

arXiv1 repo

arXiv:2109.10852

LocateAnything-3B

Transferring Knowledge from Vision to Language: How to Achieve it and how to Measure it?

arXiv1 repo

arXiv:2109.11321

Kosmos-X

Trans-Encoder: Unsupervised sentence-pair modelling through self- and mutual-distillations

arXiv1 repo

arXiv:2109.13059

trans-encoder

Robust SLAM Systems: Are We There Yet?

arXiv1 repo

arXiv:2109.13160

slambench

VoiceFixer: Toward General Speech Restoration with Neural Vocoder

arXiv1 repo

arXiv:2109.13731

voicefixer

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

arXiv1 repo

arXiv:2109.14084

fairseq

Fast Multi-Resolution Transformer Fine-tuning for Extreme Multi-label Text Classification

arXiv1 repo

arXiv:2110.00685

pecos

arXiv:2110.01200

arXiv1 repo

arXiv:2110.01200

MTP-Codebase

WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition

arXiv1 repo

arXiv:2110.03370

WenetSpeech

CLIP-Adapter: Better Vision-Language Models with Feature Adapters

arXiv1 repo

arXiv:2110.04544

BiomedCoOp

Vector-quantized Image Modeling with Improved VQGAN

arXiv1 repo

arXiv:2110.04627

echo-vqgan

Large-scale Self-Supervised Speech Representation Learning for Automatic Speaker Verification

arXiv1 repo

arXiv:2110.05777

UniSpeech

ABO: Dataset and Benchmarks for Real-World 3D Object Understanding

arXiv1 repo

arXiv:2110.06199

Cap3D

S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations

arXiv1 repo

arXiv:2110.06280

s3prl-vc

Salient Phrase Aware Dense Retrieval: Can a Dense Retriever Imitate a Sparse One?

arXiv1 repo

arXiv:2110.06918

dpr-scale

Toward Degradation-Robust Voice Conversion

arXiv1 repo

arXiv:2110.07537

RobustVC

P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

arXiv1 repo

arXiv:2110.07602

P-tuning-v2

HumBugDB: A Large-scale Acoustic Mosquito Dataset

arXiv1 repo

arXiv:2110.07607

BEANS-Zero

Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models

arXiv1 repo

arXiv:2110.08173

medlama

BBQ: A Hand-Built Bias Benchmark for Question Answering

arXiv1 repo

arXiv:2110.08193

langtest

LSA: Modeling Aspect Sentiment Coherency via Local Sentiment Aggregation

arXiv1 repo

arXiv:2110.08604

News-Sentiment-Analysis

NormFormer: Improved Transformer Pretraining with Extra Normalization

arXiv1 repo

arXiv:2110.09456

CLIP-ViT-L-14-laion2B-s32B-b82K

SSAST: Self-Supervised Audio Spectrogram Transformer

arXiv1 repo

arXiv:2110.09784

ast

Wacky Weights in Learned Sparse Representations and the Revenge of Score-at-a-Time Query Evaluation

arXiv1 repo

arXiv:2110.11540

SPLADERunner

Lhotse: a speech data representation library for the modern deep learning ecosystem

arXiv1 repo

arXiv:2110.12561

lhotse

Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations

arXiv1 repo

arXiv:2110.14513

NANSY

Distilling Relation Embeddings from Pre-trained Language Models

arXiv1 repo

arXiv:2110.15705

relbert

Chaos Engineering of Ethereum Blockchain Clients

arXiv1 repo

arXiv:2111.00221

awesome-chaos-engineering

RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and Intensity Responses

arXiv1 repo

arXiv:2111.00962

vocoder

LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

arXiv1 repo

arXiv:2111.02114

mlcd-vit-large-patch14-336

An Empirical Study of Training End-to-End Vision-and-Language Transformers

arXiv1 repo

arXiv:2111.02387

VLE

WORD: A large scale dataset, benchmark and clinical applicable study for abdominal organ segmentation from CT image

arXiv1 repo

arXiv:2111.02403

WORD

A Unified View of Relational Deep Learning for Drug Pair Scoring

arXiv1 repo

arXiv:2111.02916

chemicalx

Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

arXiv1 repo

arXiv:2111.03930

BiomedCoOp

arXiv:2111.06178

arXiv1 repo

arXiv:2111.06178

HEBO

Adding more data does not always help: A study in medical conversation summarization with PEGASUS

arXiv1 repo

arXiv:2111.07564

curai-research

XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

arXiv1 repo

arXiv:2111.09296

wav2vec2-xls-r-300m

MEDCOD: A Medically-Accurate, Emotive, Diverse, and Controllable Dialog System

arXiv1 repo

arXiv:2111.09381

curai-research

ClipCap: CLIP Prefix for Image Captioning

arXiv1 repo

arXiv:2111.09734

webapp

Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions

arXiv1 repo

arXiv:2111.10337

TimeChat-Online-139K

arXiv:2111.12704

arXiv1 repo

arXiv:2111.12704

RealBasicVSR

Replicating Monotonic Payoffs Without Oracles

arXiv1 repo

arXiv:2111.13740

osmosis

What Do You See in this Patient? Behavioral Testing of Clinical NLP Models

arXiv1 repo

arXiv:2111.15512

langtest

DKPLM: Decomposable Knowledge-enhanced Pre-trained Language Model for Natural Language Understanding

arXiv1 repo

arXiv:2112.01047

EasyNLP

BERTMap: A BERT-based Ontology Alignment System

arXiv1 repo

arXiv:2112.02682

BERTMap

Grounded Language-Image Pre-training

arXiv1 repo

arXiv:2112.03857

GLIP

arXiv:2112.05251

arXiv1 repo

arXiv:2112.05251

Kosmos-X

DistilCSE: Effective Knowledge Distillation For Contrastive Sentence Embeddings

arXiv1 repo

arXiv:2112.05638

sentemb

Margin Calibration for Long-Tailed Visual Recognition

arXiv1 repo

arXiv:2112.07225

robustlearn

Large Dual Encoders Are Generalizable Retrievers

arXiv1 repo

arXiv:2112.07899

gtr-t5-base

Learning Cross-Lingual IR from an English Retriever

arXiv1 repo

arXiv:2112.08185

DrDecr_XOR-TyDi_whitebox

StyleMC: Multi-Channel Based Fast Text-Guided Image Generation and Manipulation

arXiv1 repo

arXiv:2112.08493

StyleMC-SentenceTransformers

DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching Models

arXiv1 repo

arXiv:2112.08609

DuReader

Self-Supervised Learning for speech recognition with Intermediate layer supervision

arXiv1 repo

arXiv:2112.08778

UniSpeech

PeopleSansPeople: A Synthetic Data Generator for Human-Centric Computer Vision

arXiv1 repo

arXiv:2112.09290

humans

WebGPT: Browser-assisted question-answering with human feedback

arXiv1 repo

arXiv:2112.09332

webgpt_comparisons

Align and Prompt: Video-and-Language Pre-training with Entity Prompts

arXiv1 repo

arXiv:2112.09583

LAVIS

What are Weak Links in the npm Supply Chain?

arXiv1 repo

arXiv:2112.10165

guarddog

Mask2Former for Video Instance Segmentation

arXiv1 repo

arXiv:2112.10764

Mask2Former

Does MAML Only Work via Feature Re-use? A Data Centric Perspective

arXiv1 repo

arXiv:2112.13137

ultimate-utils

Temporally Constrained Neural Networks (TCNN): A framework for semi-supervised video semantic segmentation

arXiv1 repo

arXiv:2112.13815

Endoscapes

LeSICiN: A Heterogeneous Graph-based Approach for Automatic Legal Statute Identification from Indian Legal Documents

arXiv1 repo

arXiv:2112.14731

InLegalBERT

Towards a secure API client generator for IoT devices

arXiv1 repo

arXiv:2201.00270

openapi-generator

arXiv:2201.00487

arXiv1 repo

arXiv:2201.00487

ReferFormer

Swin UNETR: Swin Transformers for Semantic Segmentation of Brain Tumors in MRI Images

arXiv1 repo

arXiv:2201.01266

TriALS

A Survey of JSON-compatible Binary Serialization Specifications

arXiv1 repo

arXiv:2201.02089

smile-format-specification

Detecting Twenty-thousand Classes using Image-level Supervision

arXiv1 repo

arXiv:2201.02605

pixmo-count

A Benchmark of JSON-compatible Binary Serialization Specifications

arXiv1 repo

arXiv:2201.03051

smile-format-specification

DeepKE: A Deep Learning Based Knowledge Extraction Toolkit for Knowledge Base Population

arXiv1 repo

arXiv:2201.03335

DeepKE

CVSS Corpus and Massively Multilingual Speech-to-Speech Translation

arXiv1 repo

arXiv:2201.03713

UniSS

UnifiedSKG: Unifying and Multi-Tasking Structured Knowledge Grounding with Text-to-Text Language Models

arXiv1 repo

arXiv:2201.05966

UnifiedSKG

Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth

arXiv1 repo

arXiv:2201.07436

GLPDepth

Can Model Compression Improve NLP Fairness

arXiv1 repo

arXiv:2201.08542

distilgpt2

STRIDE-based Cyber Security Threat Modeling for IoT-enabled Precision Agriculture Systems

arXiv1 repo

arXiv:2201.09493

awesome-connected-things-sec

Describing Differences between Text Distributions with Natural Language

arXiv1 repo

arXiv:2201.12323

imodelsX

DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

arXiv1 repo

arXiv:2201.12329

DINO

Incorporating Commonsense Knowledge into Story Ending Generation via Heterogeneous Graph Networks

arXiv1 repo

arXiv:2201.12538

AwesomeSEG

A Dataset for Medical Instructional Video Classification and Question Answering

arXiv1 repo

arXiv:2201.12888

VPTSL

Negativity Spreads Faster: A Large-Scale Multilingual Twitter Analysis on the Role of Sentiment in Political Communication

arXiv1 repo

arXiv:2202.00396

xlm-twitter-politics-sentiment

Locally Typical Sampling

arXiv1 repo

arXiv:2202.00666

LLM-Sampling

data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

arXiv1 repo

arXiv:2202.03555

data2vec-audio-large

MaskGIT: Masked Generative Image Transformer

arXiv1 repo

arXiv:2202.04200

nanoMFM

InPars: Data Augmentation for Information Retrieval using Large Language Models

arXiv1 repo

arXiv:2202.05144

InPars

TwHIN: Embedding the Twitter Heterogeneous Information Network for Personalized Recommendation

arXiv1 repo

arXiv:2202.05387

the-algorithm-ml

A Contrastive Framework for Neural Text Generation

arXiv1 repo

arXiv:2202.06417

4th-Bookathon-The-Unbearable-Heaviness-of-GPT

arXiv:2202.06558

arXiv1 repo

arXiv:2202.06558

HEBO

Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?

arXiv1 repo

arXiv:2202.06675

clip-retrieval

Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text

arXiv1 repo

arXiv:2202.06935

awesome-nlg

Transformers in Time Series: A Survey

arXiv1 repo

arXiv:2202.07125

time-moe

arXiv:2202.07359

arXiv1 repo

arXiv:2202.07359

textlesslib

Information Extraction in Low-Resource Scenarios: Survey and Perspective

arXiv1 repo

arXiv:2202.08063

DeepKE

Probing Pretrained Models of Source Code

arXiv1 repo

arXiv:2202.08975

probings4code

Pseudo Numerical Methods for Diffusion Models on Manifolds

arXiv1 repo

arXiv:2202.09778

stable-diffusion

Phrase-Based Affordance Detection via Cyclic Bilateral Interaction

arXiv1 repo

arXiv:2202.12076

Cross-View-AG

LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding

arXiv1 repo

arXiv:2202.13669

LiLT

Combining Modular Skills in Multitask Learning

arXiv1 repo

arXiv:2202.13914

mttl

Mukayese: Turkish NLP Strikes Back

arXiv1 repo

arXiv:2203.01215

mukayese

DN-DETR: Accelerate DETR Training by Introducing Query DeNoising

arXiv1 repo

arXiv:2203.01305

DINO

Nuclei instance segmentation and classification in histopathology images with StarDist

arXiv1 repo

arXiv:2203.02284

stardist

arXiv:2203.02923

arXiv1 repo

arXiv:2203.02923

smalldiffusion

HEAR: Holistic Evaluation of Audio Representations

arXiv1 repo

arXiv:2203.03022

usad

Improving CTC-based speech recognition via knowledge transferring from pre-trained language models

arXiv1 repo

arXiv:2203.03582

ASR-Knowledge-Transferring

DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

arXiv1 repo

arXiv:2203.03605

DINO

A Simple Multi-Modality Transfer Learning Baseline for Sign Language Translation

arXiv1 repo

arXiv:2203.04287

SLRT

Temporal Difference Learning for Model Predictive Control

arXiv1 repo

arXiv:2203.04955

xwm

Conditional Prompt Learning for Vision-Language Models

arXiv1 repo

arXiv:2203.05557

BiomedCoOp

CMKD: CNN/Transformer-Based Cross-Model Knowledge Distillation for Audio Classification

arXiv1 repo

arXiv:2203.06760

ast

Block-STM: Scaling Blockchain Execution by Turning Ordering Curse to a Performance Blessing

arXiv1 repo

arXiv:2203.06871

cosmos-sdk

Formalising Decentralised Exchanges in Coq

arXiv1 repo

arXiv:2203.08016

ConCert

Surrogate Gap Minimization Improves Sharpness-Aware Training

arXiv1 repo

arXiv:2203.08065

vision_transformer

AUTOMATA: Gradient Based Data Subset Selection for Compute-Efficient Hyper-parameter Tuning

arXiv1 repo

arXiv:2203.08212

cords

Multilingual Pre-training with Language and Task Adaptation for Multilingual Text Style Transfer

arXiv1 repo

arXiv:2203.08552

multilingual_tst_copy

PosePipe: Open-Source Human Pose Estimation Pipeline for Clinical Research

arXiv1 repo

arXiv:2203.08792

PosePipeline

A Survey of Multi-Tenant Deep Learning Inference on GPU

arXiv1 repo

arXiv:2203.09040

awesome-gpu-engineering

CaRTS: Causality-driven Robot Tool Segmentation from Vision and Kinematics Data

arXiv1 repo

arXiv:2203.09475

CaRTS

Learning Affordance Grounding from Exocentric Images

arXiv1 repo

arXiv:2203.09905

Cross-View-AG

DuReader_retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine

arXiv1 repo

arXiv:2203.10232

DuReader

Mitigating Gender Bias in Distilled Language Models via Counterfactual Role Reversal

arXiv1 repo

arXiv:2203.12574

distilgpt2

Video Polyp Segmentation: A Deep Learning Perspective

arXiv1 repo

arXiv:2203.14291

VPS

Certified Mergeable Replicated Data Types

arXiv1 repo

arXiv:2203.14518

irmin

Socially Compliant Navigation Dataset (SCAND): A Large-Scale Dataset of Demonstrations for Social Navigation

arXiv1 repo

arXiv:2203.15041

UniWM_Dataset

X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval

arXiv1 repo

arXiv:2203.15086

xpool

Killing Two Birds with One Stone:Efficient and Robust Training of Face Recognition CNNs by Partial FC

arXiv1 repo

arXiv:2203.15565

insightface

Earnings-22: A Practical Benchmark for Accents in the Wild

arXiv1 repo

arXiv:2203.15591

earnings22

Parameter-efficient Model Adaptation for Vision Transformers

arXiv1 repo

arXiv:2203.16329

MoA

MMER: Multimodal Multi-task Learning for Speech Emotion Recognition

arXiv1 repo

arXiv:2203.16794

MMER

BRIO: Bringing Order to Abstractive Summarization

arXiv1 repo

arXiv:2203.16804

BRIO

Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech Data

arXiv1 repo

arXiv:2203.17113

SpeechT5

Making Pre-trained Language Models End-to-end Few-shot Learners with Contrastive Prompt Tuning

arXiv1 repo

arXiv:2204.00166

EasyNLP

PriMock57: A Dataset Of Primary Care Mock Consultations

arXiv1 repo

arXiv:2204.00333

primock57

Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation

arXiv1 repo

arXiv:2204.00447

primock57

Learning Audio-Video Modalities from Image Captions

arXiv1 repo

arXiv:2204.00679

videoCC-data

Distributional Gradient Boosting Machines

arXiv1 repo

arXiv:2204.00778

LightGBMLSS

HLDC: Hindi Legal Documents Corpus

arXiv1 repo

arXiv:2204.00806

HLDC

arXiv:2204.01715

arXiv1 repo

arXiv:2204.01715

llm_test

Temporal Alignment Networks for Long-term Video

arXiv1 repo

arXiv:2204.02968

TemporalAlignNet

BioBART: Pretraining and Evaluation of A Biomedical Generative Language Model

arXiv1 repo

arXiv:2204.03905

Fengshenbang-LM

ASQA: Factoid Questions Meet Long-Form Answers

arXiv1 repo

arXiv:2204.06092

KoPrivateGPT

WikiDiverse: A Multimodal Entity Linking Dataset with Diversified Contextual Topics and Entity Types

arXiv1 repo

arXiv:2204.06347

GEMEL

GPT-NeoX-20B: An Open-Source Autoregressive Language Model

arXiv1 repo

arXiv:2204.06745

level3_nlp_finalproject-nlp-12

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

arXiv1 repo

arXiv:2204.08387

LayoutLMv3-DocVQA

Dress Code: High-Resolution Multi-Category Virtual Try-On

arXiv1 repo

arXiv:2204.08532

dress-code

ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

arXiv1 repo

arXiv:2204.08790

GLIP

ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers

arXiv1 repo

arXiv:2204.09224

contentvec

SciTS: A Benchmark for Time-Series Databases in Scientific Experiments and Industrial Internet of Things

arXiv1 repo

arXiv:2204.09795

ClickBench

Learning to Revise References for Faithful Summarization

arXiv1 repo

arXiv:2204.10290

summary-reference-revision

TorchSparse: Efficient Point Cloud Inference Engine

arXiv1 repo

arXiv:2204.10319

torchsparse

arXiv:2204.12260

arXiv1 repo

arXiv:2204.12260

TIL-2023

DoPose-6D dataset for object segmentation and 6D pose estimation

arXiv1 repo

arXiv:2204.13613

image_agnostic_segmentation

SVTR: Scene Text Recognition with a Single Visual Model

arXiv1 repo

arXiv:2205.00159

yas

MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning

arXiv1 repo

arXiv:2205.00445

math

Sequencer: Deep LSTM for Image Classification

arXiv1 repo

arXiv:2205.01972

FM4Motor

Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition

arXiv1 repo

arXiv:2205.03433

vocalsound

Reducing Activation Recomputation in Large Transformer Models

arXiv1 repo

arXiv:2205.05198

MindSpeed-MM

Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context Learning

arXiv1 repo

arXiv:2205.05638

level3_nlp_finalproject-nlp-12

ViT5: Pretrained Text-to-Text Transformer for Vietnamese Language Generation

arXiv1 repo

arXiv:2205.06457

ViT5

Consistent Human Evaluation of Machine Translation across Language Pairs

arXiv1 repo

arXiv:2205.08533

pearmut

Summarization as Indirect Supervision for Relation Extraction

arXiv1 repo

arXiv:2205.09837

SuRE

DeepStruct: Pretraining of Language Models for Structure Prediction

arXiv1 repo

arXiv:2205.10475

trusted_ke

Deep Learning Workload Scheduling in GPU Datacenters: Taxonomy, Challenges and Vision

arXiv1 repo

arXiv:2205.11913

awesome-gpu-engineering

RetroMAE: Pre-Training Retrieval-oriented Language Models Via Masked Auto-Encoder

arXiv1 repo

arXiv:2205.12035

ember-v1

FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

arXiv1 repo

arXiv:2205.12446

fleurs

End-to-End Multimodal Fact-Checking and Explanation Generation: A Challenging Dataset and Models

arXiv1 repo

arXiv:2205.12487

Mocheg

arXiv:2205.13603

arXiv1 repo

arXiv:2205.13603

web-stable-diffusion

Multimodal Masked Autoencoders Learn Transferable Representations

arXiv1 repo

arXiv:2205.14204

m3ae_public

Controllable Text Generation with Neurally-Decomposed Oracle

arXiv1 repo

arXiv:2205.14219

constrDecoding

Multimodal Fake News Detection via CLIP-Guided Learning

arXiv1 repo

arXiv:2205.14304

Multi-fake-detective

EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

arXiv1 repo

arXiv:2205.14756

efficientvit

Prompt-aligned Gradient for Prompt Tuning

arXiv1 repo

arXiv:2205.14865

BiomedCoOp

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

arXiv1 repo

arXiv:2205.15868

CogVideo

NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages

arXiv1 repo

arXiv:2205.15960

NusaX-senti

Text2Human: Text-Driven Controllable Human Image Generation

arXiv1 repo

arXiv:2205.15996

DeepFashion-MultiModal

Squeezeformer: An Efficient Transformer for Automatic Speech Recognition

arXiv1 repo

arXiv:2206.00888

kaggle-asl-fingerspelling-1st-place-solution

Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate Progress

arXiv1 repo

arXiv:2206.01626

cleanrl

A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

arXiv1 repo

arXiv:2206.01718

sa2va_eval

Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised Learning

arXiv1 repo

arXiv:2206.02647

HistoSSLscaling

arXiv:2206.02675

arXiv1 repo

arXiv:2206.02675

HEBO

Tutel: Adaptive Mixture-of-Experts at Scale

arXiv1 repo

arXiv:2206.03382

Swin-Transformer

JuMP 1.0: Recent improvements to a modeling language for mathematical optimization

arXiv1 repo

arXiv:2206.03866

JuMP.jl

Sparse Mixture-of-Experts are Domain Generalizable Learners

arXiv1 repo

arXiv:2206.04046

MoA

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

arXiv1 repo

arXiv:2206.04615

BIG-bench

Factuality Enhanced Language Models for Open-Ended Text Generation

arXiv1 repo

arXiv:2206.04624

FasterTransformer

On Data Scaling in Masked Image Modeling

arXiv1 repo

arXiv:2206.04664

Swin-Transformer

SparseNeuS: Fast Generalizable Neural Surface Reconstruction from Sparse Views

arXiv1 repo

arXiv:2206.05737

torchsparse

The YiTrans End-to-End Speech Translation System for IWSLT 2022 Offline Shared Task

arXiv1 repo

arXiv:2206.05777

SpeechT5

GLIPv2: Unifying Localization and Vision-Language Understanding

arXiv1 repo

arXiv:2206.05836

GLIP

Semantic-Discriminative Mixup for Generalizable Sensor-based Cross-domain Activity Recognition

arXiv1 repo

arXiv:2206.06629

robustlearn

Aeneas: Rust Verification by Functional Translation

arXiv1 repo

arXiv:2206.07185

aeneas

arXiv:2206.07293

arXiv1 repo

arXiv:2206.07293

TIL-2023

PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change

arXiv1 repo

arXiv:2206.10498

LLMs-Planning

Questions Are All You Need to Train a Dense Passage Retriever

arXiv1 repo

arXiv:2206.10658

art

ProGen2: Exploring the Boundaries of Protein Language Models

arXiv1 repo

arXiv:2206.13517

jaxformer

Feature Refinement to Improve High Resolution Image Inpainting

arXiv1 repo

arXiv:2206.13644

lama

TweetNLP: Cutting-Edge Natural Language Processing for Social Media

arXiv1 repo

arXiv:2206.14774

tweetnlp

Solving Quantitative Reasoning Problems with Language Models

arXiv1 repo

arXiv:2206.14858

RLPR-Evaluation

Dissecting Self-Supervised Learning Methods for Surgical Computer Vision

arXiv1 repo

arXiv:2207.00449

Endoscapes

An Efficiency Study for SPLADE Models

arXiv1 repo

arXiv:2207.03834

LLM_Web_search

Improving Entity Disambiguation by Reasoning over a Knowledge Base

arXiv1 repo

arXiv:2207.04106

trusted_ke

ReFinED: An Efficient Zero-shot-capable Approach to End-to-End Entity Linking

arXiv1 repo

arXiv:2207.04108

trusted_ke

arXiv:2207.04296

arXiv1 repo

arXiv:2207.04296

web-stable-diffusion

A Comparative Study of Self-supervised Speech Representation Based Voice Conversion

arXiv1 repo

arXiv:2207.04356

s3prl-vc

Embedding Recycling for Language Models

arXiv1 repo

arXiv:2207.04993

EmbeddingRecycling

Next-ViT: Next Generation Vision Transformer for Efficient Deployment in Realistic Industrial Scenarios

arXiv1 repo

arXiv:2207.05501

MiDaS

OSLAT: Open Set Label Attention Transformer for Medical Entity Retrieval and Span Extraction

arXiv1 repo

arXiv:2207.05817

curai-research

ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech

arXiv1 repo

arXiv:2207.06389

Make-An-Audio

Confident Adaptive Language Modeling

arXiv1 repo

arXiv:2207.07061

mixture_of_recursions

Parameter-Efficient Prompt Tuning Makes Generalized and Calibrated Neural Text Retrievers

arXiv1 repo

arXiv:2207.07087

P-tuning-v2

CoqQ: Foundational Verification of Quantum Programs

arXiv1 repo

arXiv:2207.11350

analysis

Towards Complex Document Understanding By Discrete Reasoning

arXiv1 repo

arXiv:2207.11871

TAT-DQA

Finding smart contract vulnerabilities with ConCert's property-based testing framework

arXiv1 repo

arXiv:2208.00758

ConCert

Prompt Tuning for Generative Multimodal Pretrained Models

arXiv1 repo

arXiv:2208.02532

OFA

Investigating Efficiently Extending Transformers for Long Input Summarization

arXiv1 repo

arXiv:2208.04347

pegasus-x-base-synthsumm_open-16k

Exploring Hate Speech Detection with HateXplain and BERT

arXiv1 repo

arXiv:2208.04489

DeepLearningProject

TotalSegmentator: robust segmentation of 104 anatomical structures in CT images

arXiv1 repo

arXiv:2208.05868

TotalSegmentator

Perspective Reconstruction of Human Faces by Joint Mesh and Landmark Regression

arXiv1 repo

arXiv:2208.07142

insightface

Domain-Specific Risk Minimization for Out-of-Distribution Generalization

arXiv1 repo

arXiv:2208.08661

robustlearn

Flat Multi-modal Interaction Transformer for Named Entity Recognition

arXiv1 repo

arXiv:2208.11039

Fengshenbang-LM

Automatic music mixing with deep learning and out-of-domain data

arXiv1 repo

arXiv:2208.11428

FxNorm-automix

DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

arXiv1 repo

arXiv:2208.12242

dreambooth

Measure Construction by Extension in Dependent Type Theory with Application to Integration

arXiv1 repo

arXiv:2209.02345

analysis

What does a platypus look like? Generating customized prompts for zero-shot image classification

arXiv1 repo

arXiv:2209.03320

CLIP_benchmark

FP8 Formats for Deep Learning

arXiv1 repo

arXiv:2209.05433

ao

Pre-trained Language Models for the Legal Domain: A Case Study on Indian Law

arXiv1 repo

arXiv:2209.06049

InLegalBERT

arXiv:2209.06995

arXiv1 repo

arXiv:2209.06995

Patron

Out-of-Distribution Representation Learning for Time Series Classification

arXiv1 repo

arXiv:2209.07027

robustlearn

Monolith: Real Time Recommendation System With Collisionless Embedding Table

arXiv1 repo

arXiv:2209.07663

SIMURG

LAVIS: A Library for Language-Vision Intelligence

arXiv1 repo

arXiv:2209.09019

LAVIS

Operationalizing Machine Learning: An Interview Study

arXiv1 repo

arXiv:2209.09125

dtu_mlops

Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

arXiv1 repo

arXiv:2209.09513

ScienceQA

Generate rather than Retrieve: Large Language Models are Strong Context Generators

arXiv1 repo

arXiv:2209.10063

Ensemble-of-Retrievers

Efficient Few-Shot Learning Without Prompts

arXiv1 repo

arXiv:2209.11055

ember-v1

OLIVES Dataset: Ophthalmic Labels for Investigating Visual Eye Semantics

arXiv1 repo

arXiv:2209.11195

OLIVES_Dataset

All are Worth Words: A ViT Backbone for Diffusion Models

arXiv1 repo

arXiv:2209.12152

U-ViT

T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition

arXiv1 repo

arXiv:2209.12616

tner

WikiDes: A Wikipedia-Based Dataset for Generating Short Descriptions from Paragraphs

arXiv1 repo

arXiv:2209.13101

WikiDes

A Benchmark Comparison of Python Malware Detection Approaches

arXiv1 repo

arXiv:2209.13288

guarddog

COLO: A Contrastive Learning based Re-ranking Framework for One-Stage Summarization

arXiv1 repo

arXiv:2209.14569

TestCoLo

SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data

arXiv1 repo

arXiv:2209.15329

SpeechT5

Music-to-Text Synaesthesia: Generating Descriptive Text from Music Recordings

arXiv1 repo

arXiv:2210.00434

emotion-english-distilroberta-base

Improving Sample Quality of Diffusion Models Using Self-Attention Guidance

arXiv1 repo

arXiv:2210.00939

Fooocus

Omnigrok: Grokking Beyond Algorithmic Data

arXiv1 repo

arXiv:2210.01117

grokfast

Explaining Patterns in Data with Language Models via Interpretable Autoprompting

arXiv1 repo

arXiv:2210.01848

imodelsX

TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis

arXiv1 repo

arXiv:2210.02186

frn-50k-baseline

Decomposed Prompting: A Modular Approach for Solving Complex Tasks

arXiv1 repo

arXiv:2210.02406

DecomP-ODQA

arXiv:2210.02410

arXiv1 repo

arXiv:2210.02410

Vendi-Score

arXiv:2210.02437

arXiv1 repo

arXiv:2210.02437

MTP-Codebase

Binding Language Models in Symbolic Languages

arXiv1 repo

arXiv:2210.02875

blendsql

On Distillation of Guided Diffusion Models

arXiv1 repo

arXiv:2210.03142

i2p

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

arXiv1 repo

arXiv:2210.03347

pix2struct-large

Measuring and Narrowing the Compositionality Gap in Language Models

arXiv1 repo

arXiv:2210.03350

self-ask

SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-training

arXiv1 repo

arXiv:2210.03730

SpeechT5

ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering

arXiv1 repo

arXiv:2210.03849

ConvFinQA

DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability

arXiv1 repo

arXiv:2210.05148

DiffRoll

Foundation Transformers

arXiv1 repo

arXiv:2210.06423

Kosmos-X

InfoCSE: Information-aggregated Contrastive Learning of Sentence Embeddings

arXiv1 repo

arXiv:2210.06432

sentemb

arXiv:2210.06886

arXiv1 repo

arXiv:2210.06886

ImaginaryNet

arXiv:2210.07229

arXiv1 repo

arXiv:2210.07229

KEditVis-LLM-Editing

Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training

arXiv1 repo

arXiv:2210.08773

LAVIS

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

arXiv1 repo

arXiv:2210.09261

bbh

Revision Transformers: Instructing Language Models to Change their Values

arXiv1 repo

arXiv:2210.10332

Revision-Transformer

TabLLM: Few-shot Classification of Tabular Data with Large Language Models

arXiv1 repo

arXiv:2210.10723

TabLLM

Improving Chinese Spelling Check by Character Pronunciation Prediction: The Effects of Adaptivity and Granularity

arXiv1 repo

arXiv:2210.10996

ChineseBERT-for-csc

CEFR-Based Sentence Difficulty Annotation and Assessment

arXiv1 repo

arXiv:2210.11766

CEFR-SP

An Analysis of Fusion Functions for Hybrid Retrieval

arXiv1 repo

arXiv:2210.11934

KoPrivateGPT

arXiv:2210.12213

arXiv1 repo

arXiv:2210.12213

spabert

Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs

arXiv1 repo

arXiv:2210.12283

LeanTool

BEANS: The Benchmark of Animal Sounds

arXiv1 repo

arXiv:2210.12300

BEANS-Zero

Bootstrapping meaning through listening: Unsupervised learning of spoken sentence embeddings

arXiv1 repo

arXiv:2210.12857

spoken_sent_embedding

NVIDIA FLARE: Federated Learning from Simulation to Real-World

arXiv1 repo

arXiv:2210.13291

NVFlare

EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones

arXiv1 repo

arXiv:2210.14090

moshi

arXiv:2210.14648

arXiv1 repo

arXiv:2210.14648

TIL-2023

Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

arXiv1 repo

arXiv:2210.14712

Glot500

DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

arXiv1 repo

arXiv:2210.14896

diffusiondb

Truncation Sampling as Language Model Desmoothing

arXiv1 repo

arXiv:2210.15191

LLM-Sampling

Wespeaker: A Research and Production oriented Speaker Embedding Learning Toolkit

arXiv1 repo

arXiv:2210.17016

wespeaker

Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

arXiv1 repo

arXiv:2210.17027

SpeechT5

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

arXiv1 repo

arXiv:2211.00593

ARENA_2.0

Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022

arXiv1 repo

arXiv:2211.00815

wespeaker

DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

arXiv1 repo

arXiv:2211.01095

TCD-SDXL-LoRA

Two-Stream Network for Sign Language Recognition and Translation

arXiv1 repo

arXiv:2211.01367

SLRT

MPCFormer: fast, performant and private Transformer inference with MPC

arXiv1 repo

arXiv:2211.01452

MPCFormer

Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model

arXiv1 repo

arXiv:2211.02001

bloom

Multi-Head Adapter Routing for Cross-Task Generalization

arXiv1 repo

arXiv:2211.03831

mttl

Will we run out of data? Limits of LLM scaling based on human-generated data

arXiv1 repo

arXiv:2211.04325

gigatoken

InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions

arXiv1 repo

arXiv:2211.05778

InternImage

A Novel Sampling Scheme for Text- and Image-Conditional Image Synthesis in Quantized Latent Spaces

arXiv1 repo

arXiv:2211.07292

i2p

Towards a Mathematics Formalisation Assistant using Large Language Models

arXiv1 repo

arXiv:2211.07524

LeanAide

EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

arXiv1 repo

arXiv:2211.07636

EVA

UniRel: Unified Representation and Interaction for Joint Relational Triple Extraction

arXiv1 repo

arXiv:2211.09039

trusted_ke

Holistic Evaluation of Language Models

arXiv1 repo

arXiv:2211.09110

lmms-eval

Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks

arXiv1 repo

arXiv:2211.09808

InternImage

CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval

arXiv1 repo

arXiv:2211.10411

dpr-scale

PAL: Program-aided Language Models

arXiv1 repo

arXiv:2211.10435

gsm-hard

BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision

arXiv1 repo

arXiv:2211.10439

InternImage

CryptOpt: Verified Compilation with Randomized Program Search for Cryptographic Primitives (full version)

arXiv1 repo

arXiv:2211.10665

CryptOpt

An Empirical Study On Contrastive Search And Contrastive Decoding For Open-ended Text Generation

arXiv1 repo

arXiv:2211.10797

Adaptive-Contrastive-Search

You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model

arXiv1 repo

arXiv:2211.11152

OFA

L3Cube-MahaSBERT and HindSBERT: Sentence BERT Models and Benchmarking BERT Sentence Representations for Hindi and Marathi

arXiv1 repo

arXiv:2211.11187

bengali-sentence-similarity-sbert

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

arXiv1 repo

arXiv:2211.11256

UniMSE

VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

arXiv1 repo

arXiv:2211.11275

SpeechT5

Visual Programming: Compositional visual reasoning without training

arXiv1 repo

arXiv:2211.11559

visprog

Coreference Resolution through a seq2seq Transition-Based System

arXiv1 repo

arXiv:2211.12142

trusted_ke

Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

arXiv1 repo

arXiv:2211.12368

RAD-NeRF

EDICT: Exact Diffusion Inversion via Coupled Transformations

arXiv1 repo

arXiv:2211.12446

DOODL

G^3: Geolocation via Guidebook Grounding

arXiv1 repo

arXiv:2211.15521

diff-mining

Connecting the Dots: Floorplan Reconstruction Using Two-Level Queries

arXiv1 repo

arXiv:2211.15658

RoomFormer

On Word Error Rate Definitions and their Efficient Computation for Multi-Speaker Speech Recognition Systems

arXiv1 repo

arXiv:2211.16112

meeteval

NeuralLift-360: Lifting An In-the-wild 2D Photo to A 3D Object with 360° Views

arXiv1 repo

arXiv:2211.16431

NeuralLift-360

Rationale-Guided Few-Shot Classification to Detect Abusive Language

arXiv1 repo

arXiv:2211.17046

Rationale_predictor

Rethinking Causality-driven Robot Tool Segmentation with Temporal Constraints

arXiv1 repo

arXiv:2212.00072

CaRTS

MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition

arXiv1 repo

arXiv:2212.00500

OFA

Scaling Language-Image Pre-training via Masking

arXiv1 repo

arXiv:2212.00794

Chinese-CLIP

Moving Beyond Downstream Task Accuracy for Information Retrieval Benchmarking

arXiv1 repo

arXiv:2212.01340

ColBERT

Melody transcription via generative pre-training

arXiv1 repo

arXiv:2212.01884

SheetSage2

One-shot Implicit Animatable Avatars with Model-based Priors

arXiv1 repo

arXiv:2212.02469

ELICIT

3DGazeNet: Generalizing Gaze Estimation with Weak-Supervision from Synthetic Views

arXiv1 repo

arXiv:2212.02997

insightface

NeRDi: Single-View NeRF Synthesis with Language-Guided Diffusion as General Image Priors

arXiv1 repo

arXiv:2212.03267

zero123

Discovering Latent Knowledge in Language Models Without Supervision

arXiv1 repo

arXiv:2212.03827

language_exploration

Latent Graph Representations for Critical View of Safety Assessment

arXiv1 repo

arXiv:2212.04155

Endoscapes

Diffusion Guided Domain Adaptation of Image Generators

arXiv1 repo

arXiv:2212.04473

styleganfusion

Multi-Concept Customization of Text-to-Image Diffusion

arXiv1 repo

arXiv:2212.04488

custom-diffusion

Learning Video Representations from Large Language Models

arXiv1 repo

arXiv:2212.04501

VideoTree

Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

arXiv1 repo

arXiv:2212.05032

Structured-Diffusion-Guidance

Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints

arXiv1 repo

arXiv:2212.05055

smolMoELM-custom

Transcoding Unicode Characters with AVX-512 Instructions

arXiv1 repo

arXiv:2212.05098

simdutf

MAGVIT: Masked Generative Video Transformer

arXiv1 repo

arXiv:2212.05199

magvit

How to Backdoor Diffusion Models?

arXiv1 repo

arXiv:2212.05400

BadDiffusion

Unfolding Local Growth Rate Estimates for (Almost) Perfect Adversarial Detection

arXiv1 repo

arXiv:2212.06776

deepfake_multiLID

SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning

arXiv1 repo

arXiv:2212.07489

smacv2

Attention as a Guide for Simultaneous Speech Translation

arXiv1 repo

arXiv:2212.07850

naist-simulst

Constitutional AI: Harmlessness from AI Feedback

arXiv1 repo

arXiv:2212.08073

minihf

Transferring General Multimodal Pretrained Models to Text Recognition

arXiv1 repo

arXiv:2212.09297

OFA

Rethinking Label Smoothing on Multi-hop Question Answering

arXiv1 repo

arXiv:2212.09512

Smoothing-R3

Visconde: Multi-document QA with GPT-3 and Neural Reranking

arXiv1 repo

arXiv:2212.09656

KoPrivateGPT

Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model

arXiv1 repo

arXiv:2212.09811

nllb-pruning

Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

arXiv1 repo

arXiv:2212.10509

ircot

From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models

arXiv1 repo

arXiv:2212.10846

LAVIS

Training language models to summarize narratives improves brain alignment

arXiv1 repo

arXiv:2212.10898

brain_language_summarization

OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization

arXiv1 repo

arXiv:2212.12017

opt-iml-max-1.3b

MAUVE Scores for Generative Models: Theory and Practice

arXiv1 repo

arXiv:2212.14578

ssharoff.github.io

Argoverse 2: Next Generation Datasets for Self-Driving Perception and Forecasting

arXiv1 repo

arXiv:2301.00493

torchsparse

ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders

arXiv1 repo

arXiv:2301.00808

ConvNeXt-V2

Iterated Decomposition: Improving Science Q&A by Supervising Reasoning Processes

arXiv1 repo

arXiv:2301.01751

ice

InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval

arXiv1 repo

arXiv:2301.01820

InPars

Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU

arXiv1 repo

arXiv:2301.03598

flashinfer

Learning to Exploit Temporal Structure for Biomedical Vision-Language Processing

arXiv1 repo

arXiv:2301.04558

BiomedVLP-BioViL-T

Progress measures for grokking via mechanistic interpretability

arXiv1 repo

arXiv:2301.05217

TransformerLens

Domain Expansion of Image Generators

arXiv1 repo

arXiv:2301.05225

ControlNet

GLIGEN: Open-Set Grounded Text-to-Image Generation

arXiv1 repo

arXiv:2301.07093

GLIGEN

How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection

arXiv1 repo

arXiv:2301.07597

humanizer-skill

Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

arXiv1 repo

arXiv:2301.08243

xwm

Blind Spots: Automatically detecting ignored program inputs

arXiv1 repo

arXiv:2301.08700

publications

Set-Theoretic and Type-Theoretic Ordinals Coincide

arXiv1 repo

arXiv:2301.10696

TypeTopology

DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature

arXiv1 repo

arXiv:2301.11305

AdaDetectGPT

3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion Models

arXiv1 repo

arXiv:2301.11445

3DShape2VecSet

SEGA: Instructing Text-to-Image Models using Semantic Guidance

arXiv1 repo

arXiv:2301.12247

ControlNet

Domain Theory in Constructive and Predicative Univalent Foundations

arXiv1 repo

arXiv:2301.12405

TypeTopology

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

arXiv1 repo

arXiv:2301.12661

Make-An-Audio

Edge-guided Multi-domain RGB-to-TIR image Translation for Training Vision Tasks with Challenging Labels

arXiv1 repo

arXiv:2301.12689

sRGB-TIR

SAN: Inducing Metrizability of GAN with Discriminative Normalized Linear Layer

arXiv1 repo

arXiv:2301.12811

san

arXiv:2301.12844

arXiv1 repo

arXiv:2301.12844

HEBO

Icicle: A Re-Designed Emulator for Grey-Box Firmware Fuzzing

arXiv1 repo

arXiv:2301.13346

awesome-connected-things-sec

GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

arXiv1 repo

arXiv:2301.13430

GeneFace

In-Context Retrieval-Augmented Language Models

arXiv1 repo

arXiv:2302.00083

TinyRAG

Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization

arXiv1 repo

arXiv:2302.00275

StreetCLIP

Mixture of Diffusers for scene composition and high resolution image generation

arXiv1 repo

arXiv:2302.02412

mixture-of-diffusers

Colossal-Auto: Unified Automation of Parallelization and Activation Checkpoint for Large-scale Models

arXiv1 repo

arXiv:2302.02599

ColossalAI

Structure and Content-Guided Video Synthesis with Diffusion Models

arXiv1 repo

arXiv:2302.03011

videophy

A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations

arXiv1 repo

arXiv:2302.03025

TransformerLens

Zero-shot Image-to-Image Translation

arXiv1 repo

arXiv:2302.03027

pix2pix-zero

Exploring the Benefits of Training Expert Language Models over Instruction Tuning

arXiv1 repo

arXiv:2302.03202

ELM

Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

arXiv1 repo

arXiv:2302.03540

whisperspeech

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity

arXiv1 repo

arXiv:2302.04023

Julia_bench

GPTScore: Evaluate as You Desire

arXiv1 repo

arXiv:2302.04166

GPTScore

Will ChatGPT get you caught? Rethinking of Plagiarism Detection

arXiv1 repo

arXiv:2302.04335

verify-ai

MaskSketch: Unpaired Structure-guided Masked Image Generation

arXiv1 repo

arXiv:2302.05496

ControlNet

Level Generation Through Large Language Models

arXiv1 repo

arXiv:2302.05817

lm-pcg

Reveal the Unknown: Out-of-Knowledge-Base Mention Discovery with Entity Linking

arXiv1 repo

arXiv:2302.07189

LM-ontology-concept-placement

Cliff-Learning

arXiv1 repo

arXiv:2302.07348

scaling

How to Train Your DRAGON: Diverse Augmentation Towards Generalizable Dense Retrieval

arXiv1 repo

arXiv:2302.07452

dpr-scale

Slapo: A Schedule Language for Progressive Optimization of Large Deep Learning Model Training

arXiv1 repo

arXiv:2302.08005

slapo

Composer: Creative and Controllable Image Synthesis with Composable Conditions

arXiv1 repo

arXiv:2302.09778

videocomposer

Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels

arXiv1 repo

arXiv:2302.10586

U-ViT

FiNER-ORD: Financial Named Entity Recognition Open Research Dataset

arXiv1 repo

arXiv:2302.11157

flare-finer-ord

Guiding Large Language Models via Directional Stimulus Prompting

arXiv1 repo

arXiv:2302.11520

Directional-Stimulus-Prompting

Encoder-based Domain Tuning for Fast Personalization of Text-to-Image Models

arXiv1 repo

arXiv:2302.12228

e4t-diffusion

Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple Views

arXiv1 repo

arXiv:2302.12297

temporal-robustness

VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining

arXiv1 repo

arXiv:2302.12584

VivesDebate-Speech

FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning

arXiv1 repo

arXiv:2302.13485

robustlearn

Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolation

arXiv1 repo

arXiv:2303.00440

EMA-VFI

Do Machine Learning Models Learn Statistical Rules Inferred from Data?

arXiv1 repo

arXiv:2303.01433

sqrl

Chasing Low-Carbon Electricity for Practical and Sustainable DNN Training

arXiv1 repo

arXiv:2303.02508

zeus

Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes

arXiv1 repo

arXiv:2303.02760

HumanArt

Evaluating Parameter-Efficient Transfer Learning Approaches on SURE Benchmark for Speech Understanding

arXiv1 repo

arXiv:2303.03267

speech-adapters

CoTEVer: Chain of Thought Prompting Annotation Toolkit for Explanation Verification

arXiv1 repo

arXiv:2303.03628

CoTEVer

Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

arXiv1 repo

arXiv:2303.03926

SpeechT5

Scaling up GANs for Text-to-Image Synthesis

arXiv1 repo

arXiv:2303.05511

GigaGAN

ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation

arXiv1 repo

arXiv:2303.06458

CLFM

Query2doc: Query Expansion with Large Language Models

arXiv1 repo

arXiv:2303.07678

SPLADERunner

A Simple Framework for Open-Vocabulary Segmentation and Detection

arXiv1 repo

arXiv:2303.08131

DINO

SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference

arXiv1 repo

arXiv:2303.08308

Moonlit

VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation

arXiv1 repo

arXiv:2303.08320

MuseV

PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs

arXiv1 repo

arXiv:2303.08954

presto

TriAAN-VC: Triple Adaptive Attention Normalization for Any-to-Any Voice Conversion

arXiv1 repo

arXiv:2303.09057

TriAAN-VC

BanglaCoNER: Towards Robust Bangla Complex Named Entity Recognition

arXiv1 repo

arXiv:2303.09306

Bangla-Complex-Named-Entity-Recognition-Challenge

ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices

arXiv1 repo

arXiv:2303.09730

Moonlit

On the rise of fear speech in online social media

arXiv1 repo

arXiv:2303.10311

Fearspeech-project

AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

arXiv1 repo

arXiv:2303.10512

Continual-NExT

EVA-02: A Visual Representation for Neon Genesis

arXiv1 repo

arXiv:2303.11331

eva02_large_patch14_448.mim_m38m_ft_in1k

Reflexion: Language Agents with Verbal Reinforcement Learning

arXiv1 repo

arXiv:2303.11366

ReWOO

Stable Bias: Analyzing Societal Representations in Diffusion Models

arXiv1 repo

arXiv:2303.11408

BDM1.0

arXiv:2303.11866

arXiv1 repo

arXiv:2303.11866

LilT

Natural Language-Assisted Sign Language Recognition

arXiv1 repo

arXiv:2303.12080

SLRT

Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions

arXiv1 repo

arXiv:2303.12789

instruct-nerf2nerf

CiCo: Domain-Aware Sign Language Retrieval via Cross-Lingual Contrastive Learning

arXiv1 repo

arXiv:2303.12793

SLRT

Visual-Language Prompt Tuning with Knowledge-guided Context Optimization

arXiv1 repo

arXiv:2303.13283

BiomedCoOp

Xplainer: From X-Ray Observations to Explainable Zero-Shot Diagnosis

arXiv1 repo

arXiv:2303.13391

Fine-Grained_Features_Alignment_via_Constrastive_Learning

End-to-End Diffusion Latent Optimization Improves Classifier Guidance

arXiv1 repo

arXiv:2303.13703

DOODL

Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation

arXiv1 repo

arXiv:2303.13873

Fantasia3D

Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion Prior

arXiv1 repo

arXiv:2303.14184

Make-It-3D

Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

arXiv1 repo

arXiv:2303.14307

Llama-AVSR

Human Preference Score: Better Aligning Text-to-Image Models with Human Preference

arXiv1 repo

arXiv:2303.14420

align_sd

ChatGPT Outperforms Crowd-Workers for Text-Annotation Tasks

arXiv1 repo

arXiv:2303.15056

autolabel

Fine-grained Audible Video Description

arXiv1 repo

arXiv:2303.15616

FAVDBench

A Pilot Study of Query-Free Adversarial Attack against Stable Diffusion

arXiv1 repo

arXiv:2303.16378

Fine-Grained_Features_Alignment_via_Constrastive_Learning

Hierarchical Video-Moment Retrieval and Step-Captioning

arXiv1 repo

arXiv:2303.16406

TimeChat-Online-139K

Fairlearn: Assessing and Improving Fairness of AI Systems

arXiv1 repo

arXiv:2303.16626

lares

AnnoLLM: Making Large Language Models to Be Better Crowdsourced Annotators

arXiv1 repo

arXiv:2303.16854

autolabel

DERA: Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents

arXiv1 repo

arXiv:2303.17071

curai-research

arXiv:2303.17602

arXiv1 repo

arXiv:2303.17602

TIL-2023

Token Merging for Fast Stable Diffusion

arXiv1 repo

arXiv:2303.17604

Text2Video-Zero

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

arXiv1 repo

arXiv:2303.17760

camel

Evaluating GPT-4 and ChatGPT on Japanese Medical Licensing Examinations

arXiv1 repo

arXiv:2303.18027

JMedBench

ViMMRC 2.0 -- Enhancing Machine Reading Comprehension on Vietnamese Literature Text

arXiv1 repo

arXiv:2303.18162

ViQG

Assessing Language Model Deployment with Risk Cards

arXiv1 repo

arXiv:2303.18190

garak

arXiv:2304.01715

arXiv1 repo

arXiv:2304.01715

LVVIS

AToMiC: An Image/Text Retrieval Test Collection to Support Multimedia Content Creation

arXiv1 repo

arXiv:2304.01961

wit

Dialogue-Contextualized Re-ranking for Medical History-Taking

arXiv1 repo

arXiv:2304.01974

curai-research

DoUnseen: Tuning-Free Class-Adaptive Object Detection of Unseen Objects for Robotic Grasping

arXiv1 repo

arXiv:2304.02833

image_agnostic_segmentation

arXiv:2304.02970

arXiv1 repo

arXiv:2304.02970

Echo-ViLD

Instruction Tuning with GPT-4

arXiv1 repo

arXiv:2304.03277

Instruct-SkillMix-SDD

Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

arXiv1 repo

arXiv:2304.03279

machiavelli

Decoder-Only or Encoder-Decoder? Interpreting Language Model as a Regularized Encoder-Decoder

arXiv1 repo

arXiv:2304.04052

Partial-Attention-Language-Model

OpenAGI: When LLM Meets Domain Experts

arXiv1 repo

arXiv:2304.04370

OpenAGI

VARS: Video Assistant Referee System for Automated Soccer Decision Making from Multiple Views

arXiv1 repo

arXiv:2304.04617

thesis_automatic_faul_recognition

On the Possibilities of AI-Generated Text Detection

arXiv1 repo

arXiv:2304.04736

verify-ai

Multi-step Jailbreaking Privacy Attacks on ChatGPT

arXiv1 repo

arXiv:2304.05197

LLM-Multistep-Jailbreak

Improving Diffusion Models for Scene Text Editing with Dual Encoders

arXiv1 repo

arXiv:2304.05568

DiffSTE

Unicom: Universal and Compact Representation Learning for Image Retrieval

arXiv1 repo

arXiv:2304.05884

unicom

DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion

arXiv1 repo

arXiv:2304.06025

DreamPose

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

arXiv1 repo

arXiv:2304.06364

XVERSE-7B

HuaTuo: Tuning LLaMA Model with Chinese Medical Knowledge

arXiv1 repo

arXiv:2304.06975

Huatuo-Llama-Med-Chinese

OpenAssistant Conversations -- Democratizing Large Language Model Alignment

arXiv1 repo

arXiv:2304.07327

oasst1

ArguGPT: evaluating, understanding and identifying argumentative essays generated by GPT models

arXiv1 repo

arXiv:2304.07666

verify-ai

Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation

arXiv1 repo

arXiv:2304.07854

BiLLa

DETRs Beat YOLOs on Real-time Object Detection

arXiv1 repo

arXiv:2304.08069

rt-detr-huggingface

MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

arXiv1 repo

arXiv:2304.08247

DiagnosisCoding

The MiniPile Challenge for Data-Efficient Language Models

arXiv1 repo

arXiv:2304.08442

minipile

MasaCtrl: Tuning-Free Mutual Self-Attention Control for Consistent Image Synthesis and Editing

arXiv1 repo

arXiv:2304.08465

MasaCtrl

Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models

arXiv1 repo

arXiv:2304.08818

NeuroClips

Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents

arXiv1 repo

arXiv:2304.09542

RankGPT

Anything-3D: Towards Single-view Anything Reconstruction in the Wild

arXiv1 repo

arXiv:2304.10261

Anything-3D

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

arXiv1 repo

arXiv:2304.10592

bilingual-gpt-neox-4b-minigpt4

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

arXiv1 repo

arXiv:2304.11277

MindSpeed-MM

L3Cube-IndicSBERT: A simple approach for learning cross-lingual sentence representations using multilingual BERT

arXiv1 repo

arXiv:2304.11434

bengali-sentence-similarity-sbert

Unlocking Context Constraints of LLMs: Enhancing Context Efficiency of LLMs with Self-Information-Based Content Filtering

arXiv1 repo

arXiv:2304.12102

Selective_Context

Segment Anything in Medical Images

arXiv1 repo

arXiv:2304.12306

SAMReg

A Static Pruning Study on Sparse Neural Retrievers

arXiv1 repo

arXiv:2304.12702

splade

Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

arXiv1 repo

arXiv:2304.13731

tango

Customized Segment Anything Model for Medical Image Segmentation

arXiv1 repo

arXiv:2304.13785

TriALS

DataComp: In search of the next generation of multimodal datasets

arXiv1 repo

arXiv:2304.14108

datacomp

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

arXiv1 repo

arXiv:2304.14178

A-Lightweight-Unified-Autoregressive-MLLM

PMC-LLaMA: Towards Building Open-source Language Models for Medicine

arXiv1 repo

arXiv:2304.14454

PMC-LLaMA

MinMaxLTTB: Leveraging MinMax-Preselection to Scale LTTB

arXiv1 repo

arXiv:2305.00332

flot-downsample

The Art of the Fugue: Minimizing Interleaving in Collaborative Text Editing

arXiv1 repo

arXiv:2305.00583

loro

Self-Evaluation Guided Beam Search for Reasoning

arXiv1 repo

arXiv:2305.00633

minihf

TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion Synthesis

arXiv1 repo

arXiv:2305.00976

TMR-SOMA-RP-v1

VPGTrans: Transfer Visual Prompt Generator across LLMs

arXiv1 repo

arXiv:2305.01278

A-Lightweight-Unified-Autoregressive-MLLM

Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation

arXiv1 repo

arXiv:2305.01569

banana100-additional-iqa-models

Finding Neurons in a Haystack: Case Studies with Sparse Probing

arXiv1 repo

arXiv:2305.01610

TransformerLens

CodeGen2: Lessons for Training LLMs on Programming and Natural Languages

arXiv1 repo

arXiv:2305.02309

CodeGen

Caption Anything: Interactive Image Description with Diverse Multimodal Controls

arXiv1 repo

arXiv:2305.02677

Caption-Anything

HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

arXiv1 repo

arXiv:2305.02765

AcademiCodec

Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision

arXiv1 repo

arXiv:2305.03047

minihf

Personalize Segment Anything Model with One Shot

arXiv1 repo

arXiv:2305.03048

geti-instant-learn

Automatic Prompt Optimization with "Gradient Descent" and Beam Search

arXiv1 repo

arXiv:2305.03495

PRL-Prompts-from-Reinforcement-Learning

Exploring One-shot Semi-supervised Federated Learning with A Pre-trained Diffusion Model

arXiv1 repo

arXiv:2305.04063

FedDISC

Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

arXiv1 repo

arXiv:2305.04091

gpt-researcher

Flex-SFU: Accelerating DNN Activation Functions by Non-Uniform Piecewise Approximation

arXiv1 repo

arXiv:2305.04546

flex-sfu

WikiWeb2M: A Page-Level Multimodal Wikipedia Dataset

arXiv1 repo

arXiv:2305.05432

wit

Vision-Language Models in Remote Sensing: Current Progress and Future Trends

arXiv1 repo

arXiv:2305.05726

VRSBench

VideoChat: Chat-Centric Video Understanding

arXiv1 repo

arXiv:2305.06355

Ask-Anything

Active Retrieval Augmented Generation

arXiv1 repo

arXiv:2305.06983

FLARE

Exploiting Diffusion Prior for Real-World Image Super-Resolution

arXiv1 repo

arXiv:2305.07015

StableSR-TestSets

OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

arXiv1 repo

arXiv:2305.07895

OCRBench

Multilingual Previously Fact-Checked Claim Retrieval

arXiv1 repo

arXiv:2305.07991

claim-retrival

Instance-Aware Repeat Factor Sampling for Long-Tailed Object Detection

arXiv1 repo

arXiv:2305.08069

camie-tagger-v2

arXiv:2305.08227

arXiv1 repo

arXiv:2305.08227

TIL-2023

From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

arXiv1 repo

arXiv:2305.08283

modular_pluralism

Large Language Model Guided Tree-of-Thought

arXiv1 repo

arXiv:2305.08291

tree-of-thought-prompting

Document Understanding Dataset and Evaluation (DUDE)

arXiv1 repo

arXiv:2305.08455

sa2va_eval

Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts

arXiv1 repo

arXiv:2305.08850

Make-A-Protagonist

SatLM: Satisfiability-Aided Language Models Using Declarative Prompting

arXiv1 repo

arXiv:2305.09656

aurora-m2

Explaining black box text modules in natural language with language models

arXiv1 repo

arXiv:2305.09863

imodelsX

arXiv:2305.10037

arXiv1 repo

arXiv:2305.10037

model_swarm

UniEX: An Effective and Efficient Framework for Unified Information Extraction via a Span-extractive Perspective

arXiv1 repo

arXiv:2305.10306

Fengshenbang-LM

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

arXiv1 repo

arXiv:2305.10425

pair-preference-model-LLaMA3-8B

Statistical Knowledge Assessment for Large Language Models

arXiv1 repo

arXiv:2305.10519

DistFactAssessLM

Structural Pruning for Diffusion Models

arXiv1 repo

arXiv:2305.10924

Diff-Pruning

SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

arXiv1 repo

arXiv:2305.11000

SpeechGPT

Going Denser with Open-Vocabulary Part Segmentation

arXiv1 repo

arXiv:2305.11173

grounded-segment-any-parts

VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

arXiv1 repo

arXiv:2305.11175

VisionLLM

Self-QA: Unsupervised Knowledge Guided Language Model Alignment

arXiv1 repo

arXiv:2305.11952

XuanYuan

XuanYuan 2.0: A Large Chinese Financial Chat Model with Hundreds of Billions Parameters

arXiv1 repo

arXiv:2305.12002

XuanYuan

Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

arXiv1 repo

arXiv:2305.12031

med42

Movie101: A New Movie Understanding Benchmark

arXiv1 repo

arXiv:2305.12140

TimeChat-Online-139K

Are Your Explanations Reliable? Investigating the Stability of LIME in Explaining Text Classifiers by Marrying XAI and Adversarial Attack

arXiv1 repo

arXiv:2305.12351

XAIFooler

Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

arXiv1 repo

arXiv:2305.13035

siglip-so400m-patch14-384

RWKV: Reinventing RNNs for the Transformer Era

arXiv1 repo

arXiv:2305.13048

rwkv

Making Language Models Better Tool Learners with Execution Feedback

arXiv1 repo

arXiv:2305.13068

KnowLM

Editing Large Language Models: Problems, Methods, and Opportunities

arXiv1 repo

arXiv:2305.13172

KnowEdit

RecurrentGPT: Interactive Generation of (Arbitrarily) Long Text

arXiv1 repo

arXiv:2305.13304

Recurrent-LLM

Matcher: Segment Anything with One Shot Using All-Purpose Feature Matching

arXiv1 repo

arXiv:2305.13310

geti-instant-learn

Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach

arXiv1 repo

arXiv:2305.13579

ProFusion

LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

arXiv1 repo

arXiv:2305.13655

Omost

Aligning Large Language Models through Synthetic Feedback

arXiv1 repo

arXiv:2305.13735

minihf

Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation

arXiv1 repo

arXiv:2305.14004

Saamayik

Parts of Speech-Grounded Subspaces in Vision-Language Models

arXiv1 repo

arXiv:2305.14053

PoS-subspaces

Language Models with Rationality

arXiv1 repo

arXiv:2305.14250

llm-mysteries

Linear Cross-Lingual Mapping of Sentence Embeddings

arXiv1 repo

arXiv:2305.14256

Wikinews-multilingual

Query Rewriting for Retrieval-Augmented Large Language Models

arXiv1 repo

arXiv:2305.14283

RAG-query-rewriting

Connecting Multi-modal Contrastive Representations

arXiv1 repo

arXiv:2305.14381

20251R0136COSE40500

CGCE: A Chinese Generative Chat Evaluation Benchmark for General and Financial Domains

arXiv1 repo

arXiv:2305.14471

XuanYuan

Optimal Linear Subspace Search: Learning to Construct Fast and High-Quality Schedulers for Diffusion Models

arXiv1 repo

arXiv:2305.14677

diffSynth-studio-notes

Investigating Table-to-Text Generation Capabilities of LLMs in Real-World Information Seeking Scenarios

arXiv1 repo

arXiv:2305.14987

LLM-T2T

Gorilla: Large Language Model Connected with Massive APIs

arXiv1 repo

arXiv:2305.15334

gorilla

arXiv:2305.15717

arXiv1 repo

arXiv:2305.15717

alpaca_eval

On the Planning Abilities of Large Language Models : A Critical Investigation

arXiv1 repo

arXiv:2305.15771

LLMs-Planning

arXiv:2305.15778

arXiv1 repo

arXiv:2305.15778

awesome-ai-sre

arXiv:2305.15930

arXiv1 repo

arXiv:2305.15930

HEBO

VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

arXiv1 repo

arXiv:2305.16107

SpeechT5

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

arXiv1 repo

arXiv:2305.16213

prolificdreamer

Landmark Attention: Random-Access Infinite Context Length for Transformers

arXiv1 repo

arXiv:2305.16300

landmark-attention

IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages

arXiv1 repo

arXiv:2305.16307

IndicTrans2

Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models

arXiv1 repo

arXiv:2305.16322

Uni-ControlNet

An Empirical Comparison of LM-based Question and Answer Generation Methods

arXiv1 repo

arXiv:2305.17002

lm-question-generation

BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks

arXiv1 repo

arXiv:2305.17100

DermFM-Zero

PromptNER: Prompt Locating and Typing for Named Entity Recognition

arXiv1 repo

arXiv:2305.17104

PromptNER

arXiv:2305.17216

arXiv1 repo

arXiv:2305.17216

CoBSAT

Fine-Tuning Language Models with Just Forward Passes

arXiv1 repo

arXiv:2305.17333

MeZO

DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text

arXiv1 repo

arXiv:2305.17359

AdaDetectGPT

A Practical Toolkit for Multilingual Question and Answer Generation

arXiv1 repo

arXiv:2305.17416

lm-question-generation

SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created Through Human-Machine Collaboration

arXiv1 repo

arXiv:2305.17696

korean-safety-benchmarks

KoSBi: A Dataset for Mitigating Social Bias Risks Towards Safer Large Language Model Application

arXiv1 repo

arXiv:2305.17701

korean-safety-benchmarks

arXiv:2305.17926

arXiv1 repo

arXiv:2305.17926

alpaca_eval

ChatGPT-powered Conversational Drug Editing Using Retrieval and Domain Feedback

arXiv1 repo

arXiv:2305.18090

ChatDrug

Do Large Language Models Know What They Don't Know?

arXiv1 repo

arXiv:2305.18153

SelfAware

GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

arXiv1 repo

arXiv:2305.18752

GPT4Tools

Voice Conversion With Just Nearest Neighbors

arXiv1 repo

arXiv:2305.18975

knn-vc

Generating with Confidence: Uncertainty Quantification for Black-box Large Language Models

arXiv1 repo

arXiv:2305.19187

UQ-NLG

infoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-information

arXiv1 repo

arXiv:2305.19344

infoVerse

CryptOpt: Automatic Optimization of Straightline Code

arXiv1 repo

arXiv:2305.19586

CryptOpt

Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust

arXiv1 repo

arXiv:2305.20030

tree-ring-watermark

Let's Verify Step by Step

arXiv1 repo

arXiv:2305.20050

KOpen-platypus

Humans in 4D: Reconstructing and Tracking Humans with Transformers

arXiv1 repo

arXiv:2305.20091

4D-Humans

Preference-grounded Token-level Guidance for Language Model Fine-tuning

arXiv1 repo

arXiv:2306.00398

DenseRewardRLHF-PPO

End-to-end Knowledge Retrieval with Multi-modal Queries

arXiv1 repo

arXiv:2306.00424

ReMuQ

Masked Autoencoders with Multi-Window Local-Global Attention Are Better Audio Learners

arXiv1 repo

arXiv:2306.00561

j2a

arXiv:2306.00745

arXiv1 repo

arXiv:2306.00745

Jellyfish-13B

Continual Learning for Abdominal Multi-Organ and Tumor Segmentation

arXiv1 repo

arXiv:2306.00988

awesome-cybersecurity-agentic-ai

Towards Robust FastSpeech 2 by Modelling Residual Multimodality

arXiv1 repo

arXiv:2306.01442

ai-research-code

SACSoN: Scalable Autonomous Control for Social Navigation

arXiv1 repo

arXiv:2306.01874

UniWM_Dataset

Can Contextual Biasing Remain Effective with Whisper and GPT-2?

arXiv1 repo

arXiv:2306.01942

WhisperBiasing

Exploring the Optimal Choice for Generative Processes in Diffusion Models: Ordinary vs Stochastic Differential Equations

arXiv1 repo

arXiv:2306.02063

LakonLab

Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation

arXiv1 repo

arXiv:2306.02416

adapt_med_seg

Benchmarking Large Language Models on CMExam -- A Comprehensive Chinese Medical Exam Dataset

arXiv1 repo

arXiv:2306.03030

CMExam

Sequential Monte Carlo Steering of Large Language Models using Probabilistic Programs

arXiv1 repo

arXiv:2306.03081

llamppl

Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

arXiv1 repo

arXiv:2306.03509

MagVITS

GEO-Bench: Toward Foundation Models for Earth Monitoring

arXiv1 repo

arXiv:2306.03831

geo-bench

CL-UZH at SemEval-2023 Task 10: Sexism Detection through Incremental Fine-Tuning and Multi-Task Learning with Label Descriptions

arXiv1 repo

arXiv:2306.03907

CL-UZH-EDOS-2023

M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

arXiv1 repo

arXiv:2306.04387

M3IT

arXiv:2306.04618

arXiv1 repo

arXiv:2306.04618

OOD_NLP

On the Reliability of Watermarks for Large Language Models

arXiv1 repo

arXiv:2306.04634

lm-watermarking

Generalizable Low-Resource Activity Recognition with Diverse and Discriminative Representation Learning

arXiv1 repo

arXiv:2306.04641

robustlearn

Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts

arXiv1 repo

arXiv:2306.04723

L2D

arXiv:2306.04848

arXiv1 repo

arXiv:2306.04848

smalldiffusion

K2: A Foundation Language Model for Geoscience Knowledge Understanding and Utilization

arXiv1 repo

arXiv:2306.05064

k2

PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models

arXiv1 repo

arXiv:2306.05350

peft-ser

PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance

arXiv1 repo

arXiv:2306.05443

PIXIU

Prompt Injection attack against LLM-integrated Applications

arXiv1 repo

arXiv:2306.05499

Awesome-OpenClaw

There's Plenty of Room in the Middle: The Unsung Revolution of the Renormalization Group

arXiv1 repo

arXiv:2306.06020

Low-power-E-Paper-OS

Image Vectorization: a Review

arXiv1 repo

arXiv:2306.06441

vtracer

detrex: Benchmarking Detection Transformers

arXiv1 repo

arXiv:2306.07265

detrex

Scalable 3D Captioning with Pretrained Models

arXiv1 repo

arXiv:2306.07279

Cap3D

Resources for Brewing BEIR: Reproducible Reference Models and an Official Leaderboard

arXiv1 repo

arXiv:2306.07471

beir

TART: A plug-and-play Transformer module for task-agnostic reasoning

arXiv1 repo

arXiv:2306.07536

TART

SqueezeLLM: Dense-and-Sparse Quantization

arXiv1 repo

arXiv:2306.07629

turboquant-vllm

WebGLM: Towards An Efficient Web-Enhanced Question Answering System with Human Preferences

arXiv1 repo

arXiv:2306.07906

WebGLM

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

arXiv1 repo

arXiv:2306.07954

Rerender_A_Video

LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting

arXiv1 repo

arXiv:2306.08259

Traffic_Forecast_Benchmark

TryOnDiffusion: A Tale of Two UNets

arXiv1 repo

arXiv:2306.08276

opentryon

arXiv:2306.08645

arXiv1 repo

arXiv:2306.08645

Video-Infinity

Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations

arXiv1 repo

arXiv:2306.08658

Babel-ImageNet

LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

arXiv1 repo

arXiv:2306.09265

Multi-Modality-Arena

DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data

arXiv1 repo

arXiv:2306.09344

dreamsim

Scaling Open-Vocabulary Object Detection

arXiv1 repo

arXiv:2306.09683

owlv2-base-patch16-ensemble

arXiv:2306.09803

arXiv1 repo

arXiv:2306.09803

HEBO

ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation

arXiv1 repo

arXiv:2306.09968

ClinicalGPT-base-zh

Self-Supervised Learning for Time Series Analysis: Taxonomy, Progress, and Prospects

arXiv1 repo

arXiv:2306.10125

time-moe

MARBLE: Music Audio Representation Benchmark for Universal Evaluation

arXiv1 repo

arXiv:2306.10548

usad

Guiding Language Models of Code with Global Context using Monitors

arXiv1 repo

arXiv:2306.10763

monitors4codegen

SelfTalk: A Self-Supervised Commutative Training Diagram to Comprehend 3D Talking Faces

arXiv1 repo

arXiv:2306.10799

SelfTalk_release

Textbooks Are All You Need

arXiv1 repo

arXiv:2306.11644

cosmopedia

A Simple and Effective Pruning Approach for Large Language Models

arXiv1 repo

arXiv:2306.11695

cosyvoice3-inference-acceleration

Training Transformers with 4-bit Integers

arXiv1 repo

arXiv:2306.11987

JobList

LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models

arXiv1 repo

arXiv:2306.12420

LMFlow

Generative Multimodal Entity Linking

arXiv1 repo

arXiv:2306.12725

GEMEL

An overview on the evaluated video retrieval tasks at TRECVID 2022

arXiv1 repo

arXiv:2306.13118

ladi-overview

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

arXiv1 repo

arXiv:2306.13394

Video-MME

Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data

arXiv1 repo

arXiv:2306.13840

ultimate-utils

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models

arXiv1 repo

arXiv:2306.14048

mlx-flash

Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference

arXiv1 repo

arXiv:2306.14393

Moonlit

SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

arXiv1 repo

arXiv:2306.14610

SugarCrepe_pp

Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

arXiv1 repo

arXiv:2306.15195

shikra

arXiv:2306.15895

arXiv1 repo

arXiv:2306.15895

ReGen

UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

arXiv1 repo

arXiv:2306.16083

UnitSpeech

SE-PQA: Personalized Community Question Answering

arXiv1 repo

arXiv:2306.16261

SE-PQA

Foundation Model for Endoscopy Video Analysis via Large-scale Self-supervised Pre-train

arXiv1 repo

arXiv:2306.16741

Endo-FM

Tokenization and the Noiseless Channel

arXiv1 repo

arXiv:2306.16842

tokenizer-flores-validation

Hierarchical Neural Coding for Controllable CAD Model Generation

arXiv1 repo

arXiv:2307.00149

FlexCAD

Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language Models

arXiv1 repo

arXiv:2307.01379

SAR

Detecting Images Generated by Deep Diffusion Models using their Local Intrinsic Dimensionality

arXiv1 repo

arXiv:2307.02347

deepfake_multiLID

Jailbroken: How Does LLM Safety Training Fail?

arXiv1 repo

arXiv:2307.02483

JailbreakLab

GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

arXiv1 repo

arXiv:2307.03601

gradio-box

A Survey on Graph Neural Networks for Time Series: Forecasting, Classification, Imputation, and Anomaly Detection

arXiv1 repo

arXiv:2307.03759

time-moe

InPars Toolkit: A Unified and Reproducible Synthetic Data Generation Pipeline for Neural Information Retrieval

arXiv1 repo

arXiv:2307.04601

InPars

arXiv:2307.05222

arXiv1 repo

arXiv:2307.05222

CoBSAT

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

arXiv1 repo

arXiv:2307.06942

InternVid-Full

MGit: A Model Versioning and Management System

arXiv1 repo

arXiv:2307.07507

mgit

CoTracker: It is Better to Track Together

arXiv1 repo

arXiv:2307.07635

co-tracker

CA-LoRA: Adapting Existing LoRA for Compressed LLMs to Enable Efficient Multi-Tasking on Personal Devices

arXiv1 repo

arXiv:2307.07705

CA-LoRA

ChatDev: Communicative Agents for Software Development

arXiv1 repo

arXiv:2307.07924

ChatDev

Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models

arXiv1 repo

arXiv:2307.08487

latent-jailbreak

Generative Type Inference for Python

arXiv1 repo

arXiv:2307.09163

TypeGen

CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

arXiv1 repo

arXiv:2307.09705

COIG-CQIA

arXiv:2307.10373

arXiv1 repo

arXiv:2307.10373

TokenFlow

L-Eval: Instituting Standardized Evaluation for Long Context Language Models

arXiv1 repo

arXiv:2307.11088

Qwen

MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems

arXiv1 repo

arXiv:2307.11394

meeteval

A Change of Heart: Improving Speech Emotion Recognition through Speech-to-Text Modality Conversion

arXiv1 repo

arXiv:2307.11584

dissertation-project

Understanding the Latent Space of Diffusion Models through the Lens of Riemannian Geometry

arXiv1 repo

arXiv:2307.12868

Diffusion-Pullback

FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios

arXiv1 repo

arXiv:2307.13528

SDAK

Measuring Faithfulness in Chain-of-Thought Reasoning

arXiv1 repo

arXiv:2307.13702

CoTFaithChecker

Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image Classification

arXiv1 repo

arXiv:2307.15254

MHIM-MIL

ChatHome: Development and Evaluation of a Domain-Specific Language Model for Home Renovation

arXiv1 repo

arXiv:2307.15290

BELLE

A Dataflow Compiler for Efficient LLM Inference using Custom Microscaling Formats

arXiv1 repo

arXiv:2307.15517

mase

The Hydra Effect: Emergent Self-repair in Language Model Computations

arXiv1 repo

arXiv:2307.15771

tfmlens

Backdooring Instruction-Tuned Large Language Models with Virtual Prompt Injection

arXiv1 repo

arXiv:2307.16888

virtual-prompt-injection

Three Bricks to Consolidate Watermarks for Large Language Models

arXiv1 repo

arXiv:2308.00113

lm-watermarking

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

arXiv1 repo

arXiv:2308.00352

codel

From Sparse to Soft Mixtures of Experts

arXiv1 repo

arXiv:2308.00951

PETL_AST

Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data

arXiv1 repo

arXiv:2308.02463

RadFM

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

arXiv1 repo

arXiv:2308.02490

MM-Vet

Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning

arXiv1 repo

arXiv:2308.02533

robustlearn

Studying Large Language Model Generalization with Influence Functions

arXiv1 repo

arXiv:2308.03296

bergson

DiffSynth: Latent In-Iteration Deflickering for Realistic Video Synthesis

arXiv1 repo

arXiv:2308.03463

diffSynth-studio-notes

AgentBench: Evaluating LLMs as Agents

arXiv1 repo

arXiv:2308.03688

Awesome-OpenClaw

TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models

arXiv1 repo

arXiv:2308.03729

Multi-Modality-Arena

SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

arXiv1 repo

arXiv:2308.04430

silo-lm

Classification of Human- and AI-Generated Texts: Investigating Features for ChatGPT

arXiv1 repo

arXiv:2308.05341

verify-ai

Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance Flow

arXiv1 repo

arXiv:2308.06101

DCI-VTON-Virtual-Try-On

DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion Models

arXiv1 repo

arXiv:2308.06160

DatasetDM

Self-Alignment with Instruction Backtranslation

arXiv1 repo

arXiv:2308.06259

EasyInstruct

Foundation Model is Efficient Multimodal Multitask Model Selector

arXiv1 repo

arXiv:2308.06262

Multitask-Model-Selector

Tiny and Efficient Model for the Edge Detection Generalization

arXiv1 repo

arXiv:2308.06468

ComfyUI-Anyline

EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models

arXiv1 repo

arXiv:2308.07269

KnowEdit

LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

arXiv1 repo

arXiv:2308.07308

llm-self-defense

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

arXiv1 repo

arXiv:2308.08089

dragnuwa-pruned-safetensors

Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

arXiv1 repo

arXiv:2308.08769

Chat-Scene

Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?

arXiv1 repo

arXiv:2308.10168

DistFactAssessLM

ROSGPT_Vision: Commanding Robots Using Only Language Models' Prompts

arXiv1 repo

arXiv:2308.11236

ROSGPT_Vision

Efficient Benchmarking of Language Models

arXiv1 repo

arXiv:2308.11696

helm

arXiv:2308.11957

arXiv1 repo

arXiv:2308.11957

hashing-baseline

Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

arXiv1 repo

arXiv:2308.12038

OmniLMM-12B

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

arXiv1 repo

arXiv:2308.12067

InstructionGPT-4

MultiCapCLIP: Auto-Encoding Prompts for Zero-Shot Multilingual Visual Captioning

arXiv1 repo

arXiv:2308.13218

MultiCapCLIP

arXiv:2308.13418

arXiv1 repo

arXiv:2308.13418

nougat

Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models

arXiv1 repo

arXiv:2308.13437

pvit

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

arXiv1 repo

arXiv:2308.14508

Marathon

CLIPTrans: Transferring Visual Knowledge with Pre-trained Models for Multimodal Machine Translation

arXiv1 repo

arXiv:2308.15226

CLIPTrans

When Do Program-of-Thoughts Work for Reasoning?

arXiv1 repo

arXiv:2308.15452

EasyInstruct

ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer

arXiv1 repo

arXiv:2308.15459

ParaGuide

arXiv:2308.16361

arXiv1 repo

arXiv:2308.16361

Jellyfish-13B

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants

arXiv1 repo

arXiv:2308.16884

M-AbstainQA

TouchStone: Evaluating Vision-Language Models by Language Models

arXiv1 repo

arXiv:2308.16890

TouchStone

RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

arXiv1 repo

arXiv:2309.00267

YiVal

NLLB-CLIP -- train performant multilingual image retrieval model on a budget

arXiv1 repo

arXiv:2309.01859

CLIP_benchmark

arXiv:2309.02243

arXiv1 repo

arXiv:2309.02243

improvnet

Certifying LLM Safety against Adversarial Prompting

arXiv1 repo

arXiv:2309.02705

certified-llm-safety

Norm Tweaking: High-performance Low-bit Quantization of Large Language Models

arXiv1 repo

arXiv:2309.02784

norm-tweaking

BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network

arXiv1 repo

arXiv:2309.02836

bigvsan

Matcha-TTS: A fast TTS architecture with conditional flow matching

arXiv1 repo

arXiv:2309.03199

Matcha-TTS

Large Language Models as Optimizers

arXiv1 repo

arXiv:2309.03409

YiVal

DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

arXiv1 repo

arXiv:2309.03883

SLED

Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis

arXiv1 repo

arXiv:2309.03904

Aurora

Large-Scale Automatic Audiobook Creation

arXiv1 repo

arXiv:2309.03926

SynapseML

Knowledge-tuning Large Language Models with Structured Medical Knowledge Bases for Reliable Response Generation in Chinese

arXiv1 repo

arXiv:2309.04175

Huatuo-Llama-Med-Chinese

Don't Ignore Dual Logic Ability of LLMs while Privatizing: A Data-Intensive Analysis in Medical Domain

arXiv1 repo

arXiv:2309.04198

Huatuo-Llama-Med-Chinese

From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting

arXiv1 repo

arXiv:2309.04269

YiVal

MADLAD-400: A Multilingual And Document-Level Large Audited Dataset

arXiv1 repo

arXiv:2309.04662

madlad400-10b-mt

arXiv:2309.05519

arXiv1 repo

arXiv:2309.05519

NExT-GPT

Kani: A Lightweight and Highly Hackable Framework for Building Language Model Applications

arXiv1 repo

arXiv:2309.05542

kani

An Empirical Study of NetOps Capability of Pre-Trained Large Language Models

arXiv1 repo

arXiv:2309.05557

neteval-exam

InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation

arXiv1 repo

arXiv:2309.06380

InstaFlow

An Image Dataset for Benchmarking Recommender Systems with Raw Pixels

arXiv1 repo

arXiv:2309.06789

PixelRec

RAIN: Your Language Models Can Align Themselves without Finetuning

arXiv1 repo

arXiv:2309.07124

RAIN

PromptASR for contextualized ASR with controllable style

arXiv1 repo

arXiv:2309.07414

libriheavy

VerilogEval: Evaluating Large Language Models for Verilog Code Generation

arXiv1 repo

arXiv:2309.07544

verilog-eval

Generative AI Text Classification using Ensemble LLM Approaches

arXiv1 repo

arXiv:2309.07755

verify-ai

Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

arXiv1 repo

arXiv:2309.07875

med-safety-bench

Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates

arXiv1 repo

arXiv:2309.08125

Cornstarch

Advancing the Evaluation of Traditional Chinese Language Models: Towards a Comprehensive Benchmark Suite

arXiv1 repo

arXiv:2309.08448

TCEval-v2

Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-training and Multi-modal Tokens

arXiv1 repo

arXiv:2309.08531

Image-to-Speech

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers

arXiv1 repo

arXiv:2309.08532

PRL-Prompts-from-Reinforcement-Learning

HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform

arXiv1 repo

arXiv:2309.09493

kanade-tokenizer

Adapting Large Language Models to Domains via Reading Comprehension

arXiv1 repo

arXiv:2309.09530

KoCommercial-Dataset

Baichuan 2: Open Large-scale Language Models

arXiv1 repo

arXiv:2309.10305

Baichuan2

PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training

arXiv1 repo

arXiv:2309.10400

EasyContext

A Configurable Library for Generating and Manipulating Maze Datasets

arXiv1 repo

arXiv:2309.10498

maze-dataset

MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback

arXiv1 repo

arXiv:2309.10691

mint-bench

DISC-LawLLM: Fine-tuning Large Language Models for Intelligent Legal Services

arXiv1 repo

arXiv:2309.11325

DISC-FinLLM

You Only Look at Screens: Multimodal Chain-of-Action Agents

arXiv1 repo

arXiv:2309.11436

carl_distrl

SignBank+: Preparing a Multilingual Sign Language Dataset for Machine Translation Using Large Language Models

arXiv1 repo

arXiv:2309.11566

signbank-plus

VoiceLDM: Text-to-Speech with Environmental Context

arXiv1 repo

arXiv:2309.13664

VoiceLDM

VidChapters-7M: Video Chapters at Scale

arXiv1 repo

arXiv:2309.13952

VidChapters

UnitedHuman: Harnessing Multi-Source Data for High-Resolution Human Generation

arXiv1 repo

arXiv:2309.14335

CosmicMan

Joint Audio and Speech Understanding

arXiv1 repo

arXiv:2309.14405

Speech-IFEval

Motions in Microseconds via Vectorized Sampling-Based Planning

arXiv1 repo

arXiv:2309.14545

mdtensor

Updated Corpora and Benchmarks for Long-Form Speech Recognition

arXiv1 repo

arXiv:2309.15013

speech-datasets

RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models

arXiv1 repo

arXiv:2309.15088

pyterrier_genrank

A Content-Driven Micro-Video Recommendation Dataset at Scale

arXiv1 repo

arXiv:2309.15379

MicroLens

AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

arXiv1 repo

arXiv:2309.16058

papagei-foundation-model

Masked Autoencoders are Scalable Learners of Cellular Morphology

arXiv1 repo

arXiv:2309.16064

maes_microscopy

Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

arXiv1 repo

arXiv:2309.16240

f-divergence-dpo

MHG-GNN: Combination of Molecular Hypergraph Grammar with Graph Neural Network

arXiv1 repo

arXiv:2309.16374

materials

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

arXiv1 repo

arXiv:2309.16429

JavisBench

Vision Transformers Need Registers

arXiv1 repo

arXiv:2309.16588

geti-instant-learn

DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation

arXiv1 repo

arXiv:2309.16653

controlled-dreamgaussian

Demystifying CLIP Data

arXiv1 repo

arXiv:2309.16671

MetaCLIP

Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities

arXiv1 repo

arXiv:2309.16739

SplitFM

Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

arXiv1 repo

arXiv:2309.17002

robustlearn

Unsupervised Large Language Model Alignment for Information Retrieval via Contrastive Feedback

arXiv1 repo

arXiv:2309.17078

RLCF

Guiding Instruction-based Image Editing via Multimodal Large Language Models

arXiv1 repo

arXiv:2309.17102

ml-mgie

arXiv:2309.17444

arXiv1 repo

arXiv:2309.17444

LLM-groundedVideoDiffusion

ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

arXiv1 repo

arXiv:2309.17452

NuminaMath-7B-TIR

FELM: Benchmarking Factuality Evaluation of Large Language Models

arXiv1 repo

arXiv:2310.00741

felm

Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models

arXiv1 repo

arXiv:2310.01107

Ground-A-Video

Quantifying the Plausibility of Context Reliance in Neural Machine Translation

arXiv1 repo

arXiv:2310.01188

LRP-eXplains-Transformers

Merge, Then Compress: Demystify Efficient SMoE with Hints from Its Routing Policy

arXiv1 repo

arXiv:2310.01334

MC-SMoE

Compressing LLMs: The Truth is Rarely Pure and Never Simple

arXiv1 repo

arXiv:2310.01382

llm-kick

arXiv:2310.01405

arXiv1 repo

arXiv:2310.01405

drowse

ImagenHub: Standardizing the evaluation of conditional image generation models

arXiv1 repo

arXiv:2310.01596

ImagenHub

CAT-LM: Training Language Models on Aligned Code And Tests

arXiv1 repo

arXiv:2310.01602

CAT-LM

Fool Your (Vision and) Language Model With Embarrassingly Simple Permutations

arXiv1 repo

arXiv:2310.01651

FoolyourVLLMs

MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

arXiv1 repo

arXiv:2310.02255

MathVista

Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion

arXiv1 repo

arXiv:2310.02279

ctm

Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

arXiv1 repo

arXiv:2310.02949

shadow-alignment

DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training

arXiv1 repo

arXiv:2310.03294

EasyContext

Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

arXiv1 repo

arXiv:2310.03708

modpo

Exploiting Activation Sparsity with Dense to Dynamic-k Mixture-of-Experts Conversion

arXiv1 repo

arXiv:2310.04361

MoFEbaseD2D

Amortizing intractable inference in large language models

arXiv1 repo

arXiv:2310.04363

gfn-lm-tuning

Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

arXiv1 repo

arXiv:2310.04406

PromptingTools.jl

What's the Magic Word? A Control Theory of LLM Prompting

arXiv1 repo

arXiv:2310.04444

Magic_Words

AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

arXiv1 repo

arXiv:2310.04451

GA

LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

arXiv1 repo

arXiv:2310.04673

FunCodec

DORIS-MAE: Scientific Document Retrieval using Multi-level Aspect-based Queries

arXiv1 repo

arXiv:2310.04678

Doris-Mae-Dataset

Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages

arXiv1 repo

arXiv:2310.04799

deccp

SteerLM: Attribute Conditioned SFT as an (User-Steerable) Alternative to RLHF

arXiv1 repo

arXiv:2310.05344

Llama-3_3-Nemotron-Super-49B-GenRM

Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel Decoding

arXiv1 repo

arXiv:2310.05424

mixture_of_recursions

Generative Judge for Evaluating Alignment

arXiv1 repo

arXiv:2310.05470

JudgeBench

UAVs and Neural Networks for search and rescue missions

arXiv1 repo

arXiv:2310.05512

argus

Towards Verifiable Generation: A Benchmark for Knowledge-aware Language Model Attribution

arXiv1 repo

arXiv:2310.05634

SAFE

LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

arXiv1 repo

arXiv:2310.05736

Prompt-Compression

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

arXiv1 repo

arXiv:2310.05737

Pyramid-Flow

Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models

arXiv1 repo

arXiv:2310.06313

PCDMs

Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

arXiv1 repo

arXiv:2310.06387

llm-jailbreaking-defense

iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

arXiv1 repo

arXiv:2310.06625

frn-50k-baseline

Text Embeddings Reveal (Almost) As Much As Text

arXiv1 repo

arXiv:2310.06816

align-text-encoders

LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression

arXiv1 repo

arXiv:2310.06839

Prompt-Compression

Violation of Expectation via Metacognitive Prompting Reduces Theory of Mind Prediction Error in Large Language Models

arXiv1 repo

arXiv:2310.06983

tutor-gpt

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

arXiv1 repo

arXiv:2310.06987

Jailbreak_LLM

Composite Backdoor Attacks Against Large Language Models

arXiv1 repo

arXiv:2310.07676

CBA

MatFormer: Nested Transformer for Elastic Inference

arXiv1 repo

arXiv:2310.07707

mlx-flash

DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model

arXiv1 repo

arXiv:2310.07771

DrivingDiffusion

Large Language Models Are Zero-Shot Time Series Forecasters

arXiv1 repo

arXiv:2310.07820

llmtime

MemGPT: Towards LLMs as Operating Systems

arXiv1 repo

arXiv:2310.08560

letta-code

arXiv:2310.08588

arXiv1 repo

arXiv:2310.08588

Octopus

The Data Lakehouse: Data Warehousing and More

arXiv1 repo

arXiv:2310.08697

data-engineer-handbook

Extending Multi-modal Contrastive Representations

arXiv1 repo

arXiv:2310.08884

20251R0136COSE40500

SeqXGPT: Sentence-Level AI-Generated Text Detection

arXiv1 repo

arXiv:2310.08903

SeqXGPT

A Setwise Approach for Effective and Highly Efficient Zero-shot Ranking with Large Language Models

arXiv1 repo

arXiv:2310.09497

llm-rankers

TEQ: Trainable Equivalent Transformation for Quantization of LLMs

arXiv1 repo

arXiv:2310.10944

auto-round

Domain Generalization Using Large Pretrained Models with Mixture-of-Adapters

arXiv1 repo

arXiv:2310.11031

MoA

Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis

arXiv1 repo

arXiv:2310.11722

SDAK

A General Theoretical Paradigm to Understand Learning from Human Preferences

arXiv1 repo

arXiv:2310.12036

Reinforcement-Learning-Full-Pipeline

Enhancing High-Resolution 3D Generation through Pixel-wise Gradient Clipping

arXiv1 repo

arXiv:2310.12474

PGC-3D

arXiv:2310.12537

arXiv1 repo

arXiv:2310.12537

Jellyfish-13B

arXiv:2310.12952

arXiv1 repo

arXiv:2310.12952

Vendi-Score

GraphGPT: Graph Instruction Tuning for Large Language Models

arXiv1 repo

arXiv:2310.13023

GraphGPT

Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots

arXiv1 repo

arXiv:2310.13724

habitat-lab

Contrast Everything: A Hierarchical Contrastive Framework for Medical Time-Series

arXiv1 repo

arXiv:2310.14017

UTSD

Tree Prompting: Efficient Task Adaptation without Fine-Tuning

arXiv1 repo

arXiv:2310.14034

imodelsX

Language Model Unalignment: Parametric Red-Teaming to Expose Hidden Harms and Biases

arXiv1 repo

arXiv:2310.14303

red-instruct

arXiv:2310.14478

arXiv1 repo

arXiv:2310.14478

geolm

HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

arXiv1 repo

arXiv:2310.14566

LRV-Instruction

Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

arXiv1 repo

arXiv:2310.15110

zero123plus

Open-Set Image Tagging with Multi-Grained Text Supervision

arXiv1 repo

arXiv:2310.15200

recognize-anything

Efficient Online String Matching through Linked Weak Factors

arXiv1 repo

arXiv:2310.15711

HashChain

In-Context Learning Creates Task Vectors

arXiv1 repo

arXiv:2310.15916

icl_task_vectors

Woodpecker: Hallucination Correction for Multimodal Large Language Models

arXiv1 repo

arXiv:2310.16045

Woodpecker

MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning

arXiv1 repo

arXiv:2310.16049

llm-mysteries

Towards long-tailed, multi-label disease classification from chest X-ray: Overview of the CXR-LT challenge

arXiv1 repo

arXiv:2310.16112

Fine-Grained_Features_Alignment_via_Constrastive_Learning

Attention Lens: A Tool for Mechanistically Interpreting the Attention Head Information Retrieval Mechanism

arXiv1 repo

arXiv:2310.16270

AttentionLens

A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation

arXiv1 repo

arXiv:2310.16656

tuxemon

DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion Prior

arXiv1 repo

arXiv:2310.16818

DreamCraft3D

TD-MPC2: Scalable, Robust World Models for Continuous Control

arXiv1 repo

arXiv:2310.16828

xwm

How do Language Models Bind Entities in Context?

arXiv1 repo

arXiv:2310.17191

binding-iclr

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

arXiv1 repo

arXiv:2310.17631

JudgeBench

Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

arXiv1 repo

arXiv:2310.17884

confaide

FormalGeo: An Extensible Formalized Framework for Olympiad Geometric Problem Solving

arXiv1 repo

arXiv:2310.18021

FormalGeo

DUMA: a Dual-Mind Conversational Agent with Fast and Slow Thinking

arXiv1 repo

arXiv:2310.18075

BELLE

Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

arXiv1 repo

arXiv:2310.18235

DSG

Image Clustering Conditioned on Text Criteria

arXiv1 repo

arXiv:2310.18297

ICTC

LLMSTEP: LLM proofstep suggestions in Lean

arXiv1 repo

arXiv:2310.18457

llmstep

EHRXQA: A Multi-Modal Question Answering Dataset for Electronic Health Records with Chest X-ray Images

arXiv1 repo

arXiv:2310.18652

blendsql

Atom: Low-bit Quantization for Efficient and Accurate LLM Serving

arXiv1 repo

arXiv:2310.19102

deepcompressor

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

arXiv1 repo

arXiv:2310.19512

VideoCrafter

arXiv:2310.20145

arXiv1 repo

arXiv:2310.20145

HEBO

Integrating curation into scientific publishing to train AI models

arXiv1 repo

arXiv:2310.20440

SourceData

JADE: A Linguistics-based Safety Evaluation Platform for Large Language Models

arXiv1 repo

arXiv:2311.00286

jade-db

TopicGPT: A Prompt-based Topic Modeling Framework

arXiv1 repo

arXiv:2311.01449

turkish-complaint-topic-clustering

Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization

arXiv1 repo

arXiv:2311.01544

who_what_benchmark

Adapting Frechet Audio Distance for Generative Music Evaluation

arXiv1 repo

arXiv:2311.01616

fadtk

SAC3: Reliable Hallucination Detection in Black-Box Language Models via Semantic-aware Cross-check Consistency

arXiv1 repo

arXiv:2311.01740

sac3

AnyText: Multilingual Visual Text Generation And Editing

arXiv1 repo

arXiv:2311.03054

AnyText

GLaMM: Pixel Grounding Large Multimodal Model

arXiv1 repo

arXiv:2311.03356

groundingLMM

I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

arXiv1 repo

arXiv:2311.04145

i2vgen-xl

Holistic Evaluation of Text-To-Image Models

arXiv1 repo

arXiv:2311.04287

helm

Prompt Cache: Modular Attention Reuse for Low-Latency Inference

arXiv1 repo

arXiv:2311.04934

prompt-cache

LooGLE: Can Long-Context Language Models Understand Long Contexts?

arXiv1 repo

arXiv:2311.04939

LooGLE

BeLLM: Backward Dependency Enhanced Large Language Model for Sentence Embeddings

arXiv1 repo

arXiv:2311.05296

AnglE

Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model

arXiv1 repo

arXiv:2311.06214

DiffSplat

Knowledge-Augmented Large Language Models for Personalized Contextual Query Suggestion

arXiv1 repo

arXiv:2311.06318

azure-openai-llm-cookbook

PerceptionGPT: Effectively Fusing Visual Perception into LLM

arXiv1 repo

arXiv:2311.06612

medllm

Embarassingly Simple Dataset Distillation

arXiv1 repo

arXiv:2311.07025

PoDD

Attention-Challenging Multiple Instance Learning for Whole Slide Image Classification

arXiv1 repo

arXiv:2311.07125

AEM-dataset

arXiv:2311.07574

arXiv1 repo

arXiv:2311.07574

prismatic-vlms

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

arXiv1 repo

arXiv:2311.07575

idefics2-8b

Large Language Models can Strategically Deceive their Users when Put Under Pressure

arXiv1 repo

arXiv:2311.07590

how-to-catch-an-ai-liar

Computing Implicitizations of Multi-Graded Polynomial Maps

arXiv1 repo

arXiv:2311.07678

smash

Fair Abstractive Summarization of Diverse Perspectives

arXiv1 repo

arXiv:2311.07884

FairSumm

One-2-3-45++: Fast Single Image to 3D Objects with Consistent Multi-View Generation and 3D Diffusion

arXiv1 repo

arXiv:2311.07885

torchsparse

REST: Retrieval-Based Speculative Decoding

arXiv1 repo

arXiv:2311.08252

REST

Learning to Filter Context for Retrieval-Augmented Generation

arXiv1 repo

arXiv:2311.08377

TinyRAG

Towards Open-Ended Visual Recognition with Large Language Model

arXiv1 repo

arXiv:2311.08400

OmniScient-Model

Fine-tuning Language Models for Factuality

arXiv1 repo

arXiv:2311.08401

llm_factuality_tuning

PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models

arXiv1 repo

arXiv:2311.08590

PEMA

Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation

arXiv1 repo

arXiv:2311.08640

MCKD-semantic-segmentation

OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining

arXiv1 repo

arXiv:2311.08849

ofa

FastBlend: a Powerful Model-Free Toolkit Making Video Stylization Easier

arXiv1 repo

arXiv:2311.09265

diffSynth-studio-notes

HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

arXiv1 repo

arXiv:2311.09528

Llama-3_3-Nemotron-Super-49B-GenRM

MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and Classification

arXiv1 repo

arXiv:2311.09761

MAFALDA

TransFusion -- A Transparency-Based Diffusion Model for Anomaly Detection

arXiv1 repo

arXiv:2311.09999

awesome-cybersecurity-agentic-ai

The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation

arXiv1 repo

arXiv:2311.10057

song-describer-dataset

Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

arXiv1 repo

arXiv:2311.10702

proxy-tuning

MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

arXiv1 repo

arXiv:2311.10774

LRV-Instruction

LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching

arXiv1 repo

arXiv:2311.11284

LucidDreamer

Large Pre-trained time series models for cross-domain Time series analysis tasks

arXiv1 repo

arXiv:2311.11413

Samay

FinanceBench: A New Benchmark for Financial Question Answering

arXiv1 repo

arXiv:2311.11944

PageIndex

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

arXiv1 repo

arXiv:2311.12022

RLPR-Evaluation

T-Rex: Counting by Visual Prompting

arXiv1 repo

arXiv:2311.13596

Rex-Omni

ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs

arXiv1 repo

arXiv:2311.13600

ComfyUI-LoRA-Optimizer

MAIRA-1: A specialised large multimodal model for radiology report generation

arXiv1 repo

arXiv:2311.13668

rad-dino

SinSR: Diffusion-Based Image Super-Resolution in a Single Step

arXiv1 repo

arXiv:2311.14760

SinSR

Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity Vocoder

arXiv1 repo

arXiv:2311.14957

BigVGAN

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

arXiv1 repo

arXiv:2311.15127

videophy

GART: Gaussian Articulated Template Models

arXiv1 repo

arXiv:2311.16099

GART

Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

arXiv1 repo

arXiv:2311.16103

LanguageBind

Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability

arXiv1 repo

arXiv:2311.16484

SeeingEyeToAI

DemoFusion: Democratising High-Resolution Image Generation With No $$$

arXiv1 repo

arXiv:2311.16973

DemoFusion

Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following

arXiv1 repo

arXiv:2311.17002

Omost

arXiv:2311.17005

arXiv1 repo

arXiv:2311.17005

MVBench

Diffusion 3D Features (Diff3F): Decorating Untextured Shapes with Distilled Semantic Features

arXiv1 repo

arXiv:2311.17024

Diffusion-3D-Features

MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training

arXiv1 repo

arXiv:2311.17049

ml-mobileclip

DreamPropeller: Supercharge Text-to-3D Generation with Parallel Sampling

arXiv1 repo

arXiv:2311.17082

DreamPropeller

SEED-Bench-2: Benchmarking Multimodal Large Language Models

arXiv1 repo

arXiv:2311.17092

SEED-Bench

Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

arXiv1 repo

arXiv:2311.17117

MuseV

TaskWeaver: A Code-First Agent Framework

arXiv1 repo

arXiv:2311.17541

TaskWeaver

SyncTalk: The Devil is in the Synchronization for Talking Head Synthesis

arXiv1 repo

arXiv:2311.17590

SyncTalk

Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving

arXiv1 repo

arXiv:2311.17918

Drive-WM

VBench: Comprehensive Benchmark Suite for Video Generative Models

arXiv1 repo

arXiv:2311.17982

VBench

mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model

arXiv1 repo

arXiv:2311.18248

mPLUG-DocOwl

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

arXiv1 repo

arXiv:2311.18259

Ego4d

Splitwise: Efficient generative LLM inference using phase splitting

arXiv1 repo

arXiv:2311.18677

Nanoflow

TaskBench: Benchmarking Large Language Models for Task Automation

arXiv1 repo

arXiv:2311.18760

JARVIS

X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning

arXiv1 repo

arXiv:2311.18799

LAVIS

Distributed Global Structure-from-Motion with a Deep Front-End

arXiv1 repo

arXiv:2311.18801

gtsfm

Exploiting Diffusion Prior for Generalizable Dense Prediction

arXiv1 repo

arXiv:2311.18832

dmp

SparseGS: Sparse View Synthesis using 3D Gaussian Splatting

arXiv1 repo

arXiv:2312.00206

SparseGS

arXiv:2312.00438

arXiv1 repo

arXiv:2312.00438

Dolphins

RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback

arXiv1 repo

arXiv:2312.00849

OmniLMM-12B

DeepCache: Accelerating Diffusion Models for Free

arXiv1 repo

arXiv:2312.00858

DeepCache

Segment and Caption Anything

arXiv1 repo

arXiv:2312.00869

segment-caption-anything

arXiv:2312.01305

arXiv1 repo

arXiv:2312.01305

vivid123

arXiv:2312.01678

arXiv1 repo

arXiv:2312.01678

Jellyfish-13B

StableVITON: Learning Semantic Correspondence with Latent Diffusion Model for Virtual Try-On

arXiv1 repo

arXiv:2312.01725

StableVITON

Language-only Efficient Training of Zero-shot Composed Image Retrieval

arXiv1 repo

arXiv:2312.01998

lincir

Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

arXiv1 repo

arXiv:2312.02119

GA

QuantAttack: Exploiting Dynamic Quantization to Attack Vision Transformers

arXiv1 repo

arXiv:2312.02220

QuantAttack

Recursive Visual Programming

arXiv1 repo

arXiv:2312.02249

RVP

RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!

arXiv1 repo

arXiv:2312.02724

pyterrier_genrank

Customization Assistant for Text-to-image Generation

arXiv1 repo

arXiv:2312.03045

ProFusion

LooseControl: Lifting ControlNet for Generalized Depth Conditioning

arXiv1 repo

arXiv:2312.03079

LooseControl

DiffusionSat: A Generative Foundation Model for Satellite Imagery

arXiv1 repo

arXiv:2312.03606

DiffusionSat

Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers

arXiv1 repo

arXiv:2312.03694

PETL_AST

Graph Convolutions Enrich the Self-Attention in Transformers!

arXiv1 repo

arXiv:2312.04234

GFSA

Smooth Diffusion: Crafting Smooth Latent Spaces in Diffusion Models

arXiv1 repo

arXiv:2312.04410

Smooth-Diffusion

DreamVideo: Composing Your Dream Videos with Customized Subject and Motion

arXiv1 repo

arXiv:2312.04433

i2vgen-xl

Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation

arXiv1 repo

arXiv:2312.04483

i2vgen-xl

An LLM Compiler for Parallel Function Calling

arXiv1 repo

arXiv:2312.04511

generative-ai

Generating Illustrated Instructions

arXiv1 repo

arXiv:2312.04552

generating-illustrated-instructions-reproduction

HuRef: HUman-REadable Fingerprint for Large Language Models

arXiv1 repo

arXiv:2312.04828

HuRef

Text-to-3D Generation with Bidirectional Diffusion using both 2D and 3D priors

arXiv1 repo

arXiv:2312.04963

bidiff

SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score Distillation

arXiv1 repo

arXiv:2312.05239

SwiftBrush

SlimSAM: 0.1% Data Makes Segment Anything Slim

arXiv1 repo

arXiv:2312.05284

SAMReg

Jumpstarting Surgical Computer Vision

arXiv1 repo

arXiv:2312.05968

Endoscapes

Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

arXiv1 repo

arXiv:2312.06109

GOT-OCR2_0

Cataract-1K: Cataract Surgery Dataset for Scene Segmentation, Phase Recognition, and Irregularity Detection

arXiv1 repo

arXiv:2312.06295

Cataract-1K

EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion

arXiv1 repo

arXiv:2312.06725

EpiDiff

Honeybee: Locality-enhanced Projector for Multimodal LLM

arXiv1 repo

arXiv:2312.06742

honeybee

Encoding Surgical Videos as Latent Spatiotemporal Graphs for Object and Anatomy-Driven Reasoning

arXiv1 repo

arXiv:2312.06829

Endoscapes

Reducing Energy Bloat in Large Model Training

arXiv1 repo

arXiv:2312.06902

zeus

BIRB: A Generalization Benchmark for Information Retrieval in Bioacoustics

arXiv1 repo

arXiv:2312.07439

perch

arXiv:2312.07509

arXiv1 repo

arXiv:2312.07509

Peekaboo

FreeControl: Training-Free Spatial Control of Any Text-to-Image Diffusion Model with Any Condition

arXiv1 repo

arXiv:2312.07536

freecontrol

PaperQA: Retrieval-Augmented Generative Agent for Scientific Research

arXiv1 repo

arXiv:2312.07559

paper-qa

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

arXiv1 repo

arXiv:2312.08168

Chat-Scene

A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions

arXiv1 repo

arXiv:2312.08578

DCI

RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution

arXiv1 repo

arXiv:2312.08617

magicoder

MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning

arXiv1 repo

arXiv:2312.08636

nugie-jax-nemotron-3-nano

VideoLCM: Video Latent Consistency Model

arXiv1 repo

arXiv:2312.09109

i2vgen-xl

Marathon: A Race Through the Realm of Long Context with Large Language Models

arXiv1 repo

arXiv:2312.09542

Marathon

MobileSAMv2: Faster Segment Anything to Everything

arXiv1 repo

arXiv:2312.09579

MobileSAM

Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference

arXiv1 repo

arXiv:2312.09608

Faster-Diffusion

Retrieval-Augmented Generation for Large Language Models: A Survey

arXiv1 repo

arXiv:2312.10997

TinyRAG

G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model

arXiv1 repo

arXiv:2312.11370

R1-V

arXiv:2312.11805

arXiv1 repo

arXiv:2312.11805

CoBSAT

The Endoscapes Dataset for Surgical Scene Segmentation, Object Detection, and Critical View of Safety Assessment: Official Splits and Benchmark

arXiv1 repo

arXiv:2312.12429

Endoscapes

ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation

arXiv1 repo

arXiv:2312.13108

WorldGUI

Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting

arXiv1 repo

arXiv:2312.13271

repaint123

The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction

arXiv1 repo

arXiv:2312.13558

AIDocks

TraceFL: Interpretability-Driven Debugging in Federated Learning via Neuron Provenance

arXiv1 repo

arXiv:2312.13632

TraceFL

Typhoon: Thai Large Language Models

arXiv1 repo

arXiv:2312.13951

typhoon-7b

DUSt3R: Geometric 3D Vision Made Easy

arXiv1 repo

arXiv:2312.14132

svraster

Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models

arXiv1 repo

arXiv:2312.14197

BIPIA

TACO: Topics in Algorithmic COde generation dataset

arXiv1 repo

arXiv:2312.14852

TACO

UniHuman: A Unified Model for Editing Human Images in the Wild

arXiv1 repo

arXiv:2312.14985

UniHuman

What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning

arXiv1 repo

arXiv:2312.15685

deita-10k-v0-sft

Align on the Fly: Adapting Chatbot Behavior to Established Norms

arXiv1 repo

arXiv:2312.15907

OPO

One-Dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications

arXiv1 repo

arXiv:2312.16145

SPM

Experiential Co-Learning of Software-Developing Agents

arXiv1 repo

arXiv:2312.17025

ChatDev

DreamGaussian4D: Generative 4D Gaussian Splatting

arXiv1 repo

arXiv:2312.17142

dreamgaussian4d

On-Demand JSON: A Better Way to Parse Documents?

arXiv1 repo

arXiv:2312.17149

simdjson

4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency

arXiv1 repo

arXiv:2312.17225

4DGen

Fast Inference of Mixture-of-Experts Language Models with Offloading

arXiv1 repo

arXiv:2312.17238

mixtral-offloading

RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models

arXiv1 repo

arXiv:2401.00396

MiniCheck

Fairness in Serving Large Language Models

arXiv1 repo

arXiv:2401.00588

S-LoRA

Benchmarking Large Language Models on Controllable Generation under Diversified Instructions

arXiv1 repo

arXiv:2401.00690

CoDI-Eval

ToolEyes: Fine-Grained Evaluation for Tool Learning Capabilities of Large Language Models in Real-world Scenarios

arXiv1 repo

arXiv:2401.00741

ToolEyes

Improving the Stability and Efficiency of Diffusion Models for Content Consistent Super-Resolution

arXiv1 repo

arXiv:2401.00877

OSEDiff

TrailBlazer: Trajectory Control for Diffusion-Based Video Generation

arXiv1 repo

arXiv:2401.00896

TrailBlazer

DocLLM: A layout-aware generative language model for multimodal document understanding

arXiv1 repo

arXiv:2401.00908

DocLLM

LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

arXiv1 repo

arXiv:2401.01325

LongLM

Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting

arXiv1 repo

arXiv:2401.01339

mystreetgscar

DiffusionEdge: Diffusion Probabilistic Model for Crisp Edge Detection

arXiv1 repo

arXiv:2401.02032

DiffusionEdge

An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

arXiv1 repo

arXiv:2401.02361

mmdetection

Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss

arXiv1 repo

arXiv:2401.02677

Segmind-Vega

German Text Embedding Clustering Benchmark

arXiv1 repo

arXiv:2401.02709

mteb-1.34.14

From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models

arXiv1 repo

arXiv:2401.02777

BELLE

SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems

arXiv1 repo

arXiv:2401.03945

SpeechGPT

FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild

arXiv1 repo

arXiv:2401.04210

FunnyNet-W

U-Mamba: Enhancing Long-range Dependency for Biomedical Image Segmentation

arXiv1 repo

arXiv:2401.04722

TriALS

A Simple Baseline for Spoken Language to Sign Language Translation with 3D Avatars

arXiv1 repo

arXiv:2401.04730

SLRT

Towards Online Continuous Sign Language Recognition and Translation

arXiv1 repo

arXiv:2401.05336

SLRT

Combating Adversarial Attacks with Multi-Agent Debate

arXiv1 repo

arXiv:2401.05998

rightmind

Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models

arXiv1 repo

arXiv:2401.06102

AudioLens

EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

arXiv1 repo

arXiv:2401.06201

JARVIS

How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

arXiv1 repo

arXiv:2401.06373

llm-jailbreaking-defense

TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models

arXiv1 repo

arXiv:2401.06620

TransliCo

Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation

arXiv1 repo

arXiv:2401.06643

LLM-div-incts

Don't Rank, Combine! Combining Machine Translation Hypotheses Using Quality Estimation

arXiv1 repo

arXiv:2401.06688

qe-fusion

arXiv:2401.08417

arXiv1 repo

arXiv:2401.08417

CPO_SIMPO

Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models

arXiv1 repo

arXiv:2401.08491

acl2025-contrastive-perplexity

Scalable Pre-training of Large Autoregressive Image Models

arXiv1 repo

arXiv:2401.08541

ml-aim

Tuning Language Models by Proxy

arXiv1 repo

arXiv:2401.08565

proxy-tuning

HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance

arXiv1 repo

arXiv:2401.08772

HuixiangDou

XTable in Action: Seamless Interoperability in Data Lakes

arXiv1 repo

arXiv:2401.09621

data-engineer-handbook

Self-Rewarding Language Models

arXiv1 repo

arXiv:2401.10020

fineweb-edu

MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

arXiv1 repo

arXiv:2401.10208

MM-Interleaved

CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference

arXiv1 repo

arXiv:2401.11240

lightllm

BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU

arXiv1 repo

arXiv:2401.11324

slater

Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

arXiv1 repo

arXiv:2401.11708

RPG_models

Benchmarking Large Multimodal Models against Common Corruptions

arXiv1 repo

arXiv:2401.11943

MMCBench

CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark

arXiv1 repo

arXiv:2401.11944

Yi-34B-Chat

DITTO: Diffusion Inference-Time T-Optimization for Music Generation

arXiv1 repo

arXiv:2401.12179

TuneJury

Universal Neurons in GPT2 Language Models

arXiv1 repo

arXiv:2401.12181

gemma-scope-2b-pt-transcoders

Text Embedding Inversion Security for Multilingual Language Models

arXiv1 repo

arXiv:2401.12192

MultiVec2Text

End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization

arXiv1 repo

arXiv:2401.12850

SHARC

Lumiere: A Space-Time Diffusion Model for Video Generation

arXiv1 repo

arXiv:2401.12945

videophy

GALA: Generating Animatable Layered Assets from a Single Scan

arXiv1 repo

arXiv:2401.12979

gala

PatternPortrait: Draw Me Like One of Your Scribbles

arXiv1 repo

arXiv:2401.13001

awesome-plotters

AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents

arXiv1 repo

arXiv:2401.13178

agentboard

SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation

arXiv1 repo

arXiv:2401.13527

SpeechGPT

Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild

arXiv1 repo

arXiv:2401.13627

SUPIR

Conformal Prediction Sets Improve Human Decision Making

arXiv1 repo

arXiv:2401.13744

hitl-conformal-prediction

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

arXiv1 repo

arXiv:2401.14159

GroundingDINO

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

arXiv1 repo

arXiv:2401.14196

Reinforcement-Learning-Full-Pipeline

Demystifying Chains, Trees, and Graphs of Thoughts

arXiv1 repo

arXiv:2401.14295

rightmind

MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache

arXiv1 repo

arXiv:2401.14361

MoE-Infinity

Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

arXiv1 repo

arXiv:2401.14624

Knowledge_Pile

Spatial Transcriptomics Analysis of Zero-shot Gene Expression Prediction

arXiv1 repo

arXiv:2401.14772

SGN

FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models

arXiv1 repo

arXiv:2401.15636

FreeStyle

StableIdentity: Inserting Anybody into Anywhere at First Sight

arXiv1 repo

arXiv:2401.15975

StableIdentity

Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

arXiv1 repo

arXiv:2401.16158

MobileAgent

Diffutoon: High-Resolution Editable Toon Shading via Diffusion Models

arXiv1 repo

arXiv:2401.16224

diffSynth-studio-notes

Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning

arXiv1 repo

arXiv:2401.17186

CLFM

Proactive Detection of Voice Cloning with Localized Watermarking

arXiv1 repo

arXiv:2401.17264

audioseal

YOLO-World: Real-Time Open-Vocabulary Object Detection

arXiv1 repo

arXiv:2401.17270

YOLO-World

Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens

arXiv1 repo

arXiv:2401.17377

infini-gram

KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization

arXiv1 repo

arXiv:2401.18079

KVQuant

ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models

arXiv1 repo

arXiv:2402.00794

ReAGent

Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters

arXiv1 repo

arXiv:2402.00828

PETL_AST

EVA-GAN: Enhanced Various Audio Generation via Scalable Generative Adversarial Networks

arXiv1 repo

arXiv:2402.00892

vocoder

Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack

arXiv1 repo

arXiv:2402.01109

Vaccine

arXiv:2402.01293

arXiv1 repo

arXiv:2402.01293

CoBSAT

KTO: Model Alignment as Prospect Theoretic Optimization

arXiv1 repo

arXiv:2402.01306

Reinforcement-Learning-Full-Pipeline

Rethinking Interpretability in the Era of Large Language Models

arXiv1 repo

arXiv:2402.01761

imodelsX

When Large Language Models Meet Vector Databases: A Survey

arXiv1 repo

arXiv:2402.01763

TinyRAG

S2malloc: Statistically Secure Allocator for Use-After-Free Protection And More

arXiv1 repo

arXiv:2402.01894

mimalloc-bench

EffiBench: Benchmarking the Efficiency of Automatically Generated Code

arXiv1 repo

arXiv:2402.02037

EffiBench

arXiv:2402.02057

arXiv1 repo

arXiv:2402.02057

LookaheadDecoding

BECLR: Batch Enhanced Contrastive Few-Shot Learning

arXiv1 repo

arXiv:2402.02444

awesome-cybersecurity-agentic-ai

Verifiable evaluations of machine learning models using zkSNARKs

arXiv1 repo

arXiv:2402.02675

zkbc

Position: What Can Large Language Models Tell Us about Time Series Analysis

arXiv1 repo

arXiv:2402.02713

time-moe

arXiv:2402.02952

arXiv1 repo

arXiv:2402.02952

awesome-ai-sre

Is Mamba Capable of In-Context Learning?

arXiv1 repo

arXiv:2402.03170

is_mamba_capable_of_icl

Training-Free Consistent Text-to-Image Generation

arXiv1 repo

arXiv:2402.03286

experimental-consistory

arXiv:2402.03310

arXiv1 repo

arXiv:2402.03310

VIRL

SpecFormer: Guarding Vision Transformer Robustness via Maximum Singular Value Penalization

arXiv1 repo

arXiv:2402.03317

robustlearn

Self-Discover: Large Language Models Self-Compose Reasoning Structures

arXiv1 repo

arXiv:2402.03620

awesome-dspy

MolTC: Towards Molecular Relational Modeling In Language Models

arXiv1 repo

arXiv:2402.03781

MolTC

AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls

arXiv1 repo

arXiv:2402.04253

gbrain

LESS: Selecting Influential Data for Targeted Instruction Tuning

arXiv1 repo

arXiv:2402.04333

bergson

The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry

arXiv1 repo

arXiv:2402.04347

LLaDA-Hybrid

ScreenAI: A Vision-Language Model for UI and Infographics Understanding

arXiv1 repo

arXiv:2402.04615

transformer-final-proj

Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning

arXiv1 repo

arXiv:2402.04833

OpenHermes-2.5-1k-longest

Data-efficient Large Vision Models through Sequential Autoregression

arXiv1 repo

arXiv:2402.04841

DeLVM

Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

arXiv1 repo

arXiv:2402.05109

Hydra

More Agents Is All You Need

arXiv1 repo

arXiv:2402.05120

rightmind

Zero-Shot Clinical Trial Patient Matching with LLMs

arXiv1 repo

arXiv:2402.05125

clinical_trial_patient_matching

AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

arXiv1 repo

arXiv:2402.05602

mirage

Learning to Route Among Specialized Experts for Zero-Shot Generalization

arXiv1 repo

arXiv:2402.05859

mttl

Let Your Graph Do the Talking: Encoding Structured Data for LLMs

arXiv1 repo

arXiv:2402.05862

GNN4TaskPlan

On the Out-Of-Distribution Generalization of Multimodal Large Language Models

arXiv1 repo

arXiv:2402.06599

OOD-Generalization-of-LMMs

Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

arXiv1 repo

arXiv:2402.06619

aya_dataset

OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning

arXiv1 repo

arXiv:2402.06954

FedLLM-Bench

Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models

arXiv1 repo

arXiv:2402.07033

fiddler

Dólares or Dollars? Unraveling the Bilingual Prowess of Financial LLMs Between Spanish and English

arXiv1 repo

arXiv:2402.07405

PIXIU

Benchmarking and Building Long-Context Retrieval Models with LoCo and M2-BERT

arXiv1 repo

arXiv:2402.07440

gte-multilingual-base

T-RAG: Lessons from the LLM Trenches

arXiv1 repo

arXiv:2402.07483

LARS

A Multinomial Canonical Decomposition Model, with emphasis on the analysis of Multivariate Binary data

arXiv1 repo

arXiv:2402.07634

smash

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

arXiv1 repo

arXiv:2402.07827

aya-101

arXiv:2402.07865

arXiv1 repo

arXiv:2402.07865

prismatic-vlms

eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data

arXiv1 repo

arXiv:2402.08831

eCeLLM

Attacking Large Language Models with Projected Gradient Descent

arXiv1 repo

arXiv:2402.09154

reinforce-attacks-llms

Less is More: Fewer Interpretable Region via Submodular Subset Selection

arXiv1 repo

arXiv:2402.09164

SMDL-Attribution

OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM

arXiv1 repo

arXiv:2402.09181

OmniMedVQA

Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code

arXiv1 repo

arXiv:2402.09299

TraWiC

LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset

arXiv1 repo

arXiv:2402.09391

BioMedGPT-Mol

CodeMind: Evaluating Large Language Models for Code Reasoning

arXiv1 repo

arXiv:2402.09664

CodeMind

QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference

arXiv1 repo

arXiv:2402.10076

QUICK

OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset

arXiv1 repo

arXiv:2402.10176

OpenMathInstruct-1

BitDelta: Your Fine-Tune May Only Be Worth One Bit

arXiv1 repo

arXiv:2402.10193

BitDelta

Chain-of-Thought Reasoning Without Prompting

arXiv1 repo

arXiv:2402.10200

LLM-Sampling

A StrongREJECT for Empty Jailbreaks

arXiv1 repo

arXiv:2402.10260

karma-electric-llama31-8b

BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains

arXiv1 repo

arXiv:2402.10373

BioMistral-7B

Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs

arXiv1 repo

arXiv:2402.10517

any-precision-llm

GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models

arXiv1 repo

arXiv:2402.10744

GenRES

ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages

arXiv1 repo

arXiv:2402.10753

Awesome-OpenClaw

Contrastive Instruction Tuning

arXiv1 repo

arXiv:2402.11138

CoIN

Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents

arXiv1 repo

arXiv:2402.11208

agent-backdoor-attacks

ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs

arXiv1 repo

arXiv:2402.11235

ZeroG

KMMLU: Measuring Massive Multitask Language Understanding in Korean

arXiv1 repo

arXiv:2402.11548

KMMLU

Machine-Generated Text Localization

arXiv1 repo

arXiv:2402.11744

MGT_Localization

ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

arXiv1 repo

arXiv:2402.11753

JailbreakLab

An Empirical Evaluation of LLMs for Solving Offensive Security Challenges

arXiv1 repo

arXiv:2402.11814

LLM_CTF

Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT

arXiv1 repo

arXiv:2402.12201

circuit_backup

A Critical Evaluation of AI Feedback for Aligning Large Language Models

arXiv1 repo

arXiv:2402.12366

dpo-rlaif

Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

arXiv1 repo

arXiv:2402.12374

mlx-flash

Understanding and Mitigating the Threat of Vec2Text to Dense Retrieval Systems

arXiv1 repo

arXiv:2402.12784

vec2text-dense_retriever-threat

Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data

arXiv1 repo

arXiv:2402.12869

opendataloader-bench

RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models

arXiv1 repo

arXiv:2402.12908

RealCompo

Benchmarking Retrieval-Augmented Generation for Medicine

arXiv1 repo

arXiv:2402.13178

textbooks

Transformer tricks: Precomputing the first layer

arXiv1 repo

arXiv:2402.13388

transformer-tricks

SDXL-Lightning: Progressive Adversarial Diffusion Distillation

arXiv1 repo

arXiv:2402.13929

SDXL-Lightning

arXiv:2402.14017

arXiv1 repo

arXiv:2402.14017

RectifID

BIRCO: A Benchmark of Information Retrieval Tasks with Complex Objectives

arXiv1 repo

arXiv:2402.14151

BIRCO

T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching

arXiv1 repo

arXiv:2402.14167

T-Stitch

INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models

arXiv1 repo

arXiv:2402.14334

InstructIR

Orca-Math: Unlocking the potential of SLMs in Grade School Math

arXiv1 repo

arXiv:2402.14830

orca-math-word-problems-193k-korean

arXiv:2402.14891

arXiv1 repo

arXiv:2402.14891

LLMBind

Watermarking Makes Language Models Radioactive

arXiv1 repo

arXiv:2402.14904

watermarks-remover

MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

arXiv1 repo

arXiv:2402.14905

minimind

tinyBenchmarks: evaluating LLMs with fewer examples

arXiv1 repo

arXiv:2402.14992

SWEBench-verified-mini

The AffectToolbox: Affect Analysis for Everyone

arXiv1 repo

arXiv:2402.15195

AffectToolbox

Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency

arXiv1 repo

arXiv:2402.15481

Bias-Volatility-Framework

AgentLite: A Lightweight Library for Building and Advancing Task-Oriented LLM Agent System

arXiv1 repo

arXiv:2402.15538

AgentLite

Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing

arXiv1 repo

arXiv:2402.16192

llm-jailbreaking-defense

IR2: Information Regularization for Information Retrieval

arXiv1 repo

arXiv:2402.16200

Information-Regularization

CodeS: Towards Building Open-source Language Models for Text-to-SQL

arXiv1 repo

arXiv:2402.16347

text2sql-demo

TOTEM: TOkenized Time Series EMbeddings for General Time Series Analysis

arXiv1 repo

arXiv:2402.16412

moment

Defending LLMs against Jailbreaking Attacks via Backtranslation

arXiv1 repo

arXiv:2402.16459

llm-jailbreaking-defense

A Survey on Data Selection for Language Models

arXiv1 repo

arXiv:2402.16827

llm_project

Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step

arXiv1 repo

arXiv:2402.16906

LLMDebugger

LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language

arXiv1 repo

arXiv:2402.16929

LangGPT

MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning

arXiv1 repo

arXiv:2402.17231

MathSensei

VoCo: A Simple-yet-Effective Volume Contrastive Learning Framework for 3D Medical Image Analysis

arXiv1 repo

arXiv:2402.17300

VoCo_Downstream

OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

arXiv1 repo

arXiv:2402.17553

omniact

From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions

arXiv1 repo

arXiv:2402.17633

ytseg

arXiv:2402.17700

arXiv1 repo

arXiv:2402.17700

SAEBench

Case-Based or Rule-Based: How Do Transformers Do the Math?

arXiv1 repo

arXiv:2402.17709

Case_or_Rule

Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

arXiv1 repo

arXiv:2402.17733

TowerInstruct-7B-v0.1

The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

arXiv1 repo

arXiv:2402.17764

BitNet

A Language Model based Framework for New Concept Placement in Ontologies

arXiv1 repo

arXiv:2402.17897

LM-ontology-concept-placement

arXiv:2402.18158

arXiv1 repo

arXiv:2402.18158

qllm-eval

Bluebell: An Alliance of Relational Lifting and Independence For Probabilistic Reasoning

arXiv1 repo

arXiv:2402.18708

ArkLib

Panda-70M: Captioning 70M Videos with Multiple Cross-Modality Teachers

arXiv1 repo

arXiv:2402.19479

Panda-70M

Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

arXiv1 repo

arXiv:2403.00231

VisRAG-Ret-Train-In-domain-data

AtP*: An efficient and scalable method for localizing LLM behaviour to components

arXiv1 repo

arXiv:2403.00745

notebooks

IntactKV: Improving Large Language Model Quantization by Keeping Pivot Tokens Intact

arXiv1 repo

arXiv:2403.01241

KVQuant

Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal

arXiv1 repo

arXiv:2403.01244

SSR

KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations

arXiv1 repo

arXiv:2403.01469

KorMedMCQA

Rethinking LLM Language Adaptation: A Case Study on Chinese Mixtral

arXiv1 repo

arXiv:2403.01851

Chinese-LLaMA-Alpaca-3

An Improved Traditional Chinese Evaluation Suite for Foundation Model

arXiv1 repo

arXiv:2403.01858

tmmluplus

InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

arXiv1 repo

arXiv:2403.02691

moltshield

KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents

arXiv1 repo

arXiv:2403.03101

KnowLM

CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation

arXiv1 repo

arXiv:2403.03411

CrossNet

NoiseCollage: A Layout-Aware Text-to-Image Diffusion Model Based on Noise Cropping and Merging

arXiv1 repo

arXiv:2403.03485

noisecollage

GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

arXiv1 repo

arXiv:2403.03507

Finetune_with_GaLore

MedSafetyBench: Evaluating and Improving the Medical Safety of Large Language Models

arXiv1 repo

arXiv:2403.03744

med-safety-bench

Learning to Decode Collaboratively with Multiple Language Models

arXiv1 repo

arXiv:2403.03870

co-llm

IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators

arXiv1 repo

arXiv:2403.03894

acl2024-ircoder

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders

arXiv1 repo

arXiv:2403.03952

AMAZON-Products-2023

Symmetry Considerations for Learning Task Symmetric Robot Policies

arXiv1 repo

arXiv:2403.04359

gauss_gym

Learning to Remove Wrinkled Transparent Film with Polarized Prior

arXiv1 repo

arXiv:2403.04368

FilmRemoval

PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

arXiv1 repo

arXiv:2403.04692

PixArt-Sigma-XL-2-512-MS

Face2Diffusion for Fast and Editable Face Personalization

arXiv1 repo

arXiv:2403.05094

Face2Diffusion

LightM-UNet: Mamba Assists in Lightweight UNet for Medical Image Segmentation

arXiv1 repo

arXiv:2403.05246

TriALS

VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models

arXiv1 repo

arXiv:2403.06098

VidProM

MACE: Mass Concept Erasure in Diffusion Models

arXiv1 repo

arXiv:2403.06135

MACE

No Language is an Island: Unifying Chinese and English in Financial Large Language Models, Instruction Data, and Benchmarks

arXiv1 repo

arXiv:2403.06249

PIXIU

Editing Conceptual Knowledge for Large Language Models

arXiv1 repo

arXiv:2403.06259

ConceptEdit

CLIcK: A Benchmark Dataset of Cultural and Linguistic Intelligence in Korean

arXiv1 repo

arXiv:2403.06412

Phi-3.5-MoE-instruct

SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection

arXiv1 repo

arXiv:2403.06534

sardet_100k

Leveraging Foundation Models for Content-Based Image Retrieval in Radiology

arXiv1 repo

arXiv:2403.06567

foundation-models-for-cbmir

Real-Time Multimodal Cognitive Assistant for Emergency Medical Services

arXiv1 repo

arXiv:2403.06734

EMS-Pipeline

Accurate Spatial Gene Expression Prediction by integrating Multi-resolution features

arXiv1 repo

arXiv:2403.07592

TRIPLEX

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

arXiv1 repo

arXiv:2403.07974

code_generation_lite

A bargain for mergesorts -- How to prove your mergesort correct and stable, almost for free

arXiv1 repo

arXiv:2403.08173

stablesort

Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator

arXiv1 repo

arXiv:2403.08495

MING

Language-Grounded Dynamic Scene Graphs for Interactive Object Search with Mobile Manipulation

arXiv1 repo

arXiv:2403.08605

SmallPlan

Faster Projected GAN: Towards Faster Few-Shot Image Generation

arXiv1 repo

arXiv:2403.08778

Faster-Projected-GAN

Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset

arXiv1 repo

arXiv:2403.09029

WebSight

CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences

arXiv1 repo

arXiv:2403.09032

magicoder

RAGGED: Towards Informed Design of Scalable and Stable RAG Systems

arXiv1 repo

arXiv:2403.09040

ragged

SemanticDraw: Towards Real-Time Interactive Content Creation from Image Diffusion Models

arXiv1 repo

arXiv:2403.09055

semantic-draw

Dial-insight: Fine-tuning Large Language Models with High-Quality Domain-Specific Data Preventing Capability Collapse

arXiv1 repo

arXiv:2403.09167

BELLE

Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

arXiv1 repo

arXiv:2403.09472

easy-to-hard

Hyper-CL: Conditioning Sentence Representations with Hypernetworks

arXiv1 repo

arXiv:2403.09490

Hyper-CL

Repoformer: Selective Retrieval for Repository-Level Code Completion

arXiv1 repo

arXiv:2403.10059

Repoformer

Improving Medical Multi-modal Contrastive Learning with Expert Annotations

arXiv1 repo

arXiv:2403.10153

eCLIP

Block Verification Accelerates Speculative Decoding

arXiv1 repo

arXiv:2403.10444

Hierarchical-Speculative-Decoding

VideoAgent: Long-form Video Understanding with Large Language Model as Agent

arXiv1 repo

arXiv:2403.10517

VideoTree

S3LLM: Large-Scale Scientific Software Understanding with LLMs using Source, Metadata, and Document

arXiv1 repo

arXiv:2403.10588

S3LLM

LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

arXiv1 repo

arXiv:2403.11703

OmniLMM-12B

SV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video Diffusion

arXiv1 repo

arXiv:2403.12008

DiffSplat

One-Step Image Translation with Text-to-Image Models

arXiv1 repo

arXiv:2403.12036

img2img-turbo

Distilling Datasets Into Less Than One Image

arXiv1 repo

arXiv:2403.12040

PoDD

EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models

arXiv1 repo

arXiv:2403.12171

EasyJailbreak

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

arXiv1 repo

arXiv:2403.12895

mPLUG-DocOwl

LLMLingua-2: Data Distillation for Efficient and Faithful Task-Agnostic Prompt Compression

arXiv1 repo

arXiv:2403.12968

llmlingua-2-bert-base-multilingual-cased-meetingbank

When Do We Not Need Larger Vision Models?

arXiv1 repo

arXiv:2403.13043

scaling_on_scales

AdaptSFL: Adaptive Split Federated Learning in Resource-constrained Edge Networks

arXiv1 repo

arXiv:2403.13101

SplitFM

Mora: Enabling Generalist Video Generation via A Multi-Agent Framework

arXiv1 repo

arXiv:2403.13248

Mora

Arcee's MergeKit: A Toolkit for Merging Large Language Models

arXiv1 repo

arXiv:2403.13257

donutloop-genesis

Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding

arXiv1 repo

arXiv:2403.14174

UniSDNet

SyncTweedies: A General Generative Framework Based on Synchronized Diffusions

arXiv1 repo

arXiv:2403.14370

FlexiSyncMVD

AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

arXiv1 repo

arXiv:2403.14468

AnyV2V

Implicit Style-Content Separation using B-LoRA

arXiv1 repo

arXiv:2403.14572

B-LoRA

Foundation Models for Time Series Analysis: A Tutorial and Survey

arXiv1 repo

arXiv:2403.14735

time-moe

VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding

arXiv1 repo

arXiv:2403.14743

VURF

KeyPoint Relative Position Encoding for Face Recognition

arXiv1 repo

arXiv:2403.14852

AdaFace

Long-CLIP: Unlocking the Long-Text Capability of CLIP

arXiv1 repo

arXiv:2403.15378

Long-CLIP

LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

arXiv1 repo

arXiv:2403.15388

LLaVA-PruMerge

ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models

arXiv1 repo

arXiv:2403.16167

ESREAL

Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms

arXiv1 repo

arXiv:2403.17806

KnowledgeCircuits

Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography

arXiv1 repo

arXiv:2403.17834

CT-CLIP

2D Gaussian Splatting for Geometrically Accurate Radiance Fields

arXiv1 repo

arXiv:2403.17888

svraster

LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning

arXiv1 repo

arXiv:2403.17919

LMFlow

arXiv:2403.18118

arXiv1 repo

arXiv:2403.18118

egolifter

Annolid: Annotate, Segment, and Track Anything You Need

arXiv1 repo

arXiv:2403.18690

annolid

TextCraftor: Your Text Encoder Can be Image Quality Controller

arXiv1 repo

arXiv:2403.18978

textcraftor

LITA: Language Instructed Temporal-Localization Assistant

arXiv1 repo

arXiv:2403.19046

LITA

Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

arXiv1 repo

arXiv:2403.19114

evoeval

sDPO: Don't Use Your Data All at Once

arXiv1 repo

arXiv:2403.19270

SOLAR-10.7B-Instruct-v1.0

OAKINK2: A Dataset of Bimanual Hands-Object Manipulation in Complex Task Completion

arXiv1 repo

arXiv:2403.19417

OakInk-v2

Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

arXiv1 repo

arXiv:2403.19521

Factual-Recall-Mechanism

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

arXiv1 repo

arXiv:2403.19647

notebooks

Jamba: A Hybrid Transformer-Mamba Language Model

arXiv1 repo

arXiv:2403.19887

Jamba-v0.1

TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods

arXiv1 repo

arXiv:2403.20150

TAB

Monocular Identity-Conditioned Facial Reflectance Reconstruction

arXiv1 repo

arXiv:2404.00301

insightface

Towards Variable and Coordinated Holistic Co-Speech Motion Generation

arXiv1 repo

arXiv:2404.00368

probtalk

QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs

arXiv1 repo

arXiv:2404.00456

deepcompressor

EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories

arXiv1 repo

arXiv:2404.00599

EvoCodeBench

WavLLM: Towards Robust and Adaptive Speech Large Language Model

arXiv1 repo

arXiv:2404.00656

SpeechT5

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

arXiv1 repo

arXiv:2404.01318

JBB-Behaviors

arXiv:2404.01363

arXiv1 repo

arXiv:2404.01363

awesome-ai-sre

Are large language models superhuman chemists?

arXiv1 repo

arXiv:2404.01475

chembench

M2SA: Multimodal and Multilingual Model for Sentiment Analysis of Tweets

arXiv1 repo

arXiv:2404.01753

M2SA-multimodal-multilingual-sentiment-analysis

CameraCtrl: Enabling Camera Control for Text-to-Video Generation

arXiv1 repo

arXiv:2404.02101

CameraCtrl

Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

arXiv1 repo

arXiv:2404.02151

jailbreakbench

LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIP

arXiv1 repo

arXiv:2404.02285

BiomedCoOp

On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons

arXiv1 repo

arXiv:2404.02431

lang_neuron

Sound Borrow-Checking for Rust via Symbolic Semantics (Long Version)

arXiv1 repo

arXiv:2404.02680

aeneas

InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation

arXiv1 repo

arXiv:2404.02733

InstantStyle

Faster Diffusion via Temporal Attention Decomposition

arXiv1 repo

arXiv:2404.02747

T-GATE

arXiv:2404.02949

arXiv1 repo

arXiv:2404.02949

feud

Skeleton Recall Loss for Connectivity Conserving and Resource Efficient Segmentation of Thin Tubular Structures

arXiv1 repo

arXiv:2404.03010

TotalSegmentator

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

arXiv1 repo

arXiv:2404.03413

MiniGPT4-Video

DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sampling

arXiv1 repo

arXiv:2404.03575

DreamScene

arXiv:2404.04057

arXiv1 repo

arXiv:2404.04057

ml-sid-dit

No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance

arXiv1 repo

arXiv:2404.04125

frequency_determines_performance

Robust Depth Enhancement via Polarization Prompt Fusion Tuning

arXiv1 repo

arXiv:2404.04318

Polarization-Prompt-Fusion-Tuning

Idea23D: Collaborative LMM Agents Enable 3D Model Generation from Interleaved Multimodal Inputs

arXiv1 repo

arXiv:2404.04363

Idea23D

arXiv:2404.04475

arXiv1 repo

arXiv:2404.04475

alpaca_eval

MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems

arXiv1 repo

arXiv:2404.04735

math-solving

AI2Apps: A Visual IDE for Building LLM-based AI Agent Applications

arXiv1 repo

arXiv:2404.04902

ai2apps

Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts

arXiv1 repo

arXiv:2404.05019

colibri

DinoBloom: A Foundation Model for Generalizable Cell Embeddings in Hematology

arXiv1 repo

arXiv:2404.05022

DinoBloom

MedExpQA: Multilingual Benchmarking of Large Language Models for Medical Question Answering

arXiv1 repo

arXiv:2404.05590

PodGPT

VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain

arXiv1 repo

arXiv:2404.05659

MultiMed

SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing

arXiv1 repo

arXiv:2404.05717

swap-anything

Learning Embeddings with Centroid Triplet Loss for Object Identification in Robotic Grasping

arXiv1 repo

arXiv:2404.06277

image_agnostic_segmentation

Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models

arXiv1 repo

arXiv:2404.06309

ClipClap-GZSL

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

arXiv1 repo

arXiv:2404.06395

MiniCPM3-4B

Automated Federated Pipeline for Parameter-Efficient Fine-Tuning of Large Language Models

arXiv1 repo

arXiv:2404.06448

SplitFM

Autonomous Evaluation and Refinement of Digital Agents

arXiv1 repo

arXiv:2404.06474

Agent-Eval-Refine

Visually Descriptive Language Model for Vector Graphics Reasoning

arXiv1 repo

arXiv:2404.06479

vtracer

Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?

arXiv1 repo

arXiv:2404.06644

khayyam-challenge

arXiv:2404.06798

arXiv1 repo

arXiv:2404.06798

uMedGround

GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM Applications

arXiv1 repo

arXiv:2404.06921

gorilla

Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

arXiv1 repo

arXiv:2404.07143

InfiniTransformer

InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models

arXiv1 repo

arXiv:2404.07191

InstantMesh

RealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth Diffusion

arXiv1 repo

arXiv:2404.07199

realmdreamer

Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain

arXiv1 repo

arXiv:2404.07613

Multilingual-Medical-Corpus

rollama: An R package for using generative large language models through Ollama

arXiv1 repo

arXiv:2404.07654

rollama

High-Dimension Human Value Representation in Large Language Models

arXiv1 repo

arXiv:2404.07900

UniVaR

Taming Stable Diffusion for Text to 360° Panorama Image Generation

arXiv1 repo

arXiv:2404.07949

PanFusion

Manipulating Large Language Models to Increase Product Visibility

arXiv1 repo

arXiv:2404.07981

llm-rank-optimizer

OpenBias: Open-set Bias Detection in Text-to-Image Generative Models

arXiv1 repo

arXiv:2404.07990

OpenBias

Revisiting Feature Prediction for Learning Visual Representations from Video

arXiv1 repo

arXiv:2404.08471

xwm

Probing the 3D Awareness of Visual Foundation Models

arXiv1 repo

arXiv:2404.08636

probe3d

COCONut: Modernizing COCO Segmentation

arXiv1 repo

arXiv:2404.08639

coconut_cvpr2024

MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts

arXiv1 repo

arXiv:2404.09027

MING

arXiv:2404.09837

arXiv1 repo

arXiv:2404.09837

awesome-ai-sre

Demonstration of DB-GPT: Next Generation Data Interaction System Empowered by Large Language Models

arXiv1 repo

arXiv:2404.10209

DB-GPT

MobileNetV4 -- Universal Models for the Mobile Ecosystem

arXiv1 repo

arXiv:2404.10518

MaaAI

ViTextVQA: A Large-Scale Visual Question Answering Dataset and a Novel Multimodal Feature Fusion Method for Vietnamese Text Comprehension in Images

arXiv1 repo

arXiv:2404.10652

ViTextVQA

Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

arXiv1 repo

arXiv:2404.10719

ReaLHF

Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes

arXiv1 repo

arXiv:2404.10772

mip-splatting

TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image Editing

arXiv1 repo

arXiv:2404.11120

TiNO-Edit

LongEmbed: Extending Embedding Models for Long Context Retrieval

arXiv1 repo

arXiv:2404.12096

mteb-1.34.14

Advancing the Robustness of Large Language Models through Self-Denoised Smoothing

arXiv1 repo

arXiv:2404.12274

SelfDenoise

KV-weights are all you need for skipless transformers

arXiv1 repo

arXiv:2404.12362

transformer-tricks

Lean Copilot: Large Language Models as Copilots for Theorem Proving in Lean

arXiv1 repo

arXiv:2404.12534

LeanCopilot

Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs

arXiv1 repo

arXiv:2404.13033

LLM-Tuning

SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs

arXiv1 repo

arXiv:2404.13081

ICLR24_SuRe

Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

arXiv1 repo

arXiv:2404.13686

Hyper-SD

Guess The Unseen: Dynamic 3D Scene Reconstruction from Partial 2D Glimpses

arXiv1 repo

arXiv:2404.14410

gtu

SnapKV: LLM Knows What You are Looking for Before Generation

arXiv1 repo

arXiv:2404.14469

GUI-KV

Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous Decoding

arXiv1 repo

arXiv:2404.14600

lost-in-decoding

FlashSpeech: Efficient Zero-Shot Speech Synthesis

arXiv1 repo

arXiv:2404.14700

FlashSpeech

Setting up the Data Printer with Improved English to Ukrainian Machine Translation

arXiv1 repo

arXiv:2404.15196

dragoman

CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies

arXiv1 repo

arXiv:2404.15238

modular_pluralism

TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting

arXiv1 repo

arXiv:2404.15264

TalkingGaussian

From Parts to Whole: A Unified Reference Framework for Controllable Human Image Generation

arXiv1 repo

arXiv:2404.15267

DeepFashion-MultiModal-Parts2Whole

Metric3Dv2: A Versatile Monocular Geometric Foundation Model for Zero-shot Metric Depth and Surface Normal Estimation

arXiv1 repo

arXiv:2404.15506

Depth-Estimation

Semantic Routing for Enhanced Performance of LLM-Assisted Intent-Based 5G Core Network Management and Orchestration

arXiv1 repo

arXiv:2404.15869

semantic-router

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

arXiv1 repo

arXiv:2404.16006

MMT-Bench

GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian Splatting

arXiv1 repo

arXiv:2404.16012

GaussianTalker

arXiv:2404.16014

arXiv1 repo

arXiv:2404.16014

dictionary_learning

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

arXiv1 repo

arXiv:2404.16019

VoteSim

MoDE: CLIP Data Experts via Clustering

arXiv1 repo

arXiv:2404.16030

MetaCLIP

Validating Traces of Distributed Programs Against TLA+ Specifications

arXiv1 repo

arXiv:2404.16075

formal-web

Leveraging tropical reef, bird and unrelated sounds for superior transfer learning in marine bioacoustics

arXiv1 repo

arXiv:2404.16436

perch

TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning

arXiv1 repo

arXiv:2404.16635

mPLUG-DocOwl

LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding

arXiv1 repo

arXiv:2404.16710

mlx-flash

SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

arXiv1 repo

arXiv:2404.16790

SEED-Bench

When to Trust LLMs: Aligning Confidence with Response Quality

arXiv1 repo

arXiv:2404.17287

CONQORD

BlenderAlchemy: Editing 3D Graphics with Vision-Language Models

arXiv1 repo

arXiv:2404.17672

BlenderAlchemyOfficial

Diffusion-Aided Joint Source Channel Coding For High Realism Wireless Image Transmission

arXiv1 repo

arXiv:2404.17736

DiffJSCC

KAN: Kolmogorov-Arnold Networks

arXiv1 repo

arXiv:2404.19756

imodelsX

Clover: Regressive Lightweight Speculative Decoding with Sequential Knowledge

arXiv1 repo

arXiv:2405.00263

clover

WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting

arXiv1 repo

arXiv:2405.00823

WorkBench

On Mechanistic Knowledge Localization in Text-to-Image Generative Models

arXiv1 repo

arXiv:2405.01008

LocoGen

MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors

arXiv1 repo

arXiv:2405.01413

MiniGPT-3D

StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation

arXiv1 repo

arXiv:2405.01434

StoryDiffusion

Automatically Extracting Numerical Results from Randomized Controlled Trials with Large Language Models

arXiv1 repo

arXiv:2405.01686

llm-meta-analysis

Autoformalizing Natural Language to First-Order Logic: A Case Study in Logical Fallacy Detection

arXiv1 repo

arXiv:2405.02318

NL2FOL

Labeling supervised fine-tuning data with the scaling law

arXiv1 repo

arXiv:2405.02817

HuixiangDou

MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition

arXiv1 repo

arXiv:2405.03152

WenetSpeech-Chuan

Bridging discrete and continuous state spaces: Exploring the Ehrenfest process in time-continuous diffusion models

arXiv1 repo

arXiv:2405.03549

EhrenfestDiffusion

A Construct-Optimize Approach to Sparse View Synthesis without Camera Pose

arXiv1 repo

arXiv:2405.03659

COGS

sqlelf: a SQL-centric Approach to ELF Analysis

arXiv1 repo

arXiv:2405.03883

selfdb

Iterative Experience Refinement of Software-Developing Agents

arXiv1 repo

arXiv:2405.04219

ChatDev

xLSTM: Extended Long Short-Term Memory

arXiv1 repo

arXiv:2405.04517

xlstm

Bridging the Bosphorus: Advancing Turkish Large Language Models through Strategies for Low-Resource Language Adaptation and Benchmarking

arXiv1 repo

arXiv:2405.04685

turkish-llm

You Only Cache Once: Decoder-Decoder Architectures for Language Models

arXiv1 repo

arXiv:2405.05254

LCKV

Mirage: A Multi-Level Superoptimizer for Tensor Programs

arXiv1 repo

arXiv:2405.05751

mirage

Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

arXiv1 repo

arXiv:2405.05904

ChroKnowledge

An Investigation of Incorporating Mamba for Speech Enhancement

arXiv1 repo

arXiv:2405.06573

SEMamba

LLM-Generated Black-box Explanations Can Be Adversarially Helpful

arXiv1 repo

arXiv:2405.06800

adversarial_helpfulness

PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play Accelerator

arXiv1 repo

arXiv:2405.07510

Rectified-Diffusion

MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning

arXiv1 repo

arXiv:2405.07551

aimo-progress-prize

Characterizing virulence differences in a parasitoid wasp through comparative transcriptomic and proteomic

arXiv1 repo

arXiv:2405.07772

PusaV1

Forecasting with Hyper-Trees

arXiv1 repo

arXiv:2405.07836

Hyper-Trees

UnMarker: A Universal Attack on Defensive Image Watermarking

arXiv1 repo

arXiv:2405.08363

watermarks-remover

Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video

arXiv1 repo

arXiv:2405.08890

PDL

MTP: A Meaning-Typed Language Abstraction for AI-Integrated Programming

arXiv1 repo

arXiv:2405.08965

jac

A safety realignment framework via subspace-oriented model fusion for large language models

arXiv1 repo

arXiv:2405.09055

safety_realignment

Chameleon: Mixed-Modal Early-Fusion Foundation Models

arXiv1 repo

arXiv:2405.09818

YoChameleon

DocuMint: Docstring Generation for Python using Small Language Models

arXiv1 repo

arXiv:2405.10243

DocuMint

PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology

arXiv1 repo

arXiv:2405.10254

TITAN

One registration is worth two segmentations

arXiv1 repo

arXiv:2405.10879

SAMReg

Adhesion of a nematic elastomer cylinder

arXiv1 repo

arXiv:2405.11116

compagent

OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

arXiv1 repo

arXiv:2405.11143

selfplay-redteaming

Towards Modular LLMs by Building and Reusing a Library of LoRAs

arXiv1 repo

arXiv:2405.11157

mttl

DocReLM: Mastering Document Retrieval with Language Model

arXiv1 repo

arXiv:2405.11461

sci-bert-finetune

Training Data Attribution via Approximate Unrolled Differentiation

arXiv1 repo

arXiv:2405.12186

bergson

Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices

arXiv1 repo

arXiv:2405.12211

Slicedit

Images that Sound: Composing Images and Sounds on a Single Canvas

arXiv1 repo

arXiv:2405.12221

images-that-sound

Tagengo: A Multilingual Chat Dataset

arXiv1 repo

arXiv:2405.12612

suzume-llama-3-8B-japanese

Pytorch-Wildlife: A Collaborative Deep Learning Framework for Conservation

arXiv1 repo

arXiv:2405.12930

Depth-Estimation

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

arXiv1 repo

arXiv:2405.12981

LCKV

FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research

arXiv1 repo

arXiv:2405.13576

FlashRAG_datasets

Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation

arXiv1 repo

arXiv:2405.13622

auto-rag-eval

Irreducibility in generalized power series

arXiv1 repo

arXiv:2405.13815

conway-refinement

Automatically Identifying Local and Global Circuits with Linear Computation Graphs

arXiv1 repo

arXiv:2405.13868

circuit_backup

Focus Anywhere for Fine-grained Multi-page Document Understanding

arXiv1 repo

arXiv:2405.14295

GOT-OCR2_0

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference

arXiv1 repo

arXiv:2405.14430

xDiT

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

arXiv1 repo

arXiv:2405.14573

UI-TARS

exLong: Generating Exceptional Behavior Tests with Large Language Models

arXiv1 repo

arXiv:2405.14619

exLong

arXiv:2405.14677

arXiv1 repo

arXiv:2405.14677

RectifID

HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models

arXiv1 repo

arXiv:2405.14831

hippo-memory

Extracting Prompts by Inverting LLM Outputs

arXiv1 repo

arXiv:2405.15012

output2prompt

Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization

arXiv1 repo

arXiv:2405.15071

GrokkedTransformer

DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

arXiv1 repo

arXiv:2405.15232

DEEM

arXiv:2405.15593

arXiv1 repo

arXiv:2405.15593

MicroAdam

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

arXiv1 repo

arXiv:2405.15793

mini-swe-agent

Intruding with Words: Towards Understanding Graph Injection Attacks at the Text Level

arXiv1 repo

arXiv:2405.16405

Text-level-Graph-Attack

SpinQuant: LLM quantization with learned rotations

arXiv1 repo

arXiv:2405.16406

turboquant-vllm

arXiv:2405.16444

arXiv1 repo

arXiv:2405.16444

CacheBlend

M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

arXiv1 repo

arXiv:2405.16473

M3CoT

Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation

arXiv1 repo

arXiv:2405.16504

UnifiedImplicitAttnRepr

Tool Learning in the Wild: Empowering Language Models as Automatic Tool Agents

arXiv1 repo

arXiv:2405.16533

AutoTools

Crafting Interpretable Embeddings by Asking LLMs Questions

arXiv1 repo

arXiv:2405.16714

imodelsX

Saturn: Sample-efficient Generative Molecular Design using Memory Manipulation

arXiv1 repo

arXiv:2405.17066

sego

ReMoDetect: Reward Models Recognize Aligned LLM's Generations

arXiv1 repo

arXiv:2405.17382

ReMoDetect-deberta

Getting More Juice Out of the SFT Data: Reward Learning from Human Demonstration Improves SFT for LLM Alignment

arXiv1 repo

arXiv:2405.17888

Reward_learning_SFT

Knowledge Circuits in Pretrained Transformers

arXiv1 repo

arXiv:2405.17969

KnowledgeCircuits

Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

arXiv1 repo

arXiv:2405.18392

modded-nanogpt

Learning diverse attacks on large language models for robust red-teaming and safety tuning

arXiv1 repo

arXiv:2405.18540

red-teaming

Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack

arXiv1 repo

arXiv:2405.18641

Lisa

NeRF On-the-go: Exploiting Uncertainty for Distractor-free NeRFs in the Wild

arXiv1 repo

arXiv:2405.18715

nerfonthego-undistorted

SketchDeco: Training-Free Latent Composition for Precise Sketch Colourisation

arXiv1 repo

arXiv:2405.18716

sketchdeco-code

T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward Feedback

arXiv1 repo

arXiv:2405.18750

t2v-turbo

Can Graph Learning Improve Planning in LLM-based Agents?

arXiv1 repo

arXiv:2405.19119

GNN4TaskPlan

VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

arXiv1 repo

arXiv:2405.19209

VideoTree

ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron Pruning

arXiv1 repo

arXiv:2405.19237

BackdoorDM

TotalSegmentator MRI: Robust Sequence-independent Segmentation of Multiple Anatomic Structures in MRI

arXiv1 repo

arXiv:2405.19492

TotalSegmentator

One-Shot Safety Alignment for Large Language Models via Optimal Dualization

arXiv1 repo

arXiv:2405.19544

CAN

EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos

arXiv1 repo

arXiv:2405.19644

EgoSurgery

Grokfast: Accelerated Grokking by Amplifying Slow Gradients

arXiv1 repo

arXiv:2405.20233

grokfast

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

arXiv1 repo

arXiv:2405.20279

CV-VAE

Improving the Training of Rectified Flows

arXiv1 repo

arXiv:2405.20320

Rectified-Diffusion

From Zero to Hero: Cold-Start Anomaly Detection

arXiv1 repo

arXiv:2405.20341

ColdFusion

Unique3D: High-Quality and Efficient 3D Mesh Generation from a Single Image

arXiv1 repo

arXiv:2405.20343

Unique3D

Unraveling and Mitigating Retriever Inconsistencies in Retrieval-Augmented Large Language Models

arXiv1 repo

arXiv:2405.20680

Ensemble-of-Retrievers

Grammar-Aligned Decoding

arXiv1 repo

arXiv:2405.21047

transformers-GAD

Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction

arXiv1 repo

arXiv:2405.21069

LPCNet

Learning Manipulation by Predicting Interaction

arXiv1 repo

arXiv:2406.00439

MPI

SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources with Retrieval and Semantic Parsing

arXiv1 repo

arXiv:2406.00562

wikipedia

Invisible Backdoor Attacks on Diffusion Models

arXiv1 repo

arXiv:2406.00816

BackdoorDM

UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation

arXiv1 repo

arXiv:2406.01188

UniAnimate-DiT

TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy

arXiv1 repo

arXiv:2406.01326

TabPedia_v1.0

BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM Safeguards

arXiv1 repo

arXiv:2406.01364

BELLS

arXiv:2406.01561

arXiv1 repo

arXiv:2406.01561

ml-sid-dit

Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA

arXiv1 repo

arXiv:2406.02044

QROA

PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling

arXiv1 repo

arXiv:2406.02069

GUI-KV

Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning

arXiv1 repo

arXiv:2406.02265

RobustCap

GrootVL: Tree Topology is All You Need in State Space Model

arXiv1 repo

arXiv:2406.02395

MindOmni

The Scandinavian Embedding Benchmarks: Comprehensive Assessment of Multilingual and Monolingual Text Embedding

arXiv1 repo

arXiv:2406.02396

mteb-1.34.14

Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques

arXiv1 repo

arXiv:2406.02500

Unified-MoE-Compression

Guiding a Diffusion Model with a Bad Version of Itself

arXiv1 repo

arXiv:2406.02507

edm2

Loki: Low-rank Keys for Efficient Sparse Attention

arXiv1 repo

arXiv:2406.02542

loki

Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition

arXiv1 repo

arXiv:2406.02554

glimmer

Collision-Affording Point Trees: SIMD-Amenable Nearest Neighbors for Fast Collision Checking

arXiv1 repo

arXiv:2406.02807

capt

Efficient Minimum Bayes Risk Decoding using Low-Rank Matrix Completion Algorithms

arXiv1 repo

arXiv:2406.02832

mbrs

DenoDet: Attention as Deformable Multi-Subspace Feature Denoising for Target Detection in SAR Images

arXiv1 repo

arXiv:2406.02833

sardet_100k

EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos

arXiv1 repo

arXiv:2406.03095

EgoSurgery

Text-to-Image Rectified Flow as Plug-and-Play Priors

arXiv1 repo

arXiv:2406.03293

InstaFlow

LW-DETR: A Transformer Replacement to YOLO for Real-Time Detection

arXiv1 repo

arXiv:2406.03459

rf-detr

BLSP-Emo: Towards Empathetic Large Speech-Language Models

arXiv1 repo

arXiv:2406.03872

Speech-IFEval

CDMamba: Incorporating Local Clues into Mamba for Remote Sensing Image Binary Change Detection

arXiv1 repo

arXiv:2406.04207

Land-Change-Detection

ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language Models

arXiv1 repo

arXiv:2406.04214

ValueBench

ReNO: Enhancing One-step Text-to-Image Models through Reward-based Noise Optimization

arXiv1 repo

arXiv:2406.04312

ReNO

LawGPT: A Chinese Legal Knowledge-Enhanced Large Language Model

arXiv1 repo

arXiv:2406.04614

LaWGPT

WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

arXiv1 repo

arXiv:2406.04770

WildBench

MoE Jetpack: From Dense Checkpoints to Adaptive Mixture of Experts for Vision Tasks

arXiv1 repo

arXiv:2406.04801

Monkey

FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models

arXiv1 repo

arXiv:2406.04845

FedLLM-Bench

Hibou: A Family of Foundational Vision Transformers for Pathology

arXiv1 repo

arXiv:2406.05074

dpfm_factory

DALD: Improving Logits-based Detector without Logits from Black-box LLMs

arXiv1 repo

arXiv:2406.05232

DALD

Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis

arXiv1 repo

arXiv:2406.05478

ImprovedNAT

Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

arXiv1 repo

arXiv:2406.05551

dots.tts

PaRa: Personalizing Text-to-Image Diffusion via Parameter Rank Reduction

arXiv1 repo

arXiv:2406.05641

DEFT

WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

arXiv1 repo

arXiv:2406.05763

F5-TTS

Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents

arXiv1 repo

arXiv:2406.05870

jamming_attack

Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

arXiv1 repo

arXiv:2406.05955

PowerInfer

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

arXiv1 repo

arXiv:2406.05967

cvqa

EpiLearn: A Python Library for Machine Learning in Epidemic Modeling

arXiv1 repo

arXiv:2406.06016

EpiLearn

PowerInfer-2: Fast Large Language Model Inference on a Smartphone

arXiv1 repo

arXiv:2406.06282

PowerInfer

UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor

arXiv1 repo

arXiv:2406.06519

reranker-as-judge

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

arXiv1 repo

arXiv:2406.06525

LlamaGen

Achieving Sparse Activation in Small Language Models

arXiv1 repo

arXiv:2406.06562

Sparse-Activation

SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM

arXiv1 repo

arXiv:2406.06571

subllm

AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising

arXiv1 repo

arXiv:2406.06911

DeepCache

MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

arXiv1 repo

arXiv:2406.07057

MMTrustEval

Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter

arXiv1 repo

arXiv:2406.07096

SLU_pipeline

NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images

arXiv1 repo

arXiv:2406.07111

NeRSP

Scaling Large Language Model-based Multi-Agent Collaboration

arXiv1 repo

arXiv:2406.07155

ChatDev

EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark

arXiv1 repo

arXiv:2406.07162

EmoBox

When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models

arXiv1 repo

arXiv:2406.07368

Linearized-LLM

MINERS: Multilingual Language Models as Semantic Retrievers

arXiv1 repo

arXiv:2406.07424

miners

Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

arXiv1 repo

arXiv:2406.07522

Samba

EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech

arXiv1 repo

arXiv:2406.07803

FastSpeech2-Plus

LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning

arXiv1 repo

arXiv:2406.07969

libritts-r-filtered-speaker-descriptions

One-Step Effective Diffusion Network for Real-World Image Super-Resolution

arXiv1 repo

arXiv:2406.08177

OSEDiff

VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

arXiv1 repo

arXiv:2406.08394

VisionLLM

Real3D: Scaling Up Large Reconstruction Models with Real-World Images

arXiv1 repo

arXiv:2406.08479

Real3D

Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis

arXiv1 repo

arXiv:2406.08568

TTDS

HelpSteer2: Open-source dataset for training top-performing reward models

arXiv1 repo

arXiv:2406.08673

Llama-3_3-Nemotron-Super-49B-GenRM

Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

arXiv1 repo

arXiv:2406.08801

hallo

EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts

arXiv1 repo

arXiv:2406.09162

ELLA

An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval

arXiv1 repo

arXiv:2406.09188

lincir

MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

arXiv1 repo

arXiv:2406.09297

LCKV

WonderWorld: Interactive 3D Scene Generation from a Single Image

arXiv1 repo

arXiv:2406.09394

WonderWorld

The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments

arXiv1 repo

arXiv:2406.09494

Displace2024_baseline_updated

$S^3$ -- Semantic Signal Separation

arXiv1 repo

arXiv:2406.09556

turftopic

Large language model validity via enhanced conformal prediction methods

arXiv1 repo

arXiv:2406.09714

OLAPH

Grounding Image Matching in 3D with MASt3R

arXiv1 repo

arXiv:2406.09756

svraster

A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention

arXiv1 repo

arXiv:2406.09827

hip-ainl

Nymeria: A Massive Collection of Multimodal Egocentric Daily Motion in the Wild

arXiv1 repo

arXiv:2406.09905

nymeria_dataset

BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages

arXiv1 repo

arXiv:2406.09948

BLEnD

Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection

arXiv1 repo

arXiv:2406.10052

WhisperLiveKit

GenQA: Generating Millions of Instructions from a Handful of Prompts

arXiv1 repo

arXiv:2406.10323

GenQA

STAR: Scale-wise Text-conditioned AutoRegressive image generation

arXiv1 repo

arXiv:2406.10797

VAR

garak: A Framework for Security Probing Large Language Models

arXiv1 repo

arXiv:2406.11036

garak

WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

arXiv1 repo

arXiv:2406.11069

vision-arena

arXiv:2406.11149

arXiv1 repo

arXiv:2406.11149

GoldCoin

Liberal Entity Matching as a Compound AI Toolchain

arXiv1 repo

arXiv:2406.11255

libem

VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

arXiv1 repo

arXiv:2406.11303

VideoVista_Train

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

arXiv1 repo

arXiv:2406.11546

AudioBench-N

DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling

arXiv1 repo

arXiv:2406.11617

ComfyUI-LoRA-Optimizer

Task Me Anything

arXiv1 repo

arXiv:2406.11775

TaskMeAnything-v1-imageqa-random

Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations

arXiv1 repo

arXiv:2406.11801

safety-arithmetic

VideoLLM-online: Online Video Large Language Model for Streaming Video

arXiv1 repo

arXiv:2406.11816

videollm-online

MegaScenes: Scene-Level View Synthesis at Scale

arXiv1 repo

arXiv:2406.11819

dataset

mDPO: Conditional Preference Optimization for Multimodal Large Language Models

arXiv1 repo

arXiv:2406.11839

mDPO

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

arXiv1 repo

arXiv:2406.11931

evalplus

Transcoders Find Interpretable LLM Feature Circuits

arXiv1 repo

arXiv:2406.11944

gemma-scope-2b-pt-transcoders

SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization

arXiv1 repo

arXiv:2406.12233

gujarati-vsr

Navigating Knowledge Management Implementation Success in Government Organizations: A type-2 fuzzy approach

arXiv1 repo

arXiv:2406.12345

Recap-COCO-30K

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

arXiv1 repo

arXiv:2406.12459

humansplat

Unified Active Retrieval for Retrieval Augmented Generation

arXiv1 repo

arXiv:2406.12534

UAR_qwen

DetectBench: Can Large Language Model Detect and Piece Together Implicit Evidence?

arXiv1 repo

arXiv:2406.12641

llm-mysteries

LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation

arXiv1 repo

arXiv:2406.12832

ComfyUI-LoRA-Optimizer

Medical Spoken Named Entity Recognition

arXiv1 repo

arXiv:2406.13337

MultiMed

VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models

arXiv1 repo

arXiv:2406.13362

VisualRWKV

VDebugger: Harnessing Execution Feedback for Debugging Visual Programs

arXiv1 repo

arXiv:2406.13444

vdebugger

SpatialBot: Precise Spatial Understanding with Vision Language Models

arXiv1 repo

arXiv:2406.13642

SpatialBot

Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation

arXiv1 repo

arXiv:2406.13663

mirage

Rethinking Abdominal Organ Segmentation (RAOS) in the clinical scenario: A robustness evaluation benchmark with challenging cases

arXiv1 repo

arXiv:2406.13674

RAOS

Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented Generation

arXiv1 repo

arXiv:2406.13692

sync-ralm-faithfulness

GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

arXiv1 repo

arXiv:2406.13743

t2v_metrics

EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms

arXiv1 repo

arXiv:2406.14228

evoagent

SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

arXiv1 repo

arXiv:2406.14598

unintentional-unalignment

Direct Multi-Turn Preference Optimization for Language Agents

arXiv1 repo

arXiv:2406.14868

DMPO

Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging

arXiv1 repo

arXiv:2406.15479

Twin-Merging

PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

arXiv1 repo

arXiv:2406.15513

PKU-SafeRLHF

Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-Polygraph

arXiv1 repo

arXiv:2406.15627

SAR

video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

arXiv1 repo

arXiv:2406.15704

SALMONN

What Matters in Transformers? Not All Attention is Needed

arXiv1 repo

arXiv:2406.15786

LLM-Drop

Real-time Speech Summarization for Medical Conversations

arXiv1 repo

arXiv:2406.15888

MultiMed

Teaching LLMs to Abstain across Languages via Multilingual Feedback

arXiv1 repo

arXiv:2406.15948

M-AbstainQA

Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

arXiv1 repo

arXiv:2406.15951

modular_pluralism

AudioBench: A Universal Benchmark for Audio Large Language Models

arXiv1 repo

arXiv:2406.16020

MERaLiON-AudioLLM-Whisper-SEA-LION

Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs

arXiv1 repo

arXiv:2406.16797

lottery-ticket-adaptation

RaTEScore: A Metric for Radiology Report Generation

arXiv1 repo

arXiv:2406.16845

PMC-LLaMA

StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal

arXiv1 repo

arXiv:2406.16864

StableNormal

CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation

arXiv1 repo

arXiv:2406.17186

CLERC

A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens

arXiv1 repo

arXiv:2406.17378

Text_aligns_tokens

MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation

arXiv1 repo

arXiv:2406.17484

MING

LumberChunker: Long-Form Narrative Document Segmentation

arXiv1 repo

arXiv:2406.17526

gacha

Training-Free Exponential Context Extension via Cascading KV Cache

arXiv1 repo

arXiv:2406.17808

cascading_kv_cache

E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

arXiv1 repo

arXiv:2406.18009

F5-TTS

RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

arXiv1 repo

arXiv:2406.18284

Sonic

GaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enhanced Quality

arXiv1 repo

arXiv:2406.18462

GaussianDreamerPro

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

arXiv1 repo

arXiv:2406.18495

selfplay-redteaming

WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

arXiv1 repo

arXiv:2406.18510

wildteaming

PrExMe! Large Scale Prompt Exploration of Open Source LLMs for Machine Translation and Summarization Evaluation

arXiv1 repo

arXiv:2406.18528

PrExMe

Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

arXiv1 repo

arXiv:2406.18583

Lumina-T2X

arXiv:2406.18925

arXiv1 repo

arXiv:2406.18925

VisArgs

AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation

arXiv1 repo

arXiv:2406.18958

AnyControl

DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability

arXiv1 repo

arXiv:2406.19135

DEX-TTS

PathAlign: A vision-language model for whole slide images in histopathology

arXiv1 repo

arXiv:2406.19578

VLSA

Fine-tuning of Geospatial Foundation Models for Aboveground Biomass Estimation

arXiv1 repo

arXiv:2406.19888

granite-geospatial-biomass

arXiv:2407.00023

arXiv1 repo

arXiv:2407.00023

preble

PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration

arXiv1 repo

arXiv:2407.00203

AEM-dataset

Tarsier: Recipes for Training and Evaluating Large Video Description Models

arXiv1 repo

arXiv:2407.00634

tarsier

Preserving Multilingual Quality While Tuning Query Encoder on English Only

arXiv1 repo

arXiv:2407.00923

arxiv-negatives

Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs

arXiv1 repo

arXiv:2407.00945

EEP

SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models

arXiv1 repo

arXiv:2407.00952

SplitFM

Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs

arXiv1 repo

arXiv:2407.01082

LLM-Sampling

BERGEN: A Benchmarking Library for Retrieval-Augmented Generation

arXiv1 repo

arXiv:2407.01102

bergen

arXiv:2407.01392

arXiv1 repo

arXiv:2407.01392

SkyReels-V2-I2V-14B-720P

Retrieval-augmented generation in multilingual settings

arXiv1 repo

arXiv:2407.01463

bergen

FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

arXiv1 repo

arXiv:2407.01494

FoleyCrafter

MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

arXiv1 repo

arXiv:2407.01523

MMLongBench-Doc

fVDB: A Deep-Learning Framework for Sparse, Large-Scale, and High-Performance Spatial Intelligence

arXiv1 repo

arXiv:2407.01781

SCube

VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs

arXiv1 repo

arXiv:2407.01863

Mirage

CoIR: A Comprehensive Benchmark for Code Information Retrieval Models

arXiv1 repo

arXiv:2407.02883

coir

AgentInstruct: Toward Generative Teaching with Agentic Flows

arXiv1 repo

arXiv:2407.03502

aurora-m2

M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks

arXiv1 repo

arXiv:2407.03791

m5b_v2

ChartGemma: Visual Instruction-tuning for Chart Reasoning in the Wild

arXiv1 repo

arXiv:2407.04172

chartgemma

arXiv:2407.04292

arXiv1 repo

arXiv:2407.04292

Corki

LaRa: Efficient Large-Baseline Radiance Fields

arXiv1 repo

arXiv:2407.04699

LaRa

UltraEdit: Instruction-based Fine-Grained Image Editing at Scale

arXiv1 repo

arXiv:2407.05282

detikzify-v2-8b

Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

arXiv1 repo

arXiv:2407.05361

F5-TTS

CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

arXiv1 repo

arXiv:2407.05407

BreezyVoice

See Further for Parameter Efficient Fine-tuning by Standing on the Shoulders of Decomposition

arXiv1 repo

arXiv:2407.05417

Subspace-Tuning

SmurfCat at PAN 2024 TextDetox: Alignment of Multilingual Transformers for Text Detoxification

arXiv1 repo

arXiv:2407.05449

mt0-xl-detox-orpo

T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models

arXiv1 repo

arXiv:2407.05965

JailbreakDiffusionBench

DebUnc: Improving Large Language Model Agent Communication With Uncertainty Metrics

arXiv1 repo

arXiv:2407.06426

debunc

OffsetBias: Leveraging Debiased Data for Tuning Evaluators

arXiv1 repo

arXiv:2407.06551

offsetbias

PaliGemma: A versatile 3B VLM for transfer

arXiv1 repo

arXiv:2407.07726

detikzify-v2-8b

Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities

arXiv1 repo

arXiv:2407.07791

KnowledgeSpread

OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training

arXiv1 repo

arXiv:2407.07852

OpenDiloco

Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients

arXiv1 repo

arXiv:2407.08296

Finetune_with_GaLore

Self-training Language Models for Arithmetic Reasoning

arXiv1 repo

arXiv:2407.08400

calc-x

MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

arXiv1 repo

arXiv:2407.08739

Image-Generation-CoT

TAPFixer: Automatic Detection and Repair of Home Automation Vulnerabilities based on Negated-property Reasoning

arXiv1 repo

arXiv:2407.09095

ComfyUI-LoRA-Optimizer

ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts

arXiv1 repo

arXiv:2407.09447

ASTPrompter

arXiv:2407.09450

arXiv1 repo

arXiv:2407.09450

HEBO

Follow the Rules: Reasoning for Video Anomaly Detection with Large Language Models

arXiv1 repo

arXiv:2407.10299

AnomalyRuler

LAB-Bench: Measuring Capabilities of Language Models for Biology Research

arXiv1 repo

arXiv:2407.10362

LAB-Bench

Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

arXiv1 repo

arXiv:2407.10627

aurora-m2

PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition

arXiv1 repo

arXiv:2407.11214

open-atp

Unraveling the Truth: Do VLMs really Understand Charts? A Deep Dive into Consistency and Robustness

arXiv1 repo

arXiv:2407.11229

Robust-CQA

arXiv:2407.11325

arXiv1 repo

arXiv:2407.11325

VISA

Continuity Preserving Online CenterLine Graph Learning

arXiv1 repo

arXiv:2407.11337

CGNet

Revisiting the Impact of Pursuing Modularity for Code Generation

arXiv1 repo

arXiv:2407.11406

Revisiting-Modularity

Scaling Diffusion Transformers to 16 Billion Parameters

arXiv1 repo

arXiv:2407.11633

DiT-MoE-diffusers

Stepping Stones: A Progressive Training Strategy for Audio-Visual Semantic Segmentation

arXiv1 repo

arXiv:2407.11820

Stepping-Stones

SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge

arXiv1 repo

arXiv:2407.11906

CaRTS

EchoSight: Advancing Visual-Language Models with Wiki Knowledge

arXiv1 repo

arXiv:2407.12735

EchoSight

BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval

arXiv1 repo

arXiv:2407.12883

BRIGHT

Textualized and Feature-based Models for Compound Multimodal Emotion Recognition in the Wild

arXiv1 repo

arXiv:2407.12927

feature-vs-text-compound-emotion

Deep Time Series Models: A Comprehensive Survey and Benchmark

arXiv1 repo

arXiv:2407.13278

Time-Series-Library

T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation

arXiv1 repo

arXiv:2407.14505

TTOM

Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models

arXiv1 repo

arXiv:2407.15408

ChronAccRet

Invariance Times Transfer Properties

arXiv1 repo

arXiv:2407.15460

anchormind

LLMmap: Fingerprinting For Large Language Models

arXiv1 repo

arXiv:2407.15847

LLMmap

Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation

arXiv1 repo

arXiv:2407.17274

AVG

AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents

arXiv1 repo

arXiv:2407.17490

AMEX

Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models

arXiv1 repo

arXiv:2407.18015

GiftEvalPretrain

Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation

arXiv1 repo

arXiv:2407.18698

Adaptive-Contrastive-Search

MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

arXiv1 repo

arXiv:2407.19034

manga-ub

ByteCheckpoint: A Unified Checkpointing System for Large Foundation Model Development

arXiv1 repo

arXiv:2407.20143

VeOmni

arXiv:2407.21004

arXiv1 repo

arXiv:2407.21004

Evolver

CLEFT: Language-Image Contrastive Learning with Efficient Large Language Model and Prompt Fine-Tuning

arXiv1 repo

arXiv:2407.21011

CLEFT

arXiv:2407.21315

arXiv1 repo

arXiv:2407.21315

SpeechCueLLM

MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training

arXiv1 repo

arXiv:2407.21439

RagVL

Tora: Trajectory-oriented Diffusion Transformer for Video Generation

arXiv1 repo

arXiv:2407.21705

Tora

Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding

arXiv1 repo

arXiv:2408.00264

clover

EXAONEPath 1.0 Patch-level Foundation Model for Pathology

arXiv1 repo

arXiv:2408.00380

dpfm_factory

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

arXiv1 repo

arXiv:2408.01337

AudioBench-N

VLG-CBM: Training Concept Bottleneck Models with Vision-Language Guidance

arXiv1 repo

arXiv:2408.01432

Concept-Bottleneck-LLM

arXiv:2408.02514

arXiv1 repo

arXiv:2408.02514

Stem-JEPA

Multistain Pretraining for Slide Representation Learning in Pathology

arXiv1 repo

arXiv:2408.02859

MADELEINE

VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge

arXiv1 repo

arXiv:2408.02865

MedUMM

UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization

arXiv1 repo

arXiv:2408.05939

UniPortrait

DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation

arXiv1 repo

arXiv:2408.06010

DEEPTalk

BMX: Entropy-weighted Similarity and Semantic-enhanced Lexical Search

arXiv1 repo

arXiv:2408.06643

baguetter

Imagen 3

arXiv1 repo

arXiv:2408.07009

t2v_metrics

Post-Training Sparse Attention with Double Sparsity

arXiv1 repo

arXiv:2408.07092

DoubleSparse

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

arXiv1 repo

arXiv:2408.07199

agent-q

BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning

arXiv1 repo

arXiv:2408.07440

baple

RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation

arXiv1 repo

arXiv:2408.08067

RAGChecker

ECG-Chat: A Large ECG-Language Model for Cardiac Disease Diagnosis

arXiv1 repo

arXiv:2408.08849

ECG-Chat

xGen-MM (BLIP-3): A Family of Open Large Multimodal Models

arXiv1 repo

arXiv:2408.08872

xgen-mm-phi3-mini-instruct-interleave-r-v1.5

Selective Prompt Anchoring for Code Generation

arXiv1 repo

arXiv:2408.09121

Selective-Prompt-Anchoring

BLADE: Benchmarking Language Model Agents for Data-Driven Science

arXiv1 repo

arXiv:2408.09667

BLADE

RealCustom++: Representing Images as Real Textual Word for Real-Time Customization

arXiv1 repo

arXiv:2408.09744

RealCustom

LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain

arXiv1 repo

arXiv:2408.10343

legalbenchrag

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

arXiv1 repo

arXiv:2408.11039

zen5

Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation

arXiv1 repo

arXiv:2408.11053

verilog-eval

Ophthalmic Biomarker Detection: Highlights from the IEEE Video and Image Processing Cup 2023 Student Competition

arXiv1 repo

arXiv:2408.11170

OLIVES_Dataset

Real-Time Video Generation with Pyramid Attention Broadcast

arXiv1 repo

arXiv:2408.12588

OpenDiT

Building and better understanding vision-language models: insights and future directions

arXiv1 repo

arXiv:2408.12637

Idefics3-8B-Llama3

NanoFlow: Towards Optimal Large Language Model Serving Throughput

arXiv1 repo

arXiv:2408.12757

Nanoflow

SONICS: Synthetic Or Not -- Identifying Counterfeit Songs

arXiv1 repo

arXiv:2408.14080

bach-or-bot

A model of generation of a jet in stratified nonequilibrium plasma

arXiv1 repo

arXiv:2408.14210

SphereForge

arXiv:2408.14262

arXiv1 repo

arXiv:2408.14262

s3m-aave

CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts

arXiv1 repo

arXiv:2408.14419

acl2026-misleading-visualizations

Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations

arXiv1 repo

arXiv:2408.15232

storm

LRP4RAG: Detecting Hallucinations in Retrieval-Augmented Generation via Layer-wise Relevance Propagation

arXiv1 repo

arXiv:2408.15533

LRP-eXplains-Transformers

Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models

arXiv1 repo

arXiv:2408.15585

wespeaker

More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding

arXiv1 repo

arXiv:2408.15966

MiniGPT-3D

CogVLM2: Visual Language Models for Image and Video Understanding

arXiv1 repo

arXiv:2408.16500

cogvlm2-llama3-chat-19B

Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

arXiv1 repo

arXiv:2408.16737

aurora-m2

VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning

arXiv1 repo

arXiv:2409.01667

VProChart

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

arXiv1 repo

arXiv:2409.01704

GOT-OCR2_0

Boosting Vision-Language Models for Histopathology Classification: Predict all at once

arXiv1 repo

arXiv:2409.01883

Histo-TransCLIP

BEAVER: An Enterprise Benchmark for Text-to-SQL

arXiv1 repo

arXiv:2409.02038

beaver

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins (early version)

arXiv1 repo

arXiv:2409.02920

RoboTwin

mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

arXiv1 repo

arXiv:2409.03420

mPLUG-DocOwl

Lexicon3D: Probing Visual Foundation Models for Complex 3D Scene Understanding

arXiv1 repo

arXiv:2409.03757

Lexicon3D

BreachSeek: A Multi-Agent Automated Penetration Tester

arXiv1 repo

arXiv:2409.03789

awesome-cybersecurity-agentic-ai

Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models

arXiv1 repo

arXiv:2409.04701

late-chunking

Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling

arXiv1 repo

arXiv:2409.05395

vl_mamba

MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation

arXiv1 repo

arXiv:2409.05591

MemoRAG

FLoRA: Federated Fine-Tuning Large Language Models with Heterogeneous Low-Rank Adaptations

arXiv1 repo

arXiv:2409.05976

FederatedLLM

World-Grounded Human Motion Recovery via Gravity-View Coordinates

arXiv1 repo

arXiv:2409.06662

GVHMR

gsplat: An Open-Source Library for Gaussian Splatting

arXiv1 repo

arXiv:2409.06765

gsplat

Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem

arXiv1 repo

arXiv:2409.07123

Cross-Refine

EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance

arXiv1 repo

arXiv:2409.08091

EZIGen

arXiv:2409.08248

arXiv1 repo

arXiv:2409.08248

textboost

Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

arXiv1 repo

arXiv:2409.08264

UI-TARS

Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

arXiv1 repo

arXiv:2409.09016

CLOVER

TG-LLaVA: Text Guided LLaVA via Learnable Latent Embeddings

arXiv1 repo

arXiv:2409.09564

TG-LLaVA

Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search

arXiv1 repo

arXiv:2409.09913

RaBitQ-Library

Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT

arXiv1 repo

arXiv:2409.10103

speaker_disentangled_hubert

jina-embeddings-v3: Multilingual Embeddings With Task LoRA

arXiv1 repo

arXiv:2409.10173

jina-embeddings-v3

MusicLIME: Explainable Multimodal Music Understanding

arXiv1 repo

arXiv:2409.10496

bach-or-bot

MotIF: Motion Instruction Fine-tuning

arXiv1 repo

arXiv:2409.10683

motif

LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

arXiv1 repo

arXiv:2409.11239

KUDGE

Towards Fair RAG: On the Impact of Fair Ranking in Retrieval-Augmented Generation

arXiv1 repo

arXiv:2409.11598

Fair-RAG

Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference

arXiv1 repo

arXiv:2409.12117

nemo-nano-codec-22khz-1.89kbps-21.5fps

arXiv:2409.12147

arXiv1 repo

arXiv:2409.12147

MAgICoRE

Large Language Models are Strong Audio-Visual Speech Recognition Learners

arXiv1 repo

arXiv:2409.12319

Llama-AVSR

FlexiTex: Enhancing Texture Generation via Visual Guidance

arXiv1 repo

arXiv:2409.12431

FlexiSyncMVD

AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework

arXiv1 repo

arXiv:2409.12466

AudioEditor

JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models

arXiv1 repo

arXiv:2409.13317

JMedBench

Higher-Order Message Passing for Glycan Representation Learning

arXiv1 repo

arXiv:2409.13467

GIFFLAR

Logically Consistent Language Models via Neuro-Symbolic Integration

arXiv1 repo

arXiv:2409.13724

loco-llm

Language agents achieve superhuman synthesis of scientific knowledge

arXiv1 repo

arXiv:2409.13740

paper-qa

MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder

arXiv1 repo

arXiv:2409.14074

MultiMed

PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects

arXiv1 repo

arXiv:2409.14331

PISR

MobileViews: A Million-scale and Diverse Mobile GUI Dataset

arXiv1 repo

arXiv:2409.14337

MobileViews

arXiv:2409.14507

arXiv1 repo

arXiv:2409.14507

SAEBench

Inference-Friendly Models With MixAttention

arXiv1 repo

arXiv:2409.15012

LCKV

RAMBO: Enhancing RAG-based Repository-Level Method Body Completion

arXiv1 repo

arXiv:2409.15204

RAMBO

Making s-wave superconductors topological with magnetic field

arXiv1 repo

arXiv:2409.15266

SphereForge

OmniBench: Towards The Future of Universal Omni-Language Models

arXiv1 repo

arXiv:2409.15272

OmniInstruct_v1

Parse Trees Guided LLM Prompt Compression

arXiv1 repo

arXiv:2409.15395

Prompt-Compression

Making Text Embedders Few-Shot Learners

arXiv1 repo

arXiv:2409.15700

bge-en-icl

HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models

arXiv1 repo

arXiv:2409.16191

HelloBench

Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts

arXiv1 repo

arXiv:2409.17106

CADFusion

Internalizing ASR with Implicit Chain of Thought for Efficient Speech-to-Speech Conversational LLM

arXiv1 repo

arXiv:2409.17353

SpeechLLM

Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography

arXiv1 repo

arXiv:2409.18119

MaMA

Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction

arXiv1 repo

arXiv:2409.18121

rsrd

LangSAMP: Language-Script Aware Multilingual Pretraining

arXiv1 repo

arXiv:2409.18199

LangSAMP

MinerU: An Open-Source Solution for Precise Document Content Extraction

arXiv1 repo

arXiv:2409.18839

MinerU

FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark

arXiv1 repo

arXiv:2409.19014

FLEX

HybridFlow: A Flexible and Efficient RLHF Framework

arXiv1 repo

arXiv:2409.19256

verl-pipeline

On the Nonlinear Excitation of Phononic Frequency Combs in Molecules

arXiv1 repo

arXiv:2409.19607

Tiny-R2

Old Optimizer, New Norm: An Anthology

arXiv1 repo

arXiv:2409.20325

modded-nanogpt

The Perfect Blend: Redefining RLHF with Mixture of Judges

arXiv1 repo

arXiv:2409.20370

open-perfectblend

A Hitchhikers Guide to Fine-Grained Face Forgery Detection Using Common Sense Reasoning

arXiv1 repo

arXiv:2410.00485

HitchhikersGuide

From Code to Correctness: Closing the Last Mile of Code Generation with Hierarchical Debugging

arXiv1 repo

arXiv:2410.01215

ComfyUI-LoRA-Optimizer

Enhancement of superconductivity coexisting with charge density wave in lattice expanded $\textrm{NbTe}_2$

arXiv1 repo

arXiv:2410.01247

Titan-Memory

HelpSteer2-Preference: Complementing Ratings with Preferences

arXiv1 repo

arXiv:2410.01257

Llama-3_3-Nemotron-Super-49B-GenRM

Endless Jailbreaks with Bijection Learning

arXiv1 repo

arXiv:2410.01294

GA

CrowdCounter: A benchmark type-specific multi-target counterspeech dataset

arXiv1 repo

arXiv:2410.01400

CrowdCounter

SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

arXiv1 repo

arXiv:2410.01481

SonicSim

shapiq: Shapley Interactions for Machine Learning

arXiv1 repo

arXiv:2410.01649

shapiq

Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo

arXiv1 repo

arXiv:2410.01920

TSMC4MATH

DeepProtein: Deep Learning Library and Benchmark for Protein Sequence Learning

arXiv1 repo

arXiv:2410.02023

DeepProtein

Adversarial Decoding: Generating Readable Documents for Adversarial Objectives

arXiv1 repo

arXiv:2410.02163

adversarial_decoding

Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment

arXiv1 repo

arXiv:2410.02197

general-preference-model

Contextual Document Embeddings

arXiv1 repo

arXiv:2410.02525

cde-small-v1

HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly

arXiv1 repo

arXiv:2410.02694

HELMET

LLaVA-Video: Video Instruction Tuning With Synthetic Data

arXiv1 repo

arXiv:2410.02713

LLaVA-Video-178K

ToolGen: Unified Tool Retrieval and Calling via Generation

arXiv1 repo

arXiv:2410.03439

ToolGen-Datasets

RAFT: Realistic Attacks to Fool Text Detectors

arXiv1 repo

arXiv:2410.03658

RAFT

Learning from Committee: Reasoning Distillation from a Mixture of Teachers with Peer-Review

arXiv1 repo

arXiv:2410.03663

Learn-from-Committee

MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

arXiv1 repo

arXiv:2410.03825

MonST3R_PO-TA-S-W_ViTLarge_BaseDecoder_512_dpt

Learning Code Preference via Synthetic Evolution

arXiv1 repo

arXiv:2410.03837

llm-code-preference

SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation

arXiv1 repo

arXiv:2410.03960

Llama-3.1-SwiftKV-8B-Instruct

A Survey on LLM-based Code Generation for Low-Resource and Domain-Specific Programming Languages

arXiv1 repo

arXiv:2410.03981

agentty

IV-Mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video Synthesis

arXiv1 repo

arXiv:2410.04171

IV-mixed-Sampler

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

arXiv1 repo

arXiv:2410.04527

Casablanca

CAR: Controllable Autoregressive Modeling for Visual Generation

arXiv1 repo

arXiv:2410.04671

CAR

Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models

arXiv1 repo

arXiv:2410.04727

ForgettingCurve

arXiv:2410.05102

arXiv1 repo

arXiv:2410.05102

HEBO

GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting

arXiv1 repo

arXiv:2410.05259

GS-VTON

AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

arXiv1 repo

arXiv:2410.05295

GA

Image Watermarks are Removable Using Controllable Regeneration from Clean Noise

arXiv1 repo

arXiv:2410.05470

watermarks-remover

Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization

arXiv1 repo

arXiv:2410.06244

story-iter

TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training

arXiv1 repo

arXiv:2410.06511

torchtitan

InstantIR: Blind Image Restoration with Instant Generative Reference

arXiv1 repo

arXiv:2410.06551

InstantIR

Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures

arXiv1 repo

arXiv:2410.06672

circuit_backup

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

arXiv1 repo

arXiv:2410.07095

aideml

VHELM: A Holistic Evaluation of Vision Language Models

arXiv1 repo

arXiv:2410.07112

helm

IterGen: Iterative Semantic-aware Structured LLM Generation with Backtracking

arXiv1 repo

arXiv:2410.07295

itergen

Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

arXiv1 repo

arXiv:2410.08020

TTFT-SIFT

Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

arXiv1 repo

arXiv:2410.08102

DataFlow

Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation

arXiv1 repo

arXiv:2410.08371

DAM

Bilinear MLPs enable weight-based mechanistic interpretability

arXiv1 repo

arXiv:2410.08417

bilinear-decomposition

arXiv:2410.08709

arXiv1 repo

arXiv:2410.08709

di4c

VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model

arXiv1 repo

arXiv:2410.08792

SeeDo

Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

arXiv1 repo

arXiv:2410.08847

unintentional-unalignment

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

arXiv1 repo

arXiv:2410.09024

felonybench

When Graph meets Multimodal: Benchmarking and Meditating on Multimodal Attributed Graphs Learning

arXiv1 repo

arXiv:2410.09132

MAGB

Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

arXiv1 repo

arXiv:2410.09302

verl-pipeline

Skipping Computations in Multimodal LLMs

arXiv1 repo

arXiv:2410.09454

ima-lmms

LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

arXiv1 repo

arXiv:2410.09732

LOKI

Text4Seg: Reimagining Image Segmentation as Text Generation

arXiv1 repo

arXiv:2410.09855

Text4Seg

ChroKnowledge: Unveiling Chronological Knowledge of Language Models in Multiple Domains

arXiv1 repo

arXiv:2410.09870

ChroKnowBench

arXiv:2410.09893

arXiv1 repo

arXiv:2410.09893

RMB-Reward-Model-Benchmark

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

arXiv1 repo

arXiv:2410.10122

MuseTalk

Fed-pilot: Optimizing LoRA Allocation for Efficient Federated Fine-Tuning with Heterogeneous Clients

arXiv1 repo

arXiv:2410.10200

Fed-PLoRA

KBLaM: Knowledge Base augmented Language Model

arXiv1 repo

arXiv:2410.10450

KBLaM

Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image Classification

arXiv1 repo

arXiv:2410.10573

VLSA

LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

arXiv1 repo

arXiv:2410.10700

SafeMTData

HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

arXiv1 repo

arXiv:2410.10812

hart

Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free

arXiv1 repo

arXiv:2410.10814

MoE-Embedding

Depth Any Video with Scalable Synthetic Data

arXiv1 repo

arXiv:2410.10815

DepthAnyVideo

When Does Perceptual Alignment Benefit Vision Representations?

arXiv1 repo

arXiv:2410.10817

dreamsim

arXiv:2410.10819

arXiv1 repo

arXiv:2410.10819

Block-Sparse-Attention

Liger Kernel: Efficient Triton Kernels for LLM Training

arXiv1 repo

arXiv:2410.10989

Liger-Kernel

arXiv:2410.11163

arXiv1 repo

arXiv:2410.11163

model_swarm

Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling

arXiv1 repo

arXiv:2410.11236

Ctrl-U

Improving Long-Text Alignment for Text-to-Image Diffusion Models

arXiv1 repo

arXiv:2410.11817

LongAlign

MoH: Multi-Head Attention as Mixture-of-Head Attention

arXiv1 repo

arXiv:2410.11842

Chat-UniVi

FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation

arXiv1 repo

arXiv:2410.12266

AudioLCM

Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up

arXiv1 repo

arXiv:2410.12323

Less-is-More

Towards Neural Scaling Laws for Time Series Foundation Models

arXiv1 repo

arXiv:2410.12360

time-moe

DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception

arXiv1 repo

arXiv:2410.12628

DocLayout-YOLO

VividMed: Vision Language Model with Versatile Visual Grounding for Medicine

arXiv1 repo

arXiv:2410.12694

MMMM

Context is Key(NMF): Modelling Topical Information Dynamics in Chinese Diaspora Media

arXiv1 repo

arXiv:2410.12791

turftopic

Enterprise Benchmarks for Large Language Model Evaluation

arXiv1 repo

arXiv:2410.12857

helm

RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models

arXiv1 repo

arXiv:2410.13360

RAP-MLLM

MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures

arXiv1 repo

arXiv:2410.13754

MixEval

FiTv2: Scalable and Improved Flexible Vision Transformer for Diffusion Model

arXiv1 repo

arXiv:2410.13925

FiT-diffusers

arXiv:2410.13928

arXiv1 repo

arXiv:2410.13928

SAEBench

HiCo: Hierarchical Controllable Diffusion Model for Layout-to-image Generation

arXiv1 repo

arXiv:2410.14324

PlanGen

SNAC: Multi-Scale Neural Audio Codec

arXiv1 repo

arXiv:2410.14411

snac

A Systematic Study of Cross-Layer KV Sharing for Efficient LLM Inference

arXiv1 repo

arXiv:2410.14442

LCKV

DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

arXiv1 repo

arXiv:2410.14803

carl_distrl

A Multimodal Vision Foundation Model for Clinical Dermatology

arXiv1 repo

arXiv:2410.15038

PanDerm

Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

arXiv1 repo

arXiv:2410.15553

smoltalk2

Moonshine: Speech Recognition for Live Transcription and Voice Commands

arXiv1 repo

arXiv:2410.15608

moonshine

TimeMixer++: A General Time Series Pattern Machine for Universal Predictive Analysis

arXiv1 repo

arXiv:2410.16032

time-moe

1-bit AI Infra: Part 1.1, Fast and Lossless BitNet b1.58 Inference on CPUs

arXiv1 repo

arXiv:2410.16144

BitNet

RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style

arXiv1 repo

arXiv:2410.16184

Llama-3_3-Nemotron-Super-49B-GenRM

Improve Vision Language Model Chain-of-thought Reasoning

arXiv1 repo

arXiv:2410.16198

LLaVA-Hound-DPO

Beyond Browsing: API-Based Web Agents

arXiv1 repo

arXiv:2410.16464

guaca

TIPS: Text-Image Pretraining with Spatial awareness

arXiv1 repo

arXiv:2410.16512

tips

The Scene Language: Representing Scenes with Programs, Words, and Embeddings

arXiv1 repo

arXiv:2410.16770

scene-language

VoiceBench: Benchmarking LLM-Based Voice Assistants

arXiv1 repo

arXiv:2410.17196

voicebench

Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss

arXiv1 repo

arXiv:2410.17243

VideoLLaMA2

JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

arXiv1 repo

arXiv:2410.17250

JMMMU

Altogether: Image Captioning via Re-aligning Alt-text

arXiv1 repo

arXiv:2410.17251

MetaCLIP

Scalable Influence and Fact Tracing for Large Language Model Pretraining

arXiv1 repo

arXiv:2410.17413

bergson

ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents

arXiv1 repo

arXiv:2410.17657

MING

arXiv:2410.17736

arXiv1 repo

arXiv:2410.17736

llama2.mojo

Value Residual Learning

arXiv1 repo

arXiv:2410.17897

modded-nanogpt

FreeVS: Generative View Synthesis on Free Driving Trajectory

arXiv1 repo

arXiv:2410.18079

FreeVS

ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment

arXiv1 repo

arXiv:2410.18194

ZIP-FIT

DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation

arXiv1 repo

arXiv:2410.18666

DreamClear

How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs

arXiv1 repo

arXiv:2410.18697

prometheus-eval

arXiv:2410.18745

arXiv1 repo

arXiv:2410.18745

STRING

arXiv:2410.18798

arXiv1 repo

arXiv:2410.18798

CharXiv

Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction

arXiv1 repo

arXiv:2410.18962

GST

arXiv:2410.19278

arXiv1 repo

arXiv:2410.19278

SAEBench

NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction

arXiv1 repo

arXiv:2410.19452

NeuroClips

CoqPilot, a plugin for LLM-based generation of proofs

arXiv1 repo

arXiv:2410.19605

coqpilot

RARe: Retrieval Augmented Retrieval with In-Context Examples

arXiv1 repo

arXiv:2410.20088

RARe

SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement

arXiv1 repo

arXiv:2410.20285

moatless-tools

Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

arXiv1 repo

arXiv:2410.20526

circuit_backup

PaPaGei: Open Foundation Models for Optical Physiological Signals

arXiv1 repo

arXiv:2410.20542

papagei-foundation-model

Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA

arXiv1 repo

arXiv:2410.20672

mixture_of_recursions

Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation

arXiv1 repo

arXiv:2410.20724

SubgraphRAG

Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation

arXiv1 repo

arXiv:2410.20941

BLEUless_DocMT

Flaming-hot Initiation with Regular Execution Sampling for Large Language Models

arXiv1 repo

arXiv:2410.21236

verl-pipeline

arXiv:2410.21357

arXiv1 repo

arXiv:2410.21357

Energy-Diffusion-LLM

Are Decoder-Only Large Language Models the Silver Bullet for Code Search?

arXiv1 repo

arXiv:2410.22240

DecoderLLMs-CodeSearch

Online Detection of LLM-Generated Texts via Sequential Hypothesis Testing by Betting

arXiv1 repo

arXiv:2410.22318

online-llm-detection

arXiv:2410.22376

arXiv1 repo

arXiv:2410.22376

Rare-to-Frequent

Image2Struct: Benchmarking Structure Extraction for Vision-Language Models

arXiv1 repo

arXiv:2410.22456

helm

FlowDCN: Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution

arXiv1 repo

arXiv:2410.22655

FlowDCN

Emotional RAG: Enhancing Role-Playing Agents through Emotional Retrieval

arXiv1 repo

arXiv:2410.23041

Role-Playing-LLM-Megumin

Controlling Language and Diffusion Models by Transporting Activations

arXiv1 repo

arXiv:2410.23054

ml-lineas

On Memorization of Large Language Models in Logical Reasoning

arXiv1 repo

arXiv:2410.23123

knights-and-knaves

Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms

arXiv1 repo

arXiv:2410.23144

PD12M

$100K or 100 Days: Trade-offs when Pre-Training with Academic Resources

arXiv1 repo

arXiv:2410.23261

academic-pretraining

VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference Acceleration

arXiv1 repo

arXiv:2410.23317

GUI-KV

GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages

arXiv1 repo

arXiv:2410.23825

GlotCC-V1

Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models

arXiv1 repo

arXiv:2410.23841

InfoSearch

SelfCodeAlign: Self-Alignment for Code Generation

arXiv1 repo

arXiv:2410.24198

starcoder2-instruct-15b-v0.1

DELTA: Dense Efficient Long-range 3D Tracking for any video

arXiv1 repo

arXiv:2410.24211

DELTA_densetrack3d

Randomized Autoregressive Visual Generation

arXiv1 repo

arXiv:2411.00776

clustermark_1d-tokenizer

CycleResearcher: Improving Automated Research via Automated Review

arXiv1 repo

arXiv:2411.00816

Researcher

SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding

arXiv1 repo

arXiv:2411.01106

VisR-Bench

Sample-Efficient Alignment for LLMs

arXiv1 repo

arXiv:2411.01493

oat

How Far is Video Generation from World Model: A Physical Law Perspective

arXiv1 repo

arXiv:2411.02385

phyworld

AutoVFX: Physically Realistic Video Editing from Natural Language Instructions

arXiv1 repo

arXiv:2411.02394

autovfx

Training-free Regional Prompting for Diffusion Transformers

arXiv1 repo

arXiv:2411.02395

Regional-Prompting-FLUX

SLED: Self Logits Evolution Decoding for Improving Factuality in Large Language Models

arXiv1 repo

arXiv:2411.02433

SLED

Dr. SoW: Density Ratio of Strong-over-weak LLMs for Reducing the Cost of Human Annotation in Preference Tuning

arXiv1 repo

arXiv:2411.02481

reward_hub

ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy

arXiv1 repo

arXiv:2411.02572

rxrx3-core

On the Loss of Context-awareness in General Instruction Fine-tuning

arXiv1 repo

arXiv:2411.02688

context_awareness

ADOPT: Modified Adam Can Converge with Any $β_2$ with the Optimal Rate

arXiv1 repo

arXiv:2411.02853

adopt

SMoA: Improving Multi-agent Large Language Models with Sparse Mixture-of-Agents

arXiv1 repo

arXiv:2411.03284

rightmind

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

arXiv1 repo

arXiv:2411.03628

StreamingBench

A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning

arXiv1 repo

arXiv:2411.04105

prop-logic-transformer-circuit

Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

arXiv1 repo

arXiv:2411.04118

eval-medical-dapt

Vision Language Models are In-Context Value Learners

arXiv1 repo

arXiv:2411.04549

reward-scope

TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation

arXiv1 repo

arXiv:2411.04709

TIP-I2V

VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

arXiv1 repo

arXiv:2411.04923

VideoGLaMM

BitNet a4.8: 4-bit Activations for 1-bit LLMs

arXiv1 repo

arXiv:2411.04965

BitNet

DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning

arXiv1 repo

arXiv:2411.04983

stable-worldmodel

Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder

arXiv1 repo

arXiv:2411.05195

CLIP-Embeds

Don't Look Twice: Faster Video Transformers with Run-Length Tokenization

arXiv1 repo

arXiv:2411.05222

rlt

Using Language Models to Disambiguate Lexical Choices in Translation

arXiv1 repo

arXiv:2411.05781

Lex-Rules

GFT: Graph Foundation Model with Transferable Tree Vocabulary

arXiv1 repo

arXiv:2411.06070

GFT

M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

arXiv1 repo

arXiv:2411.06176

multimodal-docs-public

ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?

arXiv1 repo

arXiv:2411.06469

ClinicalBench

arXiv:2411.07186

arXiv1 repo

arXiv:2411.07186

NatureLM-audio

InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance

arXiv1 repo

arXiv:2411.07795

InvisMark

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

arXiv1 repo

arXiv:2411.07975

Janus

Large Language Models Can Self-Improve in Long-context Reasoning

arXiv1 repo

arXiv:2411.08147

SEALONG

The Limited Impact of Medical Adaptation of Large Language and Vision-Language Models

arXiv1 repo

arXiv:2411.08870

eval-medical-dapt

Reliable Decision Making via Calibration Oriented Retrieval Augmented Generation

arXiv1 repo

arXiv:2411.08891

calibrag

"Should I Give Up Now?" Investigating LLM Pitfalls in Software Engineering

arXiv1 repo

arXiv:2411.09916

unlazy

EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge

arXiv1 repo

arXiv:2411.10224

MLRG

Does Prompt Formatting Have Any Impact on LLM Performance?

arXiv1 repo

arXiv:2411.10541

opendataloader-bench

arXiv:2411.10557

arXiv1 repo

arXiv:2411.10557

MLAN

Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer

arXiv1 repo

arXiv:2411.10781

test-time-scaling

SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization

arXiv1 repo

arXiv:2411.10958

SageAttention

Scalable Autoregressive Monocular Depth Estimation

arXiv1 repo

arXiv:2411.11361

VAR

Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model

arXiv1 repo

arXiv:2411.12783

Med-2E3

Stylecodes: Encoding Stylistic Information For Image Generation

arXiv1 repo

arXiv:2411.12811

stylecodes

VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge

arXiv1 repo

arXiv:2411.12915

VLM

Veryl: A New Hardware Description Language as an Altarnative to SystemVerilog

arXiv1 repo

arXiv:2411.12983

veryl

VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

arXiv1 repo

arXiv:2411.13503

VBench

Find Any Part in 3D

arXiv1 repo

arXiv:2411.13550

Find3D

Hymba: A Hybrid-head Architecture for Small Language Models

arXiv1 repo

arXiv:2411.13676

hymba

Novel View Extrapolation with Video Diffusion Priors

arXiv1 repo

arXiv:2411.14208

ViewExtrapolator

arXiv:2411.14280

arXiv1 repo

arXiv:2411.14280

EasyHOI

Masala-CHAI: A Large-Scale SPICE Netlist Dataset for Analog Circuits by Harnessing AI

arXiv1 repo

arXiv:2411.14299

Masala-CHAI

DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

arXiv1 repo

arXiv:2411.14347

Rex-Omni

Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction

arXiv1 repo

arXiv:2411.14384

Open-OmniVCus

FedMLLM: Federated Fine-tuning MLLM on Multimodal Heterogeneity Data

arXiv1 repo

arXiv:2411.14717

FedMLLM

OminiControl: Minimal and Universal Control for Diffusion Transformer

arXiv1 repo

arXiv:2411.15098

Subjects200K

XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models

arXiv1 repo

arXiv:2411.15100

xgrammar

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

arXiv1 repo

arXiv:2411.15114

aideml

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement

arXiv1 repo

arXiv:2411.15115

VideoRepair

Tulu 3: Pushing Frontiers in Open Language Model Post-Training

arXiv1 repo

arXiv:2411.15124

OLMoE-1B-7B-0125-Instruct

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

arXiv1 repo

arXiv:2411.15260

VIVID-10M

Foundation Cures Personalization: Improving Personalized Models' Prompt Consistency via Hidden Foundation Knowledge

arXiv1 repo

arXiv:2411.15277

FreeCure

Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT

arXiv1 repo

arXiv:2411.15418

panspecies-dti

Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation

arXiv1 repo

arXiv:2411.16185

Fancy123

Functionality understanding and segmentation in 3D scenes

arXiv1 repo

arXiv:2411.16310

fun3du

Preference Optimization for Reasoning with Pseudo Feedback

arXiv1 repo

arXiv:2411.16345

EMPO

Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models

arXiv1 repo

arXiv:2411.16602

Chat2SVG

Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks

arXiv1 repo

arXiv:2411.16721

ASTRA

arXiv:2411.16778

arXiv1 repo

arXiv:2411.16778

uMedGround

Controllable Human Image Generation with Personalized Multi-Garments

arXiv1 repo

arXiv:2411.16801

BootComp

Probing the limitations of multimodal language models for chemistry and materials research

arXiv1 repo

arXiv:2411.16955

chembench

LiteVAR: Compressing Visual Autoregressive Modelling with Efficient Attention and Quantization

arXiv1 repo

arXiv:2411.17178

VAR

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning

arXiv1 repo

arXiv:2411.17426

TransArch

WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

arXiv1 repo

arXiv:2411.17459

Open-Sora-Plan

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

arXiv1 repo

arXiv:2411.17465

ShowUI-desktop

CHOICE: Benchmarking the Remote Sensing Capabilities of Large Vision-Language Models

arXiv1 repo

arXiv:2411.18145

CHOICE

arXiv:2411.18301

arXiv1 repo

arXiv:2411.18301

LaRender

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

arXiv1 repo

arXiv:2411.18363

Rex-Omni

FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving

arXiv1 repo

arXiv:2411.18424

lightllm

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

arXiv1 repo

arXiv:2411.18673

ac3d

Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits

arXiv1 repo

arXiv:2411.18704

open-value

arXiv:2411.18895

arXiv1 repo

arXiv:2411.18895

SAEBench

Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

arXiv1 repo

arXiv:2411.19146

Llama-3_3-Nemotron-Super-49B-v1_5

HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

arXiv1 repo

arXiv:2411.19167

ObjectForesight-HOT3D-DiT

Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

arXiv1 repo

arXiv:2412.00127

Orthus-7B-base

ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain

arXiv1 repo

arXiv:2412.00532

mteb-1.34.14

arXiv:2412.00568

arXiv1 repo

arXiv:2412.00568

the_well

Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics

arXiv1 repo

arXiv:2412.00651

UMPIRE

Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild

arXiv1 repo

arXiv:2412.00811

Vid-Morp

WAFFLE: Multimodal Floorplan Understanding in the Wild

arXiv1 repo

arXiv:2412.00955

WAFFLE

INTELLECT-1 Technical Report

arXiv1 repo

arXiv:2412.01152

prime-diloco

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

arXiv1 repo

arXiv:2412.01282

CVPR2025_Align-KD

Free Process Rewards without Process Labels

arXiv1 repo

arXiv:2412.01981

PRIME

Progress-Aware Video Frame Captioning

arXiv1 repo

arXiv:2412.02071

ProgCaptioner

CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy

arXiv1 repo

arXiv:2412.02210

CC-OCR

VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention

arXiv1 repo

arXiv:2412.02259

VideoGen-of-Thought

Pixel-level and Semantic-level Adjustable Super-resolution: A Dual-LoRA Approach

arXiv1 repo

arXiv:2412.03017

OSEDiff

TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

arXiv1 repo

arXiv:2412.03069

TokenFlow

AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning

arXiv1 repo

arXiv:2412.03248

AIM

Imagine360: Immersive 360 Video Generation from Perspective Anchor

arXiv1 repo

arXiv:2412.03552

Imagine360

PaliGemma 2: A Family of Versatile VLMs for Transfer

arXiv1 repo

arXiv:2412.03555

gemma.cpp

Reducing Tool Hallucination via Reliability Alignment

arXiv1 repo

arXiv:2412.04141

ToolHallucination

AL-QASIDA: Analyzing LLM Quality and Accuracy Systematically in Dialectal Arabic

arXiv1 repo

arXiv:2412.04193

al-qasida

Densing Law of LLMs

arXiv1 repo

arXiv:2412.04315

Ultra-FineWeb

Liquid: Language Models are Scalable and Unified Multi-modal Generators

arXiv1 repo

arXiv:2412.04332

Monkey

Sparse Voxels Rasterization: Real-time High-fidelity Radiance Field Rendering

arXiv1 repo

arXiv:2412.04459

svraster

NVILA: Efficient Frontier Visual Language Models

arXiv1 repo

arXiv:2412.04468

VILA

SleeperMark: Towards Robust Watermark against Fine-Tuning Text-to-image Diffusion Models

arXiv1 repo

arXiv:2412.04852

SleeperMark

A Practical Examination of AI-Generated Text Detectors for Large Language Models

arXiv1 repo

arXiv:2412.05139

llm-detector-eval

Multimodal Fact-Checking with Vision Language Models: A Probing Classifier based Solution with Embedding Strategies

arXiv1 repo

arXiv:2412.05155

Multimodal-Fact-Checking-with-Vision-Language-Models

DreamColour: Controllable Video Colour Editing without Training

arXiv1 repo

arXiv:2412.05180

sketchdeco-code

UniScene: Unified Occupancy-centric Driving Scene Generation

arXiv1 repo

arXiv:2412.05435

UniScene

SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

arXiv1 repo

arXiv:2412.05818

SILMM-Implementation

Text-to-3D Generation by 2D Editing

arXiv1 repo

arXiv:2412.05929

GE3D

Training Large Language Models to Reason in a Continuous Latent Space

arXiv1 repo

arXiv:2412.06769

Mirage

Maya: An Instruction Finetuned Multilingual Multimodal Model

arXiv1 repo

arXiv:2412.07112

maya

Hierarchical Split Federated Learning: Convergence Analysis and System Optimization

arXiv1 repo

arXiv:2412.07197

SplitFM

ObjCtrl-2.5D: Training-free Object Control with Camera Poses

arXiv1 repo

arXiv:2412.07721

ObjCtrl-2.5D

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

arXiv1 repo

arXiv:2412.07755

Mirage

SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

arXiv1 repo

arXiv:2412.07760

SynCamVideo-Dataset

NLPineers@ NLU of Devanagari Script Languages 2025: Hate Speech Detection using Ensembling of BERT-based models

arXiv1 repo

arXiv:2412.08163

NLPineers

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

arXiv1 repo

arXiv:2412.08503

ComfyUI-StyleStudio

Large Concept Models: Language Modeling in a Sentence Representation Space

arXiv1 repo

arXiv:2412.08821

fairseq2

FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D Reconstruction

arXiv1 repo

arXiv:2412.09573

FreeSplatter

Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion

arXiv1 repo

arXiv:2412.09593

Neural-LightRig

MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models

arXiv1 repo

arXiv:2412.09818

MERaLiON-AudioLLM-Whisper-SEA-LION

BrushEdit: All-In-One Image Inpainting and Editing

arXiv1 repo

arXiv:2412.10316

BrushEdit

SCBench: A KV Cache-Centric Analysis of Long-Context Methods

arXiv1 repo

arXiv:2412.10319

SCBench

AdvPrefix: An Objective for Nuanced LLM Jailbreaks

arXiv1 repo

arXiv:2412.10321

jailbreak-objectives

Generative AI in Medicine

arXiv1 repo

arXiv:2412.10337

llava-rad

EvalGIM: A Library for Evaluating Generative Image Models

arXiv1 repo

arXiv:2412.10604

EvalGIM

VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation

arXiv1 repo

arXiv:2412.10704

VisDoM

Glimpse: Enabling White-Box Methods to Use Proprietary Models for Zero-Shot LLM-Generated Text Detection

arXiv1 repo

arXiv:2412.11506

glimpse

ColorFlow: Retrieval-Augmented Image Sequence Colorization

arXiv1 repo

arXiv:2412.11815

ColorFlow

PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting

arXiv1 repo

arXiv:2412.12096

PanSplat

EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation

arXiv1 repo

arXiv:2412.12559

EXIT

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference

arXiv1 repo

arXiv:2412.12785

Visual-Region

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

arXiv1 repo

arXiv:2412.12932

Mirage

Improving Explainability of Sentence-level Metrics via Edit-level Attribution for Grammatical Error Correction

arXiv1 repo

arXiv:2412.13110

gec-attribute

ROMAS: A Role-Based Multi-Agent System for Database monitoring and Planning

arXiv1 repo

arXiv:2412.13520

DB-GPT

Clio: Privacy-Preserving Insights into Real-World AI Use

arXiv1 repo

arXiv:2412.13678

enabling-independent-research

Open Universal Arabic ASR Leaderboard

arXiv1 repo

arXiv:2412.13788

open_universal_arabic_asr_leaderboard

Performance Gap in Entity Knowledge Extraction Across Modalities in Vision Language Models

arXiv1 repo

arXiv:2412.14133

PopVQA

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

arXiv1 repo

arXiv:2412.14642

BioMedGPT-Mol

ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing

arXiv1 repo

arXiv:2412.14711

zen5

Scylla: Translating an Applicative Subset of C to Safe Rust

arXiv1 repo

arXiv:2412.15042

libcrux

MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark

arXiv1 repo

arXiv:2412.15194

MMLU-CF

arXiv:2412.15206

arXiv1 repo

arXiv:2412.15206

AutoTrust

OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving

arXiv1 repo

arXiv:2412.15208

OpenEMMA

EnvGS: Modeling View-Dependent Appearance with Environment Gaussian

arXiv1 repo

arXiv:2412.15215

EnvGS

Building an Explainable Graph-based Biomedical Paper Recommendation System (Technical Report)

arXiv1 repo

arXiv:2412.15229

NarrativeRecommender

Dimension Reduction with Locally Adjusted Graphs

arXiv1 repo

arXiv:2412.15426

PaCMAP

Insights into resource utilization of code small language models serving with runtime engines and execution providers

arXiv1 repo

arXiv:2412.15441

energy-ml-serving

XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation

arXiv1 repo

arXiv:2412.15529

XRAG

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

arXiv1 repo

arXiv:2412.15649

SLAM-LLM

WebLLM: A High-Performance In-Browser LLM Inference Engine

arXiv1 repo

arXiv:2412.15803

web-llm

arXiv:2412.16117

arXiv1 repo

arXiv:2412.16117

Mate

Aria-UI: Visual Grounding for GUI Instructions

arXiv1 repo

arXiv:2412.16256

Aria-UI_Data

DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

arXiv1 repo

arXiv:2412.16334

dinov3

RCAEval: A Benchmark for Root Cause Analysis of Microservice Systems with Telemetry Data

arXiv1 repo

arXiv:2412.17015

kibana

A Reality Check on Context Utilisation for Retrieval-Augmented Generation

arXiv1 repo

arXiv:2412.17031

druid

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

arXiv1 repo

arXiv:2412.17667

versa

VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

arXiv1 repo

arXiv:2412.18194

vlabench_primitive_ft_dataset

Leave-One-EquiVariant: Alleviating invariance-related information loss in contrastive music representations

arXiv1 repo

arXiv:2412.18955

GDRetriever

Jasper and Stella: distillation of SOTA embedding models

arXiv1 repo

arXiv:2412.19048

RAG-Retrieval

RAG with Differential Privacy

arXiv1 repo

arXiv:2412.19291

dp-rag

An Engorgio Prompt Makes Large Language Model Babble on

arXiv1 repo

arXiv:2412.19394

Engorgio-prompt

OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

arXiv1 repo

arXiv:2412.19723

SeeClick

Fortran2CPP: Automating Fortran-to-C++ Translation using LLMs via Multi-Turn Dialogue and Dual-Agent Integration

arXiv1 repo

arXiv:2412.19770

Fortran2Cpp

Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging

arXiv1 repo

arXiv:2412.20070

Med-MAT

Navigating Image Restoration with VAR's Distribution Alignment Prior

arXiv1 repo

arXiv:2412.21063

VAR

Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

arXiv1 repo

arXiv:2412.21117

Prometheus

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

arXiv1 repo

arXiv:2412.21187

DeepEnlighten

Titans: Learning to Memorize at Test Time

arXiv1 repo

arXiv:2501.00663

Titan-Memory

AutoPresent: Designing Structured Visuals from Scratch

arXiv1 repo

arXiv:2501.00912

AutoPresent

FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

arXiv1 repo

arXiv:2501.01005

flashinfer

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models

arXiv1 repo

arXiv:2501.01034

Multitask-National-Speech-Corpus-v1

SVFR: A Unified Framework for Generalized Video Face Restoration

arXiv1 repo

arXiv:2501.01235

SVFR

LEO-Split: A Semi-Supervised Split Learning Framework over LEO Satellite Networks

arXiv1 repo

arXiv:2501.01293

SplitFM

arXiv:2501.02531

arXiv1 repo

arXiv:2501.02531

awesome-ai-sre

REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

arXiv1 repo

arXiv:2501.03262

DeepEnlighten

The Multiple Equal-Difference Structure of Cyclotomic Cosets

arXiv1 repo

arXiv:2501.03516

ru-promptriever

PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides

arXiv1 repo

arXiv:2501.03936

PPTAgent

RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark

arXiv1 repo

arXiv:2501.04440

sardet_100k

SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images

arXiv1 repo

arXiv:2501.04689

stable-point-aware-3d

TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training

arXiv1 repo

arXiv:2501.04765

HDM-xut-340M-anime

Reproducing HotFlip for Corpus Poisoning Attacks in Dense Retrieval

arXiv1 repo

arXiv:2501.04802

hotflip_corpus_poisoning

OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

arXiv1 repo

arXiv:2501.05510

OVO-Bench

Approximate well-balanced WENO finite difference schemes using a global-flux quadrature method with multi-step ODE integrator weights

arXiv1 repo

arXiv:2501.06155

Sonus-Lab

A General Framework for Inference-time Scaling and Steering of Diffusion Models

arXiv1 repo

arXiv:2501.06848

Fk-Diffusion-Steering

BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature

arXiv1 repo

arXiv:2501.07171

open-pmc-18m

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

arXiv1 repo

arXiv:2501.07525

RadAlign

Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

arXiv1 repo

arXiv:2501.07730

clustermark_1d-tokenizer

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

arXiv1 repo

arXiv:2501.07888

tarsier

Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models

arXiv1 repo

arXiv:2501.08248

ml-icr2

MiniMax-01: Scaling Foundation Models with Lightning Attention

arXiv1 repo

arXiv:2501.08313

MMLongBench-Doc

arXiv:2501.08325

arXiv1 repo

arXiv:2501.08325

GameFactory-Dataset

Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

arXiv1 repo

arXiv:2501.08331

Go-with-the-Flow

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

arXiv1 repo

arXiv:2501.08549

VRS-HQ

GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge

arXiv1 repo

arXiv:2501.08913

raid

FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

arXiv1 repo

arXiv:2501.09213

Med-R1

Foundations of Large Language Models

arXiv1 repo

arXiv:2501.09223

Medical-Assistant

Vision-Language Models Do Not Understand Negation

arXiv1 repo

arXiv:2501.09425

negbench

Revealing the $χ_{\rm eff}$-$q$ Correlation among Coalescing Binary Black Holes and Tentative Evidence for AGN-driven Hierarchical Mergers

arXiv1 repo

arXiv:2501.09495

Titan-Memory

Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis

arXiv1 repo

arXiv:2501.09502

ViSpeak

Enhancing the De-identification of Personally Identifiable Information in Educational Data

arXiv1 repo

arXiv:2501.09765

PrivacyAI

AI-Generated Music Detection and its Challenges

arXiv1 repo

arXiv:2501.10111

deepfake-detector

ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario

arXiv1 repo

arXiv:2501.10132

ComplexFuncBench

MechIR: A Mechanistic Interpretability Framework for Information Retrieval

arXiv1 repo

arXiv:2501.10165

MechIR

The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs

arXiv1 repo

arXiv:2501.10970

mmar-freeform

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

arXiv1 repo

arXiv:2501.11325

CatV2TON

Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks

arXiv1 repo

arXiv:2501.11733

MobileAgent

Parallel Sequence Modeling via Generalized Spatial Propagation Network

arXiv1 repo

arXiv:2501.12381

GSPN

O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

arXiv1 repo

arXiv:2501.12570

O1-Pruner

Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home

arXiv1 repo

arXiv:2501.12835

AdaRAGUE

Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback

arXiv1 repo

arXiv:2501.12895

TPO

Episodic Memories Generation and Evaluation Benchmark for Large Language Models

arXiv1 repo

arXiv:2501.13121

StepDeepResearch

Design of Bayesian Clinical Trials with Clustered Data

arXiv1 repo

arXiv:2501.13218

Titan-Memory

Improving Video Generation with Human Feedback

arXiv1 repo

arXiv:2501.13918

VideoReward

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities

arXiv1 repo

arXiv:2501.13921

BreezyVoice

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

arXiv1 repo

arXiv:2501.13926

Image-Generation-CoT

Longitudinal Abuse and Sentiment Analysis of Hollywood Movie Dialogues using Language Models

arXiv1 repo

arXiv:2501.13948

sentimentanalysis-Hollywood

Constructive Ordinal Exponentiation

arXiv1 repo

arXiv:2501.14542

TypeTopology

MatAnyone: Stable Video Matting with Consistent Memory Propagation

arXiv1 repo

arXiv:2501.14677

MatAnyone

CodeMonkeys: Scaling Test-Time Compute for Software Engineering

arXiv1 repo

arXiv:2501.14723

codemonkeys

MDEval: Evaluating and Enhancing Markdown Awareness in Large Language Models

arXiv1 repo

arXiv:2501.15000

opendataloader-bench

Overview of the Amphion Toolkit (v0.2)

arXiv1 repo

arXiv:2501.15442

Vevo1.5

Distributional Surgery for Language Model Activations

arXiv1 repo

arXiv:2501.15758

OT-Intervention

SpatialVLA: Exploring Spatial Representations for Visual-Language-Action Model

arXiv1 repo

arXiv:2501.15830

SpatialVLA

PISCO: Pretty Simple Compression for Retrieval-Augmented Generation

arXiv1 repo

arXiv:2501.16075

pisco

A foundation model for human-AI collaboration in medical literature mining

arXiv1 repo

arXiv:2501.16255

DeepRetrieval

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

arXiv1 repo

arXiv:2501.16411

PhysBench

360Brew: A Decoder-only Foundation Model for Personalized Ranking and Recommendation

arXiv1 repo

arXiv:2501.16450

linkedin-skills

TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models

arXiv1 repo

arXiv:2501.16937

TAID

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

arXiv1 repo

arXiv:2501.17148

MidSteer

SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

arXiv1 repo

arXiv:2501.17161

SFTvsRL_Data

Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation

arXiv1 repo

arXiv:2501.17433

Virus

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

arXiv1 repo

arXiv:2501.17811

Janus-Pro-7B

A Video-grounded Dialogue Dataset and Metric for Event-driven Activities

arXiv1 repo

arXiv:2501.18324

VDAct

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

arXiv1 repo

arXiv:2501.18427

SANA1.5_1.6B_1024px

ExeCoder: Empowering Large Language Models with Executability Representation for Code Translation

arXiv1 repo

arXiv:2501.18460

ExeCoder

Track-On: Transformer-based Online Point Tracking with Memory

arXiv1 repo

arXiv:2501.18487

track_on

R.I.P.: Better Models by Survival of the Fittest Prompts

arXiv1 repo

arXiv:2501.18578

fairseq2

Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

arXiv1 repo

arXiv:2501.18585

unlazy

arXiv:2501.18593

arXiv1 repo

arXiv:2501.18593

Jazz

Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models

arXiv1 repo

arXiv:2501.18877

des

Visual Autoregressive Modeling for Image Super-Resolution

arXiv1 repo

arXiv:2501.18993

VARSR

Enabling Autonomic Microservice Management through Self-Learning Agents

arXiv1 repo

arXiv:2501.19056

ACV

Scalable-Softmax Is Superior for Attention

arXiv1 repo

arXiv:2501.19399

Devstral-Small-2-24B-Instruct-2512

Low-Rank Adapting Models for Sparse Autoencoders

arXiv1 repo

arXiv:2501.19406

sae_kl_finetune

arXiv:2502.00055

arXiv1 repo

arXiv:2502.00055

awesome-ai-sre

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment

arXiv1 repo

arXiv:2502.00203

Llama-3_3-Nemotron-Super-49B-v1_5

RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes

arXiv1 repo

arXiv:2502.00392

PhysicalAI-VANTAGE-Bench

arXiv:2502.00640

arXiv1 repo

arXiv:2502.00640

collabllm

COVE: COntext and VEracity prediction for out-of-context images

arXiv1 repo

arXiv:2502.01194

5pils

On Almost Surely Safe Alignment of Large Language Models at Inference-Time

arXiv1 repo

arXiv:2502.01208

inf-guard

Preference Leakage: A Contamination Problem in LLM-as-a-judge

arXiv1 repo

arXiv:2502.01534

Llama-Krikri-8B-Instruct-GGUF

Massive Values in Self-Attention Modules are the Key to Contextual Knowledge Understanding

arXiv1 repo

arXiv:2502.01563

Rope_with_LLM

SliderSpace: Decomposing the Visual Capabilities of Diffusion Models

arXiv1 repo

arXiv:2502.01639

sliderspace

arXiv:2502.01651

arXiv1 repo

arXiv:2502.01651

llama2.mojo

Layer by Layer: Uncovering Hidden Representations in Language Models

arXiv1 repo

arXiv:2502.02013

information_flow

COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation

arXiv1 repo

arXiv:2502.02589

coconut_pancap

Reconstructing 3D Flow from 2D Data with Diffusion Transformer

arXiv1 repo

arXiv:2502.02593

graphify

arXiv:2502.02716

arXiv1 repo

arXiv:2502.02716

drowse

arXiv:2502.03052

arXiv1 repo

arXiv:2502.03052

dlm-jailbreak-transfer

CARROT: A Cost Aware Rate Optimal Router

arXiv1 repo

arXiv:2502.03261

router

High-Fidelity Simultaneous Speech-To-Speech Translation

arXiv1 repo

arXiv:2502.03382

tts-1.6b-en_fr

Pre-training Epidemic Time Series Forecasters with Compartmental Prototypes

arXiv1 repo

arXiv:2502.03393

CAPE

Do Large Language Model Benchmarks Test Reliability?

arXiv1 repo

arXiv:2502.03461

mmlu-redux

Efficient Image Restoration via Latent Consistency Flow Matching

arXiv1 repo

arXiv:2502.03500

ELIR

DynVFX: Augmenting Real Videos with Dynamic Content

arXiv1 repo

arXiv:2502.03621

dynvfx

arXiv:2502.03979

arXiv1 repo

arXiv:2502.03979

Music2Emotion

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

arXiv1 repo

arXiv:2502.04320

ConceptAttention

When One LLM Drools, Multi-LLM Collaboration Rules

arXiv1 repo

arXiv:2502.04506

model_collaboration

Fast Video Generation with Sliding Tile Attention

arXiv1 repo

arXiv:2502.04507

FastVideo

Towards Cost-Effective Reward Guided Text Generation

arXiv1 repo

arXiv:2502.04517

FaRMA

arXiv:2502.04522

arXiv1 repo

arXiv:2502.04522

improvnet

Sparsity-Based Interpolation of External, Internal and Swap Regret

arXiv1 repo

arXiv:2502.04543

Tiny-R2

EigenLoRAx: Recycling Adapters to Find Principal Subspaces for Resource-Efficient Adaptation and Inference

arXiv1 repo

arXiv:2502.04700

EigenLoRA

Chest X-ray Foundation Model with Global and Local Representations Integration

arXiv1 repo

arXiv:2502.05142

CheXFound

Apple vs. Oranges: Evaluating the Apple Silicon M-Series SoCs for HPC Performance and Efficiency

arXiv1 repo

arXiv:2502.05317

mlx-vit-tune

APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding

arXiv1 repo

arXiv:2502.05431

APE

Head-Specific Intervention Can Induce Misaligned AI Coordination in Large Language Models

arXiv1 repo

arXiv:2502.05945

targeted_intervention

VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer

arXiv1 repo

arXiv:2502.05979

Omni-Effects

VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data

arXiv1 repo

arXiv:2502.06737

VersaPRM

On the Emergence of Thinking in LLMs I: Searching for the Right Intuition

arXiv1 repo

arXiv:2502.06773

OpenRLHF

Towards a Robust Framework for Multimodal Hate Detection: A Study on Video vs. Image-based Content

arXiv1 repo

arXiv:2502.07138

Video-vs-Meme-Hate

GENERator: A Long-Context Generative Genomic Foundation Model

arXiv1 repo

arXiv:2502.07272

carbon-pretraining-corpus

Enhance-A-Video: Better Generated Video for Free

arXiv1 repo

arXiv:2502.07508

Enhance-A-Video

Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting

arXiv1 repo

arXiv:2502.07608

time2lang

TransMLA: Multi-Head Latent Attention Is All You Need

arXiv1 repo

arXiv:2502.07864

TransArch

Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?

arXiv1 repo

arXiv:2502.07963

MedLitSpin

Training Sparse Mixture Of Experts Text Embedding Models

arXiv1 repo

arXiv:2502.07972

nomic-embed-text-v2-moe

Light-A-Video: Training-free Video Relighting via Progressive Light Fusion

arXiv1 repo

arXiv:2502.08590

Light-A-Video

Harnessing Vision Models for Time Series Analysis: A Survey

arXiv1 repo

arXiv:2502.08869

TS-RAG

InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU

arXiv1 repo

arXiv:2502.08910

hip-attention

CRANE: Reasoning with constrained LLM generation

arXiv1 repo

arXiv:2502.09061

CRANE

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

arXiv1 repo

arXiv:2502.09674

LRP-eXplains-Transformers

Precise Parameter Localization for Textual Generation in Diffusion Models

arXiv1 repo

arXiv:2502.09935

t2i-text-localization

STAR: Spectral Truncation and Rescale for Model Merging

arXiv1 repo

arXiv:2502.10339

ComfyUI-LoRA-Optimizer

ReStyle3D: Scene-Level Appearance Transfer with Semantic Correspondences

arXiv1 repo

arXiv:2502.10377

ReStyle3D

arXiv:2502.10385

arXiv1 repo

arXiv:2502.10385

hashing-baseline

Region-Adaptive Sampling for Diffusion Transformers

arXiv1 repo

arXiv:2502.10389

RAS

I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models

arXiv1 repo

arXiv:2502.10458

ThinkDiff

D-CIPHER: Dynamic Collaborative Intelligent Multi-Agent System with Planner and Heterogeneous Executors for Offensive Security

arXiv1 repo

arXiv:2502.10931

awesome-cybersecurity-agentic-ai

Investigating Language Preference of Multilingual RAG Systems

arXiv1 repo

arXiv:2502.11175

LanguagePreference

MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation

arXiv1 repo

arXiv:2502.11234

maskflow

MARS: Mesh AutoRegressive Model for 3D Shape Detailization

arXiv1 repo

arXiv:2502.11390

VAR

Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More

arXiv1 repo

arXiv:2502.11494

DART

Continuous Diffusion Model for Language Modeling

arXiv1 repo

arXiv:2502.11564

dlms-sinks

Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?

arXiv1 repo

arXiv:2502.11598

watermarks-remover

HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims

arXiv1 repo

arXiv:2502.11753

5pils

ChordFormer: A Conformer-Based Architecture for Large-Vocabulary Audio Chord Recognition

arXiv1 repo

arXiv:2502.11840

SheetSage2

JoLT: Joint Probabilistic Predictions on Tabular Data Using LLMs

arXiv1 repo

arXiv:2502.11877

jolt

Bitnet.cpp: Efficient Edge Inference for Ternary LLMs

arXiv1 repo

arXiv:2502.11880

BitNet

A-MEM: Agentic Memory for LLM Agents

arXiv1 repo

arXiv:2502.12110

ai-memory

Idiosyncrasies in Large Language Models

arXiv1 repo

arXiv:2502.12150

llm-idiosyncrasies

Diffusion Models without Classifier-free Guidance

arXiv1 repo

arXiv:2502.12154

classifier-free-guidance-pytorch

Independence Tests for Language Models

arXiv1 repo

arXiv:2502.12292

model-tracing

YOLOv12: Attention-Centric Real-Time Object Detectors

arXiv1 repo

arXiv:2502.12524

DINOV3-YOLOV12

CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image

arXiv1 repo

arXiv:2502.12894

CAST

SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation

arXiv1 repo

arXiv:2502.13128

SongGen

AIDE: AI-Driven Exploration in the Space of Code

arXiv1 repo

arXiv:2502.13138

aideml

REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation

arXiv1 repo

arXiv:2502.13270

REALTALK

MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification

arXiv1 repo

arXiv:2502.13383

DataFlow

GeLLMO: Generalizing Large Language Models for Multi-property Molecule Optimization

arXiv1 repo

arXiv:2502.13398

BioMedGPT-Mol

Medical Image Classification with KAN-Integrated Transformers and Dilated Neighborhood Attention

arXiv1 repo

arXiv:2502.13693

medmnistc-api

GIMMICK -- Globally Inclusive Multimodal Multitask Cultural Knowledge Benchmarking

arXiv1 repo

arXiv:2502.13766

gimmick

Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models

arXiv1 repo

arXiv:2502.13989

erase-eval

Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data

arXiv1 repo

arXiv:2502.14044

FairLLaVA

PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC

arXiv1 repo

arXiv:2502.14282

MobileAgent

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

arXiv1 repo

arXiv:2502.14420

ChatVLA_public

Tree-of-Debate: Multi-Persona Debate Trees Elicit Critical Thinking for Scientific Comparative Analysis

arXiv1 repo

arXiv:2502.14767

tree-of-debate

Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

arXiv1 repo

arXiv:2502.14768

OpenRLHF

arXiv:2502.14866

arXiv1 repo

arXiv:2502.14866

Block-Sparse-Attention

PathRAG: Pruning Graph-based Retrieval Augmented Generation with Relational Paths

arXiv1 repo

arXiv:2502.14902

sweet-search

CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs

arXiv1 repo

arXiv:2502.14925

CodePromptZip-Token-Pruning

Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models

arXiv1 repo

arXiv:2502.15451

minimind

KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

arXiv1 repo

arXiv:2502.15602

kadtk

Learning to Reason from Feedback at Test-Time

arXiv1 repo

arXiv:2502.15771

FTTT

arXiv:2502.15814

arXiv1 repo

arXiv:2502.15814

slam_scaled

A Close Look at Decomposition-based XAI-Methods for Transformer Language Models

arXiv1 repo

arXiv:2502.15886

LRP-eXplains-Transformers

arXiv:2502.15964

arXiv1 repo

arXiv:2502.15964

minions

Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare

arXiv1 repo

arXiv:2502.16051

mentat

arXiv:2502.16681

arXiv1 repo

arXiv:2502.16681

SAEBench

Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch

arXiv1 repo

arXiv:2502.17173

CheemsRM

REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective

arXiv1 repo

arXiv:2502.17254

reinforce-attacks-llms

Benchmarking Retrieval-Augmented Generation in Multi-Modal Contexts

arXiv1 repo

arXiv:2502.17297

RARE

Delta Decompression for MoE-based LLMs Compression

arXiv1 repo

arXiv:2502.17298

D2MoE

On Relation-Specific Neurons in Large Language Models

arXiv1 repo

arXiv:2502.17355

relation-specific-neurons

ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis

arXiv1 repo

arXiv:2502.17475

minimind

ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents

arXiv1 repo

arXiv:2502.18017

ViDoSeek

TextGames: Learning to Self-Play Text-Based Puzzle Games via Language Model Reasoning

arXiv1 repo

arXiv:2502.18431

SmallPlan

DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers

arXiv1 repo

arXiv:2502.18460

dpr-scale

SolEval: Benchmarking Large Language Models for Repository-level Solidity Code Generation

arXiv1 repo

arXiv:2502.18793

SolEval

FinTSB: A Comprehensive and Practical Benchmark for Financial Time Series Forecasting

arXiv1 repo

arXiv:2502.18834

FinTSB

LongEval: A Comprehensive Analysis of Long-Text Generation Through a Plan-based Paradigm

arXiv1 repo

arXiv:2502.19103

LongEval

CritiQ: Mining Data Quality Criteria from Human Preferences

arXiv1 repo

arXiv:2502.19279

CritiQ

TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding

arXiv1 repo

arXiv:2502.19400

TheoremExplainAgent

The Mighty ToRR: A Benchmark for Table Reasoning and Robustness

arXiv1 repo

arXiv:2502.19412

helm

Stay Focused: Problem Drift in Multi-Agent Debate

arXiv1 repo

arXiv:2502.19559

rightmind

Self-rewarding correction for mathematical reasoning

arXiv1 repo

arXiv:2502.19613

verl-pipeline

MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

arXiv1 repo

arXiv:2502.19634

MedVLM-R1

Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation

arXiv1 repo

arXiv:2502.20056

MLRG

R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts

arXiv1 repo

arXiv:2502.20395

R2-T2

Protecting multimodal large language models against misleading visualizations

arXiv1 repo

arXiv:2502.20503

acl2026-misleading-visualizations

Autoregressive Medical Image Segmentation via Next-Scale Mask Prediction

arXiv1 repo

arXiv:2502.20784

VAR

Adaptive Keyframe Sampling for Long Video Understanding

arXiv1 repo

arXiv:2502.21271

CRAFT

Decoupling Content and Expression: Two-Dimensional Detection of AI-Generated Text

arXiv1 repo

arXiv:2503.00258

truth-mirror

Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models

arXiv1 repo

arXiv:2503.00743

ScoreRS

DuoDecoding: Hardware-aware Heterogeneous Speculative Decoding with Dynamic Multi-Sequence Drafting

arXiv1 repo

arXiv:2503.00784

DuoDecoding

TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop

arXiv1 repo

arXiv:2503.01013

TS-RAG

DLF: Extreme Image Compression with Dual-generative Latent Fusion

arXiv1 repo

arXiv:2503.01428

Dual-generative-Latent-Fusion

Effective High-order Graph Representation Learning for Credit Card Fraud Detection

arXiv1 repo

arXiv:2503.01556

antifraud

Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations

arXiv1 repo

arXiv:2503.01623

localmod

Detecting Stylistic Fingerprints of Large Language Models

arXiv1 repo

arXiv:2503.01659

humanizer

Elliptic Loss Regularization

arXiv1 repo

arXiv:2503.02138

Qwen3.6-VL-REAP-26B-A3B

Q-Filters: Leveraging QK Geometry for Efficient KV Cache Compression

arXiv1 repo

arXiv:2503.02812

qfilters

Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation

arXiv1 repo

arXiv:2503.03492

FindTrack

Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models

arXiv1 repo

arXiv:2503.03669

parlant

Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment

arXiv1 repo

arXiv:2503.04647

Implicit-Cross-Lingual-Rewarding

Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases

arXiv1 repo

arXiv:2503.04691

MedRBench

Generating Millions Of Lean Theorems With Proofs By Exploring State Transition Graphs

arXiv1 repo

arXiv:2503.04772

re-rl

FedMABench: Benchmarking Mobile Agents on Decentralized Heterogeneous User Data

arXiv1 repo

arXiv:2503.05143

FedMABench

CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation

arXiv1 repo

arXiv:2503.05255

CMMCoT

R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

arXiv1 repo

arXiv:2503.05592

SWE-Master

Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred Skills

arXiv1 repo

arXiv:2503.05641

zen5

WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs

arXiv1 repo

arXiv:2503.05683

WikiBigEdit

Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA

arXiv1 repo

arXiv:2503.05840

transformer-tricks

Rank-R1: Enhancing Reasoning in LLM-based Document Rerankers via Reinforcement Learning

arXiv1 repo

arXiv:2503.06034

llm-rankers

X2I: Seamless Integration of Multimodal Understanding into Diffusion Transformer via Attention Distillation

arXiv1 repo

arXiv:2503.06134

X2I

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

arXiv1 repo

arXiv:2503.06749

Vision-R1

From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers

arXiv1 repo

arXiv:2503.06923

ToCa

Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition

arXiv1 repo

arXiv:2503.06984

MelQCD-main

PE3R: Perception-Efficient 3D Reconstruction

arXiv1 repo

arXiv:2503.07507

PE3R

TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster

arXiv1 repo

arXiv:2503.07649

TS-RAG

OminiControl2: Efficient Conditioning for Diffusion Transformers

arXiv1 repo

arXiv:2503.08280

OminiControl

Referring to Any Person

arXiv1 repo

arXiv:2503.08507

Rex-Omni

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

arXiv1 repo

arXiv:2503.08679

CoTFaithChecker

Seal Your Backdoor with Variational Defense

arXiv1 repo

arXiv:2503.08829

VIBE

Optimal Control of Medical Drug in a Nonlocal Model of Solid Tumor Growth

arXiv1 repo

arXiv:2503.09208

sarashina2.2-ocr

VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary

arXiv1 repo

arXiv:2503.09402

VLog

PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop

arXiv1 repo

arXiv:2503.09595

pisa-experiments

CASteer: Cross-Attention Steering for Controllable Concept Erasure

arXiv1 repo

arXiv:2503.09630

CASteer

V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video

arXiv1 repo

arXiv:2503.09631

V2M4

AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents

arXiv1 repo

arXiv:2503.09780

ai-agent-privacy

Faster Inference of LLMs using FP8 on the Intel Gaudi

arXiv1 repo

arXiv:2503.09975

neural-compressor

MaterialMVP: Illumination-Invariant Material Generation via Multi-view PBR Diffusion

arXiv1 repo

arXiv:2503.10289

MaterialMVP

Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

arXiv1 repo

arXiv:2503.10460

Light-R1

A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1

arXiv1 repo

arXiv:2503.10635

M-Attack_AdvSamples

ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition

arXiv1 repo

arXiv:2503.10673

ZeroSumEval

LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss

arXiv1 repo

arXiv:2503.10679

ml-lineas

Neighboring Autoregressive Modeling for Efficient Visual Generation

arXiv1 repo

arXiv:2503.10696

NAR

FlowTok: Flowing Seamlessly Across Text and Image Tokens

arXiv1 repo

arXiv:2503.10772

clustermark_1d-tokenizer

TxAgent: An AI Agent for Therapeutic Reasoning Across a Universe of Tools

arXiv1 repo

arXiv:2503.10970

ToolUniverse

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

arXiv1 repo

arXiv:2503.11315

MMS-LLaMA

Safe-VAR: Safe Visual Autoregressive Model for Text-to-Image Generative Watermarking

arXiv1 repo

arXiv:2503.11324

VAR

Efficient Distributed MLLM Training with Cornstarch

arXiv1 repo

arXiv:2503.11367

Cornstarch

ReCamMaster: Camera-Controlled Generative Rendering from A Single Video

arXiv1 repo

arXiv:2503.11647

MultiCamVideo-Dataset

VGGT: Visual Geometry Grounded Transformer

arXiv1 repo

arXiv:2503.11651

VGGT-1B

Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

arXiv1 repo

arXiv:2503.11832

Unlearn-Trace

AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

arXiv1 repo

arXiv:2503.12559

ACL25-AdaReTaKe

Reliable and Efficient Amortized Model-based Evaluation

arXiv1 repo

arXiv:2503.13335

helm

xLSTM 7B: A Recurrent LLM for Fast and Efficient Inference

arXiv1 repo

arXiv:2503.13427

xlstm

Why Do Multi-Agent LLM Systems Fail?

arXiv1 repo

arXiv:2503.13657

prisma

Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

arXiv1 repo

arXiv:2503.13939

Med-R1

MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding

arXiv1 repo

arXiv:2503.13964

Mdocagent-dataset

Inference-Time Intervention in Large Language Models for Reliable Requirement Verification

arXiv1 repo

arXiv:2503.14130

targeted_intervention

MoonCast: High-Quality Zero-Shot Podcast Generation

arXiv1 repo

arXiv:2503.14345

MoonCast

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

arXiv1 repo

arXiv:2503.14350

VEGGIE-VidEdit

Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

arXiv1 repo

arXiv:2503.14492

cosmos-transfer1

Measuring AI Ability to Complete Long Software Tasks

arXiv1 repo

arXiv:2503.14499

unlazy

MusicInfuser: Making Video Diffusion Listen and Dance

arXiv1 repo

arXiv:2503.14505

MusicInfuser

VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

arXiv1 repo

arXiv:2503.14939

mllm_number_sense

Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator

arXiv1 repo

arXiv:2503.15457

test-time-scaling

Object-Spatial Programming

arXiv1 repo

arXiv:2503.15812

jac

Survey on Evaluation of LLM-based Agents

arXiv1 repo

arXiv:2503.16416

StepDeepResearch

XAttention: Block Sparse Attention with Antidiagonal Scoring

arXiv1 repo

arXiv:2503.16428

RetrievalAttention

Sonata: Self-Supervised Learning of Reliable Point Representations

arXiv1 repo

arXiv:2503.16429

sonata

arXiv:2503.16611

arXiv1 repo

arXiv:2503.16611

PanoramaGenInpaint

Aligning Text-to-Music Evaluation with Human Preferences

arXiv1 repo

arXiv:2503.16669

TuneJury

Revisiting End-To-End Sparse Autoencoder Training: A Short Finetune Is All You Need

arXiv1 repo

arXiv:2503.17272

sae_kl_finetune

Won: Establishing Best Practices for Korean Financial NLP

arXiv1 repo

arXiv:2503.17963

Won-Instruct

PolarFree: Polarization-based Reflection-free Imaging

arXiv1 repo

arXiv:2503.18055

PolarFree

Distil-xLSTM: Learning Attention Mechanisms through Recurrent Structures

arXiv1 repo

arXiv:2503.18565

distil_xlstm

Any6D: Model-free 6D Pose Estimation of Novel Objects

arXiv1 repo

arXiv:2503.18673

Any6D

CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models

arXiv1 repo

arXiv:2503.18886

CFG-Zero-Star

SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

arXiv1 repo

arXiv:2503.18943

ml-unigen

Aether: Geometric-Aware Unified World Modeling

arXiv1 repo

arXiv:2503.18945

AetherV1

Understanding and Improving Information Preservation in Prompt Compression for LLMs

arXiv1 repo

arXiv:2503.19114

information-preservation-in-prompt-compression

Scaling Down Text Encoders of Text-to-Image Diffusion Models

arXiv1 repo

arXiv:2503.19897

DistillT5

Scaling Vision Pre-Training to 4K Resolution

arXiv1 repo

arXiv:2503.19903

Robopoint_Humble

Audio-centric Video Understanding Benchmark without Text Shortcut

arXiv1 repo

arXiv:2503.19951

AVUTBenchmark

RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy

arXiv1 repo

arXiv:2503.20158

rxrx3-core

UniVRSE: Unified Vision-conditioned Response Semantic Entropy for Hallucination Detection in Medical Vision-Language Models

arXiv1 repo

arXiv:2503.20504

VASE

Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models

arXiv1 repo

arXiv:2503.20752

RoboBrain2.5

Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields

arXiv1 repo

arXiv:2503.20776

Feature4X

BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology

arXiv1 repo

arXiv:2503.20880

BioX-CPath

ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation

arXiv1 repo

arXiv:2503.21729

ReaRAG-20k

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness

arXiv1 repo

arXiv:2503.21755

VBench

Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

arXiv1 repo

arXiv:2503.21758

Lumina-Image-2.0

PharMolixFM: All-Atom Foundation Models for Molecular Modeling and Generation

arXiv1 repo

arXiv:2503.21788

OpenBioMed

WMCopier: Forging Invisible Image Watermarks on Arbitrary Images

arXiv1 repo

arXiv:2503.22330

WMCopier

Text-Only Data Synthesis for Vision Language Model Training

arXiv1 repo

arXiv:2503.22655

Modality_Gap_Theory

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

arXiv1 repo

arXiv:2503.22976

SPAR-Bench

Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

arXiv1 repo

arXiv:2503.23278

awesome-mcp-security

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

arXiv1 repo

arXiv:2503.23368

VLIPP

HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment

arXiv1 repo

arXiv:2503.23907

HumanAesExpert

AI2Agent: An End-to-End Framework for Deploying AI Projects as Autonomous Agents

arXiv1 repo

arXiv:2503.23948

ai2apps

A Multi-Stage Auto-Context Deep Learning Framework for Tissue and Nuclei Segmentation and Classification in H&E-Stained Histological Images of Advanced Melanoma

arXiv1 repo

arXiv:2503.23958

PumaSubmit

Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

arXiv1 repo

arXiv:2503.24290

Open-Reasoner-Zero

Easi3R: Estimating Disentangled Motion from DUSt3R Without Training

arXiv1 repo

arXiv:2503.24391

Easi3R

Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge

arXiv1 repo

arXiv:2504.00042

knowledge-gap

Universal Zero-shot Embedding Inversion

arXiv1 repo

arXiv:2504.00147

adversarial_decoding

An Illusion of Progress? Assessing the Current State of Web Agents

arXiv1 repo

arXiv:2504.01382

UI-TARS

Diffusion-Guided Gaussian Splatting for Large-Scale Unconstrained 3D Reconstruction and Novel View Synthesis

arXiv1 repo

arXiv:2504.01960

SphereForge

T*: Re-thinking Temporal Search for Long-Form Video Understanding

arXiv1 repo

arXiv:2504.02259

LongVideoHaystack

Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

arXiv1 repo

arXiv:2504.02438

ViLAMP-llava-qwen_sig

Exploration-Driven Generative Interactive Environments

arXiv1 repo

arXiv:2504.02515

stable-retro

How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence

arXiv1 repo

arXiv:2504.02904

post-training-mechanistic-analysis

Noiser: Bounded Input Perturbations for Attributing Large Language Models

arXiv1 repo

arXiv:2504.02911

Noiser

MedSAM2: Segment Anything in 3D Medical Images and Videos

arXiv1 repo

arXiv:2504.03600

MedSAM2

MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits

arXiv1 repo

arXiv:2504.03767

awesome-mcp-security

Clinical ModernBERT: An efficient and long context encoder for biomedical text

arXiv1 repo

arXiv:2504.03964

DiagnosisCoding

UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

arXiv1 repo

arXiv:2504.04423

UniToken

VSLAM-LAB: A Comprehensive Framework for Visual SLAM Methods and Datasets

arXiv1 repo

arXiv:2504.04457

VSLAM-LAB

Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs

arXiv1 repo

arXiv:2504.04715

llm-api-audit

FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

arXiv1 repo

arXiv:2504.04842

FantasyTalking

M-Prometheus: A Suite of Open Multilingual LLM Judges

arXiv1 repo

arXiv:2504.04953

prometheus-eval

SmolVLM: Redefining small and efficient multimodal models

arXiv1 repo

arXiv:2504.05299

SmolVLM2-2.2B-Instruct

Gaussian Mixture Flow Matching Models

arXiv1 repo

arXiv:2504.05304

LakonLab

STAGE: Stemmed Accompaniment Generation through Prefix-Based Conditioning

arXiv1 repo

arXiv:2504.05690

stage

Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

arXiv1 repo

arXiv:2504.05812

EMPO

HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference

arXiv1 repo

arXiv:2504.05897

HybriMoE

AEGIS: Human Attention-based Explainable Guidance for Intelligent Vehicle Systems

arXiv1 repo

arXiv:2504.05950

AEGIS

CAI: An Open, Bug Bounty-Ready Cybersecurity AI

arXiv1 repo

arXiv:2504.06017

awesome-cybersecurity-agentic-ai

QGen Studio: An Adaptive Question-Answer Generation, Training and Evaluation Platform

arXiv1 repo

arXiv:2504.06136

qgen-studio

Large language models as uncertainty-calibrated optimizers for experimental discovery

arXiv1 repo

arXiv:2504.06265

sego

Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations

arXiv1 repo

arXiv:2504.06792

EASYEP

OSCAR: Online Soft Compression And Reranking

arXiv1 repo

arXiv:2504.07109

pisco

Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction

arXiv1 repo

arXiv:2504.07375

UniHand

Malware analysis assisted by AI with R2AI

arXiv1 repo

arXiv:2504.07574

r2ai

BoxDreamer: Dreaming Box Corners for Generalizable Object Pose Estimation

arXiv1 repo

arXiv:2504.07955

BoxDreamer

Detect Anything 3D in the Wild

arXiv1 repo

arXiv:2504.07958

DetAny3D

Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction

arXiv1 repo

arXiv:2504.07961

Geo4D

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

arXiv1 repo

arXiv:2504.08066

aideml

Out of Style: RAG's Fragility to Linguistic Variation

arXiv1 repo

arXiv:2504.08231

RAG-fragility-to-linguistic-variation

Enterprise-Grade Security for the Model Context Protocol (MCP): Frameworks and Mitigation Strategies

arXiv1 repo

arXiv:2504.08623

awesome-mcp-security

SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting

arXiv1 repo

arXiv:2504.08850

SpecEE

LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping

arXiv1 repo

arXiv:2504.08902

LatentGenerativeAnamorphoses

PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models

arXiv1 repo

arXiv:2504.08966

PACT

DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training

arXiv1 repo

arXiv:2504.09710

DUMP

Reasoning Models Can Be Effective Without Thinking

arXiv1 repo

arXiv:2504.09858

CoDE-Stop

Guiding Reasoning in Small Language Models with LLM Assistance

arXiv1 repo

arXiv:2504.09923

SMART

Aligning Anime Video Generation with Human Feedback

arXiv1 repo

arXiv:2504.10044

Index-anisora

RealHarm: A Collection of Real-World Language Model Application Failures

arXiv1 repo

arXiv:2504.10277

realharm

RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users

arXiv1 repo

arXiv:2504.10445

RealWebAssist

Efficient Process Reward Model Training via Active Learning

arXiv1 repo

arXiv:2504.10559

ActivePRM

How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients

arXiv1 repo

arXiv:2504.10766

Gradient_Unified

UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer

arXiv1 repo

arXiv:2504.11289

UniAnimate-DiT

DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

arXiv1 repo

arXiv:2504.11456

DeepMath-103K

Activated LoRA: Fine-tuned LLMs for Intrinsics

arXiv1 repo

arXiv:2504.12397

granitelib-rag-r1.0

One Model to Rig Them All: Diverse Skeleton Rigging with UniRig

arXiv1 repo

arXiv:2504.12451

UniRig

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition

arXiv1 repo

arXiv:2504.12562

ZeroSumEval

Chinese-Vicuna: A Chinese Instruction-following Llama-based Model

arXiv1 repo

arXiv:2504.12737

Chinese-Vicuna

MCP Guardian: A Security-First Layer for Safeguarding MCP-Based AI System

arXiv1 repo

arXiv:2504.12757

awesome-mcp-security

Set You Straight: Auto-Steering Denoising Trajectories to Sidestep Unwanted Concepts

arXiv1 repo

arXiv:2504.12782

MACE

SoK: Security of EMV Contactless Payment Systems

arXiv1 repo

arXiv:2504.12812

awesome-connected-things-sec

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

arXiv1 repo

arXiv:2504.13059

RoboTwin

Long Range Navigator (LRN): Extending robot planning horizons beyond metric maps

arXiv1 repo

arXiv:2504.13149

nebula2-wildos

Long-context Non-factoid Question Answering in Indic Languages

arXiv1 repo

arXiv:2504.13615

IndicGenQA

Fashion-RAG: Multimodal Fashion Image Editing via Retrieval-Augmented Generation

arXiv1 repo

arXiv:2504.14011

fashion-rag

Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale

arXiv1 repo

arXiv:2504.14225

PersonaMem-v2

LoRe: Personalizing LLMs via Low-Rank Reward Modeling

arXiv1 repo

arXiv:2504.14439

LoRe

BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation

arXiv1 repo

arXiv:2504.14538

BookWorld

TAPIP3D: Tracking Any Point in Persistent 3D Geometry

arXiv1 repo

arXiv:2504.14717

tapip3d

arXiv:2504.15071

arXiv1 repo

arXiv:2504.15071

aria-medium-embedding

EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models

arXiv1 repo

arXiv:2504.15133

WikiBigEdit

Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning

arXiv1 repo

arXiv:2504.15275

PURE

Event2Vec: Processing Neuromorphic Events Directly by Representations in Vector Space

arXiv1 repo

arXiv:2504.15371

repro-event2vec-neuromorphic-events-vector-space

VeriCoder: Enhancing LLM-Based RTL Code Generation through Functional Correctness Validation

arXiv1 repo

arXiv:2504.15659

VeriCoder

Dynamic Early Exit in Reasoning Models

arXiv1 repo

arXiv:2504.15895

CoDE-Stop

LLMSR@XLLM25: Less is More: Enhancing Structured Multi-Agent Reasoning via Quality-Guided Distillation

arXiv1 repo

arXiv:2504.16408

Less-is-More

MAGIC: Near-Optimal Data Attribution for Deep Learning

arXiv1 repo

arXiv:2504.16430

bergson

Enhancing LLM-Based Agents via Global Planning and Hierarchical Execution

arXiv1 repo

arXiv:2504.16563

agentic-awesome-skills

Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward

arXiv1 repo

arXiv:2504.16727

Visual-Variations-Robustness

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

arXiv1 repo

arXiv:2504.17207

APC-VLM

Unsupervised Corpus Poisoning Attacks in Continuous Space for Dense Retrieval

arXiv1 repo

arXiv:2504.17884

unsupervised_corpus_poisoning

Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning

arXiv1 repo

arXiv:2504.17950

mindcraft

Efficiency, Expressivity, and Extensibility in a Close-to-Metal NPU Programming Interface

arXiv1 repo

arXiv:2504.18430

IRON

Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

arXiv1 repo

arXiv:2504.19413

khms-memory

semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage

arXiv1 repo

arXiv:2504.19867

Semi-PD

Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents

arXiv1 repo

arXiv:2504.19956

awesome-cybersecurity-agentic-ai

Simplified and Secure MCP Gateways for Enterprise AI Integration

arXiv1 repo

arXiv:2504.19997

awesome-mcp-security

Learning Streaming Video Representation via Multitask Training

arXiv1 repo

arXiv:2504.20041

streamformer-timesformer

UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities

arXiv1 repo

arXiv:2504.20734

UniversalRAG

Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition

arXiv1 repo

arXiv:2504.20938

circuit_backup

Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math

arXiv1 repo

arXiv:2504.21233

Phi-4-mini-reasoning

LightEMMA: A Longitudinal Evaluation of Vision-Language Models for Autonomous Driving

arXiv1 repo

arXiv:2505.00284

LightEMMA

T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

arXiv1 repo

arXiv:2505.00703

Image-Generation-CoT

FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors

arXiv1 repo

arXiv:2505.01322

FreeInsert

Practical Efficiency of Muon for Pretraining

arXiv1 repo

arXiv:2505.02222

Muon

Connecting Independently Trained Modes via Layer-Wise Connectivity

arXiv1 repo

arXiv:2505.02604

repro-connecting-independently-trained-modes-layer-wise-low-loss-paths

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

arXiv1 repo

arXiv:2505.02922

RetrievalAttention

UCSC at SemEval-2025 Task 3: Context, Models and Prompt Optimization for Automated Hallucination Detection in LLM Output

arXiv1 repo

arXiv:2505.03030

semeval-2025-task3

RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation

arXiv1 repo

arXiv:2505.03275

gbrain

RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration

arXiv1 repo

arXiv:2505.03673

RoboOS

CoCoB: Adaptive Collaborative Combinatorial Bandits for Online Recommendation

arXiv1 repo

arXiv:2505.03840

graphify

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

arXiv1 repo

arXiv:2505.04021

prism-research

arXiv:2505.04080

arXiv1 repo

arXiv:2505.04080

llama2.mojo

Multimodal Integrated Knowledge Transfer to Large Language Models through Preference Optimization with Biomedical Applications

arXiv1 repo

arXiv:2505.05736

MINT-LLM

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA

arXiv1 repo

arXiv:2505.06356

maya

Embedding Atlas: Low-Friction, Interactive Embedding Visualization

arXiv1 repo

arXiv:2505.06386

embedding-atlas

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection

arXiv1 repo

arXiv:2505.07293

attention-influence

FLUXSynID: A Framework for Identity-Controlled Synthetic Face Generation with Document and Live Images

arXiv1 repo

arXiv:2505.07530

FLUXSynID

RAI: Flexible Agent Framework for Embodied AI

arXiv1 repo

arXiv:2505.07532

rai

OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval

arXiv1 repo

arXiv:2505.07879

OMGM

Behind Maya: Building a Multilingual Vision Language Model

arXiv1 repo

arXiv:2505.08910

maya

MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment

arXiv1 repo

arXiv:2505.09372

MAKE

Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?

arXiv1 repo

arXiv:2505.09439

Omni-R1

A large-scale evaluation of commonsense knowledge in humans and large language models

arXiv1 repo

arXiv:2505.10309

commonsense-llm-eval

UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation

arXiv1 repo

arXiv:2505.10483

UniEval

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

arXiv1 repo

arXiv:2505.10610

MMLongBench

TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation

arXiv1 repo

arXiv:2505.10696

TartanGround

Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models

arXiv1 repo

arXiv:2505.10844

figures4papers

Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere

arXiv1 repo

arXiv:2505.11029

cr4wm

DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling

arXiv1 repo

arXiv:2505.11196

DiCo

Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models

arXiv1 repo

arXiv:2505.11245

Diffusion-NPO

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference

arXiv1 repo

arXiv:2505.11329

tokenweave

HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages

arXiv1 repo

arXiv:2505.11475

Llama-3_3-Nemotron-Super-49B-GenRM

SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training

arXiv1 repo

arXiv:2505.11594

SageAttention

Self-NPO: Data-Free Diffusion Model Enhancement via Truncated Diffusion Fine-Tuning

arXiv1 repo

arXiv:2505.11777

Diffusion-NPO

TinyRS-R1: Compact Multimodal Language Model for Remote Sensing

arXiv1 repo

arXiv:2505.12099

TinyRS

Estimation of Treatment Harm Rate via Partitioning

arXiv1 repo

arXiv:2505.12209

zen5

Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning

arXiv1 repo

arXiv:2505.12370

ScreenSpot-Pro-GUI-Grounding

Harnessing the Universal Geometry of Embeddings

arXiv1 repo

arXiv:2505.12540

vec2vec

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

arXiv1 repo

arXiv:2505.12669

t2m-inferalign

arXiv:2505.12674

arXiv1 repo

arXiv:2505.12674

ml-sid-dit

MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

arXiv1 repo

arXiv:2505.13031

MindOmni

Thinkless: LLM Learns When to Think

arXiv1 repo

arXiv:2505.13379

Thinkless-1.5B-Warmup

VSA: Faster Video Diffusion with Trainable Sparse Attention

arXiv1 repo

arXiv:2505.13389

FastVideo

Krikri: Advancing Open Large Language Models for Greek

arXiv1 repo

arXiv:2505.13772

Llama-Krikri-8B-Instruct-GGUF

Let's Verify Math Questions Step by Step

arXiv1 repo

arXiv:2505.13903

DataFlow

MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem

arXiv1 repo

arXiv:2505.14148

LLM-MM-Agent

Capturing the Effects of Quantization on Trojans in Code LLMs

arXiv1 repo

arXiv:2505.14200

quantized-code-llm-security

RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection

arXiv1 repo

arXiv:2505.14318

Radar

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

arXiv1 repo

arXiv:2505.14362

DeepEyes-7B

PRL: Prompts from Reinforcement Learning

arXiv1 repo

arXiv:2505.14412

PRL-Prompts-from-Reinforcement-Learning

Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models

arXiv1 repo

arXiv:2505.14454

VidCom2

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

arXiv1 repo

arXiv:2505.14640

VideoChat-Flash

Quartet: Native FP4 Training Can Be Optimal for Large Language Models

arXiv1 repo

arXiv:2505.14669

lectures

Reward Reasoning Model

arXiv1 repo

arXiv:2505.14674

RRM-7B

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

arXiv1 repo

arXiv:2505.14682

ml-unigen

Multi-modal Traffic Scenario Generation for Autonomous Driving System Testing

arXiv1 repo

arXiv:2505.14881

TrafficComposer

PiFlow: Principle-Aware Scientific Discovery with Multi-Agent Collaboration

arXiv1 repo

arXiv:2505.15047

pigollum

lmgame-Bench: How Good are LLMs at Playing Games?

arXiv1 repo

arXiv:2505.15146

GRL

R&D-Agent-Quant: A Multi-Agent Framework for Data-Centric Factors and Model Joint Optimization

arXiv1 repo

arXiv:2505.15155

qlib

Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control

arXiv1 repo

arXiv:2505.15304

sqil

MIRB: Mathematical Information Retrieval Benchmark

arXiv1 repo

arXiv:2505.15585

mirb

Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space

arXiv1 repo

arXiv:2505.15778

Soft-Thinking

arXiv:2505.16239

arXiv1 repo

arXiv:2505.16239

DOVE

arXiv:2505.16369

arXiv1 repo

arXiv:2505.16369

xares

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

arXiv1 repo

arXiv:2505.16839

LaViDa

Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs

arXiv1 repo

arXiv:2505.16967

rlhn-680K

X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs

arXiv1 repo

arXiv:2505.16997

rightmind

Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

arXiv1 repo

arXiv:2505.17017

Image-Generation-CoT

LLM Agents for Interactive Exploration of Historical Cadastre Data: Framework and Application to Venice

arXiv1 repo

arXiv:2505.17148

venice-agents

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

arXiv1 repo

arXiv:2505.17568

JALMBench

arXiv:2505.17598

arXiv1 repo

arXiv:2505.17598

dlm-jailbreak-transfer

VIBE: Vector Index Benchmark for Embeddings

arXiv1 repo

arXiv:2505.17810

vibe

Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities

arXiv1 repo

arXiv:2505.17862

Daily-Omni

WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions

arXiv1 repo

arXiv:2505.18151

WonderPlay

arXiv:2505.18186

arXiv1 repo

arXiv:2505.18186

musicdiscovery

Think or Not? Exploring Thinking Efficiency in Large Reasoning Models via an Information-Theoretic Lens

arXiv1 repo

arXiv:2505.18237

CoDE-Stop

Dynamic Risk Assessments for Offensive Cybersecurity Agents

arXiv1 repo

arXiv:2505.18384

awesome-cybersecurity-agentic-ai

CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos

arXiv1 repo

arXiv:2505.18561

CoT-RVS

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

arXiv1 repo

arXiv:2505.19037

Speech-IFEval

Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs

arXiv1 repo

arXiv:2505.19075

UniR

Conventional Contrastive Learning Often Falls Short: Improving Dense Retrieval with Cross-Encoder Listwise Distillation and Synthetic Data

arXiv1 repo

arXiv:2505.19274

cadet-embed-base-v1

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

arXiv1 repo

arXiv:2505.19462

T5Gemma-TTS

Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks

arXiv1 repo

arXiv:2505.19472

Parallel-Hybrid-Model

Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection

arXiv1 repo

arXiv:2505.19475

VDS-TTT

STRAP: Spatio-Temporal Pattern Retrieval for Out-of-Distribution Generalization

arXiv1 repo

arXiv:2505.19547

A2TTA

TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization

arXiv1 repo

arXiv:2505.19586

TailorKV

Multi-Agent Collaboration via Evolving Orchestration

arXiv1 repo

arXiv:2505.19591

ChatDev

ErpGS: Equirectangular Image Rendering enhanced with 3D Gaussian Regularization

arXiv1 repo

arXiv:2505.19883

SphereForge

EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition

arXiv1 repo

arXiv:2505.20033

Voice-Acting-Pipeline

Iterative Self-Incentivization Empowers Large Language Models as Agentic Searchers

arXiv1 repo

arXiv:2505.20128

SearchLM

syftr: Pareto-Optimal Generative AI

arXiv1 repo

arXiv:2505.20266

syftr

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

arXiv1 repo

arXiv:2505.20279

VLM-3R

Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms

arXiv1 repo

arXiv:2505.20322

data_for_STA

SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents

arXiv1 repo

arXiv:2505.20411

SWE-rebench

DIPO: Dual-State Images Controlled Articulated Object Generation Powered by Diverse Data

arXiv1 repo

arXiv:2505.20460

EmbodiedGen

RefAV: Towards Planning-Centric Scenario Mining

arXiv1 repo

arXiv:2505.20981

RefAV

Who Reasons in the Large Language Models?

arXiv1 repo

arXiv:2505.20993

SpaceOm

SageAttention2++: A More Efficient Implementation of SageAttention2

arXiv1 repo

arXiv:2505.21136

SageAttention

Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

arXiv1 repo

arXiv:2505.21374

Video-Holmes

QuARI: Query Adaptive Retrieval Improvement

arXiv1 repo

arXiv:2505.21647

QuARI

ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge

arXiv1 repo

arXiv:2505.21906

ChatVLA_public

TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation

arXiv1 repo

arXiv:2505.22176

table-metric-study

Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design

arXiv1 repo

arXiv:2505.22179

FR-Spec

Inference-Time Scaling of Discrete Diffusion Models via Importance Weighting and Optimal Proposal Design

arXiv1 repo

arXiv:2505.22524

smc_ddm_iclr

Thinking with Generated Images

arXiv1 repo

arXiv:2505.22525

thinking-with-generated-images

A Tool for Generating Exceptional Behavior Tests With Large Language Models

arXiv1 repo

arXiv:2505.22818

exLong

RocqStar: Leveraging Similarity-driven Retrieval and Agentic Systems for Rocq generation

arXiv1 repo

arXiv:2505.22846

coqpilot

From Token to Action: State Machine Reasoning to Mitigate Overthinking in Information Retrieval

arXiv1 repo

arXiv:2505.23059

SMR

Speeding up Model Loading with fastsafetensors

arXiv1 repo

arXiv:2505.23072

fastsafetensors

Less is More: Unlocking Specialization of Time Series Foundation Models via Structured Pruning

arXiv1 repo

arXiv:2505.23195

Prune-then-Finetune

Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency Optimization

arXiv1 repo

arXiv:2505.23387

Venus

KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction

arXiv1 repo

arXiv:2505.23416

KVzip

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

arXiv1 repo

arXiv:2505.23606

Muddit

Inference-time Scaling of Diffusion Models through Classical Search

arXiv1 repo

arXiv:2505.23614

Diffusion-inference-scaling

AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora

arXiv1 repo

arXiv:2505.23628

AutoSchemaKG

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

arXiv1 repo

arXiv:2505.23656

VideoREPA

EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

arXiv1 repo

arXiv:2505.23732

EmotionRankCLAP

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

arXiv1 repo

arXiv:2505.23885

owl

TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine

arXiv1 repo

arXiv:2505.24063

TCM-Ladder

Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning

arXiv1 repo

arXiv:2505.24478

cognee

arXiv:2505.24685

arXiv1 repo

arXiv:2505.24685

misp-galaxy

LegalEval-Q: A New Benchmark for The Quality Evaluation of LLM-Generated Legal Text

arXiv1 repo

arXiv:2505.24826

minimind

MiniMax-Remover: Taming Bad Noise Helps Video Object Removal

arXiv1 repo

arXiv:2505.24873

minimax-remover

ACE-Step: A Step Towards Music Generation Foundation Model

arXiv1 repo

arXiv:2506.00045

ACE-Step

Latent Wavelet Diffusion For Ultra-High-Resolution Image Synthesis

arXiv1 repo

arXiv:2506.00433

LatentWaveletDiffusion

CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

arXiv1 repo

arXiv:2506.00519

CausalAbstain

Learning with Calibration: Exploring Test-Time Computing of Spatio-Temporal Forecasting

arXiv1 repo

arXiv:2506.00635

A2TTA

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction

arXiv1 repo

arXiv:2506.00975

NTPP

GigaAM: Efficient Self-Supervised Learner for Speech Recognition

arXiv1 repo

arXiv:2506.01192

GigaAM

DNAEdit: Direct Noise Alignment for Text-Guided Rectified Flow Editing

arXiv1 repo

arXiv:2506.01430

FlowEdit

Policy as Code, Policy as Type

arXiv1 repo

arXiv:2506.01446

extensible-mcp

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

arXiv1 repo

arXiv:2506.01844

smolvla_base

Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning

arXiv1 repo

arXiv:2506.01953

Fast-in-Slow

Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem

arXiv1 repo

arXiv:2506.02040

awesome-mcp-security

Answer Convergence as a Signal for Early Stopping in Reasoning

arXiv1 repo

arXiv:2506.02536

CoDE-Stop

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

arXiv1 repo

arXiv:2506.02557

KUEA

Solving Inverse Problems with FLAIR

arXiv1 repo

arXiv:2506.02680

FLAIR

Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning

arXiv1 repo

arXiv:2506.02738

open-pmc-18m

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

arXiv1 repo

arXiv:2506.03096

fuselip

Native-Resolution Image Synthesis

arXiv1 repo

arXiv:2506.03131

NiT-diffusers

Test-Time Scaling of Diffusion Models via Noise Trajectory Search

arXiv1 repo

arXiv:2506.03164

diffusion-tts

Robustness in Both Domains: CLIP Needs a Robust Text Encoder

arXiv1 repo

arXiv:2506.03355

LEAF

Seed-Coder: Let the Code Model Curate Data for Itself

arXiv1 repo

arXiv:2506.03524

swallow-code-v2

MiMo-VL Technical Report

arXiv1 repo

arXiv:2506.03569

MiMo-VL-7B-SFT-2508

Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games

arXiv1 repo

arXiv:2506.03610

Orak

AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance

arXiv1 repo

arXiv:2506.03828

AssetOpsBench

Pre$^3$: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation

arXiv1 repo

arXiv:2506.03887

lightllm

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning

arXiv1 repo

arXiv:2506.04034

Rex-Omni

EuroLLM-9B: Technical Report

arXiv1 repo

arXiv:2506.04079

AMALIA-9B-0626-DPO

OpenThoughts: Data Recipes for Reasoning Models

arXiv1 repo

arXiv:2506.04178

OpenThoughts-114k

Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning

arXiv1 repo

arXiv:2506.04207

Revisual-R1

Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation

arXiv1 repo

arXiv:2506.04225

HunyuanWorld-Voyager

WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning

arXiv1 repo

arXiv:2506.04363

WorldPrediction

Identity Testing for Circuits with Exponentiation Gates

arXiv1 repo

arXiv:2506.04529

mirage

Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets

arXiv1 repo

arXiv:2506.04598

CLIP_benchmark

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

arXiv1 repo

arXiv:2506.04614

MobileAgent

ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning

arXiv1 repo

arXiv:2506.04941

ArtVIP

AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models

arXiv1 repo

arXiv:2506.05140

AudioLens

arXiv:2506.05301

arXiv1 repo

arXiv:2506.05301

SeedVR2-3B

arXiv:2506.05414

arXiv1 repo

arXiv:2506.05414

SAVVY-Bench

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

arXiv1 repo

arXiv:2506.05551

MLLM-Semantic-Hallucination

BecomingLit: Relightable Gaussian Avatars with Hybrid Neural Shading

arXiv1 repo

arXiv:2506.06271

becominglit

Memory OS of AI Agent

arXiv1 repo

arXiv:2506.06326

MemoryOS

Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict

arXiv1 repo

arXiv:2506.06485

LLM-KnowledgeConflict-TaskMatters

Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text

arXiv1 repo

arXiv:2506.07001

Adversarial-Paraphrasing

Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models

arXiv1 repo

arXiv:2506.07468

selfplay-redteaming

R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation

arXiv1 repo

arXiv:2506.07826

R3D2

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

arXiv1 repo

arXiv:2506.08009

Self-Forcing

ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving

arXiv1 repo

arXiv:2506.08052

recogdrive_tome

LEANN: A Low-Storage Vector Index

arXiv1 repo

arXiv:2506.08276

LEANN

arXiv:2506.08641

arXiv1 repo

arXiv:2506.08641

TiViT

Edit Flows: Flow Matching with Edit Operations

arXiv1 repo

arXiv:2506.09018

dllm

CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation

arXiv1 repo

arXiv:2506.09109

CAIRE

Seedance 1.0: Exploring the Boundaries of Video Generation Models

arXiv1 repo

arXiv:2506.09113

Awesome-AITools

TRACE: Grounding Time Series in Context for Multimodal Embedding and Retrieval

arXiv1 repo

arXiv:2506.09114

TRACE-TimeseriesRAG-Dataset

ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning

arXiv1 repo

arXiv:2506.09513

ReasonMed

MEDUSA: A Multimodal Deep Fusion Multi-Stage Training Framework for Speech Emotion Recognition in Naturalistic Conditions

arXiv1 repo

arXiv:2506.09556

medusa

Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks

arXiv1 repo

arXiv:2506.09593

medmnistc-api

Query-Level Uncertainty in Large Language Models

arXiv1 repo

arXiv:2506.09669

query_level_uncertainty

Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation

arXiv1 repo

arXiv:2506.09736

Vision-Matters

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

arXiv1 repo

arXiv:2506.09985

vjepa2

A quantum semantic framework for natural language processing

arXiv1 repo

arXiv:2506.10077

npcpy

SoK: Evaluating Jailbreak Guardrails for Large Language Models

arXiv1 repo

arXiv:2506.10597

JailbreakGuardrailBenchmark

EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence

arXiv1 repo

arXiv:2506.10600

EmbodiedGen

SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation

arXiv1 repo

arXiv:2506.10622

sdialog

Skillful joint probabilistic weather forecasting from marginals

arXiv1 repo

arXiv:2506.10772

weathernext

CyclicReflex: Improving Reasoning Models via Cyclical Reflection Token Scheduling

arXiv1 repo

arXiv:2506.11077

CyclicReflex

Fed-HeLLo: Efficient Federated Foundation Model Fine-Tuning with Heterogeneous LoRA Allocation

arXiv1 repo

arXiv:2506.12213

Fed-PLoRA

Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

arXiv1 repo

arXiv:2506.12311

Phonikud-yi

Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding

arXiv1 repo

arXiv:2506.12336

Trust-videoLLMs

Extracting Composition-Dependent Diffusion Coefficients Over a Very Large Composition Range in NiCoFeCrMn High Entropy Alloy Following Strategic Design of Diffusion Couples and Physics Informed Neural Network Numerical Method

arXiv1 repo

arXiv:2506.12345

unlimited-ocr-benchmark

FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation

arXiv1 repo

arXiv:2506.12494

FlexRAG

Scaling Test-time Compute for LLM Agents

arXiv1 repo

arXiv:2506.12928

OAgents

MAMMA: Markerless & Automatic Multi-Person Motion Action Capture

arXiv1 repo

arXiv:2506.13040

mamma

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

arXiv1 repo

arXiv:2506.13053

ZipVoice

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy

arXiv1 repo

arXiv:2506.13284

AceReason-1.1-SFT

BUT System for the MLC-SLM Challenge

arXiv1 repo

arXiv:2506.13414

DiCoW_v3_2

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

arXiv1 repo

arXiv:2506.13585

MiniMax-AI.github.io

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs

arXiv1 repo

arXiv:2506.14003

Unlearn-Trace

SimpleDoc: Multi-Modal Document Understanding with Dual-Cue Page Retrieval and Iterative Refinement

arXiv1 repo

arXiv:2506.14035

SimpleDoc

Foundation Model Insights and a Multi-Model Approach for Superior Fine-Grained One-shot Subset Selection

arXiv1 repo

arXiv:2506.14473

RAM-APL

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks

arXiv1 repo

arXiv:2506.14512

SpaceQwen2.5-VL-3B-Instruct

Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

arXiv1 repo

arXiv:2506.14965

Reasoning360

cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree

arXiv1 repo

arXiv:2506.15655

sweet-search

OAgents: An Empirical Study of Building Effective Agents

arXiv1 repo

arXiv:2506.15741

OAgents

FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

arXiv1 repo

arXiv:2506.15742

kontext-bench

ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning

arXiv1 repo

arXiv:2506.16499

aideml

Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models

arXiv1 repo

arXiv:2506.16853

RATTPO

arXiv:2506.17055

arXiv1 repo

arXiv:2506.17055

FM-music-tagging

Simultaneous Translation with Offline Speech and LLM Models in CUNI Submission to IWSLT 2025

arXiv1 repo

arXiv:2506.17077

WhisperLiveKit

Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights

arXiv1 repo

arXiv:2506.17337

MedVLMBench

Distilling On-device Language Models for Robot Planning with Minimal Human Intervention

arXiv1 repo

arXiv:2506.17486

PRISM

Pre-Trained LLM is a Semantic-Aware and Generalizable Segmentation Booster

arXiv1 repo

arXiv:2506.18034

LLM4Seg

TAB: Unified Benchmarking of Time Series Anomaly Detection Methods

arXiv1 repo

arXiv:2506.18046

TAB

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

arXiv1 repo

arXiv:2506.18421

finepdfs

USAD: Universal Speech and Audio Representation via Distillation

arXiv1 repo

arXiv:2506.18843

usad

Benchmarking Music Generation Models and Metrics via Human Preference Studies

arXiv1 repo

arXiv:2506.19085

TuneJury

Why Do Open-Source LLMs Struggle with Data Analysis? A Systematic Empirical Study

arXiv1 repo

arXiv:2506.19794

DataMind

MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

arXiv1 repo

arXiv:2506.19835

MAM

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models

arXiv1 repo

arXiv:2506.20251

Qresafe

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content

arXiv1 repo

arXiv:2506.20331

Biomed-Enriched

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

arXiv1 repo

arXiv:2506.20920

fineweb-2

Little by Little: Continual Learning via Incremental Mixture of Rank-1 Associative Memory Experts

arXiv1 repo

arXiv:2506.21035

repro-little-by-little-continual-learning-via-incremental-mixture-of-rank-1-associative-memory-e

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

arXiv1 repo

arXiv:2506.21076

Hunyuan3D-Omni

FairyGen: Storied Cartoon Video from a Single Child-Drawn Character

arXiv1 repo

arXiv:2506.21272

FairyGen

Bridging Offline and Online Reinforcement Learning for LLMs

arXiv1 repo

arXiv:2506.21495

fairseq2

Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models

arXiv1 repo

arXiv:2506.21509

DLC

MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents

arXiv1 repo

arXiv:2506.21605

Membench

The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements

arXiv1 repo

arXiv:2506.22419

aideml

arXiv:2506.22557

arXiv1 repo

arXiv:2506.22557

dlm-jailbreak-transfer

Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment

arXiv1 repo

arXiv:2506.22685

Embedding-Adjustment

MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering

arXiv1 repo

arXiv:2506.22900

MOTOR

arXiv:2506.23329

arXiv1 repo

arXiv:2506.23329

IR3D-bench

Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop

arXiv1 repo

arXiv:2506.23351

RoboTwin

PathDiff: Histopathology Image Synthesis with Unpaired Text and Mask Conditions

arXiv1 repo

arXiv:2506.23440

PathDiff

ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models

arXiv1 repo

arXiv:2507.00898

ONLY

Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks

arXiv1 repo

arXiv:2507.01297

compactds-retrieval

arXiv:2507.01663

arXiv1 repo

arXiv:2507.01663

TransferQueue

RoboBrain 2.0 Technical Report

arXiv1 repo

arXiv:2507.02029

RoboSpatial-Eval

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

arXiv1 repo

arXiv:2507.02259

MemAgent

AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench

arXiv1 repo

arXiv:2507.02554

aideml

AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models

arXiv1 repo

arXiv:2507.02664

AIGI-Holmes

VLAI: A RoBERTa-Based Model for Automated Vulnerability Severity Classification

arXiv1 repo

arXiv:2507.03607

vulnerability-severity-classification-chinese-macbert-base

Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models

arXiv1 repo

arXiv:2507.03916

SwanLab

SeqTex: Generate Mesh Textures in Video Sequence

arXiv1 repo

arXiv:2507.04285

ComfyUI-SeqTex

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

arXiv1 repo

arXiv:2507.04447

DreamVLA

Unveiling the Potential of Diffusion Large Language Model in Controllable Generation

arXiv1 repo

arXiv:2507.04504

dLLM-CtrlGen

any4: Learned 4-bit Numeric Representation for LLMs

arXiv1 repo

arXiv:2507.04610

any4

Spatio-Temporal LLM: Reasoning about Environments and Actions

arXiv1 repo

arXiv:2507.05258

Spatio-Temporal-LLM

Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model

arXiv1 repo

arXiv:2507.05513

Eagle

An autonomous agent for auditing and improving the reliability of clinical AI models

arXiv1 repo

arXiv:2507.05755

medmnistc-api

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance

arXiv1 repo

arXiv:2507.06272

Monkey

Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation

arXiv1 repo

arXiv:2507.06607

Phi-4-mini-flash-reasoning

Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

arXiv1 repo

arXiv:2507.07095

MotionHub

SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

arXiv1 repo

arXiv:2507.07610

Spatial-Visualization-Benchmark

THUNDER: Tile-level Histopathology image UNDERstanding benchmark

arXiv1 repo

arXiv:2507.07860

thunder

Predicting and generating antibiotics against future pathogens with ApexOracle

arXiv1 repo

arXiv:2507.07862

ApexOracle

M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning

arXiv1 repo

arXiv:2507.08306

M2-Reasoning

InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes

arXiv1 repo

arXiv:2507.08416

InstaScene

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity

arXiv1 repo

arXiv:2507.08771

FR-Spec

From One to More: Contextual Part Latents for 3D Generation

arXiv1 repo

arXiv:2507.08772

partverse

Exploiting Leaderboards for Large-Scale Distribution of Malicious Models

arXiv1 repo

arXiv:2507.08983

trojanclimb

Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition

arXiv1 repo

arXiv:2507.09116

WenetSpeech-Chuan

arXiv:2507.09264

arXiv1 repo

arXiv:2507.09264

walrus

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching

arXiv1 repo

arXiv:2507.09318

ZipVoice

The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs

arXiv1 repo

arXiv:2507.11097

DIJA

EXAONE 4.0: Unified Large Language Models Integrating Non-reasoning and Reasoning Modes

arXiv1 repo

arXiv:2507.11407

EXAONE-4.0

Are Vision Foundation Models Ready for Out-of-the-Box Medical Image Registration?

arXiv1 repo

arXiv:2507.11569

Foundation-based-reg

General Modular Harness for LLM Agents in Multi-Turn Gaming Environments

arXiv1 repo

arXiv:2507.11633

stable-retro

A Survey of Deep Learning for Geometry Problem Solving

arXiv1 repo

arXiv:2507.11936

VisNumBench

Kevin: Multi-Turn RL for Generating CUDA Kernels

arXiv1 repo

arXiv:2507.11948

TritonForge

Characterizing State Space Model and Hybrid Language Model Performance with Long Context

arXiv1 repo

arXiv:2507.12442

SSM-Scope

SpatialTrackerV2: 3D Point Tracking Made Easy

arXiv1 repo

arXiv:2507.12462

SpaTrackerV2

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

arXiv1 repo

arXiv:2507.12705

AudioJudge

DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment

arXiv1 repo

arXiv:2507.12796

DeQA-Doc

DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

arXiv1 repo

arXiv:2507.13985

DreamScene

Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

arXiv1 repo

arXiv:2507.15028

video-tt

OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities

arXiv1 repo

arXiv:2507.15085

OCRGenBench

Solving Formal Math Problems by Decomposition and Iterative Reflection

arXiv1 repo

arXiv:2507.15225

Seed-Prover

MEETI: A Multimodal ECG Dataset from MIMIC-IV-ECG with Signals, Images, Features and Interpretations

arXiv1 repo

arXiv:2507.15255

GEM

RDMA: Cost Effective Agent-Driven Rare Disease Mining from Electronic Health Records

arXiv1 repo

arXiv:2507.15867

RDMA

SDBench: A Comprehensive Benchmark Suite for Speaker Diarization

arXiv1 repo

arXiv:2507.16136

OpenBench

Morpheus: A Neural-driven Animatronic Face with Hybrid Actuation and Diverse Emotion Control

arXiv1 repo

arXiv:2507.16645

Morpheus-Software

AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation

arXiv1 repo

arXiv:2507.16940

AURA

InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation

arXiv1 repo

arXiv:2507.17520

VLA_Instruction_Tuning

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice

arXiv1 repo

arXiv:2507.17527

UniSS

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training

arXiv1 repo

arXiv:2507.17634

Ling-1T

Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment

arXiv1 repo

arXiv:2507.19002

banana100-additional-iqa-models

RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow

arXiv1 repo

arXiv:2507.19280

RemoteReasoner

MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks

arXiv1 repo

arXiv:2507.19634

MCIF

Fine-structure Preserved Real-world Image Super-resolution via Transfer VAE Training

arXiv1 repo

arXiv:2507.20291

OSEDiff

Meta CLIP 2: A Worldwide Scaling Recipe

arXiv1 repo

arXiv:2507.22062

MetaCLIP

CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography

arXiv1 repo

arXiv:2507.22953

CADS-dataset

Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models

arXiv1 repo

arXiv:2507.23159

Full-Duplex-Bench

Unveiling Super Experts in Mixture-of-Experts Large Language Models

arXiv1 repo

arXiv:2507.23279

reap

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding

arXiv1 repo

arXiv:2507.23478

3D-R1

Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving

arXiv1 repo

arXiv:2507.23726

Seed-Prover

CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks

arXiv1 repo

arXiv:2507.23751

synthetic-data

Phi-Ground Tech Report: Advancing Perception in GUI Grounding

arXiv1 repo

arXiv:2507.23779

Phi-Ground

From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media

arXiv1 repo

arXiv:2508.00497

SocialAlign

Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation

arXiv1 repo

arXiv:2508.00912

Reasoning_Length_Prediction

SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy

arXiv1 repo

arXiv:2508.01188

SwanLab

OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets

arXiv1 repo

arXiv:2508.01630

healthadvocate

Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time

arXiv1 repo

arXiv:2508.02037

STIM

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

arXiv1 repo

arXiv:2508.02317

VeOmni

Efficient Agents: Building Effective Agents While Reducing Cost

arXiv1 repo

arXiv:2508.02694

OAgents

AgentSight: System-Level Observability for AI Agents Using eBPF

arXiv1 repo

arXiv:2508.02736

agentsight

CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic Data

arXiv1 repo

arXiv:2508.02879

CauKer2M

LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking

arXiv1 repo

arXiv:2508.03440

Soft-Thinking

Unravelling the Probabilistic Forest: Arbitrage in Prediction Markets

arXiv1 repo

arXiv:2508.03474

CloddsBot

Agent Lightning: Train ANY AI Agents with Reinforcement Learning

arXiv1 repo

arXiv:2508.03680

agent-lightning

HPSv3: Towards Wide-Spectrum Human Preference Score

arXiv1 repo

arXiv:2508.03789

banana100-additional-iqa-models

TempFlow-GRPO: When Timing Matters for GRPO in Flow Models

arXiv1 repo

arXiv:2508.04324

EditScore

Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization

arXiv1 repo

arXiv:2508.04796

tokenizer-flores-validation

Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment

arXiv1 repo

arXiv:2508.04865

MultiPL-E

R-Zero: Self-Evolving Reasoning LLM from Zero Data

arXiv1 repo

arXiv:2508.05004

R-Zero

Reasoning through Exploration: A Reinforcement Learning Framework for Robust Function Calling

arXiv1 repo

arXiv:2508.05118

AWorld

CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL

arXiv1 repo

arXiv:2508.05242

SwanLab

Learning to Reason for Factuality

arXiv1 repo

arXiv:2508.05618

fairseq2

NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

arXiv1 repo

arXiv:2508.05835

nemo-nano-codec-22khz-1.89kbps-21.5fps

More Is Better: A MoE-Based Emotion Recognition Framework with Human Preference Alignment

arXiv1 repo

arXiv:2508.06036

MER2025-MRAC25

arXiv:2508.06098

arXiv1 repo

arXiv:2508.06098

Resonate

VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding

arXiv1 repo

arXiv:2508.07493

VisR-Bench

arXiv:2508.07647

arXiv1 repo

arXiv:2508.07647

LaRender

Beyond Ten Turns: Unlocking Long-Horizon Agentic Search with Large-Scale Asynchronous RL

arXiv1 repo

arXiv:2508.07976

ASearcher-train-data

Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing

arXiv1 repo

arXiv:2508.09192

d3LLM

MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models

arXiv1 repo

arXiv:2508.09779

MoIIE

Profile-Aware Maneuvering: A Dynamic Multi-Agent System for Robust GAIA Problem Solving by AWorld

arXiv1 repo

arXiv:2508.09889

AWorld

MOC: Meta-Optimized Classifier for Few-Shot Whole Slide Image Classification

arXiv1 repo

arXiv:2508.09967

MOC

TexVerse: A Universe of 3D Objects with High-Resolution Textures

arXiv1 repo

arXiv:2508.10868

TexVerse

VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip

arXiv1 repo

arXiv:2508.10931

VSF

arXiv:2508.11131

arXiv1 repo

arXiv:2508.11131

Block-Sparse-Attention

TinyTim: A Family of Language Models for Divergent Generation

arXiv1 repo

arXiv:2508.11607

npcpy

QuarkMed Medical Foundation Model Technical Report

arXiv1 repo

arXiv:2508.11894

MedXpertQA

Improving Densification in 3D Gaussian Splatting for High-Fidelity Rendering

arXiv1 repo

arXiv:2508.12313

SphereForge

MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security

arXiv1 repo

arXiv:2508.12538

awesome-mcp-security

Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing

arXiv1 repo

arXiv:2508.12631

router

Cryfish: On deep audio analysis with Large Language Models

arXiv1 repo

arXiv:2508.12666

VoicePersonification

arXiv:2508.13009

arXiv1 repo

arXiv:2508.13009

Matrix-Game-2.0

OptimalThinkingBench: Evaluating Over and Underthinking in LLMs

arXiv1 repo

arXiv:2508.13141

unlazy

NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

arXiv1 repo

arXiv:2508.14444

NVIDIA-Nemotron-Nano-9B-v2

Mobile-Agent-v3: Fundamental Agents for GUI Automation

arXiv1 repo

arXiv:2508.15144

MobileAgent

Intern-S1: A Scientific Multimodal Foundation Model

arXiv1 repo

arXiv:2508.15763

Intern-S1-Pro

TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill and Decode Inference

arXiv1 repo

arXiv:2508.15881

TransArch

Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping

arXiv1 repo

arXiv:2508.15904

PathPT

Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning

arXiv1 repo

arXiv:2508.16929

circuit_backup

EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems

arXiv1 repo

arXiv:2508.17623

emo-reasoning

The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis

arXiv1 repo

arXiv:2508.17627

CoDE-Stop

ST-Raptor: LLM-Powered Semi-Structured Table Question Answering

arXiv1 repo

arXiv:2508.18190

ST-Raptor

Wan-S2V: Audio-Driven Cinematic Video Generation

arXiv1 repo

arXiv:2508.18621

Wan2.2-S2V-14B

MobileCLIP2: Improving Multi-Modal Reinforced Training

arXiv1 repo

arXiv:2508.20691

ml-mobileclip

rStar2-Agent: Agentic Reasoning Technical Report

arXiv1 repo

arXiv:2508.20722

Open-AgentRL

On the Theoretical Limitations of Embedding-Based Retrieval

arXiv1 repo

arXiv:2508.21038

vestige

Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning

arXiv1 repo

arXiv:2508.21048

Veritas

AHELM: A Holistic Evaluation of Audio-Language Models

arXiv1 repo

arXiv:2508.21376

helm

MobiAgent: A Systematic Framework for Customizable Mobile Agents

arXiv1 repo

arXiv:2509.00531

MobiAgent

CCE: Confidence-Consistency Evaluation for Time Series Anomaly Detection

arXiv1 repo

arXiv:2509.01098

CCE

Reinforced Visual Perception with Tools

arXiv1 repo

arXiv:2509.01656

REVPT-data

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

arXiv1 repo

arXiv:2509.02020

FireRedTTS2

arXiv:2509.02398

arXiv1 repo

arXiv:2509.02398

Resonate

GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning

arXiv1 repo

arXiv:2509.02492

GRAM-RR-TrainingData

Jointly Reinforcing Diversity and Quality in Language Model Generations

arXiv1 repo

arXiv:2509.02534

darling

Planning with Reasoning using Vision Language World Model

arXiv1 repo

arXiv:2509.02722

cr4wm

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

arXiv1 repo

arXiv:2509.04474

SpecTTS-Bench

Why Language Models Hallucinate

arXiv1 repo

arXiv:2509.04664

generative-ai

Hunyuan-MT Technical Report

arXiv1 repo

arXiv:2509.05209

Hunyuan-MT-7B-GGUF

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

arXiv1 repo

arXiv:2509.06155

Verse-Bench

Text4Seg++: Advancing Image Segmentation via Generative Language Modeling

arXiv1 repo

arXiv:2509.06321

Text4Seg

Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet

arXiv1 repo

arXiv:2509.06861

unlazy

Interleaving Reasoning for Better Text-to-Image Generation

arXiv1 repo

arXiv:2509.06945

Vision-R1

RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use

arXiv1 repo

arXiv:2509.06980

RL-Factory

veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD

arXiv1 repo

arXiv:2509.07003

veScale

arXiv:2509.07447

arXiv1 repo

arXiv:2509.07447

trajgaze

Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates

arXiv1 repo

arXiv:2509.09550

neucodec

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives

arXiv1 repo

arXiv:2509.09838

stable-retro

Dynamic Vulnerability Patching for Heterogeneous Embedded Systems Using Stack Frame Reconstruction

arXiv1 repo

arXiv:2509.10213

awesome-connected-things-sec

GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography

arXiv1 repo

arXiv:2509.10344

GLAM

Towards Understanding Visual Grounding in Visual Language Models

arXiv1 repo

arXiv:2509.10345

gui-agent

MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models

arXiv1 repo

arXiv:2509.10569

watermarks-remover

Trading-R1: Financial Trading with LLM Reasoning via Reinforcement Learning

arXiv1 repo

arXiv:2509.11420

TradingAgents

UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning

arXiv1 repo

arXiv:2509.11543

MobileAgent

HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking

arXiv1 repo

arXiv:2509.11552

hichunk

NeuroStrike: Neuron-Level Attacks on Aligned LLMs

arXiv1 repo

arXiv:2509.11864

JailbreakLab

MMORE: Massive Multimodal Open RAG & Extraction

arXiv1 repo

arXiv:2509.11937

mmore

RailSafeNet: Visual Scene Understanding for Tram Safety

arXiv1 repo

arXiv:2509.12125

RailSafeNet

Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection

arXiv1 repo

arXiv:2509.12546

Agent4FaceForgery

Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder

arXiv1 repo

arXiv:2509.12883

lego-edit

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

arXiv1 repo

arXiv:2509.13282

ChartGaze

Do Activation Verbalization Methods Convey Privileged Information?

arXiv1 repo

arXiv:2509.13316

verb_faithfulness

OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft

arXiv1 repo

arXiv:2509.13347

VeOmni

BWCache: Accelerating Video Diffusion Transformers through Block-Wise Caching

arXiv1 repo

arXiv:2509.13789

ltx2-vidgen-skill

LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures

arXiv1 repo

arXiv:2509.14252

mlx-tune

arXiv:2509.14427

arXiv1 repo

arXiv:2509.14427

hashing-baseline

Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis

arXiv1 repo

arXiv:2509.14579

X-Voice

ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data

arXiv1 repo

arXiv:2509.15221

ScaleCUA-Data

DiffusionNFT: Online Diffusion Reinforcement with Forward Process

arXiv1 repo

arXiv:2509.16117

GRPO

Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models

arXiv1 repo

arXiv:2509.16696

decoding_uncertainty

The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

arXiv1 repo

arXiv:2509.16972

q-frame

Evolution of Concepts in Language Model Pre-Training

arXiv1 repo

arXiv:2509.17196

circuit_backup

AI Pangaea: Unifying Intelligence Islands for Adapting Myriad Tasks

arXiv1 repo

arXiv:2509.17460

awdemos

Qwen3-Omni Technical Report

arXiv1 repo

arXiv:2509.17765

Qwen3-Omni

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

arXiv1 repo

arXiv:2509.18004

WSChuan-Train

GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction

arXiv1 repo

arXiv:2509.18090

GeoSVR

The Illusion of Readiness in Health AI

arXiv1 repo

arXiv:2509.18234

PeruMedQA

Track-On2: Enhancing Online Point Tracking with Memory

arXiv1 repo

arXiv:2509.19115

track_on

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

arXiv1 repo

arXiv:2509.19244

LaViDa

Formal Verification of Minimax Algorithms

arXiv1 repo

arXiv:2509.20138

sunfish

Play by the Type Rules: Inferring Constraints for LLM Functions in Declarative Programs

arXiv1 repo

arXiv:2509.20208

blendsql

Uncovering Graph Reasoning in Decoder-only Transformers with Circuit Tracing

arXiv1 repo

arXiv:2509.20336

GraphGhost

EmbeddingGemma: Powerful and Lightweight Text Representations

arXiv1 repo

arXiv:2509.20354

embeddinggemma-300m

Revisiting Data Challenges of Computational Pathology: A Pack-based Multiple Instance Learning Training Framework

arXiv1 repo

arXiv:2509.20923

CPathPatchFeature

Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution

arXiv1 repo

arXiv:2509.21072

AWorld

NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics

arXiv1 repo

arXiv:2509.21309

NewtonGen

d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation

arXiv1 repo

arXiv:2509.21474

d2

Self-Speculative Biased Decoding for Faster Re-Translation

arXiv1 repo

arXiv:2509.21740

NoLanguageLeftWaiting

Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis

arXiv1 repo

arXiv:2509.22167

Semantic-DACVAE-Japanese-32dim

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

arXiv1 repo

arXiv:2509.22186

MinerU

Enabling Approximate Joint Sampling in Diffusion LMs

arXiv1 repo

arXiv:2509.22738

ParallelBench

A Capacity-Based Rationale for Multi-Head Attention

arXiv1 repo

arXiv:2509.22840

repro-a-capacity-based-rationale-for-multi-head-attention

Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding

arXiv1 repo

arXiv:2509.23050

understanding_lp

ToolUniverse: An open platform for democratizing AI scientists

arXiv1 repo

arXiv:2509.23426

ToolUniverse

EfficientMIL: Efficient Linear-Complexity MIL Method for WSI Classification

arXiv1 repo

arXiv:2509.23640

EfficientMIL

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

arXiv1 repo

arXiv:2509.23661

LLaVA-OneVision-1.5-4B-Instruct

Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation

arXiv1 repo

arXiv:2509.23866

dart-gui

RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization

arXiv1 repo

arXiv:2509.23991

SphereForge

SparseD: Sparse Attention for Diffusion Language Models

arXiv1 repo

arXiv:2509.24014

SparseD

UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities

arXiv1 repo

arXiv:2509.24391

UniFlow-Audio

Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting

arXiv1 repo

arXiv:2509.24789

tsfmx

Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws

arXiv1 repo

arXiv:2509.24914

repro-single-head-attention-inductive-bias-spectral-highdim

Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

arXiv1 repo

arXiv:2509.25050

GRPO

Scaling Generalist Data-Analytic Agents

arXiv1 repo

arXiv:2509.25084

DataMind

arXiv:2509.25127

arXiv1 repo

arXiv:2509.25127

ml-sid-dit

Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events

arXiv1 repo

arXiv:2509.25146

fast-feature-fields

DepthLM: Metric Depth From Vision Language Models

arXiv1 repo

arXiv:2509.25413

angleLLM

A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments

arXiv1 repo

arXiv:2509.25609

abxlab

More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models

arXiv1 repo

arXiv:2509.25848

VAPO

IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance

arXiv1 repo

arXiv:2509.26231

IMG-Multimodal-Diffusion-Alignment

MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation

arXiv1 repo

arXiv:2509.26391

MotionRAG

SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From

arXiv1 repo

arXiv:2509.26404

SeedPrints

fev-bench: A Realistic Benchmark for Time Series Forecasting

arXiv1 repo

arXiv:2509.26468

fev

dParallel: Learnable Parallel Decoding for dLLMs

arXiv1 repo

arXiv:2509.26488

d3LLM

Entropy After </Think> for reasoning model early exiting

arXiv1 repo

arXiv:2509.26522

CoDE-Stop

OceanGym: A Benchmark Environment for Underwater Embodied Agents

arXiv1 repo

arXiv:2509.26536

OceanGPT

LongCodeZip: Compress Long Context for Code Language Models

arXiv1 repo

arXiv:2510.00446

LongCodeZip

GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness

arXiv1 repo

arXiv:2510.00536

GUI-KV

On Predictability of Reinforcement Learning Dynamics for Large Language Models

arXiv1 repo

arXiv:2510.00553

ReproAlphaRL

ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning

arXiv1 repo

arXiv:2510.01010

banana100-additional-iqa-models

Aristotle: IMO-level Automated Theorem Proving

arXiv1 repo

arXiv:2510.01346

open-atp

GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation

arXiv1 repo

arXiv:2510.02186

GeoPurify

FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation

arXiv1 repo

arXiv:2510.02315

FOCUS

RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents

arXiv1 repo

arXiv:2510.02609

RedCodeAgent

To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration

arXiv1 repo

arXiv:2510.02676

ecf8

Consolidating Reinforcement Learning for Multimodal Discrete Diffusion Models

arXiv1 repo

arXiv:2510.02880

MaskGRPO

Product-Quantised Image Representation for High-Quality Image Synthesis

arXiv1 repo

arXiv:2510.03191

WeavePrompt

Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer

arXiv1 repo

arXiv:2510.03342

RoboSpatial-Eval

Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning

arXiv1 repo

arXiv:2510.04213

wespeaker

RAP: 3D Rasterization Augmented End-to-End Planning

arXiv1 repo

arXiv:2510.04333

RAP_ckpts

When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA

arXiv1 repo

arXiv:2510.04849

PsiloQA

RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection

arXiv1 repo

arXiv:2510.04885

rl-injector

Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization

arXiv1 repo

arXiv:2510.05038

test-time-hybrid-retrieval

Agentic Misalignment: How LLMs Could Be Insider Threats

arXiv1 repo

arXiv:2510.05179

ogham-mcp

NorMuon: Making Muon more efficient and scalable

arXiv1 repo

arXiv:2510.05491

modded-nanogpt

TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis

arXiv1 repo

arXiv:2510.06063

TelecomTS

TokenChain: A Discrete Speech Chain via Semantic Token Modeling

arXiv1 repo

arXiv:2510.06201

TokenChain

Conditional Denoising Diffusion Model-Based Robust MR Image Reconstruction from Highly Undersampled Data

arXiv1 repo

arXiv:2510.06335

WeavePrompt

StyleKeeper: Prevent Content Leakage using Negative Visual Query Guidance

arXiv1 repo

arXiv:2510.06827

StyleKeeper

Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods

arXiv1 repo

arXiv:2510.07143

DART

Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner

arXiv1 repo

arXiv:2510.07838

Full-Duplex-Bench

LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?

arXiv1 repo

arXiv:2510.07962

LightReasoner

Reinforcing Diffusion Models by Direct Group Preference Optimization

arXiv1 repo

arXiv:2510.08425

GRPO

ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping

arXiv1 repo

arXiv:2510.08457

Revisual-R1

ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation

arXiv1 repo

arXiv:2510.08551

ARTDECO

GraphGhost: Tracing Structures Behind Large Language Models

arXiv1 repo

arXiv:2510.08613

GraphGhost

Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight

arXiv1 repo

arXiv:2510.08713

UniWM_Dataset

When to Reason: Semantic Router for vLLM

arXiv1 repo

arXiv:2510.08731

semantic-router

Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation

arXiv1 repo

arXiv:2510.08807

humanoid-everyday

Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy

arXiv1 repo

arXiv:2510.09012

ARsample

The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections

arXiv1 repo

arXiv:2510.09023

Palisade

SynthID-Image: Image watermarking at internet scale

arXiv1 repo

arXiv:2510.09263

gemini-watermark-and-synthid-remover

Ctrl-World: A Controllable Generative World Model for Robot Manipulation

arXiv1 repo

arXiv:2510.10125

emboviz

WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

arXiv1 repo

arXiv:2510.10726

HunyuanWorld-Mirror

FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec

arXiv1 repo

arXiv:2510.10785

FAC-FACodec

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

arXiv1 repo

arXiv:2510.11098

VCB-Bench

ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding

arXiv1 repo

arXiv:2510.11498

ArtifactsBenchmark

PhySIC: Physically Plausible 3D Human-Scene Interaction and Contact from a Single Image

arXiv1 repo

arXiv:2510.11649

Phy-SIC

Diffusion Transformers with Representation Autoencoders

arXiv1 repo

arXiv:2510.11690

RAE

Scaling Language-Centric Omnimodal Representation Learning

arXiv1 repo

arXiv:2510.11693

LCO-Embedding

QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs

arXiv1 repo

arXiv:2510.11696

QeRL

Demystifying Reinforcement Learning in Agentic Reasoning

arXiv1 repo

arXiv:2510.11701

Open-AgentRL

Ax-Prover: A Deep Reasoning Agentic Framework for Theorem Proving in Mathematics and Quantum Physics

arXiv1 repo

arXiv:2510.12787

lean-lsp-mcp

Detect Anything via Next Point Prediction

arXiv1 repo

arXiv:2510.12798

LocateAnything-3B

RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge

arXiv1 repo

arXiv:2510.13590

post-graph-rag

The Art of Scaling Reinforcement Learning Compute for LLMs

arXiv1 repo

arXiv:2510.13786

OpenRLHF

Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs

arXiv1 repo

arXiv:2510.13795

Honey-Data-15M

LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimization

arXiv1 repo

arXiv:2510.13907

prompt-ops

xLLM Technical Report

arXiv1 repo

arXiv:2510.14686

xllm

LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training

arXiv1 repo

arXiv:2510.14969

UI-Simulator

From Pixels to Words -- Towards Native Vision-Language Primitives at Scale

arXiv1 repo

arXiv:2510.14979

NEO

RankSEG-RMA: An Efficient Segmentation Algorithm via Reciprocal Moment Approximation

arXiv1 repo

arXiv:2510.15362

rankseg

Attention Sinks in Diffusion Language Models

arXiv1 repo

arXiv:2510.15731

dlms-sinks

BLIP3o-NEXT: Next Frontier of Native Image Generation

arXiv1 repo

arXiv:2510.15857

BLIP3o

How Good Are LLMs at Processing Tool Outputs?

arXiv1 repo

arXiv:2510.15955

toolJSONprocessing

Demystifying Transition Matching: When and Why It Can Beat Flow Matching

arXiv1 repo

arXiv:2510.17991

TransitionFlowMatching

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

arXiv1 repo

arXiv:2510.18795

ProCLIP

Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring

arXiv1 repo

arXiv:2510.18817

FailureSensorIQ

Search Self-play: Pushing the Frontier of Agent Capability without Supervision

arXiv1 repo

arXiv:2510.18821

SSP

LightMem: Lightweight and Efficient Memory-Augmented Generation

arXiv1 repo

arXiv:2510.18866

LightMem

Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting

arXiv1 repo

arXiv:2510.18874

retaining-by-doing

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

arXiv1 repo

arXiv:2510.18876

Grasp-Any-Region

Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues

arXiv1 repo

arXiv:2510.19028

SCRIPTS

Addressing the Depth-of-Field Constraint: A New Paradigm for High Resolution Multi-Focus Image Fusion

arXiv1 repo

arXiv:2510.19581

VAEAEDOF

Learning Affordances at Inference-Time for Vision-Language-Action Models

arXiv1 repo

arXiv:2510.19752

roboeval

UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement

arXiv1 repo

arXiv:2510.20441

unified-audio

From Masks to Worlds: A Hitchhiker's Guide to World Models

arXiv1 repo

arXiv:2510.20668

Lumina-DiMOO

AlphaFlow: Understanding and Improving MeanFlow Models

arXiv1 repo

arXiv:2510.20771

alphaflow

HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives

arXiv1 repo

arXiv:2510.20822

HoloCine

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

arXiv1 repo

arXiv:2510.21311

Fines

The Principles of Diffusion Models

arXiv1 repo

arXiv:2510.21890

PyTorchTutorial

Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy

arXiv1 repo

arXiv:2510.22215

ViMDoc

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

arXiv1 repo

arXiv:2510.22603

Llama-AVSR

Correctness Forensics for Batch Speculative Decoding: Diagnosing the Ragged Tensor Problem

arXiv1 repo

arXiv:2510.22876

spec_dec

Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)

arXiv1 repo

arXiv:2510.22954

srt-hivemind

ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation

arXiv1 repo

arXiv:2510.23306

ReconViaGen

Dexbotic: Open-Source Vision-Language-Action Toolbox

arXiv1 repo

arXiv:2510.23511

dexbotic

JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence

arXiv1 repo

arXiv:2510.23538

ArtifactsBenchmark

MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection

arXiv1 repo

arXiv:2510.23727

MUStReason

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

arXiv1 repo

arXiv:2510.24563

OSWorld-MCP

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

arXiv1 repo

arXiv:2510.24693

STAR-Bench

Generative View Stitching

arXiv1 repo

arXiv:2510.24718

generative_view_stitching

RNAGenScape: Property-Guided, Optimized Generation of mRNA Sequences with Manifold Langevin Dynamics

arXiv1 repo

arXiv:2510.24736

figures4papers

Formalization of Auslander--Buchsbaum--Serre criterion in Lean4

arXiv1 repo

arXiv:2510.24818

FLT

DINO-YOLO: Self-Supervised Pre-training for Data-Efficient Object Detection in Civil Engineering Applications

arXiv1 repo

arXiv:2510.25140

DINOV3-YOLOV12

RT-DETRv4: Painlessly Furthering Real-Time Object Detection with Vision Foundation Models

arXiv1 repo

arXiv:2510.25257

RT-DETRv4

ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents

arXiv1 repo

arXiv:2510.25668

ALDEN

The Wiegold problem and free products of left-orderable groups

arXiv1 repo

arXiv:2510.26073

superhuman

UniTok-Audio: A Unified Audio Generation Framework via Generative Modeling on Discrete Codec Tokens

arXiv1 repo

arXiv:2510.26372

unified-audio

Kimi Linear: An Expressive, Efficient Attention Architecture

arXiv1 repo

arXiv:2510.26692

GatedDeltaNet-2

Category-Aware Semantic Caching for Heterogeneous LLM Workloads

arXiv1 repo

arXiv:2510.26835

semantic-router

The Denario project: Deep knowledge AI agents for scientific discovery

arXiv1 repo

arXiv:2510.26887

Denario

E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources

arXiv1 repo

arXiv:2510.27135

Nitro-E

Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMs

arXiv1 repo

arXiv:2510.27246

ogham-mcp

World Simulation with Video Foundation Models for Physical AI

arXiv1 repo

arXiv:2511.00062

Cosmos-Predict2.5-14B

CompAgent: An Agentic Framework for Visual Compliance Verification

arXiv1 repo

arXiv:2511.00171

compagent

ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation

arXiv1 repo

arXiv:2511.00511

OpenS2V-Nexus

Erasing 'Ugly' from the Internet: Propagation of the Beauty Myth in Text-Image Models

arXiv1 repo

arXiv:2511.00749

BeautyStandards

Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering

arXiv1 repo

arXiv:2511.01090

FineWeb2-Ro

Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play

arXiv1 repo

arXiv:2511.01261

speech_drame

UniREditBench: A Unified Reasoning-based Image Editing Benchmark

arXiv1 repo

arXiv:2511.01295

UniREdit-Data-100K

arXiv:2511.01833

arXiv1 repo

arXiv:2511.01833

evalscope

iFlyBot-VLA Technical Report

arXiv1 repo

arXiv:2511.01914

iFlyBot-VLA

TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models

arXiv1 repo

arXiv:2511.02802

TabTune

Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities

arXiv1 repo

arXiv:2511.02817

oolong

FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty Levels

arXiv1 repo

arXiv:2511.02872

open-atp

audio2chart: End to End Audio Transcription into playable Guitar Hero charts

arXiv1 repo

arXiv:2511.03337

audio2chart

NVIDIA Nemotron Nano V2 VL

arXiv1 repo

arXiv:2511.03929

Eagle

Submanifold Sparse Convolutional Networks for Automated 3D Segmentation of Kidneys and Kidney Tumours in Computed Tomography

arXiv1 repo

arXiv:2511.04334

ai_cancer_research

V-Thinker: Interactive Thinking with Images

arXiv1 repo

arXiv:2511.04460

V-Thinker

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

arXiv1 repo

arXiv:2511.04555

Evo-1

Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts

arXiv1 repo

arXiv:2511.04655

VSI-Bench

arXiv:2511.05171

arXiv1 repo

arXiv:2511.05171

naturelm-audio

TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models

arXiv1 repo

arXiv:2511.05275

TwinVLA

CGCE: Classifier-Guided Concept Erasure in Generative Models

arXiv1 repo

arXiv:2511.05865

CGCE

The Station: An Open-World Environment for AI-Driven Discovery

arXiv1 repo

arXiv:2511.06309

station

DIMO: Diverse 3D Motion Generation for Arbitrary Objects

arXiv1 repo

arXiv:2511.07409

DIMO

SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models

arXiv1 repo

arXiv:2511.08379

som-refusal-directions

Structured RAG for Answering Aggregative Questions

arXiv1 repo

arXiv:2511.08505

aggregative_questions

LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics

arXiv1 repo

arXiv:2511.08544

mlx-tune

TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models

arXiv1 repo

arXiv:2511.08667

TabPFN

RF-DETR: Neural Architecture Search for Real-Time Detection Transformers

arXiv1 repo

arXiv:2511.09554

rf-detr

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

arXiv1 repo

arXiv:2511.09690

fairseq2

SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control

arXiv1 repo

arXiv:2511.09715

SliderEdit

Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning

arXiv1 repo

arXiv:2511.10037

Beyond-React

ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs

arXiv1 repo

arXiv:2511.10240

ProgRAG

ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference

arXiv1 repo

arXiv:2511.10645

qwen38-27b-exl3

Depth Anything 3: Recovering the Visual Space from Any Views

arXiv1 repo

arXiv:2511.10647

DA3-BASE

Effective Brascamp-Lieb inequalities

arXiv1 repo

arXiv:2511.11091

numina-lean-agent

Moirai 2.0: When Less Is More for Time Series Forecasting

arXiv1 repo

arXiv:2511.11698

moirai-2.0-R-small

Uni-Hand: Universal Hand Motion Forecasting in Egocentric Views

arXiv1 repo

arXiv:2511.12878

UniHand

PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching

arXiv1 repo

arXiv:2511.12998

PerTouch

Distribution Matching Distillation Meets Reinforcement Learning

arXiv1 repo

arXiv:2511.13649

Z-Image-Turbo

Scaling Spatial Intelligence with Multimodal Foundation Models

arXiv1 repo

arXiv:2511.13719

SenseNova-SI-1.1-Qwen2.5-VL-7B

OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

arXiv1 repo

arXiv:2511.14582

OmniZip

UniGen-1.5: Enhancing Image Generation and Editing through Reward Unification in Reinforcement Learning

arXiv1 repo

arXiv:2511.14760

ml-unigen

IPR-1: Interactive Physical Reasoner

arXiv1 repo

arXiv:2511.15407

stable-retro

GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI

arXiv1 repo

arXiv:2511.15658

geo-bench

First Frame Is the Place to Go for Video Content Customization

arXiv1 repo

arXiv:2511.15700

FFGO-Video-Customization

Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn

arXiv1 repo

arXiv:2511.15738

Weave

Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning

arXiv1 repo

arXiv:2511.16043

Agent0

SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent

arXiv1 repo

arXiv:2511.16108

SkyRL

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

arXiv1 repo

arXiv:2511.16449

VLA-Pruner

Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs

arXiv1 repo

arXiv:2511.16664

NVIDIA-Nemotron-3-Nano-4B-GGUF

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

arXiv1 repo

arXiv:2511.16757

UTS

Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers

arXiv1 repo

arXiv:2511.17209

SPECTRE

METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model

arXiv1 repo

arXiv:2511.17366

RoboBrain_Dex

RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation

arXiv1 repo

arXiv:2511.17441

RoboCOIN

Planning with Sketch-Guided Verification for Physics-Aware Video Generation

arXiv1 repo

arXiv:2511.17450

SketchVerify

ROVER: Regulator-Driven Robust Temporal Verification of Black-Box Robot Policies

arXiv1 repo

arXiv:2511.17781

stable-retro

NeAR: Coupled Neural Asset-Renderer Stack

arXiv1 repo

arXiv:2511.18600

NeAR

arXiv:2511.18870

arXiv1 repo

arXiv:2511.18870

HunyuanVideo-1.5

Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models

arXiv1 repo

arXiv:2511.18890

Nemotron-Flash-1B

SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation

arXiv1 repo

arXiv:2511.19558

spqr

RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models

arXiv1 repo

arXiv:2511.19704

RADIO

Leveraging Foundation Models for Histological Grading in Cutaneous Squamous Cell Carcinoma using PathFMTools

arXiv1 repo

arXiv:2511.19751

PathFMTools

GigaWorld-0: World Models as Data Engine to Empower Embodied AI

arXiv1 repo

arXiv:2511.19861

Giga-World-1-Toydata

Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning

arXiv1 repo

arXiv:2511.19900

Agent0

Boosting Reasoning in Large Multimodal Models via Activation Replay

arXiv1 repo

arXiv:2511.19972

replay

WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving

arXiv1 repo

arXiv:2511.20022

WaymoQA

PRADA: Probability-Ratio-Based Attribution and Detection of Autoregressive-Generated Images

arXiv1 repo

arXiv:2511.20068

prada

NVIDIA Nemotron Parse 1.1

arXiv1 repo

arXiv:2511.20478

NVIDIA-Nemotron-Parse-v1.1

BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents

arXiv1 repo

arXiv:2511.20597

guaca

PixelDiT: Pixel Diffusion Transformers for Image Generation

arXiv1 repo

arXiv:2511.20645

PixelDiT-ImageNet

FaithFusion: Harmonizing Reconstruction and Generation via Pixel-wise Information Gain

arXiv1 repo

arXiv:2511.21113

FaithFusion

When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models

arXiv1 repo

arXiv:2511.21192

UPA-RFAS

Exploring Fusion Strategies for Multimodal Vision-Language Systems

arXiv1 repo

arXiv:2511.21889

Multimodal-Fusion-Strategies

Geometrically-Constrained Agent for Spatial Reasoning

arXiv1 repo

arXiv:2511.22659

gca

Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

arXiv1 repo

arXiv:2511.22677

Z-Image-Turbo

Test-time scaling of diffusions with flow maps

arXiv1 repo

arXiv:2511.22688

UniGenBench

Captain Safari: A World Engine with Pose-Aligned 3D Memory

arXiv1 repo

arXiv:2511.22815

Captain-Safari

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

arXiv1 repo

arXiv:2511.23269

OctoMed-7B

TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition

arXiv1 repo

arXiv:2512.01248

TRivia-3B

EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking Humans

arXiv1 repo

arXiv:2512.01340

LongCat-Video-Avatar

H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs

arXiv1 repo

arXiv:2512.01797

karma-electric-llama31-8b

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

arXiv1 repo

arXiv:2512.02014

tuna-2

CLEF: Clinically-Guided Contrastive Learning for Electrocardiogram Foundation Models

arXiv1 repo

arXiv:2512.02180

ecg-foundation-model

Guided Self-Evolving LLMs with Minimal Human Supervision

arXiv1 repo

arXiv:2512.02472

R-Zero

dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

arXiv1 repo

arXiv:2512.02498

dots.ocr

Spatially-Grounded Document Retrieval via Patch-to-Region Relevance Propagation

arXiv1 repo

arXiv:2512.02660

Snappy

A Hierarchical Tree-based approach for creating Configurable and Static Deep Research Agent (Static-DRA)

arXiv1 repo

arXiv:2512.03887

workers-personal-agent

OD-MoE: On-Demand Expert Loading for Cacheless Edge-Distributed MoE Inference

arXiv1 repo

arXiv:2512.03927

colibri

BioMedGPT-Mol: Multi-task Learning for Molecular Understanding and Generation

arXiv1 repo

arXiv:2512.04629

BioMedGPT-Mol

YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases

arXiv1 repo

arXiv:2512.04793

YingMusic-SVC

Intrinsically Interpretable Attention via Sparse Post-Training

arXiv1 repo

arXiv:2512.05865

CLT-Forge

LightSearcher: Efficient DeepSearch via Experiential Memory

arXiv1 repo

arXiv:2512.06653

MemoryOS

Group Representational Position Encoding

arXiv1 repo

arXiv:2512.07805

GRAPE

OmniPSD: Layered PSD Generation with Diffusion Transformer

arXiv1 repo

arXiv:2512.09247

OmniPSD

Hierarchy-Aware Multimodal Unlearning for Medical AI

arXiv1 repo

arXiv:2512.09867

MedForget

MotionEdit: Benchmarking and Learning Motion-Centric Image Editing

arXiv1 repo

arXiv:2512.10284

motionedit

Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning

arXiv1 repo

arXiv:2512.10691

RadVLM-GRPO

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction

arXiv1 repo

arXiv:2512.10888

granite-4.0-3b-vision

FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos

arXiv1 repo

arXiv:2512.10927

FoundationMotion

Position: Universal Aesthetic Alignment Narrows Artistic Expression

arXiv1 repo

arXiv:2512.11883

icml2026_position

MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models

arXiv1 repo

arXiv:2512.12121

MixtureKit

arXiv:2512.12218

arXiv1 repo

arXiv:2512.12218

journey-before-destination

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

arXiv1 repo

arXiv:2512.12772

JointAVBench

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

arXiv1 repo

arXiv:2512.12799

DrivePI

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

arXiv1 repo

arXiv:2512.13278

Open-AgentRL

RecTok: Reconstruction Distillation along Rectified Flow

arXiv1 repo

arXiv:2512.13421

RecTok

Adapting MLLMs for Nuanced Video Retrieval

arXiv1 repo

arXiv:2512.13511

TARA

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

arXiv1 repo

arXiv:2512.14008

LaViDa

GLM-TTS Technical Report

arXiv1 repo

arXiv:2512.14291

GLM-TTS

Native and Compact Structured Latents for 3D Generation

arXiv1 repo

arXiv:2512.14692

TRELLIS.2-4B

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

arXiv1 repo

arXiv:2512.15420

flowbind

Corrective Diffusion Language Models

arXiv1 repo

arXiv:2512.15596

ParallelBench

DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models

arXiv1 repo

arXiv:2512.15713

DiffusionVL

DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI

arXiv1 repo

arXiv:2512.16676

DataFlow

arXiv:2512.16899

arXiv1 repo

arXiv:2512.16899

EditReward

Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation

arXiv1 repo

arXiv:2512.16913

SphereForge

StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors

arXiv1 repo

arXiv:2512.16915

StereoPilot

Seed-Prover 1.5: Mastering Undergraduate-Level Theorem Proving via Learning from Experience

arXiv1 repo

arXiv:2512.17260

Seed-Prover

ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration

arXiv1 repo

arXiv:2512.17298

ProCache

Xiaomi MiMo-VL-Miloco Technical Report

arXiv1 repo

arXiv:2512.17436

xiaomi-mimo-vl-miloco

PathBench-MIL: A Comprehensive AutoML and Benchmarking Framework for Multiple Instance Learning in Histopathology

arXiv1 repo

arXiv:2512.17517

PathBench-MIL

The HydroGym Reinforcement Learning Platform for Fluid Dynamics

arXiv1 repo

arXiv:2512.17534

hydrogym

Dexterous World Models

arXiv1 repo

arXiv:2512.17907

dwm

arXiv:2512.17909

arXiv1 repo

arXiv:2512.17909

EditReward

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

arXiv1 repo

arXiv:2512.19433

Lumina-DiMOO

How Much 3D Do Video Foundation Models Encode?

arXiv1 repo

arXiv:2512.19949

VidFM3D

MolAct: An Agentic RL Framework for Molecular Editing and Property Optimization

arXiv1 repo

arXiv:2512.20135

SwanLab

QuarkAudio Technical Report

arXiv1 repo

arXiv:2512.20151

unified-audio

Step-DeepResearch Technical Report

arXiv1 repo

arXiv:2512.20491

StepDeepResearch

Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs

arXiv1 repo

arXiv:2512.20573

specdiff_aoi

Quantifying Laziness, Decoding Suboptimality, and Context Degradation in Large Language Models

arXiv1 repo

arXiv:2512.20662

unlazy

How important is Recall for Measuring Retrieval Quality?

arXiv1 repo

arXiv:2512.20854

retrieval-response

NVIDIA Nemotron 3: Efficient and Open Intelligence

arXiv1 repo

arXiv:2512.20856

NVIDIA-Nemotron-3-Nano-4B-GGUF

Streaming Video Instruction Tuning

arXiv1 repo

arXiv:2512.21334

Streamo

AstraNav-Memory: Contexts Compression for Long Memory

arXiv1 repo

arXiv:2512.21627

AstraNav-Memory

AstraNav-World: World Model for Foresight Control and Consistency

arXiv1 repo

arXiv:2512.21714

AstraNav-World

Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees

arXiv1 repo

arXiv:2512.21857

ADT-Tree

MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs

arXiv1 repo

arXiv:2512.22219

mirage

RollArt: Disaggregated Multi-Task Agentic RL Training at Scale

arXiv1 repo

arXiv:2512.22560

Mooncake

Visual Autoregressive Modelling for Monocular Depth Estimation

arXiv1 repo

arXiv:2512.22653

VAR-Depth

Reverse Personalization

arXiv1 repo

arXiv:2512.22984

reverse-personalization

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility

arXiv1 repo

arXiv:2512.23365

spatial_mosaic_vqa

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models

arXiv1 repo

arXiv:2512.23578

SLM-Style-Amnesia

Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models

arXiv1 repo

arXiv:2512.24618

Youtu-LLM-2B

mHC: Manifold-Constrained Hyper-Connections

arXiv1 repo

arXiv:2512.24880

maxtext

GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction

arXiv1 repo

arXiv:2512.25073

MVGenMaster

STELLAR: A Search-Based Testing Framework for Large Language Model Applications

arXiv1 repo

arXiv:2601.00497

STELLAR

RoboReward: General-Purpose Vision-Language Reward Models for Robotics

arXiv1 repo

arXiv:2601.00675

reward-scope

Early-Stage Prediction of Review Effort in AI-Generated Pull Requests

arXiv1 repo

arXiv:2601.00753

circuit-breaker

CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving

arXiv1 repo

arXiv:2601.01874

cogflow_code

360-GeoGS: Geometrically Consistent Feed-Forward 3D Gaussian Splatting Reconstruction for 360 Images

arXiv1 repo

arXiv:2601.02102

SphereForge

Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

arXiv1 repo

arXiv:2601.02356

talk2move

Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning

arXiv1 repo

arXiv:2601.02970

ReASC

AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

arXiv1 repo

arXiv:2601.03191

anatomix

IndexTTS 2.5 Technical Report

arXiv1 repo

arXiv:2601.03888

index-tts

PosterVerse: A Full-Workflow Framework for Commercial-Grade Poster Generation with HTML-Based Scalable Typography

arXiv1 repo

arXiv:2601.03993

PosterVerse

Re-Rankers as Relevance Judges

arXiv1 repo

arXiv:2601.04455

reranker-as-judge

SmartSearch: Process Reward-Guided Query Refinement for Search Agents

arXiv1 repo

arXiv:2601.04888

SmartSearch

Atlas 2 -- Foundation models for clinical deployment

arXiv1 repo

arXiv:2601.05148

PathoROB

RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation

arXiv1 repo

arXiv:2601.05241

RoboVIP_VDM

LaST$_{0}$: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

arXiv1 repo

arXiv:2601.05248

last0

Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

arXiv1 repo

arXiv:2601.05251

Mesh4D

Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs

arXiv1 repo

arXiv:2601.05635

SoE

Do Sparse Autoencoders Identify Reasoning Features in Language Models?

arXiv1 repo

arXiv:2601.05679

reasoning-probing

LayerGS: Decomposition and Inpainting of Layered 3D Human Avatars via 2D Gaussian Splatting

arXiv1 repo

arXiv:2601.05853

LayerGS

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning

arXiv1 repo

arXiv:2601.06803

laser

Categorize Early, Integrate Late: Divergent Processing Strategies in Automatic Speech Recognition

arXiv1 repo

arXiv:2601.06972

categorize-early-asr

Solar Open Technical Report

arXiv1 repo

arXiv:2601.07022

Solar-Open-100B

Dr. Zero: Self-Evolving Search Agents without Training Data

arXiv1 repo

arXiv:2601.07055

drzero

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

arXiv1 repo

arXiv:2601.07372

maxtext

MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head

arXiv1 repo

arXiv:2601.07832

MHLA

Training Free Zero-Shot Visual Anomaly Localization via Diffusion Inversion

arXiv1 repo

arXiv:2601.08022

DIVAD

Ministral 3

arXiv1 repo

arXiv:2601.08584

Ministral-3-8B-Instruct-2512-GGUF

TranslateGemma Technical Report

arXiv1 repo

arXiv:2601.09012

trans-gemma

A.X K1 Technical Report

arXiv1 repo

arXiv:2601.09200

A.X-K1

SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

arXiv1 repo

arXiv:2601.09385

SLAM-LLM

Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale

arXiv1 repo

arXiv:2601.10338

Palisade

Reasoning Models Generate Societies of Thought

arXiv1 repo

arXiv:2601.10825

h2aichat

FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning

arXiv1 repo

arXiv:2601.11141

FlashLabs-Chroma

ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models

arXiv1 repo

arXiv:2601.11404

AI539_NLP

UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation

arXiv1 repo

arXiv:2601.11522

UniX

ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents

arXiv1 repo

arXiv:2601.12294

ToolPRMBench

Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models

arXiv1 repo

arXiv:2601.12626

linear-mech-vlms

Think3D: Thinking with Space for Spatial Reasoning

arXiv1 repo

arXiv:2601.13029

spagent

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

arXiv1 repo

arXiv:2601.14235

Denario

RoboBrain 2.5: Depth in Sight, Time in Mind

arXiv1 repo

arXiv:2601.14352

RoboBrain2.5-4B

Next Generation Active Learning: Mixture of LLMs in the Loop

arXiv1 repo

arXiv:2601.15773

MoLLIA

Endless Terminals: Scaling RL Environments for Terminal Agents

arXiv1 repo

arXiv:2601.16443

SkyRL

OnlineSI: Taming Large Language Model for Online 3D Understanding and Grounding

arXiv1 repo

arXiv:2601.16538

online-spatial-intelligence

HapticMatch: An Exploration for Generative Material Haptic Simulation and Interaction

arXiv1 repo

arXiv:2601.16639

GelSight_Texture_Dataset

VidLaDA: Bidirectional Diffusion Large Language Models for Efficient Video Understanding

arXiv1 repo

arXiv:2601.17868

VidLaDA

DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints

arXiv1 repo

arXiv:2601.18137

deepclause-sdk

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction

arXiv1 repo

arXiv:2601.18537

minimind

Arithmetic volumes of moduli stacks of Shtukas

arXiv1 repo

arXiv:2601.18557

superhuman

A Hybrid Discriminative and Generative System for Universal Speech Enhancement

arXiv1 repo

arXiv:2601.19113

unified-audio

Native LLM and MLLM Inference at Scale on Apple Silicon

arXiv1 repo

arXiv:2601.19139

vllm-ios

Self-Distillation Enables Continual Learning

arXiv1 repo

arXiv:2601.19897

OpenClaw-RL

CPiRi: Channel Permutation-Invariant Relational Interaction for Multivariate Time Series Forecasting

arXiv1 repo

arXiv:2601.20318

CPiRi

AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

arXiv1 repo

arXiv:2601.20524

AnomalyVFM

DeepSeek-OCR 2: Visual Causal Flow

arXiv1 repo

arXiv:2601.20552

DeepSeek-OCR-2

Reinforcement Learning via Self-Distillation

arXiv1 repo

arXiv:2601.20802

OpenClaw-RL

asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation

arXiv1 repo

arXiv:2601.20992

asr_eval

AI-Assisted Engineering Should Track the Epistemic Status and Temporal Validity of Architectural Decisions

arXiv1 repo

arXiv:2601.21116

keep-the-why

MoCo: A One-Stop Shop for Model Collaboration Research

arXiv1 repo

arXiv:2601.21257

model_collaboration

Irrationality of rapidly converging series: a problem of Erdős and Graham

arXiv1 repo

arXiv:2601.21442

superhuman

OCRVerse: Towards Holistic OCR in End-to-End Vision-Language Models

arXiv1 repo

arXiv:2601.21639

OCRVerse

Mechanistic Data Attribution: Tracing the Training Origins of Interpretable LLM Units

arXiv1 repo

arXiv:2601.21996

Mechanistic-Data-Attribution

Where Do the Joules Go? Diagnosing Inference Energy Consumption

arXiv1 repo

arXiv:2601.22076

zeus

Semi-Autonomous Mathematics Discovery with Gemini: A Case Study on the Erdős Problems

arXiv1 repo

arXiv:2601.22401

superhuman

dgMARK: Decoding-Guided Watermarking for Diffusion Language Models

arXiv1 repo

arXiv:2601.22985

dgmark-watermarking

Strongly Polynomial Time Complexity of Policy Iteration for $L_\infty$ Robust MDPs

arXiv1 repo

arXiv:2601.23229

superhuman

Eigenweights for arithmetic Hirzebruch Proportionality

arXiv1 repo

arXiv:2601.23245

superhuman

Dispersion Loss Counteracts Embedding Condensation and Improves Generalization in Small Language Models

arXiv1 repo

arXiv:2602.00217

figures4papers

VoxServe: Streaming-Centric Serving System for Speech Language Models

arXiv1 repo

arXiv:2602.00269

vox-serve

From Observations to States: Latent Time Series Forecasting

arXiv1 repo

arXiv:2602.00297

LatentTSF

DecompressionLM: Deterministic, Diagnostic, and Zero-Shot Concept Graph Extraction from Language Models

arXiv1 repo

arXiv:2602.00377

decompressionlm

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

arXiv1 repo

arXiv:2602.00807

repro-any3d-vla-enhancing-vla-robustness-via-diverse-point-clouds

arXiv:2602.01382

arXiv1 repo

arXiv:2602.01382

EditReward

Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models

arXiv1 repo

arXiv:2602.01849

self-rewarding-smc

Show, Don't Tell: Morphing Latent Reasoning into Image Generation

arXiv1 repo

arXiv:2602.02227

LatentMorph

Kimi K2.5: Visual Agentic Intelligence

arXiv1 repo

arXiv:2602.02276

Kimi-K2.5

Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning

arXiv1 repo

arXiv:2602.02431

repro-full-batch-gd-outperforms-one-pass-sgd-sample-complexity-separation

Lower bounds for multivariate independence polynomials and their generalisations

arXiv1 repo

arXiv:2602.02450

superhuman

AgentRx: Diagnosing AI Agent Failures from Execution Trajectories

arXiv1 repo

arXiv:2602.02475

AgentRx

RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System

arXiv1 repo

arXiv:2602.02488

Open-AgentRL

PixelGen: Improving Pixel Diffusion with Perceptual Supervision

arXiv1 repo

arXiv:2602.02493

PixelGen-diffusers

Norm Anchors Make Model Edits Last

arXiv1 repo

arXiv:2602.02543

ComfyUI-LoRA-Optimizer

Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis

arXiv1 repo

arXiv:2602.03139

anima-turbo-4step

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection

arXiv1 repo

arXiv:2602.03216

repro-token-sparse-attention-long-context-interleaved-selection

SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

arXiv1 repo

arXiv:2602.03411

SWE-World

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

arXiv1 repo

arXiv:2602.03733

RegionReasoner

HY3D-Bench: Generation of 3D Assets

arXiv1 repo

arXiv:2602.03907

HY3D-Bench

arXiv:2602.04315

arXiv1 repo

arXiv:2602.04315

GeneralVLA

VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image

arXiv1 repo

arXiv:2602.04349

VecSet-Edit

EMA Policy Gradient: Taming Reinforcement Learning for LLMs with EMA Anchor and Top-k KL

arXiv1 repo

arXiv:2602.04417

ema-pg

Skin Tokens: A Learned Compact Representation for Unified Autoregressive Rigging

arXiv1 repo

arXiv:2602.04805

SkinTokens

PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation

arXiv1 repo

arXiv:2602.04876

PerpetualWonder

The Single-Multi Evolution Loop for Self-Improving Model Collaboration Systems

arXiv1 repo

arXiv:2602.05182

model_collaboration

Learning to Inject: Automated Prompt Injection via Reinforcement Learning

arXiv1 repo

arXiv:2602.05746

AutoInject

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

arXiv1 repo

arXiv:2602.06053

personaplex

Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making

arXiv1 repo

arXiv:2602.06570

Baichuan-M3-235B

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

arXiv1 repo

arXiv:2602.07026

Modality_Gap_Theory

DeepPrep: An LLM-Powered Agentic System for Autonomous Data Preparation

arXiv1 repo

arXiv:2602.07371

DeepAnalyze

When Is Enough Not Enough? Illusory Completion in Search Agents

arXiv1 repo

arXiv:2602.07549

illusory_completion

Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning

arXiv1 repo

arXiv:2602.07605

Finedefics_ICLR2025

Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research

arXiv1 repo

arXiv:2602.08387

modalities

Beyond Transcripts: A Renewed Perspective on Audio Chaptering

arXiv1 repo

arXiv:2602.08979

ytseg

Atlas: Enabling Cross-Vendor Authentication for IoT

arXiv1 repo

arXiv:2602.09263

awesome-connected-things-sec

UniVTAC: A Unified Simulation Platform for Visuo-Tactile Manipulation Data Generation, Learning, and Benchmarking

arXiv1 repo

arXiv:2602.10093

UniVTAC

A Vision-Language Foundation Model for Zero-shot Clinical Collaboration and Automated Concept Discovery in Dermatology

arXiv1 repo

arXiv:2602.10624

DermFM-Zero

Training and Benchmarking Code Generation for Physics-Inspired Animations

arXiv1 repo

arXiv:2602.10840

AgentFly

Embedding Inversion via Conditional Masked Diffusion Language Models

arXiv1 repo

arXiv:2602.11047

embedding-inversion-demo

It's TIME: Towards the Next Generation of Time Series Forecasting Benchmarks

arXiv1 repo

arXiv:2602.12147

TIME

dVoting: Fast Voting for dLLMs

arXiv1 repo

arXiv:2602.12153

dVoting

arXiv:2602.12322

arXiv1 repo

arXiv:2602.12322

foreact

Flow-Factory: A Unified Framework for Reinforcement Learning in Flow-Matching Models

arXiv1 repo

arXiv:2602.12529

GRPO

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

arXiv1 repo

arXiv:2602.12705

MedXpertQA

RAT-Bench: A Comprehensive Benchmark for Text Anonymization

arXiv1 repo

arXiv:2602.12806

rat-bench

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System

arXiv1 repo

arXiv:2602.13692

ThunderAgent

Towards Spatial Transcriptomics-driven Pathology Foundation Models

arXiv1 repo

arXiv:2602.14177

SEAL

Revisiting the Platonic Representation Hypothesis: An Aristotelian View

arXiv1 repo

arXiv:2602.14486

Aristotelian

Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions

arXiv1 repo

arXiv:2602.14878

tool-definition-quality-score

EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing

arXiv1 repo

arXiv:2602.15031

EditCtrl

ToaSt: Token Channel Selection and Structured Pruning for Efficient ViT

arXiv1 repo

arXiv:2602.15720

repro-toast-token-channel-selection-and-structured-pruning-for-efficient-vit

SAM 3D Body: Robust Full-Body Human Mesh Recovery

arXiv1 repo

arXiv:2602.15989

sam-3d-body

Fast KV Compaction via Attention Matching

arXiv1 repo

arXiv:2602.16284

attention-matching-rl

MerLean: An Agentic Framework for Autoformalization in Quantum Computation

arXiv1 repo

arXiv:2602.16554

lean-lsp-mcp

ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models

arXiv1 repo

arXiv:2602.16609

ColBERT-Zero

Towards a Science of AI Agent Reliability

arXiv1 repo

arXiv:2602.16666

harness-evals

DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning

arXiv1 repo

arXiv:2602.16742

DeepVision-103K

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

arXiv1 repo

arXiv:2602.16855

MobileAgent

Heterogeneous Federated Fine-Tuning with Parallel One-Rank Adaptation

arXiv1 repo

arXiv:2602.16936

Fed-PLoRA

Arcee Trinity Large Technical Report

arXiv1 repo

arXiv:2602.17004

donutloop-genesis

M2F: Automated Formalization of Mathematical Literature at Scale

arXiv1 repo

arXiv:2602.17016

lean-lsp-mcp

PartRAG: Retrieval-Augmented Part-Level 3D Generation and Editing

arXiv1 repo

arXiv:2602.17033

PartRAG

Leveraging Contrastive Learning for a Similarity-Guided Tampered Document Data Generation Pipeline

arXiv1 repo

arXiv:2602.17322

RealText-V2-Syn25k

arXiv:2602.17868

arXiv1 repo

arXiv:2602.17868

MantisV2Experiments

Improving Topic Modeling by Distilling Soft Labels from Language Models

arXiv1 repo

arXiv:2602.17907

MCompassRAG

GrandTour: A Legged Robotics Dataset in the Wild for Multi-Modal Perception and State Estimation

arXiv1 repo

arXiv:2602.18164

grand_tour_box

Improving Sampling for Masked Diffusion Models via Information Gain

arXiv1 repo

arXiv:2602.18176

Information-Gain-Sampler

CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation

arXiv1 repo

arXiv:2602.18424

CapNav

From Docs to Descriptions: Smell-Aware Evaluation of MCP Server Descriptions

arXiv1 repo

arXiv:2602.18914

tool-definition-quality-score

TimeRadar: A Domain-Rotatable Foundation Model for Time Series Anomaly Detection

arXiv1 repo

arXiv:2602.19068

TimeRadar

Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling

arXiv1 repo

arXiv:2602.19089

ani3dhuman

AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG

arXiv1 repo

arXiv:2602.19127

DataFlow

Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation

arXiv1 repo

arXiv:2602.19161

flash-vaed

SkillOrchestra: Learning to Route Agents via Skill Transfer

arXiv1 repo

arXiv:2602.19672

SkillOrchestra

Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation

arXiv1 repo

arXiv:2602.19778

ChordMiniApp

Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion

arXiv1 repo

arXiv:2602.20577

MVLAD-AD

On Data Engineering for Scaling LLM Terminal Capabilities

arXiv1 repo

arXiv:2602.21193

Nemotron-Terminal-Corpus

CARE: A Molecular-Guided Foundation Model with Adaptive Region Modeling for Whole Slide Image Analysis

arXiv1 repo

arXiv:2602.21637

CARE

Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion

arXiv1 repo

arXiv:2602.21646

smt-9b-hf

Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models

arXiv1 repo

arXiv:2602.22246

Diffusion_Self_Purification

veScale-FSDP: Flexible and High-Performance FSDP at Scale

arXiv1 repo

arXiv:2602.22437

veScale

Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language

arXiv1 repo

arXiv:2602.23201

Generalized-Neural-Memory

LK Losses: Direct Acceptance Rate Optimization for Speculative Decoding

arXiv1 repo

arXiv:2602.23881

SpecForge

Enhancing Spatial Understanding in Image Generation via Reward Modeling

arXiv1 repo

arXiv:2602.24233

UniGenBench

Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

arXiv1 repo

arXiv:2603.00431

Finedefics_ICLR2025

DreamWorld: Unified World Modeling in Video Generation

arXiv1 repo

arXiv:2603.00466

VideoREPA

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

arXiv1 repo

arXiv:2603.00667

HistoSelect

CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling

arXiv1 repo

arXiv:2603.01205

CoSMo3D

Constructive and Predicative Locale Theory in Univalent Foundations

arXiv1 repo

arXiv:2603.01308

TypeTopology

SimRecon: SimReady Compositional Scene Reconstruction from Real Videos

arXiv1 repo

arXiv:2603.02133

SimRecon

Adaptive Personalized Federated Learning via Multi-task Averaging of Kernel Mean Embeddings

arXiv1 repo

arXiv:2603.02233

repro-adaptive-personalized-fl-multi-task-averaging-kernel-mean-embedding

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

arXiv1 repo

arXiv:2603.02277

felonybench

Track4World: Feedforward World-centric Dense 3D Tracking of All Pixels

arXiv1 repo

arXiv:2603.02573

Track4World

Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement

arXiv1 repo

arXiv:2603.02641

RE-USE

SorryDB: Can AI Provers Complete Real-World Lean Theorems?

arXiv1 repo

arXiv:2603.02668

SorryDB

EvoSkill: Automated Skill Discovery for Multi-Agent Systems

arXiv1 repo

arXiv:2603.02766

EvoSkill

NeuroSkill(tm): Proactive Real-Time Agentic System Capable of Modeling Human State of Mind

arXiv1 repo

arXiv:2603.03212

neuroloop

PlaneCycle: Training-Free 2D-to-3D Lifting of Foundation Models Without Adapters

arXiv1 repo

arXiv:2603.04165

PlaneCycle

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

arXiv1 repo

arXiv:2603.04205

Real5-OmniDocBench

Toward Real-world Infrared Image Super-Resolution: A Unified Autoregressive Framework and Benchmark Dataset

arXiv1 repo

arXiv:2603.04745

FLIR-IISR

BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning

arXiv1 repo

arXiv:2603.04918

BandPO

LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution

arXiv1 repo

arXiv:2603.05947

LucidFlux

arXiv:2603.06007

arXiv1 repo

arXiv:2603.06007

MASFactory

CHMv2: Improvements in Global Canopy Height Mapping using DINOv3

arXiv1 repo

arXiv:2603.06382

dinov3

Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning

arXiv1 repo

arXiv:2603.06688

Narrative-Weaver

WaDi: Weight Direction-aware Distillation for One-step Image Synthesis

arXiv1 repo

arXiv:2603.08258

WaDi

How Far Can Unsupervised RLVR Scale LLM Training?

arXiv1 repo

arXiv:2603.08660

TTRL

SVG-EAR: Parameter-Free Linear Compensation for Sparse Video Generation via Error-aware Routing

arXiv1 repo

arXiv:2603.08982

Sparse-VideoGen

arXiv:2603.09877

arXiv1 repo

arXiv:2603.09877

GenEditEvalKit

OpenClaw-RL: Train Any Agent Simply by Talking

arXiv1 repo

arXiv:2603.10165

OpenClaw-RL

GLM-OCR Technical Report

arXiv1 repo

arXiv:2603.10910

GLM-OCR

Detect Anything in Real Time: From Single-Prompt Segmentation to Multi-Class Detection

arXiv1 repo

arXiv:2603.11441

DART

ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping

arXiv1 repo

arXiv:2603.11841

wespeaker

LMEB: Long-horizon Memory Embedding Benchmark

arXiv1 repo

arXiv:2603.12572

seahorse

When Drafts Evolve: Speculative Decoding Meets Online Learning

arXiv1 repo

arXiv:2603.12617

OnlineSPEC

DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs

arXiv1 repo

arXiv:2603.12996

ParallelBench

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

arXiv1 repo

arXiv:2603.13026

PISmith

Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach

arXiv1 repo

arXiv:2603.13056

CVPRW-26

Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents

arXiv1 repo

arXiv:2603.13151

Awesome-OpenClaw

Diffusion Reinforcement Learning via Centered Reward Distillation

arXiv1 repo

arXiv:2603.14128

GRPO

MDM-Prime-v2: Binary Encoding and Index Shuffling Enable Scaling of Diffusion Language Models

arXiv1 repo

arXiv:2603.16077

mdm-prime-v2

Omnilingual MT: Machine Translation for 1,600 Languages

arXiv1 repo

arXiv:2603.16309

pearmut

MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild

arXiv1 repo

arXiv:2603.17187

MetaClaw

Detecting the Machine: A Comprehensive Benchmark of AI-Generated Text Detectors Across Architectures, Domains, and Adversarial Conditions

arXiv1 repo

arXiv:2603.17522

sloptotal

CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents

arXiv1 repo

arXiv:2603.17829

SkyRL

Versatile Editing of Video Content, Actions, and Dynamics without Training

arXiv1 repo

arXiv:2603.17989

FlowEdit

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

arXiv1 repo

arXiv:2603.18178

Cosmos-Reason2-32B

NymeriaPlus: Enriching Nymeria Dataset with Additional Annotations and Data

arXiv1 repo

arXiv:2603.18496

nymeria_dataset

Scaling Sim-to-Real Reinforcement Learning for Robot VLAs with Generative 3D Worlds

arXiv1 repo

arXiv:2603.18532

EmbodiedGen

The Simplicity of the Hodge Bundle

arXiv1 repo

arXiv:2603.19052

superhuman

arXiv:2603.19637

arXiv1 repo

arXiv:2603.19637

UniBioTransfer

MOSS-TTSD: Text to Spoken Dialogue Generation

arXiv1 repo

arXiv:2603.19739

MOSS-TTS

Herding CATs: ALARA for Agent Harness Engineering in Portable Composable Multi-Agent Teams

arXiv1 repo

arXiv:2603.20380

npcpy

The production of meaning in the processing of natural language

arXiv1 repo

arXiv:2603.20381

npcpy

RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based Optimization

arXiv1 repo

arXiv:2603.20527

repro-rmnp-row-momentum-normalized-preconditioning-for-scalable-matrix-based-optimization

CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs

arXiv1 repo

arXiv:2603.21014

CLT-Forge

Mechanisms of Introspective Awareness

arXiv1 repo

arXiv:2603.21396

introspection-mechanisms

Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain

arXiv1 repo

arXiv:2603.21693

CEBaG

Efficient Universal Perception Encoder

arXiv1 repo

arXiv:2603.22387

EUPE

The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models

arXiv1 repo

arXiv:2603.22728

usad

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

arXiv1 repo

arXiv:2603.23067

MLLM-HWSI

Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment

arXiv1 repo

arXiv:2603.23114

contextual_moralchoice

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training

arXiv1 repo

arXiv:2603.23885

HunyuanOCR

MMTIT-Bench: A Multilingual and Multi-Scenario Benchmark with Cognition-Perception-Reasoning Guided Text-Image Machine Translation

arXiv1 repo

arXiv:2603.23896

HunyuanOCR

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

arXiv1 repo

arXiv:2603.25040

Intern-S1-Pro

S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation

arXiv1 repo

arXiv:2603.25702

S2D2

Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models

arXiv1 repo

arXiv:2603.25750

sommelier

Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

arXiv1 repo

arXiv:2603.25767

UTS

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

arXiv1 repo

arXiv:2603.26661

GaussianGPT

TAPS: Task Aware Proposal Distributions for Speculative Sampling

arXiv1 repo

arXiv:2603.27027

TAPS-Datasets

MOOZY: A Patient-First Foundation Model for Computational Pathology

arXiv1 repo

arXiv:2603.27048

MOOZY

Codebase-Memory: Tree-Sitter-Based Knowledge Graphs for LLM Code Exploration via MCP

arXiv1 repo

arXiv:2603.27277

codebase-memory-mcp

Falcon Perception

arXiv1 repo

arXiv:2603.27365

Falcon-OCR

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

arXiv1 repo

arXiv:2603.27460

Project-Imaging-X

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM

arXiv1 repo

arXiv:2603.27507

Chat-Scene

V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models

arXiv1 repo

arXiv:2603.27650

VidCom2

ITQ3_S: High-Fidelity 3-bit LLM Inference via Interleaved Ternary Quantization with Rotation-Domain Smoothing

arXiv1 repo

arXiv:2603.27914

turboquant-vllm

JaWildText: A Benchmark for Vision-Language Models on Japanese Scene Text Understanding

arXiv1 repo

arXiv:2603.27942

jawildtext

Meta-Harness: End-to-End Optimization of Model Harnesses

arXiv1 repo

arXiv:2603.28052

meta-harness

MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions

arXiv1 repo

arXiv:2603.28086

MOSS-TTS

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

arXiv1 repo

arXiv:2603.28458

TransArch

Rethinking Language Model Scaling under Transferable Hypersphere Optimization

arXiv1 repo

arXiv:2603.28743

ohara

Gen-Searcher: Reinforcing Agentic Search for Image Generation

arXiv1 repo

arXiv:2603.28767

KnowGen-Bench

MemFactory: Unified Inference & Training Framework for Agent Memory

arXiv1 repo

arXiv:2603.29493

SwanLab

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models

arXiv1 repo

arXiv:2604.00688

OmniVoice

CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery

arXiv1 repo

arXiv:2604.01658

time-series

T5Gemma-TTS Technical Report

arXiv1 repo

arXiv:2604.01760

T5Gemma-TTS

Lifting Unlabeled Internet-level Data for 3D Scene Understanding

arXiv1 repo

arXiv:2604.01907

SceneVersepp

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

arXiv1 repo

arXiv:2604.02546

UniScene3D

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

arXiv1 repo

arXiv:2604.02947

Internal-Safety-Collapse

StoryScope: Investigating idiosyncrasies in AI fiction

arXiv1 repo

arXiv:2604.03136

storyscope

VOSR: A Vision-Only Generative Model for Image Super-Resolution

arXiv1 repo

arXiv:2604.03225

OSEDiff

Version Control System for Data with MatrixOne

arXiv1 repo

arXiv:2604.03927

matrixone

AURA: Always-On Understanding and Real-Time Assistance via Video Streams

arXiv1 repo

arXiv:2604.04184

AURA

3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image

arXiv1 repo

arXiv:2604.04406

EmbodiedGen

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

arXiv1 repo

arXiv:2604.04771

MinerU

Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency

arXiv1 repo

arXiv:2604.04847

Full-Duplex-Bench

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression

arXiv1 repo

arXiv:2604.04921

triattention-main

Early Stopping for Large Reasoning Models via Confidence Dynamics

arXiv1 repo

arXiv:2604.04930

CoDE-Stop

StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing

arXiv1 repo

arXiv:2604.05014

starVLA

PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing

arXiv1 repo

arXiv:2604.05018

academic-research-skills

This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA

arXiv1 repo

arXiv:2604.05051

LLMHealthFramingEffect

TRACE: Capability-Targeted Agentic Training

arXiv1 repo

arXiv:2604.05336

TRACE

WikiSeeker: Rethinking the Role of Vision-Language Models in Knowledge-Based Visual Question Answering

arXiv1 repo

arXiv:2604.05818

WikiSeeker

CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training

arXiv1 repo

arXiv:2604.05821

CLEAR

From Debate to Decision: Conformal Social Choice for Safe Multi-Agent Deliberation

arXiv1 repo

arXiv:2604.07667

rightmind

SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

arXiv1 repo

arXiv:2604.08377

SkillClaw

PIArena: A Platform for Prompt Injection Evaluation

arXiv1 repo

arXiv:2604.08499

PIArena

RewardFlow: Generate Images by Optimizing What You Reward

arXiv1 repo

arXiv:2604.08536

RewardFlow

Enhancing LLM Problem Solving via Tutor-Student Multi-Agent Interaction

arXiv1 repo

arXiv:2604.08931

rightmind

Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence

arXiv1 repo

arXiv:2604.09057

Tora

DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio

arXiv1 repo

arXiv:2604.09344

Sidon

Conflicts Make Large Reasoning Models Vulnerable to Attacks

arXiv1 repo

arXiv:2604.09750

ConflictHarm

Pioneer Agent: Continual Improvement of Small Language Models in Production

arXiv1 repo

arXiv:2604.09791

gliner2-privacy-filter-PII-multi

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

arXiv1 repo

arXiv:2604.09860

RoboLab

When More Thinking Hurts: Overthinking in LLM Test-Time Compute Scaling

arXiv1 repo

arXiv:2604.10739

unlazy

TInR: Exploring Tool-Internalized Reasoning in Large Language Models

arXiv1 repo

arXiv:2604.10788

TInR

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

arXiv1 repo

arXiv:2604.12012

tips

Nucleus-Image: Sparse MoE for Image Generation

arXiv1 repo

arXiv:2604.12163

Nucleus-Image

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

arXiv1 repo

arXiv:2604.14268

HY-World-2.0

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

arXiv1 repo

arXiv:2604.14548

VoxSafeBench

Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models

arXiv1 repo

arXiv:2604.14591

MLN

ProtoTTA: Prototype-Guided Test-Time Adaptation

arXiv1 repo

arXiv:2604.15494

ProtoTTA

DIRT: Database-Integrated Random Testing

arXiv1 repo

arXiv:2604.16373

turso

Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy

arXiv1 repo

arXiv:2604.16923

LAPD

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

arXiv1 repo

arXiv:2604.17243

RemoteShield

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling

arXiv1 repo

arXiv:2604.18556

Qwen3.8-27B-GSQ-RCO-GGUF

Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs

arXiv1 repo

arXiv:2604.18576

time-series

Debating the Unspoken: Role-Anchored Multi-Agent Reasoning for Half-Truth Detection

arXiv1 repo

arXiv:2604.19005

rightmind

arXiv:2604.19417

arXiv1 repo

arXiv:2604.19417

MERTools

GRAFT: Geometric Refinement and Fitting Transformer for Human Scene Reconstruction

arXiv1 repo

arXiv:2604.19624

graft

Semantic-Fast-SAM: Efficient Semantic Segmenter

arXiv1 repo

arXiv:2604.20169

Semantic-Fast-SAM

Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization

arXiv1 repo

arXiv:2604.21072

BloomBee

DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

arXiv1 repo

arXiv:2604.21507

diarizen-tutorial

StructMem: Structured Memory for Long-Horizon Behavior in LLMs

arXiv1 repo

arXiv:2604.21748

LightMem

Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents

arXiv1 repo

arXiv:2604.22085

memanto

FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting

arXiv1 repo

arXiv:2604.22328

Energy_Benchmark_TSFM_pub

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

arXiv1 repo

arXiv:2604.23321

MMEB-V3

Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior in Generative Retrieval

arXiv1 repo

arXiv:2604.23396

lost-in-decoding

IndustryAssetEQA: A Neurosymbolic Operational Intelligence System for Embodied Question Answering in Industrial Asset Maintenance

arXiv1 repo

arXiv:2604.23446

AssetOpsBench

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis

arXiv1 repo

arXiv:2604.24198

DataMind

arXiv:2604.24622

arXiv1 repo

arXiv:2604.24622

CF-VLA

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

arXiv1 repo

arXiv:2604.24763

tuna-2

MAIC-UI: Making Interactive Courseware with Generative UI

arXiv1 repo

arXiv:2604.25806

MAIC-UI

DeepTutor: Towards Agentic Personalized Tutoring

arXiv1 repo

arXiv:2604.26962

DeepTutor

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

arXiv1 repo

arXiv:2604.28130

MoCapAnythingV2-weights

Model Compression with Exact Budget Constraints via Riemannian Manifolds

arXiv1 repo

arXiv:2605.00649

Qwen3.8-27B-GSQ-RCO-GGUF

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

arXiv1 repo

arXiv:2605.00658

UniVidX

Co-Generative De Novo Functional Protein Design

arXiv1 repo

arXiv:2605.00948

OpenBioMed

Shao: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation

arXiv1 repo

arXiv:2605.01790

Shao

Stochastic Sparse Attention for Memory-Bound Inference

arXiv1 repo

arXiv:2605.01910

repro-stochastic-sparse-attention-for-memory-bound-inference

Break the Block: Dynamic-size Reasoning Blocks for Diffusion Large Language Models via Monotonic Entropy Descent with Reinforcement Learning

arXiv1 repo

arXiv:2605.02263

Block-R1

Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

arXiv1 repo

arXiv:2605.02757

Seeing-Realism-from-Simulation

Compositional Neural-Cyber-Physical System Verification in the Interactive Theorem Prover of Your Choice

arXiv1 repo

arXiv:2605.02790

analysis

Contrastive Privacy: A Semantic Approach to Measuring Privacy of AI-based Sanitization

arXiv1 repo

arXiv:2605.02977

contrastive-privacy

Ortho-Hydra: Orthogonalized Experts for DiT LoRA

arXiv1 repo

arXiv:2605.03252

anima_lora

Segmenting Human-LLM Co-authored Text via Change Point Detection

arXiv1 repo

arXiv:2605.03723

DetectLLMSegmentation

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

arXiv1 repo

arXiv:2605.03937

minimind-o

arXiv:2605.05115

arXiv1 repo

arXiv:2605.05115

drowse

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

arXiv1 repo

arXiv:2605.05185

OpenSearch-VL

MidSteer: Optimal Affine Framework for Steering Generative Models

arXiv1 repo

arXiv:2605.05220

MidSteer

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

arXiv1 repo

arXiv:2605.05611

X-Voice

Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions

arXiv1 repo

arXiv:2605.06058

CoExVQA

Continuous-Time Distribution Matching for Few-Step Diffusion Distillation

arXiv1 repo

arXiv:2605.06376

anima-turbo-4step

Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models

arXiv1 repo

arXiv:2605.06388

semantic-wm

SparseForge: Efficient Semi-Structured LLM Sparsification via Annealing of Hessian-Guided Soft-Mask

arXiv1 repo

arXiv:2605.06402

SparseForge

Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models

arXiv1 repo

arXiv:2605.06510

tfmlens

MIND: Monge Inception Distance for Generative Models Evaluation

arXiv1 repo

arXiv:2605.06797

RAEv2

Narrow Secret Loyalty Dodges Black-Box Audits

arXiv1 repo

arXiv:2605.06846

whitebox-affordance-ladder

Don't Retrain, Align: Adapting Autoregressive LMs to Diffusion LMs via Representation Alignment

arXiv1 repo

arXiv:2605.06885

Open-dLLM

MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference

arXiv1 repo

arXiv:2605.07363

TransArch

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

arXiv1 repo

arXiv:2605.07723

academic-research-skills

Benchmarking Foundation Models for Renal Lesion Stratification in CT

arXiv1 repo

arXiv:2605.07749

RenalVision

Anisotropic Modality Align

arXiv1 repo

arXiv:2605.07825

Modality_Gap_Theory

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

arXiv1 repo

arXiv:2605.08029

ml-starflow

Normalizing Trajectory Models

arXiv1 repo

arXiv:2605.08078

ml-starflow

SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

arXiv1 repo

arXiv:2605.08366

big-pickle-swe-atlas

arXiv:2605.08703

arXiv1 repo

arXiv:2605.08703

EditReward

Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal

arXiv1 repo

arXiv:2605.09203

watermarks-remover

Attention Drift: What Autoregressive Speculative Decoding Models Learn

arXiv1 repo

arXiv:2605.09992

Attention-Drift

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

arXiv1 repo

arXiv:2605.09998

prime-agent

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation

arXiv1 repo

arXiv:2605.10912

WildClawBench

ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

arXiv1 repo

arXiv:2605.11086

exploitgym

Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models

arXiv1 repo

arXiv:2605.11726

Block-R1

Chronicles-OCR: A Cross-Temporal Perception Benchmark for the Evolutionary Trajectory of Chinese Characters

arXiv1 repo

arXiv:2605.11960

HunyuanOCR

L2P: Unlocking Latent Potential for Pixel Generation

arXiv1 repo

arXiv:2605.12013

L2P

Task-Adaptive Embedding Refinement via Test-time LLM Guidance

arXiv1 repo

arXiv:2605.12487

task-aware-embedding-refinement

AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward

arXiv1 repo

arXiv:2605.12495

AlphaGRPO

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

arXiv1 repo

arXiv:2605.12500

NEO

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

arXiv1 repo

arXiv:2605.12501

Phi-Ground-Any

Certified Robustness under Heterogeneous Perturbations via Hybrid Randomized Smoothing

arXiv1 repo

arXiv:2605.12876

icml26-hybrid-randomized-smoothing

Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics

arXiv1 repo

arXiv:2605.13171

formal-conjectures

LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics

arXiv1 repo

arXiv:2605.13412

RAB-Cred

arXiv:2605.13782

arXiv1 repo

arXiv:2605.13782

LMPath

TabPFN-3: Technical Report

arXiv1 repo

arXiv:2605.13986

TabPFN

SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasks

arXiv1 repo

arXiv:2605.14051

AssetOpsBench

Turning Stale Gradients into Stable Gradients: Coherent Coordinate Descent with Implicit Landscape Smoothing for Lightweight Zeroth-Order Optimization

arXiv1 repo

arXiv:2605.14373

repro-turning-stale-gradients-into-stable-gradients-coherent-coordinate-descent

Nexus : An Agentic Framework for Time Series Forecasting

arXiv1 repo

arXiv:2605.14389

time-series

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making

arXiv1 repo

arXiv:2605.14403

DermAgent

GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding

arXiv1 repo

arXiv:2605.15250

TransArch

Tweedie's Formula and Score-Driven Updating

arXiv1 repo

arXiv:2605.15902

skaters

Generative 3D Gaussians with Learned Density Control

arXiv1 repo

arXiv:2605.16355

TripoSplat

TriALS: Triphasic-Aided Liver Lesion Segmentation Benchmark in Non-Contrast CT

arXiv1 repo

arXiv:2605.16572

TriALS

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models

arXiv1 repo

arXiv:2605.16842

Lumina-DiMOO

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers

arXiv1 repo

arXiv:2605.16941

WINO-DLLM

AgentModernize: Preserving Business Logic in Legacy Modernization with Multi-Agent LLMs and Behavioral Specification Graphs

arXiv1 repo

arXiv:2605.17535

rebuild-dossier

Stable Audio 3

arXiv1 repo

arXiv:2605.17991

stable-audio-3

Improved Baselines with Representation Autoencoders

arXiv1 repo

arXiv:2605.18324

RAEv2

SAME: A Semantically-Aligned Music Autoencoder

arXiv1 repo

arXiv:2605.18613

same-l-decoder-lora

Pocket Foundation Models: Distilling TFMs into CPU-Ready Gradient-Boosted Trees

arXiv1 repo

arXiv:2605.18654

sqlite-predict

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting

arXiv1 repo

arXiv:2605.18810

SpecForge

arXiv:2605.19130

arXiv1 repo

arXiv:2605.19130

egobabyvlm

PhyWorld: Physics-Faithful World Model for Video Generation

arXiv1 repo

arXiv:2605.19242

PhyWorld

Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVR

arXiv1 repo

arXiv:2605.19282

Pion

Understanding Wacky Weights: A Dissection of SPLADE's Learned Term Importance

arXiv1 repo

arXiv:2605.19628

understanding-wacky-weights

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

arXiv1 repo

arXiv:2605.20035

SEATS

Toto 2.0: Time Series Forecasting Enters the Scaling Era

arXiv1 repo

arXiv:2605.20119

toto

Tiny-Engram: Trigger-Indexed Concept Tables for Generative Vision

arXiv1 repo

arXiv:2605.20309

TinyEngram

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

arXiv1 repo

arXiv:2605.20630

AssetOpsBench

Latent-space Attacks for Refusal Evasion in Language Models

arXiv1 repo

arXiv:2605.21706

latent-evasion

Swift Sampling: Selecting Temporal Surprises via Taylor Series

arXiv1 repo

arXiv:2605.22678

SwiftSampling

Advancing Mathematics Research with AI-Driven Formal Proof Search

arXiv1 repo

arXiv:2605.22763

formal-conjectures

Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

arXiv1 repo

arXiv:2605.22791

GatedDeltaNet-2

Cambrian-P: Pose-Grounded Video Understanding

arXiv1 repo

arXiv:2605.22819

cambrian-p

FastKernels: Benchmarking GPU Kernel Generation in Production

arXiv1 repo

arXiv:2605.23215

fastkernels

From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills

arXiv1 repo

arXiv:2605.23899

darwin-skill

In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models

arXiv1 repo

arXiv:2605.23908

picbreeder-vlm-archive

Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content

arXiv1 repo

arXiv:2605.24421

gaslit-aisoc

ECHO: Terminal Agents Learn World Models for Free

arXiv1 repo

arXiv:2605.24517

SkyRL

Towards Multi-Turn Dialog Systems for Industrial Asset Operations and Maintenance

arXiv1 repo

arXiv:2605.24953

AssetOpsBench

Leveraging Gauge Freedom for Learning Non-Gradient Population Dynamics of Stochastic Systems

arXiv1 repo

arXiv:2605.25107

repro-leveraging-gauge-freedom-for-non-gradient-population-dynamics

Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning

arXiv1 repo

arXiv:2605.26292

BiomedCoOp

Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations

arXiv1 repo

arXiv:2605.26874

AssetOpsBench

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

arXiv1 repo

arXiv:2605.27062

sopro

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

arXiv1 repo

arXiv:2605.27881

SwanLab

ABot-OCR Technical Report

arXiv1 repo

arXiv:2605.27978

ocr

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

arXiv1 repo

arXiv:2605.28091

Qwen-Image-Flash

Rethinking Memory as Continuously Evolving Connectivity

arXiv1 repo

arXiv:2605.28773

LightMem

From Pixels to Words -- Towards Native One-Vision Models at Scale

arXiv1 repo

arXiv:2605.28820

NEO

Draft-OPD: On-Policy Distillation for Speculative Draft Models

arXiv1 repo

arXiv:2605.29343

Draft-OPD

Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting

arXiv1 repo

arXiv:2605.29401

time-series

DiffSpot: Can VLMs Spot Fine-Grained Visual Differences in Web Interfaces?

arXiv1 repo

arXiv:2605.29615

layoutlens

VikingMem: A Memory Base Management System for Stateful LLM-based Applications

arXiv1 repo

arXiv:2605.29640

OpenViking

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

arXiv1 repo

arXiv:2605.29659

opir-multilang-onnx

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

arXiv1 repo

arXiv:2605.29707

SpecForge

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis

arXiv1 repo

arXiv:2605.30434

DataMind

APE: Agentic Prompt Enhancer for Image Generation and Editing

arXiv1 repo

arXiv:2606.00204

UniGenBench

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

arXiv1 repo

arXiv:2606.00206

2x-3090-GA102-300-A1-sglang-inference

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

arXiv1 repo

arXiv:2606.01348

HunyuanOCR

UniD$^3$: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning

arXiv1 repo

arXiv:2606.01394

UniD3

Bridging the Last Mile of Time Series Forecasting with LLM Agents

arXiv1 repo

arXiv:2606.02497

time-series

Cosmos 3: Omnimodal World Models for Physical AI

arXiv1 repo

arXiv:2606.02800

UniGenBench

From Long News to Accurate Forecast: Importance-Aware Fusion and PRM-Guided Reflection for Time Series Forecasting

arXiv1 repo

arXiv:2606.03097

time-series

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

arXiv1 repo

arXiv:2606.03303

superhuman

KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks

arXiv1 repo

arXiv:2606.03458

KVarN

AlignAtt4LLM: Fast AlignAtt for Decoder-Only LLMs at IWSLT 2026 Simultaneous Speech Translation Task

arXiv1 repo

arXiv:2606.03967

WhisperLiveKit

Self-Distilled Policy Gradient

arXiv1 repo

arXiv:2606.04036

halo

Who Needs Labels? Adapting Vision Foundation Models With the Metadata You Already Have

arXiv1 repo

arXiv:2606.05107

dinov3

GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

arXiv1 repo

arXiv:2606.05160

PhysicalAI-Robotics-Locomanipulation-GRAIL

AsyncWebRL: Efficient Asynchronous Reinforcement Learning for Multi-Step Visual Web Agents

arXiv1 repo

arXiv:2606.05597

webgym

T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation

arXiv1 repo

arXiv:2606.05975

T-FunS3D

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios

arXiv1 repo

arXiv:2606.06177

ouvia

RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

arXiv1 repo

arXiv:2606.06256

RedKnot

Unsupervised Skill Discovery for Agentic Data Analysis

arXiv1 repo

arXiv:2606.06416

DataMind

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

arXiv1 repo

arXiv:2606.06444

usad

VoxCPM2 Technical Report

arXiv1 repo

arXiv:2606.06928

VoxCPM

arXiv:2606.09516

arXiv1 repo

arXiv:2606.09516

SwiftVR

End-to-End Context Compression at Scale

arXiv1 repo

arXiv:2606.09659

LCLM

Collaborative Human-Agent Protocol (CHAP)

arXiv1 repo

arXiv:2606.09751

chap

Rethinking the Divergence Regularization in LLM RL

arXiv1 repo

arXiv:2606.09821

UniRL

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

arXiv1 repo

arXiv:2606.10968

UniRL

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models

arXiv1 repo

arXiv:2606.11025

UniRL

M*: A Modular, Extensible, Serving System for Multimodal Models

arXiv1 repo

arXiv:2606.12688

mstar

World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible

arXiv1 repo

arXiv:2606.13652

world-tracing-demo

FastContext: Training Efficient Repository Explorer for Coding Agents

arXiv1 repo

arXiv:2606.14066

caveman

Classifying by Proxy: Explainable and Reproducible Ensemble of Proxy Tasks for Child Sexual Abuse Imagery Classification

arXiv1 repo

arXiv:2606.15993

EISPCSAI

Scaling Human and G2P Supervision for Robust Phonetic Transcription

arXiv1 repo

arXiv:2606.16019

ML

Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents

arXiv1 repo

arXiv:2606.16038

Open-SWE-Traces

arXiv:2606.17404

arXiv1 repo

arXiv:2606.17404

ELSA

MagicSim: A Unified Infrastructure for Executable Embodied Interaction

arXiv1 repo

arXiv:2606.17511

robodojo_long

Vision-language models for chest radiography do not always need the image

arXiv1 repo

arXiv:2606.17710

causal

HumanoidArena: Benchmarking Egocentric Hierarchical Whole-body Learning

arXiv1 repo

arXiv:2606.17833

HumanoidArena

Uncertainty Quantification for Flow-Based Vision-Language-Action Models

arXiv1 repo

arXiv:2606.18043

uq_vla

REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection

arXiv1 repo

arXiv:2606.19881

REDACT-PII-Benchmark

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

arXiv1 repo

arXiv:2606.20177

NeFo

TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation

arXiv1 repo

arXiv:2606.20709

TeleStyleV2

Tmax: A simple recipe for terminal agents

arXiv1 repo

arXiv:2606.23321

tmax

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

arXiv1 repo

arXiv:2606.23344

PP-DocLayoutV3_safetensors

DiffusionBench: On Holistic Evaluation of Diffusion Transformers

arXiv1 repo

arXiv:2606.24888

diffusion-bench

AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents

arXiv1 repo

arXiv:2606.24893

agentodyssey

BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents

arXiv1 repo

arXiv:2606.25556

SwanLab

Towards an Interactive Evidence-RAG Peer-Review Workspace for the Journal of Digital History

arXiv1 repo

arXiv:2606.25837

journal-of-digital-history-evidence-rag

Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs

arXiv1 repo

arXiv:2606.26387

figures4papers

NaviCache: Test-Time Self-Calibration Caching for Video Generation

arXiv1 repo

arXiv:2606.26795

HunyuanVideo

Scaling Multi-Reference Image Generation with Dynamic Reward Optimization

arXiv1 repo

arXiv:2606.26947

DyRef

A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts

arXiv1 repo

arXiv:2606.27881

temporal-ner

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

arXiv1 repo

arXiv:2606.28128

PhysisForcing

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

arXiv1 repo

arXiv:2606.28276

SimFoundry

When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning

arXiv1 repo

arXiv:2606.29354

LSF_MDia

StrucTab: A Structured Optimization Framework for Table Parsing

arXiv1 repo

arXiv:2606.29905

HunyuanOCR

Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks

arXiv1 repo

arXiv:2606.31074

triospect

DA-Studio: An Agentic System for End-to-End Data Analysis

arXiv1 repo

arXiv:2606.31423

DeepAnalyze

Large Databases Need Small, Open-Weight Language Models

arXiv1 repo

arXiv:2606.31808

blendsql

DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation

arXiv1 repo

arXiv:2606.32028

DVG-WM

EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation

arXiv1 repo

arXiv:2607.01147

EquiSteer

TiRex-2: Generalizing TiRex to Multivariate Data and Streaming

arXiv1 repo

arXiv:2607.01204

tirex

AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models

arXiv1 repo

arXiv:2607.02269

AnyGroundBench

GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training

arXiv1 repo

arXiv:2607.02486

Geomix

arXiv:2607.03502

arXiv1 repo

arXiv:2607.03502

Filler_Token

arXiv:2607.03788

arXiv1 repo

arXiv:2607.03788

tensor-train

ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog

arXiv1 repo

arXiv:2607.04438

ppt-master

Unified Audio Intelligence Without Regressing on Text Intelligence

arXiv1 repo

arXiv:2607.05196

Nemotron-Labs-Audex-30B-A3B

Vision Pretraining for Dense Spatial Perception

arXiv1 repo

arXiv:2607.05247

lingbot-vision

LLM-as-a-Verifier: A General-Purpose Verification Framework

arXiv1 repo

arXiv:2607.05391

dsh-plugin-llm-verifier

FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games

arXiv1 repo

arXiv:2607.06514

stable-retro

EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

arXiv1 repo

arXiv:2607.07459

EmbodiedGen

UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing

arXiv1 repo

arXiv:2607.08646

UltraX-Preview

StereoSplat+: Feed-Forward Stereo Gaussian Splatting with Diffusion-Assisted Progressive Inference

arXiv1 repo

arXiv:2607.08808

StereoSplat_Plus

Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI

arXiv1 repo

arXiv:2607.09324

crowdsourced-data-tools

Index SLM Technical Report

arXiv1 repo

arXiv:2607.09885

Index-1.9B

GigaAM Multilingual: Foundation Model for Underrepresented Languages

arXiv1 repo

arXiv:2607.10371

GigaAM

SETA: Scaling Environments for Terminal Agents

arXiv1 repo

arXiv:2607.10891

seta

Higher-Order Cell Tracking Transformer

arXiv1 repo

arXiv:2607.11754

rlx-models

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

arXiv1 repo

arXiv:2607.11886

AlphaGRPO

Let RGB Be the Language of Vision

arXiv1 repo

arXiv:2607.12450

RINO

Self-Improvements in Modern Agentic Systems: A Survey

arXiv1 repo

arXiv:2607.13104

academic-research-skills

UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets

arXiv1 repo

arXiv:2607.13586

UniPhys-Bench

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

arXiv1 repo

arXiv:2607.14846

asr-benchmark-optimization

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

arXiv1 repo

arXiv:2607.18218

prov-gigapath

Patch Policy: Efficient Embodied Control via Dense Visual Representations

arXiv1 repo

arXiv:2607.18236

patch_policy

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

arXiv1 repo

arXiv:2607.19191

ABot-World-Explorer-500h

PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning

arXiv1 repo

arXiv:2607.20064

PRO-LONG

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

arXiv1 repo

arXiv:2607.20709

labs-OO-Agents

arXiv:2607.21557

arXiv1 repo

arXiv:2607.21557

OpenForge-RL

RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention

arXiv1 repo

arXiv:2607.21927

riskernel

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model

arXiv1 repo

arXiv:2607.22083

Nanbeige4.2-3B

Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents

arXiv1 repo

arXiv:2607.22563

AssetOpsBench

Modeling Local Exploit Hazard - A Bayesian Framework for Quantifying Exploit Risk and Operational Efficiency

arXiv1 repo

arXiv:2607.24618

researches

Specula: Scaling formal specifications for autonomous model checking of system code

arXiv1 repo

arXiv:2607.25333

Specula

Coevolution of epidemic dynamics and network topology driven by disease fatality and waning immunity

arXiv1 repo

arXiv:2607.25475

sci-ssci-skills

A Distributional Robustness Margin For Pathology Foundation Models

arXiv1 repo

arXiv:2607.25497

croma

Revisiting the Algebraic Foundation of Relational Data

arXiv1 repo

arXiv:2607.26356

prela

Contrastive ESA: Human Evaluation of Multiple Translations at Once

arXiv1 repo

arXiv:2607.26640

pearmut

StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

arXiv1 repo

arXiv:2607.26754

stable-retro

Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences

arXiv1 repo

arXiv:2607.26973

sat-bundleadjust

VETO: Towards Protecting Images From Frontier AI Editing

arXiv1 repo

arXiv:2607.27292

VetoBench

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

arXiv1 repo

arXiv:2607.28625

ACE-Data-0

TokTier: Exact Stateful CPU+GPU Tokenization for Agentic LLM Serving

arXiv1 repo

arXiv:2607.29678

toktier

DiffusionGemma Technical Report

arXiv1 repo

arXiv:2608.00146

awesome-gemma

TEngineDB-V: An OLAP-Native Vector Search System for Large-$k$ Workloads at Tencent

arXiv1 repo

arXiv:2608.00650

parqdb

FATE: Frame-Level Audio-Visual Temporal Embedding

arXiv1 repo

arXiv:2608.01310

FATE

arXiv:2608.01507

arXiv1 repo

arXiv:2608.01507

ripwire

Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance

arXiv1 repo

arXiv:2608.03403

ReMe

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

arXiv1 repo

arXiv:2608.03979

Vision-DeepResearch

When does training on downscaled images yield the same gradients?

arXiv1 repo

arXiv:2608.04448

anima-turbo-4step

K-EXAONE 2.0 Technical Report

arXiv1 repo

arXiv:2608.04505

K-EXAONE-2.0-750B-A37B

TS-RAG: Retrieval Augmented Generation for Time Series Forecasting

arXiv1 repo

arXiv:2608.06223

raggy

Retrofitting Linear Attention into Diffusion Language Models

arXiv1 repo

arXiv:2608.06628

LLaDA-Hybrid

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

arXiv1 repo

arXiv:2608.06867

xRouteBench

Vernata: Self-Supervised Learning of LiDAR Point Representations

arXiv1 repo

arXiv:2608.06919

vernata

The Spectral Neuron

arXiv1 repo

arXiv:2608.08003

spectral_neuron_paper

Motif 3: Technical Report

arXiv1 repo

arXiv:2608.09119

Motif-3

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

arXiv1 repo

arXiv:2608.10438

continuous-interaction-diffusion

AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

arXiv1 repo

arXiv:2608.11123

AlbumentationsX

HandEdit: A Unified Benchmark for Egocentric Human-to-Robot Dexterous Hand Image Editing

arXiv1 repo

arXiv:2608.12122

HandEdit

Alaya-EVOKE: From Linear-Scaling Supervision to Endless World

arXiv1 repo

arXiv:2608.13546

Evoke

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

arXiv1 repo

arXiv:2608.13966

Qwen3.8-27B-QUASAR-NVFP4

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

arXiv1 repo

arXiv:2608.14577

Internal-Safety-Collapse

Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI

arXiv1 repo

arXiv:2608.16319

relarena

Le Critique: Privileged Value Functions for LLM Reinforcement Learning

arXiv1 repo

arXiv:2608.16739

prime-values

Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision

arXiv1 repo

arXiv:2608.16812

ConceptEdit-12M

Agent Lightning v1.0: Towards Harnessed Agentic RL

arXiv1 repo

arXiv:2608.17528

agent-lightning

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

arXiv1 repo

arXiv:2608.19338

observerbench

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

arXiv1 repo

arXiv:2608.21381

PersonaMem-v2

Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era

arXiv1 repo

arXiv:2608.22602

accel-sim-framework

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

arXiv1 repo

arXiv:2608.23041

AutoSaddler

Prime Agent: A Self-Improving RLM Harness

arXiv1 repo

arXiv:2608.23552

prime-agent

Rebuild Dossier: Mechanically-Enforced Specs for Agentic App Rebuilds, and What Model-Tier Failures Reveal

arXiv1 repo

arXiv:2608.23616

rebuild-dossier

Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

arXiv1 repo

arXiv:2608.23691

station

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

arXiv1 repo

arXiv:2608.23873

semantic-overlays

pigzpp: Fast, Parallel, Portable Compression for the Whole Stack

arXiv1 repo

arXiv:2608.24153

pigzpp

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

arXiv1 repo

arXiv:2608.24188

paritok-4b-v1

Meta$^n$: Recursive Self-Improvement through Emergent Depth

arXiv1 repo

arXiv:2608.24735

meta-n

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

arXiv1 repo

arXiv:2608.24845

BVD

post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding at Synthesis

arXiv1 repo

arXiv:2608.24921

post-graph-rag

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

arXiv1 repo

arXiv:2608.25218

turnbench

Extending TotalSegmentator: Predicting Patient and Acquisition Characteristics from CT and MR Images

arXiv1 repo

arXiv:2608.29348

TotalSegmentator

SPARK: Skeleton-Guided Reasoning Synthesis from Large-Scale Scientific Literature

arXiv1 repo

arXiv:2608.30214

Spark-234K

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

arXiv1 repo

arXiv:2609.00111

Qwen-Drive-1.0-4B

EM^2Mem: Event-Centric Multimodal Memory for Large Language Models

arXiv1 repo

arXiv:2609.00551

LightMem

Bandits in Prod: Hyperparameter Optimization at Inference Time

arXiv1 repo

arXiv:2609.01335

IMABO

H3-World: Turning Language Understanding into World Control

arXiv1 repo

arXiv:2609.01560

H3-World

VibeVoice-ASR-Streaming Technical Report

arXiv1 repo

arXiv:2609.02812

VibeVoice-ASR-Streaming-7B

Last Translation Benchmark

arXiv1 repo

arXiv:2609.04173

last-translation-benchmark

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

arXiv1 repo

arXiv:2609.04196

Puffin-16M

Search in Power-Law Networks

arXiv1 repo

arXiv:cs/0103016

littleballoffur

Adaptive Interaction Using the Adaptive Agent Oriented Software Architecture (AAOSA)

arXiv1 repo

arXiv:cs/9812015

neuro-san-studio

Google AI beats top human players at strategy game StarCraft II

Nature1 repo

Nature:d41586-019-03298-6

learn-claude-code

Human-level control through deep reinforcement learning

Nature1 repo

Nature:nature14236

learn-claude-code

The 4D nucleome project

Nature1 repo

Nature:nature23884

alphagenome_research

Nature:nbt

Nature1 repo

Nature:nbt

nextflow

Nature:ng

Nature1 repo

Nature:ng

Histopathology_Benchmark

Circularly and elliptically polarized light under water and the Umov effect

Nature1 repo

Nature:s41377-019-0143-0

Awesome-Polarization

Nature-inspired chiral metasurfaces for circular polarization detection and full-Stokes polarimetric measurements

Nature1 repo

Nature:s41377-019-0184-4

Awesome-Polarization

Flavonoid intake is associated with lower mortality in the Danish Diet Cancer and Health Cohort

Nature1 repo

Nature:s41467-019-11622-x

HowToLiveLonger

The 4D Nucleome Data Portal as a resource for searching and visualizing curated nucleomics data

Nature1 repo

Nature:s41467-022-29697-4

alphagenome_research

Out-of-distribution generalization for learning quantum dynamics

Nature1 repo

Nature:s41467-023-39381-w

Awesome-Out-Of-Distribution-Detection

Augmenting interpretable models with large language models during training

Nature1 repo

Nature:s41467-023-43713-1

imodelsX

Sopa: a technology-invariant pipeline for analyses of image-based spatial omics

Nature1 repo

Nature:s41467-024-48981-z

sopa

Towards building multilingual language model for medicine

Nature1 repo

Nature:s41467-024-52417-z

MMedLM

Quantifying the reasoning abilities of LLMs on clinical cases

Nature1 repo

Nature:s41467-025-64769-1

MedRBench

A generalizable pathology foundation model using a unified knowledge distillation pretraining framework

Nature1 repo

Nature:s41551-025-01488-4

GPFM

A new era in functional genomics screens

Nature1 repo

Nature:s41576-021-00409-w

awesome-CRISPR

Highly accurate protein structure prediction with AlphaFold

Nature1 repo

Nature:s41586-021-03819-2

paper-reading

Highly accurate protein structure prediction for the human proteome

Nature1 repo

Nature:s41586-021-03828-1

alphafold3

Advancing mathematics by guiding human intuition with AI

Nature1 repo

Nature:s41586-021-04086-x

paper-reading

Solving olympiad geometry without human demonstrations

Nature1 repo

Nature:s41586-023-06747-5

superhuman

Accurate structure prediction of biomolecular interactions with AlphaFold 3

Nature1 repo

Nature:s41586-024-07487-w

alphafold3

A multimodal generative AI copilot for human pathology

Nature1 repo

Nature:s41586-024-07618-3

TITAN

A Drosophila computational brain model reveals sensorimotor processing

Nature1 repo

Nature:s41586-024-07763-9

fruit-fly-simulation

Larger and more instructable language models become less reliable

Nature1 repo

Nature:s41586-024-07930-y

llm-reliability

Scalable watermarking for identifying large language model outputs

Nature1 repo

Nature:s41586-024-08025-4

watermarks-remover

Functional evaluation and clinical classification of BRCA2 variants

Nature1 repo

Nature:s41586-024-08388-8

carbon

Mastering diverse control tasks through world models

Nature1 repo

Nature:s41586-025-08744-2

UI-TARS

Eye structure shapes neuron function in Drosophila motion vision

Nature1 repo

Nature:s41586-025-09276-5

fly

Advancing regulatory variant effect prediction with AlphaGenome

Nature1 repo

Nature:s41586-025-10014-0

alphagenome

Operational tropical cyclone forecasting with AI

Nature1 repo

Nature:s41586-026-10953-2

weathernext

A DNA language model based on multispecies alignment predicts the effects of genome-wide variants

Nature1 repo

Nature:s41587-024-02511-w

clinvar-vep-final

A guide to deep learning in healthcare

Nature1 repo

Nature:s41591-018-0316-z

Awesome-AI4Med

Evaluation and accurate diagnoses of pediatric diseases using artificial intelligence

Nature1 repo

Nature:s41591-018-0335-9

Awesome-AI4Med

Magnitude, demographics and dynamics of the effect of the first wave of the COVID-19 pandemic on all-cause mortality in 21 industrialized countries

Nature1 repo

Nature:s41591-020-1112-0

HowToLiveLonger

Optimized glycemic control of type 2 diabetes with reinforcement learning: a proof-of-concept trial

Nature1 repo

Nature:s41591-023-02552-9

DiagnosisCoding

Transparent medical image AI via an image–text foundation model grounded in medical literature

Nature1 repo

Nature:s41591-024-02887-x

DermFM-Zero

A multimodal vision foundation model for clinical dermatology

Nature1 repo

Nature:s41591-025-03747-y

PanDerm

Holistic evaluation of large language models for medical tasks with MedHELM

Nature1 repo

Nature:s41591-025-04151-2

helm

nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation

Nature1 repo

Nature:s41592-020-01008-z

TriALS

Alevin-fry unlocks rapid, accurate and memory-frugal quantification of single-cell RNA-seq data

Nature1 repo

Nature:s41592-022-01408-3

alevin-fry

Cellpose 2.0: how to train your own model

Nature1 repo

Nature:s41592-022-01663-4

cellpose

A cellular segmentation algorithm with fast customization

Nature1 repo

Nature:s41592-022-01664-3

cellpose

Cellpose3: one-click image restoration for improved cellular segmentation

Nature1 repo

Nature:s41592-025-02595-5

cellpose

A visual–omics foundation model to bridge histopathology with spatial transcriptomics

Nature1 repo

Nature:s41592-025-02707-1

AtlasPatch

Novae: a graph-based foundation model for spatial transcriptomics data

Nature1 repo

Nature:s41592-025-02899-6

sopa

High-parameter spatial multi-omics through histology-anchored integration

Nature1 repo

Nature:s41592-025-02926-6

SpatialEx

The ENERTALK dataset, 15 Hz electricity consumption data from 22 houses in Korea

Nature1 repo

Nature:s41597-019-0212-5

awesome-nilm

MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports

Nature1 repo

Nature:s41597-019-0322-0

rad-dino

A voltage and current measurement dataset for plug load appliance identification in households

Nature1 repo

Nature:s41597-020-0389-7

awesome-nilm

Kvasir-Capsule, a video capsule endoscopy dataset

Nature1 repo

Nature:s41597-021-00920-z

Endo-FM

The IDEAL household energy dataset, electricity, gas, contextual sensor data and survey data for 255 UK homes

Nature1 repo

Nature:s41597-021-00921-y

awesome-nilm

PAPILA: Dataset with fundus images and clinical data of both eyes of the same patient for glaucoma assessment

Nature1 repo

Nature:s41597-022-01388-1

FairMedFM

The heterogeneous pharmacological medical biochemical network PharMeBINet

Nature1 repo

Nature:s41597-022-01510-3

pykeen

BRAX, Brazilian labeled chest x-ray dataset

Nature1 repo

Nature:s41597-022-01608-8

rad-dino

Solar and wind power data from the Chinese State Grid Renewable Energy Generation Forecasting Competition

Nature1 repo

Nature:s41597-022-01696-6

Energy-EVA

MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification

Nature1 repo

Nature:s41597-022-01721-8

medmnistc-api

Three Dimensional Lower Extremity Musculoskeletal Geometry of the Visible Human Female and Male

Nature1 repo

Nature:s41597-022-01905-2

Awesome-Medical-Dataset

A dataset for medical instructional video classification and question answering

Nature1 repo

Nature:s41597-023-02036-y

Awesome-AI4Med

A Spitzoid Tumor dataset with clinical metadata and Whole Slide Images for Deep Learning models

Nature1 repo

Nature:s41597-023-02585-2

Awesome-Medical-Dataset

Open Fundus Photograph Dataset with Pathologic Myopia Recognition and Anatomical Structure Annotation

Nature1 repo

Nature:s41597-024-02911-2

Awesome-Medical-Dataset

MAPLES-DR: MESSIDOR Anatomical and Pathological Labels for Explainable Screening of Diabetic Retinopathy

Nature1 repo

Nature:s41597-024-03739-6

Awesome-Medical-Dataset

AMD-SD: An Optical Coherence Tomography Image Dataset for wet AMD Lesions Segmentation

Nature1 repo

Nature:s41597-024-03844-6

Awesome-Medical-Dataset

FARFUM-RoP, A dataset for computer-aided detection of Retinopathy of Prematurity

Nature1 repo

Nature:s41597-024-03897-7

Awesome-Medical-Dataset

An open-access lumbosacral spine MRI dataset with enhanced spinal nerve root structure resolution

Nature1 repo

Nature:s41597-024-03919-4

Awesome-Medical-Dataset

LungHist700: A dataset of histological images for deep learning in pulmonary pathology

Nature1 repo

Nature:s41597-024-03944-3

Awesome-Medical-Dataset

Dual Leap Motion Controller 2: A Robust Dataset for Multi-view Hand Pose Recognition

Nature1 repo

Nature:s41597-024-03968-9

Awesome-Medical-Dataset

Investigating the Quality of DermaMNIST and Fitzpatrick17k Dermatological Image Datasets

Nature1 repo

Nature:s41597-025-04382-5

medmnistc-api

The “Podcast” ECoG dataset for modeling neural activity during natural language comprehension

Nature1 repo

Nature:s41597-025-05462-2

automated-brain-explanations

A Custom Annotated Dataset for Segmentation of Pulmonary Veins, Arteries, and Airways

Nature1 repo

Nature:s41597-025-06074-6

TotalSegmentator

Learning a Health Knowledge Graph from Electronic Medical Records

Nature1 repo

Nature:s41598-017-05778-z

Awesome-AI4Med

Automated Gleason grading of prostate cancer tissue microarrays via deep learning

Nature1 repo

Nature:s41598-018-30535-1

Histopathology_Benchmark

Pessimism is associated with greater all-cause and cardiovascular mortality, but optimism is not protective

Nature1 repo

Nature:s41598-020-69388-y

HowToLiveLonger

A large dataset of white blood cells containing cell locations and types, along with segmented nuclei and cytoplasm

Nature1 repo

Nature:s41598-021-04426-x

DinoBloom

Bioacoustic classification of avian calls from raw sound waveforms with an open-source deep learning architecture

Nature1 repo

Nature:s41598-021-95076-6

NIPS4Bplus

iCodon customizes gene expression based on the codon composition

Nature1 repo

Nature:s41598-022-15526-7

CodonBERT

Contrastive language and vision learning of general fashion concepts

Nature1 repo

Nature:s41598-022-23052-9

Cool-GenAI-Fashion-Papers

Extracting and visualizing hidden activations and computational graphs of PyTorch models with TorchLens

Nature1 repo

Nature:s41598-023-40807-0

torchlens

PetBERT: automated ICD-11 syndromic disease coding for outbreak detection in first opinion veterinary electronic health records

Nature1 repo

Nature:s41598-023-45155-7

DiagnosisCoding

Global birdsong embeddings enable superior transfer learning for bioacoustic classification

Nature1 repo

Nature:s41598-023-49989-z

perch

Data-driven blood glucose level prediction in type 1 diabetes: a comprehensive comparative analysis

Nature1 repo

Nature:s41598-024-70277-x

nocturnal-hypo-gly-prob-forecast

A multiscale model for multivariate time series forecasting

Nature1 repo

Nature:s41598-024-82417-4

Time-Series-Library

Multiple model visual feature embedding and selection method for an efficient ocular disease classification

Nature1 repo

Nature:s41598-024-84922-y

Project-Imaging-X

Asymmetric dual pathway fusion for histology driven spatial transcriptomic prediction

Nature1 repo

Nature:s41598-026-63426-x

STP-Bench

VetTag: improving automated veterinary diagnosis coding via large-scale language modeling

Nature1 repo

Nature:s41746-019-0113-1

DiagnosisCoding

A large language model for electronic health records

Nature1 repo

Nature:s41746-022-00742-2

DiagnosisCoding

PatchSorter: a high throughput deep learning digital pathology tool for object labeling

Nature1 repo

Nature:s41746-024-01150-4

PatchSorter

Towards evaluating and building versatile large language models for medicine

Nature1 repo

Nature:s41746-024-01390-4

MedS-Ins

HONeYBEE: enabling scalable multimodal AI in oncology through foundation model-driven embeddings

Nature1 repo

Nature:s41746-025-02003-4

HoneyBee

Embeddings of clinical codes enable knowledge-grounded AI in medicine

Nature1 repo

Nature:s41746-026-02664-9

ClinVec

Deep learning models for predicting RNA degradation via dual crowdsourcing

Nature1 repo

Nature:s42256-022-00571-8

CodonBERT

Evaluation of post-hoc interpretability methods in time-series classification

Nature1 repo

Nature:s42256-023-00620-w

Awesome-Time-Series-Explainability

Parameter-efficient fine-tuning of large-scale pre-trained language models

Nature1 repo

Nature:s42256-023-00626-4

LLMSurvey

Defending ChatGPT against jailbreak attack via self-reminders

Nature1 repo

Nature:s42256-023-00765-8

llm-jailbreaking-defense

Exploring scalable medical image encoders beyond text supervision

Nature1 repo

Nature:s42256-024-00965-w

rad-dino

ImmunoStruct enables multimodal deep learning for immunogenicity prediction

Nature1 repo

Nature:s42256-025-01163-y

figures4papers

Large language models as uncertainty-calibrated optimizers for experimental discovery

Nature1 repo

Nature:s42256-026-01283-z

gollum

High-content CRISPR screening

Nature1 repo

Nature:s43586-021-00093-4

awesome-CRISPR

On-chip solar power source for self-powered smart microsensors in bulk CMOS process

Nature1 repo

Nature:s44172-025-00358-w

solarcircuits

PodGPT: an audio-augmented large language model for research and education

Nature1 repo

Nature:s44385-025-00022-0

PodGPT

Nature:sdata20157

Nature1 repo

Nature:sdata20157

awesome-nilm

Nature:sdata2018251

Nature1 repo

Nature:sdata2018251

RefinedVision

Multi-class texture analysis in colorectal cancer histology

Nature1 repo

Nature:srep27988

Histopathology_Benchmark

Titles are fetched live from arxiv.org and Crossref (for Nature and ACM articles), then cached — a bare id means that lookup hasn't completed yet.