0

All papers

Every paper that carries a page of its own here. The Papers page leads with what is trending and lets you filter the recent catalog; this is the plain index of the rest.

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesarXiv 2022NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and MappingICCV 2023 1GPT4RoI: Instruction Tuning Large Language Model on Region-of-InterestarXiv 2023Best-of-N JailbreakingarXiv 2024PatchmatchNet: Learned Multi-View Patchmatch StereoCVPR 2021 1Modeling Context in Referring ExpressionsarXiv 2016Llumnix: Dynamic Scheduling for Large Language Model ServingarXiv 2024CValues: Measuring the Values of Chinese Large Language Models from Safety to ResponsibilityarXiv 2023Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolationtrain-short-test-long-attention-with-linear-1Mask Transfiner for High-Quality Instance SegmentationCVPR 2022 1Self-labelling via simultaneous clustering and representation learningICLR 2020 1SampleRNN: An Unconditional End-to-End Neural Audio Generation ModelarXiv 2016Diffusion Models in Low-Level Vision: A SurveyarXiv 2024PINA: Leveraging Side Information in eXtreme Multi-label Classification via Predicted Instance Neighborhood AggregationarXiv 2023Equivariant Diffusion for Molecule Generation in 3DarXiv 2022ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation TransferACL 2021 5Junction Tree Variational Autoencoder for Molecular Graph Generationjunction-tree-variational-autoencoder-for-1ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image GenerationICCV 2023 1MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningarXiv 2024Generalizable and Animatable Gaussian Head AvatararXiv 2024Towards a Unified View of Parameter-Efficient Transfer Learningtowards-a-unified-view-of-parameter-efficientDenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingCVPR 2022 1DocRes: A Generalist Model Toward Unifying Document Image Restoration TasksCVPR 2024 1Mengzi: Towards Lightweight yet Ingenious Pre-trained Models for ChinesearXiv 2021Foundational Models Defining a New Era in Vision: A Survey and OutlookarXiv 2023Escaping the Big Data Paradigm with Compact TransformersarXiv 2021LAMBDA: A Large Model Based Data AgentarXiv 2024Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networksenhancing-the-reliability-of-out-of-1StyleSwin: Transformer-based GAN for High-resolution Image Generationstyleswin-transformer-based-gan-for-highFreeInit: Bridging Initialization Gap in Video Diffusion ModelsarXiv 2023Unleashing Text-to-Image Diffusion Models for Visual Perceptionunleashing-text-to-image-diffusion-models-forUsing Sequences of Life-events to Predict Human LivesarXiv 2023Pandora: Towards General World Model with Natural Language Actions and Video StatesarXiv 2024StyleSDF: High-Resolution 3D-Consistent Image and Geometry GenerationCVPR 2022 1Rewriting a Deep Generative ModelECCV 2020 8ControlNet++: Improving Conditional Controls with Efficient Consistency FeedbackarXiv 2024A Survey on All-in-One Image Restoration: Taxonomy, Evaluation and Future TrendsarXiv 2024Mass-Editing Memory in a TransformerarXiv 2022DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language ModelsarXiv 2023XrayGPT: Chest Radiographs Summarization using Medical Vision-Language ModelsarXiv 2023Dawn of the transformer era in speech emotion recognition: closing the valence gaparXiv 2022InternLM2.5-StepProver: Advancing Automated Theorem Proving via Expert Iteration on Large-Scale LEAN ProblemsarXiv 2024Can large language models provide useful feedback on research papers? A large-scale empirical analysisarXiv 2023Learning-Rate-Free Learning by D-AdaptationarXiv 2023StyleGANEX: StyleGAN-Based Manipulation Beyond Cropped Aligned FacesICCV 2023 1PeRFlow: Piecewise Rectified Flow as Universal Plug-and-Play AcceleratorarXiv 2024TinySAM: Pushing the Envelope for Efficient Segment Anything ModelarXiv 2023AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly DetectionarXiv 20233D Registration with Maximal Cliques3d-registration-with-maximal-cliquesEDGE: Editable Dance Generation From MusicCVPR 2023 1Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengtharXiv 2024Deep Video Inpaintingdeep-video-inpainting-1Neural 3D Scene Reconstruction with the Manhattan-world AssumptionCVPR 2022 1SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian SplattingCVPR 2024 1Break-A-Scene: Extracting Multiple Concepts from a Single ImagearXiv 2023Orb: A Fast, Scalable Neural Network PotentialarXiv 2024ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustnessimagenet-trained-cnns-are-biased-towards-1A Hierarchical Representation Network for Accurate and Detailed Face Reconstruction from In-The-Wild ImagesCVPR 2023 1Understanding URDF: A Dataset and AnalysisarXiv 2023Stable-Hair: Real-World Hair Transfer via Diffusion ModelarXiv 2024AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder LossarXiv 2019PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image EditorarXiv 2023Deep Unsupervised Learning using Nonequilibrium ThermodynamicsarXiv 2015Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersarXiv 2024Hungry Hungry Hippos: Towards Language Modeling with State Space ModelsarXiv 2022ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single ImageCVPR 2024 1Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D GenerationCVPR 2023 1Dynamic Snake Convolution based on Topological Geometric Constraints for Tubular Structure SegmentationICCV 2023 1A Survey of Knowledge-Enhanced Text GenerationarXiv 2020Mimic before Reconstruct: Enhancing Masked Autoencoders with Feature MimickingarXiv 2023HybVIO: Pushing the Limits of Real-time Visual-inertial OdometryarXiv 2021LettuceDetect: A Hallucination Detection Framework for RAG ApplicationsarXiv 2025OmniDrones: An Efficient and Flexible Platform for Reinforcement Learning in Drone ControlarXiv 2023Deep Residual Learning for Small-Footprint Keyword SpottingarXiv 2017G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question AnsweringarXiv 2024Visual In-Context PromptingCVPR 2024 1DialogStudio: Towards Richest and Most Diverse Unified Dataset Collection for Conversational AIarXiv 2023FewCLUE: A Chinese Few-shot Learning Evaluation BenchmarkarXiv 2021StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice ConversionarXiv 2021RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language ModelsarXiv 2023Joint 2D-3D-Semantic Data for Indoor Scene UnderstandingarXiv 2017LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QAarXiv 2024Aggregated Contextual Transformations for High-Resolution Image InpaintingarXiv 2021LLaMA Pro: Progressive LLaMA with Block ExpansionarXiv 2024E(n) Equivariant Graph Neural NetworksarXiv 2021GMAN: A Graph Multi-Attention Network for Traffic PredictionarXiv 2019MotionClone: Training-Free Motion Cloning for Controllable Video GenerationarXiv 2024HoloLens 2 Research Mode as a Tool for Computer Vision ResearcharXiv 2020GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answeringgqa-a-new-dataset-for-real-world-visualPredRNN: A Recurrent Neural Network for Spatiotemporal Predictive LearningarXiv 2021Model scale versus domain knowledge in statistical forecasting of chaotic systemsarXiv 2023The CLRS-Text Algorithmic Reasoning Language BenchmarkarXiv 2024Dual Aggregation Transformer for Image Super-ResolutionICCV 2023 1Depth Any Video with Scalable Synthetic DataarXiv 2024ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion ModelsarXiv 2023Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video UnderstandingarXiv 2025Wasserstein Auto-Encoderswasserstein-auto-encoders-1M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language ModelsarXiv 2023Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision ModelsarXiv 2024ReVersion: Diffusion-Based Relation Inversion from ImagesarXiv 2023TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task TokenizationCVPR 2025 1CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation AlignmentarXiv 2022Large Language Model Instruction Following: A Survey of Progresses and ChallengesarXiv 2023The All-Seeing Project V2: Towards General Relation Comprehension of the Open WorldarXiv 2024Deep Exemplar-based ColorizationarXiv 2018DragAnything: Motion Control for Anything using Entity RepresentationarXiv 2024StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-trainingarXiv 2023Pervasive Attention: 2D Convolutional Neural Networks for Sequence-to-Sequence Predictionpervasive-attention-2d-convolutional-neuralPretraining is All You Need for Image-to-Image TranslationarXiv 2022Your ViT is Secretly an Image Segmentation Modelyour-vit-is-secretly-an-image-segmentationGraphSAINT: Graph Sampling Based Inductive Learning MethodICLR 2020 1Efficient Wait-k Models for Simultaneous Machine TranslationarXiv 2020Parameter-Efficient Transfer Learning for NLParXiv 2019Git Re-Basin: Merging Models modulo Permutation SymmetriesarXiv 2022AudioDec: An Open-source Streaming High-fidelity Neural Audio CodecarXiv 2023Latent Video Diffusion Models for High-Fidelity Long Video GenerationarXiv 2022KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven Conversationkdconv-a-chinese-multi-domain-dialogue-1Measuring Coding Challenge Competence With APPSarXiv 2021Rotation-invariant convolutional neural networks for galaxy morphology predictionarXiv 2015HyperHuman: Hyper-Realistic Human Generation with Latent Structural DiffusionarXiv 2023GPT-GNN: Generative Pre-Training of Graph Neural NetworksarXiv 2020Manifold Mixup: Better Representations by Interpolating Hidden StatesICLR 2019 5Fine-Tuning Image-Conditional Diffusion Models is Easier than You ThinkarXiv 2024CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and GenerationarXiv 2021KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPsICCV 2021 10LW-DETR: A Transformer Replacement to YOLO for Real-Time DetectionarXiv 2024GLTR: Statistical Detection and Visualization of Generated Textgltr-statistical-detection-and-visualization-1SMASH: One-Shot Model Architecture Search through HyperNetworkssmash-one-shot-model-architecture-search-1Augraphy: A Data Augmentation Library for Document ImagesarXiv 2022Vision-LSTM: xLSTM as Generic Vision BackbonearXiv 2024Parameter Prediction for Unseen Deep ArchitecturesNeurIPS 2021 12BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose EstimationarXiv 2025Improving Retrieval-Augmented Generation in Medicine with Iterative Follow-up QuestionsarXiv 2024A Survey on the Memory Mechanism of Large Language Model based AgentsarXiv 2024MaxViT: Multi-Axis Vision TransformerarXiv 2022Declarative Experimentation in Information Retrieval using PyTerrierarXiv 2020Denoising Diffusion Models for Plug-and-Play Image RestorationarXiv 2023FloWaveNet : A Generative Flow for Raw AudioarXiv 2018Decoupled Attention Network for Text RecognitionarXiv 2019VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingarXiv 2024MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based FinetuningarXiv 2021HumanRF: High-Fidelity Neural Radiance Fields for Humans in MotionarXiv 2023DiGress: Discrete Denoising diffusion for graph generationarXiv 2022AI Challenger : A Large-scale Dataset for Going Deeper in Image UnderstandingarXiv 2017GaMeS: Mesh-Based Adapting and Modification of Gaussian SplattingarXiv 2024Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningarXiv 2024ViTMatte: Boosting Image Matting with Pretrained Plain Vision TransformersarXiv 2023CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language ModelsarXiv 2023PersFormer: 3D Lane Detection via Perspective Transformer and the OpenLane Benchmarkpersformer-3d-lane-detection-via-perspectiveBigBIO: A Framework for Data-Centric Biomedical Natural Language ProcessingarXiv 2022Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual ConceptsarXiv 2021WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement LearningarXiv 2024Unsupervised Deep Embedding for Clustering AnalysisarXiv 2015MiVOLO: Multi-input Transformer for Age and Gender EstimationarXiv 2023Grounding Language Models to Images for Multimodal Inputs and OutputsarXiv 2023Continual Learning of Large Language Models: A Comprehensive SurveyarXiv 2024Passage Re-ranking with BERTarXiv 2019Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware FusionCVPR 2021 1JADE: A Linguistics-based Safety Evaluation Platform for Large Language ModelsarXiv 2023Dank Learning: Generating Memes Using Deep Neural NetworksarXiv 2018Chain of Ideas: Revolutionizing Research Via Novel Idea Development with LLM AgentsarXiv 2024Advancing Plain Vision Transformer Towards Remote Sensing Foundation ModelarXiv 2022Sample-Efficient Neural Architecture Search by Learning Action SpacearXiv 2019Deep Networks with Stochastic DeptharXiv 2016Deep Learning for Multivariate Time Series Imputation: A SurveyarXiv 2024LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language ModelsarXiv 2023NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating Systemnl2bash-a-corpus-and-semantic-parser-for-2OmniThink: Expanding Knowledge Boundaries in Machine Writing through ThinkingarXiv 2025Pop2Piano : Pop Audio-based Piano Cover GenerationarXiv 2022AgentHarm: A Benchmark for Measuring Harmfulness of LLM AgentsarXiv 2024Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and FuturearXiv 2023Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebatearXiv 2023Your Diffusion Model is Secretly a Zero-Shot ClassifierICCV 2023 1ReLiK: Retrieve and LinK, Fast and Accurate Entity Linking and Relation Extraction on an Academic BudgetarXiv 2024PromptIR: Prompting for All-in-One Blind Image RestorationarXiv 2023A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and ChallengesarXiv 2025ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasksalfred-a-benchmark-for-interpreting-grounded-1ChemBERTa: Large-Scale Self-Supervised Pretraining for Molecular Property PredictionarXiv 2020ByteTransformer: A High-Performance Transformer Boosted for Variable-Length InputsarXiv 2022Generalizable Humanoid Manipulation with 3D Diffusion PoliciesarXiv 2024A Contrastive Framework for Neural Text GenerationarXiv 2022Semantic Image Inversion and Editing using Rectified Stochastic Differential EquationsarXiv 2024ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control SystemsarXiv 2023DKM: Dense Kernelized Feature Matching for Geometry EstimationarXiv 2022UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-TrainingarXiv 2021Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representationmichelangelo-conditional-3d-shape-generationOpenResearcher: Unleashing AI for Accelerated Scientific ResearcharXiv 2024Trellis Networks for Sequence Modelingtrellis-networks-for-sequence-modeling-1Counter-Strike Deathmatch with Large-Scale Behavioural CloningarXiv 2021Omnidata: A Scalable Pipeline for Making Multi-Task Mid-Level Vision Datasets from 3D Scansomnidata-a-scalable-pipeline-for-making-multiImage-based table recognition: data, model, and evaluationECCV 2020 8On the Use of ArXiv as a DatasetarXiv 2019GenAD: Generative End-to-End Autonomous DrivingarXiv 2024Benchmarking Uncertainty Quantification Methods for Large Language Models with LM-PolygrapharXiv 2024Visual Style Prompting with Swapping Self-AttentionarXiv 2024FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive LearningarXiv 2023The Matrix Calculus You Need For Deep LearningarXiv 2018Generating Images with Multimodal Language ModelsNeurIPS 2023 11Learning Disentangled Joint Continuous and Discrete Representationslearning-disentangled-joint-continuous-and-1What Does BERT Look At? An Analysis of BERT's Attentionwhat-does-bert-look-at-an-analysis-of-berts-1

Back to Papers