arXiv cs.CV
Computer Vision and Pattern Recognition daily digest.
52 past editions, oldest first by date. Atom feed → · all categories
01 — Past editions
most recent first- 2026-07-19 MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators 19 papers · No. 58
- 2026-07-18 MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators 19 papers · No. 57
- 2026-07-17 MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators 19 papers · No. 56
- 2026-07-16 Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study 23 papers · No. 55
- 2026-07-15 HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and... 18 papers · No. 54
- 2026-07-14 Evidence-Backed Video Question Answering 21 papers · No. 53
- 2026-07-13 Scalable Visual Pretraining for Language Intelligence 29 papers · No. 52
- 2026-07-12 OpenCoF: Learning to Reason Through Video Generation 20 papers · No. 51
- 2026-07-11 OpenCoF: Learning to Reason Through Video Generation 20 papers · No. 50
- 2026-07-10 OpenCoF: Learning to Reason Through Video Generation 20 papers · No. 49
- 2026-07-09 MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal... 21 papers · No. 48
- 2026-07-08 ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation 23 papers · No. 47
- 2026-07-07 From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model 20 papers · No. 46
- 2026-07-06 Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning 18 papers · No. 45
- 2026-07-05 Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning 18 papers · No. 44
- 2026-07-04 Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning 18 papers · No. 43
- 2026-07-03 Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning 18 papers · No. 42
- 2026-07-02 World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video 35 papers · No. 41
- 2026-07-01 FLORA: A deep learning approach to predict forest attributes from... 24 papers · No. 40
- 2026-06-30 Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware... 21 papers · No. 39
- 2026-06-29 Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly... 30 papers · No. 38
- 2026-06-28 DanceOPD: On-Policy Generative Field Distillation 20 papers · No. 37
- 2026-06-27 DanceOPD: On-Policy Generative Field Distillation 20 papers · No. 36
- 2026-06-26 DanceOPD: On-Policy Generative Field Distillation 20 papers · No. 35
- 2026-06-25 A cross-process welding penetration status prediction algorithm based on... 24 papers · No. 34
- 2026-06-24 FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse... 24 papers · No. 33
- 2026-06-23 Semantic Browsing: Controllable Diversity for Image Generation 20 papers · No. 32
- 2026-06-22 UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning 22 papers · No. 31
- 2026-06-21 UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning 22 papers · No. 30
- 2026-06-20 UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning 22 papers · No. 29
- 2026-06-19 UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning 22 papers · No. 28
- 2026-06-18 Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation 12 papers · No. 27
- 2026-06-17 Adaptive Volumetric Mechanical Property Fields Invariant to Resolution 25 papers · No. 26
- 2026-06-16 The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim... 15 papers · No. 25
- 2026-06-15 Gaze Heads: How VLMs Look at What They Describe 17 papers · No. 24
- 2026-06-14 SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 20 papers · No. 23
- 2026-06-13 SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 20 papers · No. 22
- 2026-06-12 SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning 20 papers · No. 21
- 2026-06-11 Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models 29 papers · No. 20
- 2026-06-10 FADA: Accessible fetal ultrasound interpretation and annotation with a... 9 papers · No. 19
- 2026-06-09 OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics 18 papers · No. 18
- 2026-06-08 MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding... 28 papers · No. 17
- 2026-06-07 A Vision-language Framework for Comparative Reasoning in Radiology 10 papers · No. 16
- 2026-06-06 A Vision-language Framework for Comparative Reasoning in Radiology 10 papers · No. 15
- 2026-06-05 A Vision-language Framework for Comparative Reasoning in Radiology 10 papers · No. 14
- 2026-06-02 Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via... 29 papers · No. 13
- 2026-05-30 VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion 16 papers · No. 12
- 2026-05-28 AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning 10 papers · No. 11
- 2026-05-27 LocateAnything: Fast and High-Quality Vision-Language Grounding with... 14 papers · No. 10
- 2026-05-25 ETCHR: Editing To Clarify and Harness Reasoning 29 papers · No. 9
- 2026-05-23 AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild 32 papers · No. 8
- 2026-05-22 AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild 32 papers · No. 7
Colophon
Set in Georgia, with system sans for interface chrome and a monospaced stack for code and paper identifiers. Sole accent: amber
#D99C5E. Built and served on an always-free VM. The masthead is set 14% letterspaced because newspapers do that and it works.