Multi-Faceted Analysis for Video Understanding Workshop @ WACV 2027
About
We are witnessing a paradigm shift in computer vision, driven by the explosion of video data from diverse platforms and the rise of multimodal foundation models. This workshop is dedicated to the multi-faceted analysis of video, exploring how advanced technologies—specifically Forensic Search, Video Question Answering (VQA), and Long Video Understanding—can unlock deep insights from varying perspectives. We focus on three critical and distinct modalities: Aerial/Drone Vision, Surveillance Systems, and Egocentric Vision (personal glasses and body-worn devices).
Each of these modalities offers a unique vantage point on the world, from the broad, contextual overview of a drone to the persistent monitoring of surveillance cameras and the first-person, intent-driven perspective of body-worn devices. The core theme of this workshop is to promote research into flexible interactions with these diverse video sources. By leveraging natural language queries (VQA) and advanced forensic search capabilities, we can transform passive video archives into interactive intelligence systems. This approach allows users to intuitively navigate and reason about complex events across different scales and viewpoints, bridging the gap between raw pixel data and high-level semantic understanding.
This workshop serves as a premier forum for researchers to present novel methodologies that handle the specific nuances of each modality while pushing the boundaries of interactive video understanding. By fostering discussion on these multi-faceted approaches, we aim to accelerate the development of robust, real-world applications in public safety, autonomous systems, smart cities, and next-generation media intelligence.
Call for Papers
We invite original research contributions in (but not limited to) the following areas:
• Open-vocabulary scene understanding
• Multimodal drone video analysis
• Aerial human action recognition & crowd analysis
• Efficient UAV object detection
Submissions can include short papers (up to 4 pages including references) or full papers (up to 8 pages excluding references) in the WACV main conference format. Accepted full papers will be included in the WACV proceedings. When deciding whether a submission should be a short paper or a full paper, please consider the significance and novelty of the contributions. If the proposed approach is well-developed with sufficient theoretical and/or empirical justification, consider submitting a full paper. If the work is a simple extension or summary of published work in other venues or is in the proof-of-concept stage, a short paper will provide a good basis for discussion and feedback. We welcome papers that propose a new technical approach for any of the above, or claim to take a position regarding challenging and open-ended questions that address multi-facet analysis for aerial and remote sensing video understanding
Important Dates
• Paper Submission Deadline - October 20, 2026 11:59 PM PST
• Decision Notification to Authors - October 30, 2026
• Camera Ready Submission Deadline (as per main conference) - Nov 20, 2026 11:59 PM PST
Submission Preparation Instructions
Please follow the main conference format and submission guidelines to prepare your papers. Check WACV Submission Guidelines here
Submission site
We will use OpenReview for submissions. Link to the submission portal
All authors need to have an OpenReview profile. Please plan ahead as it can take up to two weeks.
Keynote Speakers
🗣️ Speaker 1: Dr. Dinesh Manocha
Bio
Talk Title: Towards General Audio-Visual Intelligence
Abstract: Audio-Visual General Intelligence—the capacity of AI agents to deeply understand and reason about all types of auditory and visual inputs, including speech, environmental sounds, and music, and to combine them with images and videos—is crucial for enabling AI to interact seamlessly and naturally with our world. Despite this importance, audio understanding has traditionally lagged advancements in vision and language processing. This gap stems from significant challenges, including limited datasets, the complexity of audio signals, and a shortage of advanced neural architectures and effective training methods explicitly tailored for audio. In this talk, we overview our work on audio and audio-visual understanding, and the development of audio and audio-visual LLMs. We provide an overview of audio large language models (ALLMs) used for advanced audio perception and complex reasoning. This includes GAMA, which is built with a specialized architecture, optimized audio encoding, and a novel alignment dataset. We also introduce ReCLAP, a state-of-the-art audio-language encoder, and CompA, one of the first projects to tackle compositional reasoning in audio-language models—a critical challenge given the inherently compositional nature of audio. We also discuss the Audio Flamingo series and Music Flamingo, which are our open ALLM models that provide advanced long-audio understanding and reasoning capabilities across speech, sound, and music. Finally, we present audio-visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed to understand and reason over long, complex real-world (audio-visual) videos, and we highlight its performance across different scenarios.
Bio: Dinesh Manocha is the Paul Chrisman-Iribe Chair in Computer Science & ECE and a Distinguished University Professor at the University of Maryland, College Park. His research interests include virtual environments, physically-based modeling, and robotics. His group has developed numerous software packages that are standard and licensed to over 60 commercial vendors. He has published more than 900 papers & supervised 69 PhD dissertations. His group has received more than 22 best paper and test-of-time awards at leading conferences in computer graphics, solid modeling, multimedia, VR, and robotics. Manocha is a Fellow of AAAI, AAAS, ACM, IEEE, and NAI, a member of ACM SIGGRAPH and IEEE VR Academies, and a recipient of the Bézier Award from the Solid Modeling Association. He received the Distinguished Alumni Award from IIT Delhi and the Distinguished Career in Computer Science Award from the Washington Academy of Sciences. He co-founded Impulsonic, a physics-based audio simulation technology developer, which Valve Inc. acquired in November 2016.
🗣️ Speaker 2: Dr. Xiaoming Liu
Bio
Talk Title: **
Abstract:
🗣️ Speaker 3: Dr. Yang Liu
Bio
Talk Title: **
Abstract:
Program Schedule
Date:
Time: 09:00 – 18:30
| Time | Event |
|---|---|
| 09:00 – 09:10 | Opening Remarks |
| 09:10 – 10:00 | Invited Talk 1 |
| 10:00 – 10:50 | Invited Talk 2 |
| 10:50 – 11:10 | Coffee Break |
| 11:10 – 12:00 | Invited Talk 3 |
| 12:00 – 13:30 | Lunch Break and Poster Session 1 |
| 13:30 – 14:20 | Invited Talk 4 |
| 14:20 – 15:10 | Invited Talk 5 |
| 15:10 – 15:40 | Coffee Break |
| 15:40 – 16:40 | AI Challenge Session: overview, top solutions, and awards |
| 16:40 – 17:30 | Poster Session and Industry Demos |
| 17:30 – 18:15 | Panel Discussion: “The Future of Video Intelligence: From Forensics to Autonomy” |
| 18:15 – 18:30 | Best Paper Awards and Closing Remarks |
Organizers
Contact
For questions about the workshop, please contact the organizers at mavu-wacv27-organizers@googlegroups.com.
