Venue
The CVAVM workshop is part of ICCV 2017. It will take place in Venice, Italy, on 23 October 2017, at the same venue as the main ICCV conference: Palazzo del Cinema – Venice Convention Center (Lungomare Guglielmo Marconi, 3030126 Lido di Venezia – Venice, see google map). Please see the ICCV webpage for more information on venue, accommodations, and other details.

Important dates
Paper registration (title, abstract and authors): July 19, July 21 2017 (due to CMT issue)
Full paper submission: July 21, 2017
Acceptance notification: August 11, 2017
Camera-ready paper due: August 25, 2017
Workshop date: October 23, 2017 (morning)
ICCV main conference date: October 24-27, 2017
Paper Submission
Our CVAVM workshop invites paper submissions on any applications and algorithms that combine visual and audio information. See the list of topics below.
Paper submissions are handled through the workshop’s CMT website: https://cmt3.research.microsoft.com/CVAVM2017. If you have any issues or questions, do not hesitate to contact us (bazinjc AT kaist.ac.kr).
The paper registration deadline (title, abstract and authors) is July 21 and the paper submission deadline is July 21, see dates. The paper submission is similar to the ICCV main conference, see guidelines and template on the ICCV webpage. Papers are limited to 8 pages (excluding references), including figures and tables. The reviewing will be double-blind, and each submission will be reviewed by at least two reviewers. Papers that are not blind, or do not use the template, or have more than 8 pages (excluding references) will be rejected without review. All the accepted papers will be published in the ICCV workshop proceedings.
Topics include (but are not limited to):
– multi-modal learning and deep learning
– automatic video captioning
– joint audio-visual processing
– scene/action recognition, and video classification
– 3D reconstruction and tracking
– video segmentation and saliency
– speaker identification
– speech recognition in videos
– virtual/augmented reality and tele-presence
– human-computer interaction
– robotics
– automatic generation of videos
– trailer generation
– video and movie manipulation
– video synchronization
– image sonification
– video-to-music alignment
– joint audio-video retargeting
Schedule
The workshop will be on October 23, 2017. See venue information above.
| Time |
|
| 08:40 – 08:50 |
Welcome and Opening Remarks |
| 08:50 – 09:35 |
Invited keynote1 by Rif A. Saurous (Google) |
| 09:35 – 09:50 |
Oral1: “Improving Speaker Turn Embedding by Crossmodal
Transfer Learning From Face Embedding” by Nam Le and Jean-Marc Odobez |
| 09:50 – 10:05 |
Oral2: “Unsupervised Cross-Modal Deep-Model Adaptation for Audio-Visual Re-Identification With Wearable Cameras”, by Alessio Brutti and Andrea Cavallaro |
| 10:05 – 10:20 |
Oral3: “Exploiting the Complementarity of Audio and Visual Data in Multi-Speaker Tracking”, by Yutong Ban, Laurent Girin, Xavier Alameda-Pineda and Radu Horaud |
| 10:20 – 10:40 |
Coffee break |
| 10:40 – 11:25 |
Invited keynote2 by Rémi Ronfard (INRIA) |
| 11:25 – 11:40 |
Oral4: “Improved Speech Reconstruction From Silent Video”, by Ariel Ephrat, Tavi Halperin and Shmuel Peleg |
| 11:40 – 11:55 |
Oral5: “Visual Music Transcription of Clarinet Video Recordings Trained With Audio-Based Labelled Data”, by Pablo Zinemanas, Pablo Arias, Gloria Haro and Emilia Gómez |
| 11:55 – 12:40 |
Invited keynote3 by Josh McDermott (MIT) |
| 12:40 – 12:45 |
Closing Remarks |
Keynote speakers
Josh McDermott, MIT, USA
Rémi Ronfard, INRIA, France
Rif A. Saurous, Google, USA
Workshop chairs
Jean-Charles Bazin, KAIST
Zhengyou Zhang, Microsoft Research
William T. Freeman, MIT
Prof. Jean-Charles Bazin
Zhengyou Zhang
Prof. William T. Freeman
Committee members
TBA
Sponsors
