Apple is presenting new research at the annual conference on IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), which takes place in person in Nashville, Tennessee from June 11 to June 15. We are proud to sponsor the conference, which brings together the scientific and industrial research communities in computer vision and pattern recognition. Below is an overview of Apple’s participation at CVPR 2025.
Jump to a section:
Schedule
Stop by the Apple booth in the Music City Center, booth #1217, during exhibition hours. All times listed in CDT (Nashville time):
- Friday, June 13: 10:00am – 6:30pm
- Saturday, June 14: 10:00am – 6:30pm
- Sunday, June 15: 10:00am – 3:00pm
Wednesday, June 11
- WORKSHOP
- Computer Vision for Metaverse Workshop (CV4Metaverse) 2025
- 8:10am – 12:20pm, Room 107 A
-
- POSTER
- “A Stereo Image Quality Predictor for AR/VR”
- Netanel Tamir (Weizmann Institute of Science), Shir Amir, Ranel Itzhaky, Noam Atia (Tel Aviv University), Shobhita Sundaram (Massachusetts Institute of Technology), Stephanie Fu (Massachusetts Institute of Technology), Miriam Farber, Ron Sokolovsky, Richard Zhang (Independent researcher), Tali Dekel (Weizmann Institute of Science), Phillip Isola (Massachusetts Institute of Technology)
- WORKSHOP
- Fine-Grained Visual Categorization (FGVC12) 2025
- 9:00am – 5:15pm, Room 104 E
-
- POSTER
- “Rethinking Semi-Supervised Domain Adaptation for Semantic Segmentation with Semi-Supervised Learning in the Foundation Model Era”
- Joshua Kurien (University of Waterloo), Bavesh Balaji (University of Waterloo), Henry Lai, Pablo Guerrero Vela, C Thomas, Alex Wong, Sirisha Rambhatla
Thursday, June 12
- WORKSHOP
- Women in Computer Vision (WiCV)
- 8:30am – 1:00pm (Workshop), Room 105 B
- 6:00pm – 8:00pm (Mentorship Dinner), Room 202 C
- Fazilet Gokbudak, Jess Knowles, and Michael Kirchhof will be representing Apple at the WiCV Mentorship Dinner
Friday, June 13
- HIGHLIGHT POSTER
- Multimodal Autoregressive Pre-Training of Large Vision Encoders
- 4:00pm – 6:00pm, #407, Poster Session 2, Exhibit Hall D
- Enrico Fini, Mustafa Shukor (Sorbonne University), Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Louis Béthune, Zhe Gan, Victor Turrisi, Alexander Toshev, Marcin Eichner, Yinfei Yang, Moin Nabi, Josh Susskind, Alaaeldin El-Nouby
Saturday, June 14
- ORAL PRESENTATION
- From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
- 9:00am – 10:15am, Presentation #5, Oral Session 3, Davidson Ballroom
- Andrew Szot (Georgia Institute of Technology), Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira (Georgia Institute of Technology), Alexander Toshev
- POSTER
- From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons
- 10:30am – 12:30pm, #329, Poster Session 3, Exhibit Hall D
- Andrew Szot (Georgia Institute of Technology), Bogdan Mazoure, Omar Attia, Aleksei Timofeev, Harsh Agrawal, Devon Hjelm, Zhe Gan, Zsolt Kira (Georgia Institute of Technology), Alexander Toshev
- HIGHLIGHT POSTER
- Matrix3D: Large Photogrammetry Model All-in-One
- 10:30am – 12:30pm, #57, Poster Session 3, Exhibit Hall D
- Yuanxun Lu (Nanjing University), Jingyang Zhang, Tian Fang, Danny Nahmias, Yanghai Tsin, Long Quan (Hong Kong University of Science and Technology), Xun Cao (Nanjing University), Yao Yao (Nanjing University), Shiwei Li
- POSTER
- FastVLM: Efficient Vision Encoding for Vision Language Models
- 5:00pm – 7:00pm, #378, Poster Session 4, Exhibit Hall D
- Pavan Kumar Anasosalu Vasu, Fartash Faghri, Chun-Liang Li, Cem Koc, Nate True, Albert Antony, Gokul Santhanam, James Gabriel, Peter Grasch, Oncel Tuzel, Hadi Pour Ansari
Sunday, June 15
Booth Programming & Demos
Visit Apple’s booth at Music City Center, Booth #1217, during exhibition hours.
Featured Research Sessions
Technical Demos
- DEMO
- FastVLM
- FastVLM is a family of mobile-friendly vision language models.These models use a mix of CNN and Transformer architectures for vision encoding designed specifically for processing high-resolution images. Together, they deliver the best balance between accuracy and speed.
- Friday, June 13:10:00am – 12:30pm, 2:30pm – 4:30pm
- Saturday, June 14: 10:00am – 12:30pm, 2:30pm – 4:30pm
- Sunday, June 15: 10:00am – 12:30pm
Accepted Papers
- FastVLM: Efficient Vision Encoding for Vision Language Models
- Pavan Kumar Anasosalu Vasu, Fartash Faghri, Chun-Liang Li, Cem Koc, Nate True, Albert Antony, Gokul Santhanam, James Gabriel, Peter Grasch, Oncel Tuzel, Hadi Pour Ansari
- Matrix3D: Large Photogrammetry Model All-in-One
- Yuanxun Lu (Nanjing University), Jingyang Zhang, Tian Fang, Danny Nahmias, Yanghai Tsin, Long Quan (Hong Kong University of Science and Technology), Xun Cao (Nanjing University), Yao Yao (Nanjing University), Shiwei Li
- Multimodal Autoregressive Pre-training of Large Vision Encoders
- Enrico Fini, Mustafa Shukor (Sorbonne University), Xiujun Li, Philipp Dufter, Michal Klein, David Haldimann, Sai Aitharaju, Louis Béthune, Zhe Gan, Victor Turrisi, Alexander Toshev, Marcin Eichner, Yinfei Yang, Moin Nabi, Josh Susskind, Alaaeldin El-Nouby
Acknowledgements
Jack Langerman is Workshop Co-Organizer for the Workshop on Urban Scene Modeling: Where Vision Meets Photogrammetry and Graphics at CVPR.
Jeff Bigham is Workshop Co-Organizer for the VizWiz Grand Challenge Workshop at CVPR.
Qi Shan is Session Chair for CVPR.
Alex Colburn, Fartash Faghri, Hadi Pour Ansari, Mingze Xu, and Oncel Tuzel are Area Chairs for CVPR.
Amin Karimi Monsefi, Andrew Szot, Guandao Yang, Harsh Agrawal, Helisa Dhamo, Huangjie Zheng, Jack Langerman, Jiatao Gu, Liangchen Song, Michael Kirchhof, Marcin Eichner, Noam Elata, Pavan Kumar Anasosalu Vasu, Peter Fu, Raviteja Vemulapalli, Shaobo Fang, Rick Chang, Xiaoming Zhao, and Xudong Liu are Reviewers for CVPR.