{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:59:22Z","timestamp":1778083162712,"version":"3.51.4"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T00:00:00Z","timestamp":1701734400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,12,5]]},"abstract":"<jats:p>\n            DSLR cameras can achieve multiple zoom levels via shifting lens distances or swapping lens types. However, these techniques are not possible on smart-phone devices due to space constraints. Most smartphone manufacturers adopt a hybrid zoom system: commonly a Wide (\n            <jats:bold>W<\/jats:bold>\n            ) camera at a low zoom level and a Telephoto (\n            <jats:bold>T<\/jats:bold>\n            ) camera at a high zoom level. To simulate zoom levels between\n            <jats:bold>W<\/jats:bold>\n            and\n            <jats:bold>T<\/jats:bold>\n            , these systems crop and digitally upsample images from\n            <jats:bold>W<\/jats:bold>\n            , leading to significant detail loss. In this paper, we propose an efficient system for hybrid zoom super-resolution on mobile devices, which captures a synchronous pair of W and\n            <jats:bold>T<\/jats:bold>\n            shots and leverages machine learning models to align and transfer details from\n            <jats:bold>T<\/jats:bold>\n            to\n            <jats:bold>W.<\/jats:bold>\n            We further develop an adaptive blending method that accounts for depth-of-field mismatches, scene occlusion, flow uncertainty, and alignment errors. To minimize the domain gap, we design a dual-phone camera rig to capture real-world inputs and ground-truths for supervised training. Our method generates a 12-megapixel image in 500ms on a mobile platform and compares favorably against state-of-the-art methods under extensive evaluation on real-world scenarios.\n          <\/jats:p>","DOI":"10.1145\/3618362","type":"journal-article","created":{"date-parts":[[2023,12,5]],"date-time":"2023-12-05T10:20:48Z","timestamp":1701771648000},"page":"1-12","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Efficient Hybrid Zoom Using Camera Fusion on Mobile Phones"],"prefix":"10.1145","volume":"42","author":[{"given":"Xiaotong","family":"Wu","sequence":"first","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei-Sheng","family":"Lai","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yichang","family":"Shih","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Charles","family":"Herrmann","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael","family":"Krainin","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deqing","family":"Sun","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chia-Kai","family":"Liang","sequence":"additional","affiliation":[{"name":"Google, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,12,5]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-007-0041-4"},{"key":"e_1_2_2_2_1","volume-title":"Wireless software synchronization of multiple distributed cameras","author":"Ansari Sameer","unstructured":"Sameer Ansari, Neal Wadhwa, Rahul Garg, and Jiawen Chen. 2019. Wireless software synchronization of multiple distributed cameras. In ICCP. IEEE, Tokyo, Japan, 1--9."},{"key":"e_1_2_2_3_1","volume-title":"GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution","author":"Chan Kelvin C.K.","year":"2021","unstructured":"Kelvin C.K. Chan, Xintao Wang, Xiangyu Xu, Jinwei Gu, and Chen Change Loy. 2021. GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution. In CVPR. IEEE, Virtual\/Online, 14245--14254."},{"key":"e_1_2_2_4_1","unstructured":"Ferenc Huszar Jose Caballero Andrew Cunningham Alejandro Acosta Andrew Aitken Alykhan Tejani Johannes Totz Zehan Wang Wenzhe Shi Christian Ledig Lucas Theis. 2017. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. In CVPR."},{"key":"e_1_2_2_5_1","unstructured":"Xiaodong Cun and Chi-Man Pun. 2020. Defocus blur detection via depth distillation. In ECCV."},{"key":"e_1_2_2_6_1","volume-title":"Kaiming He, and Xiaoou Tang.","author":"Dong Chao","year":"2014","unstructured":"Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. 2014. Learning a deep convolutional network for image super-resolution. In ECCV."},{"key":"e_1_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Jochen Gast and Stefan Roth. 2018. Lightweight probabilistic deep networks. In ICCV.","DOI":"10.1109\/CVPR.2018.00355"},{"key":"e_1_2_2_8_1","unstructured":"Jinjin Gu Yujun Shen and Bolei Zhou. 2020. Image processing using multi-code GAN prior. In CVPR."},{"key":"e_1_2_2_9_1","volume-title":"Burst photography for high dynamic range and low-light imaging on mobile cameras. ACM TOG","author":"Hasinoff Samuel W","year":"2016","unstructured":"Samuel W Hasinoff, Dillon Sharlet, Ryan Geiss, Andrew Adams, Jonathan T Barron, Florian Kainz, Jiawen Chen, and Marc Levoy. 2016. Burst photography for high dynamic range and low-light imaging on mobile cameras. ACM TOG (2016)."},{"key":"e_1_2_2_10_1","unstructured":"Jingwen He Wu Shi Kai Chen Lean Fu and Chao Dong. 2022. GCFSR: a Generative and Controllable Face Super Resolution Method Without Facial and GAN Priors. In CVPR."},{"key":"e_1_2_2_11_1","unstructured":"HonorMagic 2023. Honor Magic4 Ultimate Camera test. https:\/\/www.dxomark.com\/honor-magic4-ultimate-camera-test-retested\/. Accessed: 2023-03-07."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00584"},{"key":"e_1_2_2_13_1","volume-title":"Xintao Wang, Chen Change Loy, and Ziwei Liu.","author":"Jiang Yuming","year":"2021","unstructured":"Yuming Jiang, Kelvin CK Chan, Xintao Wang, Chen Change Loy, and Ziwei Liu. 2021. Robust Reference-based Super-Resolution via C2-Matching. In CVPR."},{"key":"e_1_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Justin Johnson Alexandre Alahi and Li Fei-Fei. 2016. Perceptual losses for real-time style transfer and super-resolution. In ECCV.","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_2_2_15_1","volume-title":"Jung Kwon Lee, and Kyoung Mu Lee","author":"Kim Jiwon","year":"2016","unstructured":"Jiwon Kim, Jung Kwon Lee, and Kyoung Mu Lee. 2016. Accurate image super-resolution using very deep convolutional networks. In CVPR."},{"key":"e_1_2_2_16_1","unstructured":"Wei-Sheng Lai Jia-Bin Huang Narendra Ahuja and Ming-Hsuan Yang. 2017. Deep Laplacian pyramid networks for fast and accurate super-resolution. In CVPR."},{"key":"e_1_2_2_17_1","volume-title":"Face deblurring using dual camera fusion on mobile phones. ACM TOG","author":"Lai Wei-Sheng","year":"2022","unstructured":"Wei-Sheng Lai, Yichang Shih, Lun-Cheng Chu, Xiaotong Wu, Sung-Fang Tsai, Michael Krainin, Deqing Sun, and Chia-Kai Liang. 2022. Face deblurring using dual camera fusion on mobile phones. ACM TOG (2022)."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01730"},{"key":"e_1_2_2_19_1","unstructured":"Junyong Lee Sungkil Lee Sunghyun Cho and Seungyong Lee. 2019. Deep defocus map estimation using domain adaptation. In CVPR."},{"key":"e_1_2_2_20_1","unstructured":"Liying Lu Wenbo Li Xin Tao Jiangbo Lu and Jiaya Jia. 2021. MASA-SR: Matching acceleration and spatial adaptation for reference-based image super-resolution. In CVPR."},{"key":"e_1_2_2_21_1","unstructured":"Roey Mechrez Itamar Talmi and Lihi Zelnik-Manor. 2018. The contextual loss for"},{"key":"e_1_2_2_22_1","unstructured":"image transformation with non-aligned data. In ECCV."},{"key":"e_1_2_2_23_1","volume-title":"Pulse: Self-supervised photo upsampling via latent space exploration of generative models. In CVPR.","author":"Menon Sachit","year":"2020","unstructured":"Sachit Menon, Alexandru Damian, Shijia Hu, Nikhil Ravi, and Cynthia Rudin. 2020. Pulse: Self-supervised photo upsampling via latent space exploration of generative models. In CVPR."},{"key":"e_1_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Marco Pesavento Marco Volino and Adrian Hilton. 2021. Attention-based multi-reference learning for image super-resolution. In CVPR.","DOI":"10.1109\/ICCV48922.2021.01443"},{"key":"e_1_2_2_25_1","volume-title":"Film: Frame interpolation for large motion. In ECCV.","author":"Reda Fitsum","year":"2022","unstructured":"Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, and Brian Curless. 2022. Film: Frame interpolation for large motion. In ECCV."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/38.946629"},{"key":"e_1_2_2_27_1","volume-title":"RAISR: rapid and accurate image super resolution","author":"Romano Yaniv","year":"2016","unstructured":"Yaniv Romano, John Isidoro, and Peyman Milanfar. 2016. RAISR: rapid and accurate image super resolution. IEEE TCI (2016)."},{"key":"e_1_2_2_28_1","volume-title":"U-net: Convolutional networks for biomedical image segmentation. In MICCAI.","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. In MICCAI."},{"key":"e_1_2_2_29_1","doi-asserted-by":"crossref","unstructured":"Edward Rosten and Tom Drummond. 2006. Machine learning for high-speed corner detection. In ECCV.","DOI":"10.1007\/11744023_34"},{"key":"e_1_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Gyumin Shim Jinsun Park and In So Kweon. 2020. Robust reference-based super-resolution with similarity-aware deformable convolution. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00845"},{"key":"e_1_2_2_31_1","unstructured":"Deqing Sun Charles Herrmann Fitsum Reda Michael Rubinstein David J. Fleet and William T Freeman. 2022. Disentangling Architecture and Training for Optical Flow. In ECCV."},{"key":"e_1_2_2_32_1","unstructured":"Deqing Sun Daniel Vlasic Charles Herrmann Varun Jampani Michael Krainin Huiwen Chang Ramin Zabih William T Freeman and Ce Liu. 2021. AutoFlow: Learning a Better Training Set for Optical Flow. In CVPR."},{"key":"e_1_2_2_33_1","volume-title":"Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR.","author":"Sun Deqing","year":"2018","unstructured":"Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. 2018. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In CVPR."},{"key":"e_1_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Libin Sun and James Hays. 2012. Super-resolution from internet-scale scene matching. In ICCP.","DOI":"10.1109\/ICCPhot.2012.6215221"},{"key":"e_1_2_2_35_1","volume-title":"Computer vision: algorithms and applications","author":"Szeliski Richard","unstructured":"Richard Szeliski. 2022. Computer vision: algorithms and applications. Springer Nature."},{"key":"e_1_2_2_36_1","volume-title":"Defusionnet: Defocus blur detection via recurrently fusing and refining multi-scale deep features. In CVPR.","author":"Tang Chang","year":"2019","unstructured":"Chang Tang, Xinzhong Zhu, Xinwang Liu, Lizhe Wang, and Albert Zomaya. 2019. Defusionnet: Defocus blur detection via recurrently fusing and refining multi-scale deep features. In CVPR."},{"key":"e_1_2_2_37_1","volume-title":"Raft: Recurrent all-pairs field transforms for optical flow. In ECCV.","author":"Teed Zachary","year":"2020","unstructured":"Zachary Teed and Jia Deng. 2020. Raft: Recurrent all-pairs field transforms for optical flow. In ECCV."},{"key":"e_1_2_2_38_1","unstructured":"Robert Triggs. 2023. All the new HUAWEI P40 camera technology explained. https:\/\/www.androidauthority.com\/huawei-p40-camera-explained-1097350\/. Accessed: 2023-03-07."},{"key":"e_1_2_2_39_1","volume-title":"Florian Kainz, and Janne Kontkanen.","author":"Trinidad Marc Comino","year":"2019","unstructured":"Marc Comino Trinidad, Ricardo Martin Brualla, Florian Kainz, and Janne Kontkanen. 2019. Multi-view image fusion. In CVPR."},{"key":"e_1_2_2_40_1","doi-asserted-by":"crossref","unstructured":"Tengfei Wang Jiaxin Xie Wenxiu Sun Qiong Yan and Qifeng Chen. 2021. Dual-camera super-resolution with aligned attention modules. In CVPR.","DOI":"10.1109\/ICCV48922.2021.00201"},{"key":"e_1_2_2_41_1","volume-title":"ESRGAN: Enhanced super-resolution generative adversarial networks. In ECCV.","author":"Wang Xintao","year":"2018","unstructured":"Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. 2018. ESRGAN: Enhanced super-resolution generative adversarial networks. In ECCV."},{"key":"e_1_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Yufei Wang Zhe Lin Xiaohui Shen Radomir Mech Gavin Miller and Garrison W Cottrell. 2016. Event-specific image importance. In CVPR.","DOI":"10.1109\/CVPR.2016.520"},{"key":"e_1_2_2_43_1","unstructured":"Pengxu Wei Ziwei Xie Hannan Lu Zongyuan Zhan Qixiang Ye Wangmeng Zuo and Liang Lin. 2020. Component divide-and-conquer for real-world image super-resolution. In ECCV."},{"key":"e_1_2_2_44_1","volume-title":"Handheld multi-frame super-resolution. ACM TOG","author":"Wronski Bartlomiej","year":"2019","unstructured":"Bartlomiej Wronski, Ignacio Garcia-Dorado, Manfred Ernst, Damien Kelly, Michael Krainin, Chia-Kai Liang, Marc Levoy, and Peyman Milanfar. 2019. Handheld multi-frame super-resolution. ACM TOG (2019)."},{"key":"e_1_2_2_45_1","doi-asserted-by":"crossref","unstructured":"Bin Xia Yapeng Tian Yucheng Hang Wenming Yang Qingmin Liao and Jie Zhou. 2022. Coarse-to-Fine Embedded PatchMatch and Multi-Scale Dynamic Aggregation for Reference-based Super-Resolution. In AAAI.","DOI":"10.1609\/aaai.v36i3.20180"},{"key":"e_1_2_2_46_1","unstructured":"Yanchun Xie Jimin Xiao Mingjie Sun Chao Yao and Kaizhu Huang. 2020. Feature representation matters: End-to-end learning for reference-based image super-resolution. In ECCV."},{"key":"e_1_2_2_47_1","unstructured":"Shumian Xin Neal Wadhwa Tianfan Xue Jonathan T Barron Pratul P Srinivasan Jiawen Chen Ioannis Gkioulekas and Rahul Garg. 2021. Defocus map estimation and deblurring from a single dual-pixel image. In ICCV."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00881"},{"key":"e_1_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Fuzhi Yang Huan Yang Jianlong Fu Hongtao Lu and Baining Guo. 2020. Learning texture transformer network for image super-resolution. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00583"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00388"},{"key":"e_1_2_2_51_1","doi-asserted-by":"crossref","unstructured":"Xindong Zhang Hui Zeng Shi Guo and Lei Zhang. 2022b. Efficient Long-Range Attention Network for Image Super-resolution. In ECCV.","DOI":"10.1007\/978-3-031-19790-1_39"},{"key":"e_1_2_2_52_1","doi-asserted-by":"crossref","unstructured":"Yulun Zhang Kunpeng Li Kai Li Lichen Wang Bineng Zhong and Yun Fu. 2018. Image super-resolution using very deep residual channel attention networks. In ECCV.","DOI":"10.1007\/978-3-030-01234-2_18"},{"key":"e_1_2_2_53_1","doi-asserted-by":"crossref","unstructured":"Zhilu Zhang Ruohao Wang Hongzhi Zhang Yunjin Chen and Wangmeng Zuo. 2022a. Self-Supervised Learning for Real-World Super-Resolution from Dual Zoomed Observations. In ECCV.","DOI":"10.1007\/978-3-031-19797-0_35"},{"key":"e_1_2_2_54_1","doi-asserted-by":"crossref","unstructured":"Zhifei Zhang Zhaowen Wang Zhe Lin and Hairong Qi. 2019b. Image super-resolution by neural texture transfer. In CVPR.","DOI":"10.1109\/CVPR.2019.00817"},{"key":"e_1_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Wenda Zhao Bowen Zheng Qiuhua Lin and Huchuan Lu. 2019. Enhancing diversity of defocus blur detectors via cross-ensemble network. In CVPR.","DOI":"10.1109\/CVPR.2019.00911"},{"key":"e_1_2_2_56_1","volume-title":"Crossnet: An end-to-end reference-based super resolution network using cross-scale warping. In ECCV.","author":"Zheng Haitian","year":"2018","unstructured":"Haitian Zheng, Mengqi Ji, Haoqian Wang, Yebin Liu, and Lu Fang. 2018. Crossnet: An end-to-end reference-based super resolution network using cross-scale warping. In ECCV."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618362","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3618362","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T10:45:38Z","timestamp":1755773138000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3618362"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,5]]},"references-count":56,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,12,5]]}},"alternative-id":["10.1145\/3618362"],"URL":"https:\/\/doi.org\/10.1145\/3618362","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,5]]},"assertion":[{"value":"2023-12-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}