Cross-Modal Image-to-Image Translation: Techniques, Applications and Challenges (CMI2IT)

Co-located with 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)

July 5, 2026 Bangkok, Thailand

About the Workshop

This workshop focuses on cross-modal image-to-image translation, where one sensing modality is synthesized from another (e.g., visible–infrared, visible–SAR, RGB–depth, CT–MRI). It welcomes contributions across the full research pipeline, including model architectures, learning under limited or unpaired data, image quality assessment and enhancement, physics- and sensor-aware constraints, and studies of downstream tasks and applications enabled by cross-modal translation. Representative applications include remote sensing and Earth observation, autonomous driving and intelligent transportation, surveillance and security, robotics and unmanned systems, medical imaging, agriculture and environmental monitoring, industrial inspection, and night-vision enhancement.

Call for Papers

Cross-modal image-to-image translation aims to synthesize one sensing modality from another (e.g., visible–infrared, visible–SAR, RGB–depth, CT–MRI), enabling robust perception under adverse conditions and across heterogeneous sensors. Recent advances in diffusion models, GANs, and other generative frameworks have significantly improved the fidelity and diversity of generated images, yet important challenges remain: enforcing physical and sensor consistency, coping with limited or unpaired data, and meeting real-time and reliability requirements in safety-critical scenarios.

This workshop seeks original, unpublished contributions on all aspects of cross-modal image-to-image translation, spanning model design, learning paradigms, evaluation, and applications across different modalities and domains. We particularly welcome work that leverages modern generative models (e.g., diffusion models, GANs, autoregressive or flow-based models) for cross-modal translation in real-world settings. Topics of interest include, but are not limited to:

Paper Submission

We invite submissions of original research papers addressing but not limited to the topics as listed above. Submissions should adhere to the IEEE ICME 2026 formatting guidelines and will undergo a rigorous peer-review process. Accepted papers will be presented at the workshop and included in the workshop proceeding of the conference*. We also welcome submissions of demos, datasets, and position papers that contribute to the workshop's themes.

Papers have to be submitted via: https://cmt3.research.microsoft.com/IEEEICMEW2026.

Author information and submission instructions are available here: https://2026.ieeeicme.org/author-information-and-submission-instructions/

*All accepted papers need to pay the register fees.

Important Dates

Due to various requests, the submission deadline has been extended.

Program

Date and time: July 5, 2026, 2:00 PM – 5:00 PM

Location: Thai Chakkraphat 1

Session Title Speaker(s) / Author(s) Time Duration
Opening Opening Remarks Workshop Chair(s) 2:00–2:05 PM 5 min
Keynote Self-Perception Network Structure Optimization Prof. Jiao Shi (Northwestern Polytechnical University) 2:05–2:45 PM 40 min
Keynote From Metameric Dilemma to Trustworthy RGB-to-HSI Dr. Xingxing Yang (Hong Kong Baptist University) 2:45–3:25 PM 40 min
Break (3:25–3:30 PM, 5 min)
Oral Presentations
1 Multi-modal Conditioned Bridge Diffusion for Pathology Image Synthesis and Augmentation Yanan Zhang; Xiangzhi Bai 3:30–3:45 PM 15 min
2 SCISR-TTA: Test Time Adaptation for Screen Content Image Super Resolution Ajeet Kumar Verma; Vinit Jakhetiya; Badri N Subudhi 3:45–4:00 PM 15 min
3 Diverse Optical-to-SAR Image Translation with Enhanced Feature Disentanglement Zixiang Ye; Zonghao Han; Kun Hu; Yongping Zhai; Shaohui Mei 4:00–4:15 PM 15 min
4 Visible-to-Infrared Image Translation via Hierarchical Semantic Embeddings Fangqi Shi; Zonghao Han; Zixiang Ye; Zhengtao Chen; Kun Hu; Shaohui Mei 4:15–4:30 PM 15 min
5 AD-Relight: Training-Free Banner Relighting via Illumination Translation with Diffusion Priors Rameshwar Mishra; A V Subramanyam 4:30–4:45 PM 15 min
6 CATS-Diff: Channel-Adaptive Timestep Selection for Efficient Diffusion Model Training Muhammad Ridha Agam; Mas Nurul Achmadiah; Hui-Kai Su; Chi-Chia Sun; Jun-Wei Hsieh 4:45–5:00 PM 15 min

Organizers

Prof. Shaohui Mei
Prof. Shaohui Mei

JiaoNorthwestern Polytechnical University, China

Prof. Ying Fu
Prof. Ying Fu

Beijing Institute of Technology, China

Dr. Kun Hu
Dr. Kun Hu

Edith Cowan University, Australia

Prof. Zhengxia Zou
Prof. Zhengxia Zou

Beihang University, China

Prof. Jiao Shi
Prof. Jiao Shi

Northwestern Polytechnical University, China

Prof. Zhiyong Wang
Prof. Zhiyong Wang

The University of Sydney, Australia

Contact Us

If you have any questions about the workshop, please feel free to reach out to us:

Email: meish@nwpu.edu.cn