Research Scientist · Computer Vision

Muhammad Akhtar Munir

PhD Computer Science

I am a Research Scientist in Computer Vision at MBZUAI, developing reliable multimodal AI systems with a focus on vision-language models, foundation models, model calibration, and tool-augmented agents.

My work covers model evaluation and adaptation, robust prediction under domain shift, and agentic visual reasoning, with applications in object detection, geospatial intelligence, and environmental modeling. My research has appeared at CVPR, ICCV, NeurIPS, ICLR, ECCV, and TPAMI, including Highlight papers and Best Paper awards.

Focus Reliable AI · Real-world visual intelligence · Geospatial AI
Research AI agents · Vision-language models · Foundation models
Published at CVPR · ICCV · NeurIPS · ICLR · ECCV · TPAMI

Research interests

My work is organized around three connected areas: building capable models, evaluating them rigorously, and making their confidence useful in practice.

01

Geospatial AI

Foundation models, multimodal benchmarks, and agentic models for satellite imagery, weather, and air-quality forecasting.

Remote sensing · Climate AI · Foundation models

02

Reliable multimodal learning

Evaluation and adaptation methods for vision-language models, with emphasis on calibration, uncertainty, and behavior under distribution shift.

Model calibration · Domain shift · VLM adaptation

03

Agentic models

Tool-augmented models that reason over visual evidence and can be evaluated on realistic, multi-step tasks.

Tool use · Visual reasoning · Benchmarking

Representative projects

A selection illustrating the progression from calibrated perception to multimodal and agentic Earth-intelligence models.

ThinkGeo evaluation framework
2026Geospatial agentsBest Paper · MONTI-CVPR

ThinkGeo

A benchmark for evaluating whether tool-augmented agents can solve complex, multi-step remote-sensing tasks using visual evidence and external tools.

TerraFM multisensor Earth-observation model
2026Foundation modelsICLR

TerraFM

A scalable foundation model that learns shared representations across heterogeneous Earth-observation sensors and downstream tasks.

GEOBench-VLM benchmark
2025Vision-language modelsHighlight · ICCV

GEOBench-VLM

A systematic benchmark for measuring the capabilities and limitations of general-purpose vision-language models on geospatial tasks.

Cal-DETR calibration results
2023Model calibrationNeurIPS

Cal-DETR

A calibrated detection transformer designed to improve the alignment between prediction confidence and object-detection performance.

Selected publications

Full author lists and links are retained below. For citation counts and the complete record, see Google Scholar ↗.

16 publications

2026

CVPR 2026

Towards Calibrating Prompt Tuning of Vision-Language Models

Ashshak Sharifdeen, Fahad Shamshad, Muhammad Akhtar Munir, Abhishek Basu, Mohamed Insaf Ismithdeen, Jeyapriyan Jeyamohan, Chathurika Sewwandi Silva, Karthik Nandakumar, Muhammad Haris Khan

PaperCode

2026

MONTI-CVPR · Best Paper

ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks

Akashah Shabbir*, Muhammad Akhtar Munir*, Akshay Dudhane*, Muhammad Umer Sheikh, Muhammad Haris Khan, Paolo Fraccaro, Juan Bernabe Moreno, Fahad Shahbaz Khan, Salman Khan · *Equal contribution

PaperProject

2026

ECCV 2026

OpenEarthAgent: A Unified Framework for Tool-Augmented Geospatial Agents

Akashah Shabbir, Muhammad Umer Sheikh, Muhammad Akhtar Munir, Hiyam Debary, Mustansar Fiaz, Muhammad Zaigham Zaheer, Paolo Fraccaro, Fahad Shahbaz Khan, Muhammad Haris Khan, Xiao Xiang Zhu, Salman Khan

PaperCode

2026

ICLR 2026

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

Muhamad Sohail Danish, Muhammad Akhtar Munir, Syed Roshaan Ali Shah, Muhammad Haris Khan, Rao Muhammad Anwer, Jorma Laaksonen, Fahad Shahbaz Khan, Salman Khan

PaperCode

2026

npj CleanAir

Synergistic Neural Forecasting of Air Pollution with Stochastic Sampling

Yohan Abeysinghe, Muhammad Akhtar Munir, Sanoojan Baliah, Ron Sarafian, Fahad Shahbaz Khan, Yinon Rudich, Salman Khan

PaperCode

2025

ICCV 2025 · Highlight

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

Muhamad Sohail Danish*, Muhammad Akhtar Munir*, Syed Roshaan Ali Shah, Kartik Kuckreja, Fahad Shahbaz Khan, Paolo Fraccaro, Alexandre Lacoste, Salman Khan · *Equal contribution

PaperCode

2025

TerraBytes-ICML · Best Paper

AirCast: Improving Air Pollution Forecasting Through Multi-Variable Data Alignment

Vishal Nedungadi, Muhammad Akhtar Munir, Marc Rußwurm, Ron Sarafian, Ioannis N. Athanasiadis, Yinon Rudich, Fahad Shahbaz Khan, Salman Khan

PaperCode

2025

CVPR 2025 · Highlight

O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models

Ashshak Sharifdeen, Muhammad Akhtar Munir, Sanoojan Baliah, Salman Khan, Muhammad Haris Khan

PaperCode

2025

CVPR 2025

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D. Watson, Levente J. Klein, Fahad Shahbaz Khan, Salman Khan

PaperCode

2024

CCAI-NeurIPS 2024

Efficient Localized Adaptation of Neural Weather Forecasting: A Case Study in the MENA Region

Muhammad Akhtar Munir, Fahad Shahbaz Khan, Salman Khan

PaperCode

2024

CVPR 2024

Improving Single Domain-Generalized Object Detection: A Focus on Diversification and Alignment

Sohail Danish, Muhammad Haris Khan, Muhammad Akhtar Munir, Saquib Sarfraz, Mohsen Ali

PaperCode

2023

NeurIPS 2023

Cal-DETR: Calibrated Detection Transformer

Muhammad Akhtar Munir, Salman Khan, Muhammad Haris Khan, Mohsen Ali, Fahad Shahbaz Khan

PaperCode

2023

IEEE TPAMI 2023

Domain Adaptive Object Detection via Balancing between Self-Training and Adversarial Learning

Muhammad Akhtar Munir, Muhammad Haris Khan, Muhammad Saquib Sarfraz, Mohsen Ali

PaperProject

2023

CVPR 2023

Bridging Precision and Confidence: A Train-Time Loss for Calibrating Object Detection

Muhammad Akhtar Munir, Muhammad Haris Khan, Salman Khan, Fahad Shahbaz Khan

PaperCode

2022

NeurIPS 2022

Towards Improving Calibration in Object Detection under Domain Shift

Muhammad Akhtar Munir, Muhammad Haris Khan, Muhammad Saquib Sarfraz, Mohsen Ali

PaperCode

2021

NeurIPS 2021

SSAL: Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection

Muhammad Akhtar Munir, Muhammad Haris Khan, Muhammad Saquib Sarfraz, Mohsen Ali

PaperProject

Collaborators and mentees

I have been fortunate to work with students and early-career researchers across remote sensing, multimodal learning, calibration, and climate AI.

Advisors and senior collaborators: Salman Khan, Muhammad Haris Khan, Saquib Sarfraz, Fahad Shahbaz Khan, and Mohsen Ali.

Muhammad Sohail Danish

Muhammad Sohail Danish

Remote-sensing benchmarks and foundation models

Akashah Shabbir

Akashah Shabbir

Agentic pipelines for remote sensing

Vishal Nedungadi

Vishal Nedungadi

Air-pollution forecasting and climate AI

Ashshak Sharifdeen

Ashshak Sharifdeen

Calibration for vision-language models

Yohan Abeysinghe

Yohan Abeysinghe

Generative methods for climate modeling

Syed Roshaan Ali Shah

Syed Roshaan Ali Shah

Remote-sensing benchmarks and foundation models

Muhammad Umer Sheikh

Muhammad Umer Sheikh

Tool-augmented geospatial agents

Research community

Selected recognition, reviewing, teaching, and conference activity.

Recognition

  • Best Paper — ICML TerraBytes Workshop, 2025
  • Best Paper — CVPR MONTI Workshop, 2026
  • Highlight Papers — CVPR and ICCV, 2025
  • Outstanding Reviewer — CVPR, 2025
  • NeurIPS Scholar Award — 2022 and 2023
  • CVPR DEI Award — 2023
  • UAE AI Award finalist — EarthDial, Scientific Research category, 2025

Professional service

  • Reviewer: CVPR, ICCV, NeurIPS, ICLR, ECCV, EMNLP, TPAMI, and IJCV
  • Guest lecturer: CV805 — Life-long Learning Agents for Vision, MBZUAI
  • Invited speaker: University of Central Florida
  • Conference presentations: CVPR, ICCV, and NeurIPS
  • Research fellowship: Information Technology University, Pakistan

Research conversations and collaboration

I welcome conversations about reliable multimodal AI, geospatial intelligence, vision-language models, and related research collaborations.