Portrait of Rao Muhammad Anwer

Faculty / Computer Vision

Rao Muhammad Anwer

Assistant Professor of Computer Vision

Mohamed bin Zayed University of Artificial Intelligence

Research in visual recognition, multimodal learning, open-world perception, and efficient, robust deep learning for detailed scene understanding.

Citations9,100+
H-index41
Scientific papers130+
Researchers supervised30+

Biography

Rao Muhammad Anwer is an Assistant Professor of Computer Vision at the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI).

His research spans visual object recognition, multimodal learning, visual reasoning, open-world perception, and efficient and robust deep learning. He has published more than 130 scientific papers in leading AI and computer vision venues.

Before joining MBZUAI, he was a Research Scientist at the Inception Institute of Artificial Intelligence in Abu Dhabi. Earlier, he was a postdoctoral researcher at Aalto University under Finland’s Centre of Excellence in Computational Inference Research.

Assistant Professor · Jun 2020–present
Mohamed bin Zayed University of Artificial Intelligence, UAE
Research Scientist · Aug 2018–Jun 2020
Inception Institute of Artificial Intelligence, UAE
Postdoctoral Researcher · Feb 2014–Jul 2018
Aalto University, Finland
PhD · Computer Vision · 2008–2013
Autonomous University of Barcelona, Spain
MSc · Computer Vision and Artificial Intelligence · 2007–2008
Autonomous University of Barcelona, Spain
MSc · Intelligent Systems Design · 2005–2007
Chalmers University of Technology, Sweden

News

Recent research publications, funded projects, and intellectual property highlights.

  1. New papers at AAAI, CVPR, and ECCV 2026

    “Think Before You Segment” appears at AAAI; “Thinking Beyond Labels” at CVPR; and “Paying More Attention to Visual Tokens” and “Ground3D-LMM” at ECCV.

  2. TerraFM and Agent-X at ICLR 2026

    Two projects on multisensor Earth-observation foundation models and vision-centric agentic reasoning appear at ICLR 2026.

  3. New Arabic and agriculture benchmarks

    DuwatBench appears in ACL Findings, while ARB and AgriChain appear at LREC 2026.

  4. Medical vision research at WACV and MIDL

    MedROV appears at WACV 2026, and UniNuc appears at Medical Imaging with Deep Learning 2026.

  5. EnergyGPT collaboration

    A US$400K collaborative project with MBZUAI IFM, Inception, and ADNOC is developing a long-context multimodal model for well-log interpretation and petrophysical reasoning.

  6. Four U.S. patents granted

    Patents cover video instance segmentation, attention-aware person search, the EdgeNeXt mobile-vision architecture, and open-world semi-supervised satellite object detection.

Research

Developing visual systems that recognize, understand, and reason about complex real-world scenes.

  1. Large Multimodal Models and Vision-Language Reasoning

    Large multimodal models, grounded dialogue, visual agents, and step-by-step reasoning across images, video, and language.

    GLaMM · Agent-X · LLaVA-o1 · PALO
  2. Pixel-Level and Open-Vocabulary Scene Understanding

    Pixel-level grounding, open-vocabulary detection and segmentation, person search, and detailed 2D and 3D scene understanding.

    Open-YOLO 3D · OpenSeg-R · SipMask · PSTR
  3. Efficient and Robust Computer Vision

    Resource-efficient architectures, robust representation learning, limited supervision, domain generalization, and open-world learning.

    EdgeNeXt · Energy-based Latent Aligner · D2Det
  4. Video Understanding, Segmentation, and Action Recognition

    Video instance segmentation, tracking, composed retrieval, action recognition, and temporal visual reasoning.

    OW-VISFormer · CoVR · SipMask
  5. Domain Foundation Models: Agriculture, Healthcare, Arabic and Culturally-Aware AI

    Foundation models and multimodal systems tailored to specialist knowledge, languages, cultural contexts, and high-impact domains.

    AgroGPT · BiMediX · AIN · CAMEL-Bench · TerraFM

Selected projects

Representative research projects spanning multimodal reasoning, detailed scene understanding, efficient vision, video, and domain foundation models. Paper, project, code, and model links are included where publicly available.

ICLR 2026 · Multimodal reasoning

Agent-X

A benchmark and framework for evaluating deep multimodal reasoning in vision-centric agentic tasks.

CVPR 2024 · Pixel grounding

GLaMM

A large multimodal model that connects natural-language responses with precise pixel-level visual grounding.

ICLR 2025 Oral · Open-vocabulary 3D

Open-YOLO 3D

Fast open-vocabulary 3D instance segmentation using class-agnostic proposals and efficient 2D detections.

ECCV 2022 · Efficient vision

EdgeNeXt

An efficiently amalgamated CNN–Transformer architecture designed for accurate visual recognition on mobile and edge devices.

IJCV 2025 · Video understanding

OW-VISFormer

Open-world video instance segmentation for recognizing known objects, identifying unknowns, and learning new categories over time.

WACV 2025 · Agriculture

AgroGPT

An expert-tuned vision-language model for fine-grained agricultural understanding and multimodal conversation.

EMNLP 2024 Findings · Healthcare

BiMediX

A bilingual Arabic–English medical mixture-of-experts model for healthcare question answering and multi-turn interaction.

2025 · Arabic multimodal AI

AIN

An inclusive Arabic–English large multimodal model covering visual understanding, cultural context, OCR, and specialist domains.

Funding & leadership

Principal and co-principal investigator on interdisciplinary AI programs, with academic service across leading computer vision and AI venues.

Principal Investigator · US$800K

Deciphering Brain Diseases by AI

MBZUAI–WIS research on quantifying subcellular organelle dynamics in control and diseased brain organoid models.

Co-Principal Investigator · 2025–2026 · US$400K

EnergyGPT

A long-context multimodal model for well-log interpretation and petrophysical reasoning, developed with IFM, Inception, and ADNOC.

Principal Investigator · 2024 · AED215K

WAM-GPT

MBZUAI Seed Fund collaboration with the Emirates News Agency.

Co-Principal Investigator · 2024 · US$20K

Google Research Award

A climate-change and sustainability-tailored Arabic large language model.

Selected academic service

Program Chair, ACM Multimedia Asia 2026 · Senior Program Committee, AAAI 2026 and AAAI 2027 · Reviewer/committee roles for CVPR, ICCV, ECCV, WACV, NeurIPS, ICLR, ICML, and ACL ARR · Lead Guest Editor, CVIU special issue · Workshop organizer for ICCV, NeurIPS, CVPR, and ICME workshops.

Programs & awards

MBZUAI Start-up Fund, 2020–present · Co-PI on two US$800K MBZUAI–WIS programs in embryo development and biomimetic low-resolution recognition · PicSOM team member for TRECVID 2014–2015 · Honourable Mention as part of the UAB team in the 2012 PASCAL VOC image-classification challenge.

Selected publications

2026 AAAI

Think Before You Segment: An Object-Aware Reasoning Agent for Referring Audio-Visual Segmentation

J. Zhou, Y. M. Zhou, M. Han, T. Wang, X. Chang, H. Cholakkal, R. M. Anwer

2026 CVPR

Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs

D. Demidov, M. Z. Zaheer, Z. Han, O. Thawakar, R. M. Anwer

2026 ECCV

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

S. Venkatraman, R. Thawkar, O. Thawakar, R. M. Anwer, et al.

2026 ECCV

Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

A. Harsh, Z. Han, J. Lahoud, Y. Liu, R. M. Anwer, et al.

2026 ICLR

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

M. S. Danish, M. A. Munir, S. R. A. Shah, et al., R. M. Anwer, et al.

2026 ICLR

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

T. Ashraf, A. Saqib, H. Ghani, et al., R. M. Anwer, S. Khan

2025 TPAMI

Foundation Models Defining a New Era in Vision: A Survey and Outlook

M. Awais, M. Naseer, S. Khan, R. M. Anwer, et al.

2025 ICLR

Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

M. E. A. Boudjoghra, A. Dai, J. Lahoud, H. Cholakkal, R. M. Anwer, et al.

2024 CVPR

GLaMM: Pixel Grounding Large Multimodal Model

H. Rasheed, M. Maaz, S. Shaji, A. Shaker, et al., R. M. Anwer, et al.

2022 ECCV

EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications

M. Maaz, A. Shaker, H. Cholakkal, S. Khan, R. M. Anwer, F. Khan

2020 ECCV

SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation

J. Cao, R. M. Anwer, H. Cholakkal, F. Khan, Y. Pang, L. Shao, M. Shah

Patents

Six granted US patents and four patent applications spanning video, detection, efficient architectures, generation, and medical AI.

  1. Granted · 2025 · US 12,499,573

    Video instance segmentation via multi-scale spatio-temporal split attention transformer

  2. Granted · 2025 · US 12,260,674

    Attention-aware relation mixer for person search

  3. Granted · 2025 · US 12,373,672

    Efficiently amalgamated CNN-transformer architecture for mobile vision applications

  4. Granted · 2025 · US 12,380,677

    Open-world semi-supervised satellite object detection

  5. Granted · 2023 · US 11,756,244

    Handwriting generation

  6. Granted · 2022 · US 11,244,188

    Dense and discriminative neural network architectures for object detection and instance segmentation

Teaching

Course design, coordination, and instruction across core AI, computer vision, deep learning, and large multimodal models at MBZUAI.

2021–present · Instructor

Visual Object Recognition and Detection

Course lecturer and designer; offered annually since 2021.

2026 · Examiner & Instructor

Advanced Topics in Large Multimodal Models

PhD elective covering visual grounding, hallucination, bias, open-vocabulary perception, and reasoning.

2025 · Examiner & Instructor

Selected Topics in Computer Vision

Designed from scratch; course and instructor evaluations both 5.0/5.0.

2022–2024 · Examiner & Instructor

Foundations of Artificial Intelligence

Co-developed for MBZUAI’s Master’s cohort across CV, ML, and NLP.

2022 · Instructor

Deep Learning

Co-developed as an elective for students across three departments.

2021 · Examiner & Instructor

Research Communication and Dissemination

Designed from scratch for MBZUAI’s PhD and MSc cohorts.

People

More than 30 postdoctoral researchers, graduate students, and visiting researchers supervised or co-supervised.

See all people →

Postdoctoral researchers

  • Jean Lahoud Jean Lahoud2020–present
  • Zongyan Han Zongyan Han2024–present
  • Jinxing Zhou Jinxing Zhou2025–present

PhD students

  • Omkar Thawakar Omkar Thawakar2023–2026
  • Sara Ghaboura Sara Ghaboura2024–2026

Master’s students

  • Ketan More Ketan More2025–2027
  • Ritesh Thawkar Ritesh Thawkar2025–2027
  • Abdulla Alshehhi 2025–2027

Research interns & visitors

  • Amandeep Kumar Amandeep KumarResearch Assistant · 2022–2024
  • Tajamul Ashraf Tajamul AshrafResearch Assistant · 2024–2025
  • Xinyu Yan Xinyu YanTianjin University · 2025–2026

Selected Collaborations and Previous Affiliations

Contact

Email rao.anwer@mbzuai.ac.ae
Address MBZUAI, Masdar City
Abu Dhabi, United Arab Emirates Download curriculum vitae ↓