About Me

I am a PhD student in the Computer Vision Group at the University of Bonn, supervised by Professor Dr. Jürgen Gall. My research is on video understanding that looks forward rather than backward: anticipating what happens next in a video and, increasingly, generating it. Current work centres on reinforcement-learning post-training for generative video and world models, and on how to give a large multimodal model a new capability from little data without losing what it already knows.

Before the PhD I was an Associate Researcher at the Intelligent Visual Analytics Lab (IVAL) at MBZUAI, working with Dr. Salman Khan on efficient video recognition, adapting vision-language models to video, and open-vocabulary spatio-temporal grounding. Several of those models have since been adopted as backbones and baselines in subsequent video research.

I did my master’s in Image Processing and Computer Vision (IPCV) on an Erasmus Mundus scholarship, with my thesis at EPFL’s CVLAB under Dr. Mathieu Salzmann and an internship at the Empathic Computing Lab under Dr. Mark Billinghurst. My undergraduate degree is in Electrical Engineering from Habib University in Karachi, Pakistan.

I’m glad to hear from prospective students and collaborators interested in any of the topics below.

Research Interests

My work follows one line: from perceiving what is in a video, to anticipating what happens next, to generating what could happen. Three things carry through every step: multimodal alignment, computational efficiency, and adapting foundation models without erasing what they already know.

  • Anticipation and world models: long-horizon and open-vocabulary action anticipation; generative video and world models; unified models that both understand and generate
  • Post-training foundation models: reinforcement-learning post-training (GRPO) as targeted property transfer; preserving prior capabilities; self-rewarding formulations that rule out reward hacking
  • Spatio-temporal grounding and segmentation: open-vocabulary video grounding; memory-efficient long-video object segmentation; referring localization in vision-language models
  • Efficient video architectures: state-space models for images and video; encoder-free video-language alignment
  • Multimodal adaptation: prompt learning and regularized adaptation of vision-language models for zero-shot and open-vocabulary generalization

News

  • Jul 2026
    Our paper “Open-Vocabulary Long-term Action Anticipation” is accepted in ECCV 2026.
  • Mar 2026
    Our paper “RedSage: A Cybersecurity Generalist LLM” is accepted in ICLR 2026.
  • Jul 2025
    Our paper “MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation” is accepted in ICCV 2025.
  • Feb 2025
    Three of our papers (Video-Panda, GroupMamba, and STING-BEE) have been accepted in CVPR 2025.
  • Dec 2024
    Our paper “Efficient Video Object Segmentation via Modulated Cross-Attention Memory” is accepted in WACV 2025.

Publications

  • ECCV

    Open-Vocabulary Long-term Action Anticipation

    Syed Talal Wasim, Jinhui Yi, Hamid Suleman, Ahmad Javed, Yanan Luo, Muzammal Naseer and Juergen Gall

    ECCV, 2026

  • ICLR

    RedSage: A Cybersecurity Generalist LLM

    Naufal Suryanto, Muzammal Naseer, Pengfei Li, Syed Talal Wasim, Jinhui Yi, Juergen Gall, Paolo Ceravolo, Ernesto Damiani

    ICLR, 2026

  • ICCV

    MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation

    Syed Talal Wasim, Hamid Suleman, Olga Zatsarynna, Muzammal Naseer and Juergen Gall

    ICCV, 2025

  • CVPR

    Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models

    Jinhui Yi*, Syed Talal Wasim*, Yanan Luo*, Muzammal Naseer and Juergen Gall

    CVPR, 2025

  • CVPR

    GroupMamba: Parameter-Efficient and Accurate Group Visual State Space Model

    Abdelrahman Shaker, Syed Talal Wasim, Salman Khan, and Fahad Shahbaz Khan

    CVPR, 2025

  • CVPR

    STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection

    D. Velayudhan, A. Ahmed, M. Alansari, N. Gour, A. Behouch, T. Hassan, Syed Talal Wasim, N. Maalej, M. Naseer, J. Gall, M. Bennamoun, E. Damiani and N. Werghi

    CVPR, 2025

  • WACV

    Efficient Video Object Segmentation via Modulated Cross-Attention Memory

    Abdelrahman Shaker, Syed Talal Wasim, Martin Danelljan, Salman Khan, Ming-Hsuan Yang and Fahad Shahbaz Khan

    WACV, 2025

  • IJCV

    Distillation-free Scaling of Large State-Space Models for Images and Videos

    Hamid Suleman*, Syed Talal Wasim*, Muzammal Naseer and Juergen Gall

    IJCV

  • CVPR

    Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding

    Syed Talal Wasim, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang and Fahad Shahbaz Khan

    CVPR, 2024

  • VISAPP

    AR-VPT: Simple Auto-Regressive Prompts for Adapting Frozen ViTs to Videos

    Muhammad Zain Yousuf, Syed Talal Wasim, Syed Nouman Hasany and Muhammad Farhan

    VISAPP, 2024

  • NeurIPS

    Hardware Resilience Properties of Text-Guided Image Classifiers

    Syed Talal Wasim, Kabila Haile Soboka, Abdulrahman Mahmoud, Salman Khan, David Brooks and Gu-Yeon Wei

    NeurIPS, 2023

  • ICCV

    Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition

    Syed Talal Wasim*, Muhammad Uzair Khattak*, Muzammal Naseer, Salman Khan, Mubarak Shah and Fahad Shahbaz Khan

    ICCV, 2023

  • ICCV

    Self-regulating Prompts: Foundational Model Adaptation without Forgetting

    Muhammad Uzair Khattak*, Syed Talal Wasim*, Muzammal Naseer, Salman Khan, Ming-Hsuan Yang and Fahad Shahbaz Khan

    ICCV, 2023

  • CVPR

    Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting

    Syed Talal Wasim, Muzammal Naseer, Salman Khan, Fahad Shahbaz Khan and Mubarak Shah

    CVPR, 2023

  • JDMDH

    Toward Automatic Typography Analysis: Serif Classification and Font Similarities

    Syed Talal Wasim, Romain Collaud, Lara Défayes, Nicolas Henchoz, Mathieu Salzmann and Delphine Ribes

    Journal of Data Mining in Digital Humanities, 2023

  • Frontiers

    Using Facial Micro-Expressions in Combination With EEG and Physiological Signals for Emotion Recognition

    Nastaran Saffaryazdi, Syed Talal Wasim, Kuldeep Dileep, Alireza Farrokhi Nia, Suranga Nanayakkara, Elizabeth Broadbent and Mark Billinghurst

    Frontiers in Psychology, 2022

  • CVPRW

    Sim-to-Real Transfer for Object Detection and Localization on Animals

    Syed Talal Wasim, Syed N. Hasany, Kainat Abbasi, Huda Feroz, Anisa A. Ahmed, Mudasir H. Shaikh and Muhammad Farhan

    CV4Animals Workshop, CVPR 2021

Services

Journal reviewer

  • Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
  • Transactions on Neural Networks and Learning Systems (TNNLS)
  • Transactions on Image Processing (TIP)
  • Transactions on Machine Learning Research (TMLR)
  • International Journal of Computer Vision (IJCV)
  • Pattern Recognition

Conference reviewer

  • Computer vision (CVPR, ICCV, ECCV, WACV, ACCV)
  • Artificial intelligence and machine learning (NeurIPS, ICLR, ICML, AAAI)

Project supervision

  • Co-supervise undergraduate projects in computer vision at Habib University
  • Co-supervise high-school students in Pakistan for the Intel International Science and Engineering Fair (ISEF)

BibTeX