About Me
I am a PhD student in the Computer Vision Group at the University of Bonn, supervised by Professor Dr. Jürgen Gall. My research is on video understanding that looks forward rather than backward: anticipating what happens next in a video and, increasingly, generating it. Current work centres on reinforcement-learning post-training for generative video and world models, and on how to give a large multimodal model a new capability from little data without losing what it already knows.
Before the PhD I was an Associate Researcher at the Intelligent Visual Analytics Lab (IVAL) at MBZUAI, working with Dr. Salman Khan on efficient video recognition, adapting vision-language models to video, and open-vocabulary spatio-temporal grounding. Several of those models have since been adopted as backbones and baselines in subsequent video research.
I did my master’s in Image Processing and Computer Vision (IPCV) on an Erasmus Mundus scholarship, with my thesis at EPFL’s CVLAB under Dr. Mathieu Salzmann and an internship at the Empathic Computing Lab under Dr. Mark Billinghurst. My undergraduate degree is in Electrical Engineering from Habib University in Karachi, Pakistan.
I’m glad to hear from prospective students and collaborators interested in any of the topics below.
Research Interests
My work follows one line: from perceiving what is in a video, to anticipating what happens next, to generating what could happen. Three things carry through every step: multimodal alignment, computational efficiency, and adapting foundation models without erasing what they already know.
- Anticipation and world models: long-horizon and open-vocabulary action anticipation; generative video and world models; unified models that both understand and generate
- Post-training foundation models: reinforcement-learning post-training (GRPO) as targeted property transfer; preserving prior capabilities; self-rewarding formulations that rule out reward hacking
- Spatio-temporal grounding and segmentation: open-vocabulary video grounding; memory-efficient long-video object segmentation; referring localization in vision-language models
- Efficient video architectures: state-space models for images and video; encoder-free video-language alignment
- Multimodal adaptation: prompt learning and regularized adaptation of vision-language models for zero-shot and open-vocabulary generalization
News
-
Jul 2026
Our paper “Open-Vocabulary Long-term Action Anticipation” is accepted in ECCV 2026.
-
Mar 2026
Our paper “RedSage: A Cybersecurity Generalist LLM” is accepted in ICLR 2026.
-
Jul 2025
Our paper “MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation” is accepted in ICCV 2025.
-
Feb 2025
Three of our papers (Video-Panda, GroupMamba, and STING-BEE) have been accepted in CVPR 2025.
-
Dec 2024
Our paper “Efficient Video Object Segmentation via Modulated Cross-Attention Memory” is accepted in WACV 2025.
-
Oct 2024
New preprint released titled “Distillation-free Scaling of Large SSMs for Images and Videos”.
-
Mar 2024
Our paper “VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding” is accepted in CVPR 2024.
-
Feb 2024
My student Muhammad Zain Yousuf’s bachelor thesis titled “AR-VPT: Simple Auto-Regressive Prompts for Adapting Frozen ViTs to Videos” is accepted in VISAPP 2024.
-
Jan 2024
I started a PhD at the University of Bonn, Germany working on Long-Term Multimodal Video Understanding, under the supervision of Professor Dr. Juergen Gall.
-
Oct 2023
Our paper “Hardware Resilience Properties of Text-Guided Image Classifiers” is accepted in NeurIPS 2023.
-
Aug 2023
Our paper “Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition” is accepted in ICCV 2023.
-
Aug 2023
Our paper “Self-regulating Prompts: Foundational Model Adaptation without Forgetting” is accepted in ICCV 2023.
-
Jun 2023
Our paper “Toward Automatic Typography Analysis: Serif Classification and Font Similarities” is accepted in the Journal of Data Mining in Digital Humanities (JDMDH).
-
Mar 2023
Our paper “Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting” is accepted in CVPR 2023.
-
Jun 2022
Our paper “Using Facial Micro-Expressions in Combination With EEG and Physiological Signals for Emotion Recognition” is accepted in Frontiers in Psychology.
-
Apr 2022
I started working as a researcher at MBZUAI. I was supervised by Dr. Salman Khan, working on multimodal video understanding.
-
Jul 2021
I was accepted in the ETH Robotics Summer School and Symposium.
-
Jun 2021
I defended my master’s thesis and graduated from the IPCV master’s program.
-
May 2021
Our paper on synthetic data for object detection is accepted to CVPR 2021 CV4Animals workshop.
-
Feb 2021
I started my master’s thesis in the CVLAB at EPFL supervised by Dr. Mathieu Salzmann. I worked on automated typography analysis on figurative content.
-
Jul 2020
I started a remote research internship at the Empathic Computing Lab supervised by Dr. Mark Billinghurst.
-
Sep 2019
I started my master’s degree in Image Processing and Computer Vision (IPCV) funded by the Erasmus Mundus Joint Master’s Degree (EMJMD) scholarship program.
-
Jun 2019
I completed my undergraduate degree in Electrical Engineering with a Minor in computer science. Graduated first in class with the Dean’s Medal.
Publications
2026
-
ECCV Open-Vocabulary Long-term Action Anticipation
ECCV, 2026
Paper Project page@inproceedings{wasim2026openVocabAnt, title={Open-Vocabulary Long-term Action Anticipation}, author={Syed Talal Wasim and Jinhui Yi and Hamid Suleman and Ahmad Javed and Yanan Luo and Muzammal Naseer and Juergen Gall}, booktitle={ECCV}, year={2026} } -
ICLR RedSage: A Cybersecurity Generalist LLM
ICLR, 2026
Paper Project page@inproceedings{suryanto2026redsage, title={RedSage: A Cybersecurity Generalist LLM}, author={Naufal Suryanto and Muzammal Naseer and Pengfei Li and Syed Talal Wasim and Jinhui Yi and Juergen Gall and Paolo Ceravolo and Ernesto Damiani}, booktitle={ICLR}, year={2026} } 2025
-
ICCV MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation
ICCV, 2025
Paper Project page@inproceedings{wasim2025mixant, title={MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation}, author={Syed Talal Wasim and Hamid Suleman and Olga Zatsarynna and Muzammal Naseer and Juergen Gall}, booktitle={ICCV}, year={2025} } -
CVPR Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
CVPR, 2025
Paper Project page@inproceedings{yi2025vpanda, title={Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models}, author={Jinhui Yi* and Syed Talal Wasim* and Yanan Luo* and Muzammal Naseer and Juergen Gall}, booktitle={CVPR}, year={2025} } -
CVPR GroupMamba: Parameter-Efficient and Accurate Group Visual State Space Model
CVPR, 2025
-
CVPR STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection
CVPR, 2025
Paper Code@inproceedings{velayudhan2024sting, title={STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection}, author={Divya Velayudhan and Abdelfatah Ahmed and Mohamad Alansari and Neha Gour and Abderaouf Behouch and Taimur Hassan and Syed Talal Wasim and Nabil Maalej and Muzammal Naseer and Juergen Gall and Mohammed Bennamoun and Ernesto Damiani and Naoufel Werghi}, booktitle={CVPR}, year={2025} } -
WACV Efficient Video Object Segmentation via Modulated Cross-Attention Memory
WACV, 2025
-
IJCV Distillation-free Scaling of Large State-Space Models for Images and Videos
IJCV
2024
-
CVPR Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
CVPR, 2024
-
VISAPP AR-VPT: Simple Auto-Regressive Prompts for Adapting Frozen ViTs to Videos
VISAPP, 2024
2023
-
NeurIPS Hardware Resilience Properties of Text-Guided Image Classifiers
NeurIPS, 2023
-
ICCV Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition
ICCV, 2023
-
ICCV Self-regulating Prompts: Foundational Model Adaptation without Forgetting
ICCV, 2023
-
CVPR Vita-CLIP: Video and text adaptive CLIP via Multimodal Prompting
CVPR, 2023
-
JDMDH Toward Automatic Typography Analysis: Serif Classification and Font Similarities
Journal of Data Mining in Digital Humanities, 2023
Paper Code@article{wasim2023gest, title={Toward automatic typography analysis: serif classification and font similarities}, author={Syed Talal Wasim and Romain Collaud and Lara Défayes and Nicolas Henchoz and Mathieu Salzmann and Delphine Ribes}, journal={Journal of Data Mining in Digital Humanities (JDMDH)}, year={2023} } 2022
-
Frontiers Using Facial Micro-Expressions in Combination With EEG and Physiological Signals for Emotion Recognition
Frontiers in Psychology, 2022
Paper@article{wasim2022ecl, title={Using facial micro-expressions in combination with EEG and physiological signals for emotion recognition}, author={Nastaran Saffaryazdi and Syed Talal Wasim and Kuldeep Dileep and Alireza Farrokhi Nia and Suranga Nanayakkara and Elizabeth Broadbent and Mark Billinghurst}, journal={Frontiers in Psychology}, year={2022} } 2021
-
CVPRW Sim-to-Real Transfer for Object Detection and Localization on Animals
CV4Animals Workshop, CVPR 2021
Poster@inproceedings{wasim2021cv4animals, title={Sim-to-Real Transfer for Object Detection and Localization on Animals}, author={Syed Talal Wasim and Syed N. Hasany and Kainat Abbasi and Huda Feroz and Anisa A. Ahmed and Mudasir H. Shaikh and Muhammad Farhan}, booktitle={CV4Animals CVPR Workshop}, year={2021} }
Services
Journal reviewer
- Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
- Transactions on Neural Networks and Learning Systems (TNNLS)
- Transactions on Image Processing (TIP)
- Transactions on Machine Learning Research (TMLR)
- International Journal of Computer Vision (IJCV)
- Pattern Recognition
Conference reviewer
- Computer vision (CVPR, ICCV, ECCV, WACV, ACCV)
- Artificial intelligence and machine learning (NeurIPS, ICLR, ICML, AAAI)
Project supervision
- Co-supervise undergraduate projects in computer vision at Habib University
- Co-supervise high-school students in Pakistan for the Intel International Science and Engineering Fair (ISEF)