Ali Vosoughi, multimodal AI researcher at the University of Rochester
Ali Vosoughi
University of Rochester | PhD research complete
Egocentric and multimodal AI. Scene understanding, vision-language agents, and evaluation for real-time, on-device systems.
Headsets, robots, and vehicles share one problem: a first-person stream from a moving platform that must be understood and acted on within a latency budget. Ali Vosoughi builds multimodal AI for this problem and the evaluation methods that decide whether such systems ship. Over seven years of AI and ML research, his work spans egocentric perception, visual agents, video understanding, conversational speech agents, world models, and audiovisual learning, across four industry labs, Apple, Microsoft Research, Smule, and Bosch AI Research, and federally supported programs including DARPA, NIH, NSF, and NNSA. His methods include post-training of multimodal models with SFT, DPO, and reinforcement learning, and statistically rigorous evaluation with significance testing, bootstrap confidence intervals, and human-in-the-loop protocols, built in Python and PyTorch.
AAAI 2026 Best Demonstration Award Runner-up co-inventor, US patent application US 2025/0124292 A1 COLM Β· NeurIPS Β· NAACL Β· ACM MM Β· ICASSP
Egocentric Perception Scene Understanding Multimodal Agents Video Understanding with LLMs World Models Audiovisual AI Speech and Conversational AI Spatial Audio AR and Autonomous Systems Efficient and On-Device ML Medical and Scientific Imaging AI
πŸ“§ ali.vosoughi@rochester.edu
πŸ“ CS Department, Wegmans Hall 3211
🍎 Apple
Machine Learning Intern
Agentic Multimodal AI
🎡 Smule AI
Research Scientist Intern
Spatial Generation
🏒 Microsoft Research
Research Intern
Video/Audio LLM and Video Understanding
πŸš— Bosch AI Research
Research Intern
LLM and Counterfactual Learning
πŸ›‘οΈ DARPA PTG
Graduate Researcher
Autonomous Multimodal Perception and AR
πŸŽ“ University of Rochester
Graduate Research Assistant, 2019 to present
Multimodal AI across DARPA, NIH, NSF, and NNSA programs
πŸ†
AAAI 2026 Best Demonstration Award Runner-up
Caption Anything in Video (Spatiotemporal Multimodal Prompting)
πŸ“Ή
Video Understanding with LLMs
Comprehensive survey with 241+ citations (IEEE TCSVT 2025)
πŸ”¬
PW-VQA
Causal debiasing for visual question answering with 50+ citations (IEEE TMM 2024)
πŸ†
First counterfactual audio methods
ICASSP 2024 + US patent application US 2025/0124292 A1 (co-inventor, published Jan 2025)
πŸ”Š
PromptReverb
First text-to-spatial generation at 48kHz (ICASSP 2026)
🎬
AVVA
Unified audiovisual foundation model with LLM curation (EUSIPCO 2025)
🀝
Autonomous multimodal copilot
Real-time audiovisual AR demonstrations (DARPA)
πŸ“Š
VERIFY benchmark
Stage-aware evaluation of multimodal reasoning fidelity (COLM 2026)
🧠
Video LMM Post-Training
Deep dive into video reasoning with large multimodal models
πŸ“¦
AVE-2 Dataset
Open audiovisual benchmark for cross-modal event understanding

Recent News & Updates

08/2026
πŸ“„ SPIE Medical Imaging 2027: four papers submitted on medical vision-language evaluation and causal inference methods
06/2026
πŸ“„ IΒ²G released in the IEEE/CVF CVPR 2026 Workshop Proceedings: Generating Instructional Illustrations via Text-Conditioned Diffusion
05/2026
πŸ“„ ICASSP 2026 paper accepted as oral presentation: PromptReverb (Text-to-Spatial-Audio Generation at 48kHz)
12/2025
πŸ“„ NeurIPS 2025 paper accepted: MMPerspective (Multimodal LLM Reasoning, Video and Visual Perception)
09/2025
βœ… Completed research internship at Smule AI (Spatial Audio Generation and Synthesis)
06/2025
🎡 Started research internship at Smule AI (Spatial Audio Generation and Immersive Computing)
10/2024
🎀 Presented at SANE 2024 at Google, Cambridge, MA (Audio Understanding, Video LLMs, and Spatial Audio)
10/2024
πŸ“„ ACM Multimedia 2024: EAGLE (Egocentric Video Understanding and Language Generation)
08/2024
πŸ’Ό Research presentation at Microsoft Research, Seattle (Audiovisual LLM, Video and Audio Understanding)
03/2024
πŸ“„ NAACL 2024: OSCaR (Video Object State Captioning, Autonomous Video Perception)
02/2024
πŸ“„ IEEE Transactions on Multimedia 2024: PW-VQA (Causal Visual Question Answering, Video Reasoning)
08/2023
🎯 Two ICCV 2023 papers accepted (Audiovisual Sound Separation and Autonomous AR Perception System)

Publications

EAGLE egocentric video understanding engine with DARPA PTG, ACM Multimedia 2024

EAGLE: Egocentric AGgregated Language-video Engine
ACM International Conference on Multimedia (ACM MM) 2024
[Paper][DARPA PTG Demo]

VERIFY benchmark for multimodal reasoning evaluation, COLM 2026

VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
Conference on Language Modeling (COLM) 2026
[Paper][Website][πŸ€— Hugging Face]

OSCaR object state captioning for accessible video understanding, NAACL 2024

OSCaR: Object State Captioning and State Change Representation
North American Chapter of the Association for Computational Linguistics (NAACL) 2024
Supported in part by an NIH R01 on accessible video description for blind and low-vision users (R01EY034562)
[Paper][Code]

PW-VQA causal visual question answering by Ali Vosoughi, IEEE Transactions on Multimedia 2024

PW-VQA: Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA
IEEE Transactions on Multimedia (TMM) 2024

[Paper][Code][Website]

PromptReverb text to spatial audio generation at 48 kHz by Ali Vosoughi, ICASSP 2026 oral

PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2026, oral
[Paper][Website]

AVVA audio video alignment foundation model with LLM curation, EUSIPCO 2025

Quality Over Quantity? LLM-Based Curation for a Data-Efficient Audio-Video Foundation Model
European Signal Processing Conference (EUSIPCO) 2025
[Paper][Website]

Video understanding with large language models survey, IEEE TCSVT 2025

Video Understanding with Large Language Models: A Survey
IEEE Transactions on Circuits and Systems for Video Technology (TCSVT) 2025
[Paper][Code]

Counterfactual audio language learning by Ali Vosoughi, ICASSP 2024

Learning Audio Concepts from Counterfactual Natural Language
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
[Paper][Code][Patent]

AVSA-Sep audiovisual scene aware sound separation, ICCV 2023 workshop

AVSA-Sep: Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation
IEEE/CVF International Conference on Computer Vision (ICCV) 2023: ICCV AV4D Workshop
[Paper]

MISAR multimodal instructional system with augmented reality, ICCV 2023 workshop

MISAR: A Multimodal Instructional System with Augmented Reality
IEEE/CVF International Conference on Computer Vision (ICCV) 2023: ICCV AV4D Workshop
[Paper][Code][Video]

Nonlinear relation discovery in large scale settings, ICASSP 2022

Relation Discovery in Nonlinearly Related Large-scale Settings
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2022
[Paper][Code]

Pre-image based nonlinear relationship discovery, EUSIPCO 2021

Leveraging Pre-Images to Discover Nonlinear Relationships in Multivariate Environments
European Signal Processing Conference (EUSIPCO) 2021
[Paper]

Large scale nonlinear Granger causality by Ali Vosoughi, Nature Scientific Reports 2021

Large-scale Nonlinear Granger Causality for Inferring Directed Dependence from Short Multivariate Time-series Data
Scientific Reports, Nature Publishing Group (Nature) 2021
[Paper][Code]


Personal Gallery

Ali Vosoughi, multimodal AI researcher
Ali Vosoughi
Ali Vosoughi at the University of Rochester
Ali Vosoughi