Ali Vosoughi
University of Rochester | PhD research complete
Egocentric and multimodal AI. Scene understanding, vision-language agents, and evaluation for real-time, on-device systems.
Headsets, robots, and vehicles share one problem: a first-person stream from a moving platform that must be understood and acted on within a latency budget. Ali Vosoughi builds multimodal AI for this problem and the evaluation methods that decide whether such systems ship. Over seven years of AI and ML research, his work spans egocentric perception, visual agents, video understanding, conversational speech agents, world models, and audiovisual learning, across four industry labs,
Apple,
Microsoft Research,
Smule, and
Bosch AI Research,
and federally supported programs including
DARPA, NIH, NSF, and NNSA. His methods include post-training of multimodal models with SFT, DPO, and reinforcement learning, and statistically rigorous evaluation with significance testing, bootstrap confidence intervals, and human-in-the-loop protocols, built in Python and PyTorch.
AAAI 2026 Best Demonstration Award Runner-up
co-inventor, US patent application US 2025/0124292 A1
COLM Β· NeurIPS Β· NAACL Β· ACM MM Β· ICASSP
π§ ali.vosoughi@rochester.edu
π CS Department, Wegmans Hall 3211
π Apple
Machine Learning Intern
Agentic Multimodal AI
Agentic Multimodal AI
π΅ Smule AI
Research Scientist Intern
Spatial Generation
Spatial Generation
π’ Microsoft Research
Research Intern
Video/Audio LLM and Video Understanding
Video/Audio LLM and Video Understanding
π Bosch AI Research
Research Intern
LLM and Counterfactual Learning
LLM and Counterfactual Learning
π‘οΈ DARPA PTG
Graduate Researcher
Autonomous Multimodal Perception and AR
Autonomous Multimodal Perception and AR
π University of Rochester
Graduate Research Assistant, 2019 to present
Multimodal AI across DARPA, NIH, NSF, and NNSA programs
Multimodal AI across DARPA, NIH, NSF, and NNSA programs
AAAI 2026 Best Demonstration Award Runner-up
Caption Anything in Video (Spatiotemporal Multimodal Prompting)
Caption Anything in Video (Spatiotemporal Multimodal Prompting)
Video Understanding with LLMs
Comprehensive survey with 241+ citations (IEEE TCSVT 2025)
Comprehensive survey with 241+ citations (IEEE TCSVT 2025)
PW-VQA
Causal debiasing for visual question answering with 50+ citations (IEEE TMM 2024)
Causal debiasing for visual question answering with 50+ citations (IEEE TMM 2024)
First counterfactual audio methods
ICASSP 2024 + US patent application US 2025/0124292 A1 (co-inventor, published Jan 2025)
ICASSP 2024 + US patent application US 2025/0124292 A1 (co-inventor, published Jan 2025)
PromptReverb
First text-to-spatial generation at 48kHz (ICASSP 2026)
First text-to-spatial generation at 48kHz (ICASSP 2026)
AVVA
Unified audiovisual foundation model with LLM curation (EUSIPCO 2025)
Unified audiovisual foundation model with LLM curation (EUSIPCO 2025)
Autonomous multimodal copilot
Real-time audiovisual AR demonstrations (DARPA)
Real-time audiovisual AR demonstrations (DARPA)
VERIFY benchmark
Stage-aware evaluation of multimodal reasoning fidelity (COLM 2026)
Stage-aware evaluation of multimodal reasoning fidelity (COLM 2026)
Video LMM Post-Training
Deep dive into video reasoning with large multimodal models
Deep dive into video reasoning with large multimodal models
AVE-2 Dataset
Open audiovisual benchmark for cross-modal event understanding
Open audiovisual benchmark for cross-modal event understanding
Recent News & Updates
08/2026
π COLM 2026 acceptance: VERIFY, a 600-item visual reasoning benchmark with human-annotated reasoning paths and stage-aware evaluation
08/2026
π SPIE Medical Imaging 2027: four papers submitted on medical vision-language evaluation and causal inference methods
06/2026
π IΒ²G released in the IEEE/CVF CVPR 2026 Workshop Proceedings: Generating Instructional Illustrations via Text-Conditioned Diffusion
05/2026
π ICASSP 2026 paper accepted as oral presentation: PromptReverb (Text-to-Spatial-Audio Generation at 48kHz)
02/2026
π AAAI 2026 Best Demonstration Award Runner-up: Caption Anything in Video (Spatiotemporal Video Understanding and Multimodal Prompting)
12/2025
π NeurIPS 2025 paper accepted: MMPerspective (Multimodal LLM Reasoning, Video and Visual Perception)
09/2025
β
Completed research internship at Smule AI (Spatial Audio Generation and Synthesis)
06/2025
π΅ Started research internship at Smule AI (Spatial Audio Generation and Immersive Computing)
03/2025
π Published VERIFY benchmark (Multimodal Reasoning Verification for Video and Vision LLMs)
10/2024
π€ Presented at SANE 2024 at Google, Cambridge, MA (Audio Understanding, Video LLMs, and Spatial Audio)
10/2024
π ACM Multimedia 2024: EAGLE (Egocentric Video Understanding and Language Generation)
08/2024
πΌ Research presentation at Microsoft Research, Seattle (Audiovisual LLM, Video and Audio Understanding)
03/2024
π NAACL 2024: OSCaR (Video Object State Captioning, Autonomous Video Perception)
02/2024
π IEEE Transactions on Multimedia 2024: PW-VQA (Causal Visual Question Answering, Video Reasoning)
08/2023
π― Two ICCV 2023 papers accepted (Audiovisual Sound Separation and Autonomous AR Perception System)
04/2023
π’ Started internship at Bosch Center for AI (Audio Language Models and Counterfactual Reasoning)
Publications

EAGLE: Egocentric AGgregated Language-video Engine
ACM International Conference on Multimedia (ACM MM) 2024
[Paper][DARPA PTG Demo]

VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
Conference on Language Modeling (COLM) 2026
[Paper][Website][π€ Hugging Face]







AVSA-Sep: Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation
IEEE/CVF International Conference on Computer Vision (ICCV) 2023: ICCV AV4D Workshop
[Paper]



Leveraging Pre-Images to Discover Nonlinear Relationships in Multivariate Environments
European Signal Processing Conference (EUSIPCO) 2021
[Paper]

Personal Gallery

