Research Impact
M. Ali Vosoughi (محمد علی وثوقی) • University of Rochester | PhD research complete • alivosoughi.com
Publications & Citations
- 1,600+ total citations (Google Scholar, September 2026)
- h-index: 20 across computer vision, audio processing, and multimodal AI
- i10-index: 43
- First-author publications: IEEE TMM, ICASSP, EUSIPCO
- Survey leadership: Video Understanding with LLMs (IEEE TCSVT 2025)
Industry Research Collaborations
- Apple (2026): Agentic multimodal AI, machine learning internship
- Microsoft Research (2024): Unified audiovisual encoder → EUSIPCO 2025
- Bosch AI Research (2023): Counterfactual audio-language → ICASSP 2024 + US patent application
- Smule AI (2025): Spatial audio generation for real-time applications
- Clinical AI research collaborations: Aidoc, Carestream, and Gleamer (radiology AI studies with the University of Rochester radiology group)
Open Research Contributions
- Datasets: AVE-2 audiovisual dataset
- Benchmarks: VERIFY (COLM 2026), PW-VQA (causal VQA), MMPerspective (NeurIPS 2025)
- Code: GitHub repositories with reproducible implementations
- Demos: PromptReverb spatial audio demo
Recognition
- AAAI 2026 Best Demonstration Award Runner-up: Caption Anything in Video
- US National Interest Waiver (2023): Permanent residency for contributions to AI research
- Patent application: co-inventor, US 2025/0124292 A1 (audio-language learning)
- DARPA PTG: real-time multimodal AI development and AR demonstrations
Research focus: Egocentric and multimodal AI. Scene understanding, vision-language agents, and evaluation for real-time, on-device systems.