BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:For all the promise of “global” multimodal AI\, most systems still speak to a narrow slice of the world. This talk traces a path to genuinely multilingual\, culturally grounded models through three steps: first\, ALM-Bench reframes evaluation by testing 100 languages across culturally situated tasks\, revealing where today’s LMMs falter\, especially on low-resource scripts. Next\, ViMUL moves beyond images to video\, pairing a diverse 14-language\, 15-domain benchmark with a balanced baseline to show how training and evaluation can align for robust multilingual video understanding. Finally\, we examine language itself as a causal factor: a cross-lingual T2I study where grammatical gender shifts visual outputs\, surfacing a new axis of bias. Together\, these pieces offer a story and a blueprint for inclusive\, reliable multimodal systems.\n\nI am an MSc. student in the College of Engineering and Computer Science department at the University of Central Florida. I am a member of the Center for Research in Computer Vision (CRCV) Lab advised by Prof. Mubarak Shah.\n\nPreviously\, I was a Research Engineer in the Computer Vision Department\, affiliated with the IVAL-Lab at Mohamed bin Zayed University of Artificial Intelligence (MBZUAI).\n\n\n\n------\n\nPowered by addevent.com \nShare your next event with us!\n
X-ALT-DESC;FMTTYPE=text/html:For all the promise of “global” multimodal AI, most systems still speak to a narrow slice of the world. This talk traces a path to genuinely multilingual, culturally grounded models through three steps: first, ALM-Bench reframes evaluation by testing 100 languages across culturally situated tasks, revealing where today’s LMMs falter, especially on low-resource scripts. Next, ViMUL moves beyond images to video, pairing a diverse 14-language, 15-domain benchmark with a balanced baseline to show how training and evaluation can align for robust multilingual video understanding. Finally, we examine language itself as a causal factor: a cross-lingual T2I study where grammatical gender shifts visual outputs, surfacing a new axis of bias. Together, these pieces offer a story and a blueprint for inclusive, reliable multimodal systems.<br><br>I am an MSc. student in the College of Engineering and Computer Science department at the University of Central Florida. I am a member of the Center for Research in Computer Vision (CRCV) Lab advised by Prof. Mubarak Shah.<br><br>Previously, I was a Research Engineer in the Computer Vision Department, affiliated with the IVAL-Lab at Mohamed bin Zayed University of Artificial Intelligence (MBZUAI).<br><br><br /><br />------<br /><br />Powered by addevent.com <br>Share your next event with us!<br>
UID:6776713c9ce94d0585726c5cbbf348ddaddeventcom
SUMMARY:Ashmal Vayani - Seeing the World as It Speaks: Multilingual\, Culturally Aware Multimodal AI (Asia)
DTSTART;TZID=America/Los_Angeles:20251015T090000
DTEND;TZID=America/Los_Angeles:20251015T100000
DTSTAMP:20260730T201220Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/yhv-tiir-ava?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR