BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Users expect modern multimodal LLMs (MLLMs) to operate over spoken and video inputs in multiple languages\, flexibly addressing diverse task requests such as transcription\, translation\, summarization\, and question answering. Despite the central role of this instruction-following ability\, existing evaluations often remain limited to text-only and monolingual settings\, failing to reflect the complexity of real-world scenarios. As a solution\, this talk describes the crosslingual evaluation of general-purpose instruction-following models (including SpeechLLMs\, VideoLLMs\, and MLLMs) through MCIF\, a novel benchmark built from scientific talks. Developed entirely through human annotation\, MCIF has been designed to assess how the instruction-following abilities vary across languages\, modalities\, and tasks. Through extensive benchmarking of 23 state-of-the-art models\, MCIF reveals critical performance gaps that monolingual single-modality benchmarks cannot capture\, such as significant limitations in joint speech-video integration\, long-form processing\, and summarization\, establishing a foundation for advancing truly multimodal\, crosslingual instruction-following systems.\n\nSara Papi is an AI Researcher at Fondazione Bruno Kessler (FBK)\, Italy\, where she works on speech processing and multimodal LLMs within the MEETWEEN and DVPS Horizon European projects. She received her PhD cum laude in Information Engineering and Computer Science from the University of Trento in 2024\, with a focus on simultaneous speech translation and subtitling. Her research interests span multimodal and crosslingual instruction-following models\, speech foundation models\, and LLMs. Her work has been recognized with awards\, including the Best PhD Graduate 2024 Award in Information and Communication Technology from the University of Trento\, an Outstanding Paper and SAC Award at ACL 2024\, and a Social Impact Paper Award at EMNLP 2024. She actively contributes to the community as an organizer of the IWSLT Evaluation Campaign and as an Area Chair or reviewer for major conferences in speech and NLP\, such as *ACL and Interspeech.\n\n------\n\nPowered by addevent.com \nShare your next event with us!\n
X-ALT-DESC;FMTTYPE=text/html:Users expect modern multimodal LLMs (MLLMs) to operate over spoken and video inputs in multiple languages, flexibly addressing diverse task requests such as transcription, translation, summarization, and question answering. Despite the central role of this instruction-following ability, existing evaluations often remain limited to text-only and monolingual settings, failing to reflect the complexity of real-world scenarios. As a solution, this talk describes the crosslingual evaluation of general-purpose instruction-following models (including SpeechLLMs, VideoLLMs, and MLLMs) through MCIF, a novel benchmark built from scientific talks. Developed entirely through human annotation, MCIF has been designed to assess how the instruction-following abilities vary across languages, modalities, and tasks. Through extensive benchmarking of 23 state-of-the-art models, MCIF reveals critical performance gaps that monolingual single-modality benchmarks cannot capture, such as significant limitations in joint speech-video integration, long-form processing, and summarization, establishing a foundation for advancing truly multimodal, crosslingual instruction-following systems.<br><br>Sara Papi is an AI Researcher at Fondazione Bruno Kessler (FBK), Italy, where she works on speech processing and multimodal LLMs within the MEETWEEN and DVPS Horizon European projects. She received her PhD cum laude in Information Engineering and Computer Science from the University of Trento in 2024, with a focus on simultaneous speech translation and subtitling. Her research interests span multimodal and crosslingual instruction-following models, speech foundation models, and LLMs. Her work has been recognized with awards, including the Best PhD Graduate 2024 Award in Information and Communication Technology from the University of Trento, an Outstanding Paper and SAC Award at ACL 2024, and a Social Impact Paper Award at EMNLP 2024. She actively contributes to the community as an organizer of the IWSLT Evaluation Campaign and as an Area Chair or reviewer for major conferences in speech and NLP, such as *ACL and Interspeech.<br /><br />------<br /><br />Powered by addevent.com <br>Share your next event with us!<br>
UID:f2befef1a3b74c6c8ca4af584ef23363addeventcom
SUMMARY:Sara Papi - Towards Crosslingual Evaluation of General-Purpose Instruction-Following Models Across Text\, Speech\, and Vision (Multimodal)
DTSTART;TZID=America/Los_Angeles:20260123T070000
DTEND;TZID=America/Los_Angeles:20260123T080000
DTSTAMP:20260811T122519Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/kbv-mqso-kxy?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR