BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Large language models are trained on massive amounts of untrusted data from the web and third-party vendors. This talk explores how adversaries can compromise LLMs by poisoning training data at different stages. We first show that carefully mislabeled RLHF examples can create "universal backdoors" that let attackers bypass safety measures after the model is deployed. Since reinforcement learning data is usually carefully curated\, we also explore a more realistic threat model: posting poisoned data on the web so that it ends up in the pre-training dataset. We found that by corrupting even a tiny fraction (0.1% or less) of pre-training data\, attackers can introduce persistent vulnerabilities that survive through fine-tuning and alignment\, enabling attacks like denial-of-service and belief manipulation. Our findings highlight how the reliance on untrusted data\, particularly during pre-training\, creates critical security vulnerabilities.\n\nJavier Rando is a PhD student at ETH Zurich under the supervision of Prof. Florian Tramèr. His research focuses on identifying vulnerabilities in state-of-the-art AI models\, particularly large language models (LLMs)\, to understand potential risks in real-world applications. In the summer of 2024\, he interned with Meta's GenAI Safety & Trust team\, where he conducted research on poisoning attacks and participated in the robustness analysis of the latest multimodal LLaMA Guard. Prior to his PhD\, Javier obtained a MSc in Computer Science at ETH Zürich and was a visiting researcher at New York University\, working on language model truthfulness under the supervision of Prof. He He.\n\n------\n\nCreate your own Add to Calendar links with addevent.com/r/a \n
X-ALT-DESC;FMTTYPE=text/html:Large language models are trained on massive amounts of untrusted data from the web and third-party vendors. This talk explores how adversaries can compromise LLMs by poisoning training data at different stages. We first show that carefully mislabeled RLHF examples can create "universal backdoors" that let attackers bypass safety measures after the model is deployed. Since reinforcement learning data is usually carefully curated, we also explore a more realistic threat model: posting poisoned data on the web so that it ends up in the pre-training dataset. We found that by corrupting even a tiny fraction (0.1% or less) of pre-training data, attackers can introduce persistent vulnerabilities that survive through fine-tuning and alignment, enabling attacks like denial-of-service and belief manipulation. Our findings highlight how the reliance on untrusted data, particularly during pre-training, creates critical security vulnerabilities.<br><br>Javier Rando is a PhD student at ETH Zurich under the supervision of Prof. Florian Tramèr. His research focuses on identifying vulnerabilities in state-of-the-art AI models, particularly large language models (LLMs), to understand potential risks in real-world applications. In the summer of 2024, he interned with Meta's GenAI Safety &amp; Trust team, where he conducted research on poisoning attacks and participated in the robustness analysis of the latest multimodal LLaMA Guard. Prior to his PhD, Javier obtained a MSc in Computer Science at ETH Zürich and was a visiting researcher at New York University, working on language model truthfulness under the supervision of Prof. He He.<br /><br />------<br /><br />Create your own Add to Calendar links with addevent.com/r/a <br>
UID:94fa7770ea634254b06f55a1344870ebaddeventcom
SUMMARY:C4AI - Javier Rando - Poisoned Training Data Can Compromise LLMs (Safety/Align)
DTSTART;TZID=America/Los_Angeles:20250123T090000
DTEND;TZID=America/Los_Angeles:20250123T100000
DTSTAMP:20261001T162521Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/gkq-yrbx-rns?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR