BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Large language models demonstrate the intriguing ability to perform unseen tasks via in-context learning. However\, it remains unclear what mechanisms inside the model drive such task-level generalization. In this work\, we approach this question through the lens of off-by-one addition (i.e.\, 1+1=3\, 2+2=5\, 3+3=?)\, a two-step\, counterfactual task with an unexpected +1 function as a second step. Leveraging circuit-style interpretability techniques such as path patching\, we analyze the models' internal computations behind their notable performance and present three key findings. First\, we uncover a function induction mechanism that explains the model's generalization from standard addition to off-by-one addition. This mechanism resembles the structure of the induction head mechanism found in prior work and elevates it to a higher level of abstraction. Second\, we show that the induction of the +1 function is governed by multiple attention heads in parallel\, each of which emits a distinct piece of the +1 function. Finally\, we find that this function induction mechanism is reused in a broader range of tasks\, including synthetic tasks such as shifted multiple-choice QA and algorithmic tasks such as base-8 addition. Overall\, our findings offer deeper insights into how reusable and composable structures within language models enable task-level generalization.\n\nQinyuan Ye recently completed her Ph.D. in Computer Science at University of Southern California. Her research centers on enabling NLP and AI systems to learn in a data-efficient and proactive manner\, with an emphasis on meta-learning\, in-context learning\, and instruction tuning. She co-organized the workshop on Instruction Tuning and Instruction Following at NeurIPS 2023 and co-presented a tutorial on LLM-driven Instruction Following at EMNLP 2023.\n\n------\n\nCreate your own Add to Calendar links with addevent.com/r/a \n
X-ALT-DESC;FMTTYPE=text/html:Large language models demonstrate the intriguing ability to perform unseen tasks via in-context learning. However, it remains unclear what mechanisms inside the model drive such task-level generalization. In this work, we approach this question through the lens of off-by-one addition (i.e., 1+1=3, 2+2=5, 3+3=?), a two-step, counterfactual task with an unexpected +1 function as a second step. Leveraging circuit-style interpretability techniques such as path patching, we analyze the models' internal computations behind their notable performance and present three key findings. First, we uncover a function induction mechanism that explains the model's generalization from standard addition to off-by-one addition. This mechanism resembles the structure of the induction head mechanism found in prior work and elevates it to a higher level of abstraction. Second, we show that the induction of the +1 function is governed by multiple attention heads in parallel, each of which emits a distinct piece of the +1 function. Finally, we find that this function induction mechanism is reused in a broader range of tasks, including synthetic tasks such as shifted multiple-choice QA and algorithmic tasks such as base-8 addition. Overall, our findings offer deeper insights into how reusable and composable structures within language models enable task-level generalization.<br><br>Qinyuan Ye recently completed her Ph.D. in Computer Science at University of Southern California. Her research centers on enabling NLP and AI systems to learn in a data-efficient and proactive manner, with an emphasis on meta-learning, in-context learning, and instruction tuning. She co-organized the workshop on Instruction Tuning and Instruction Following at NeurIPS 2023 and co-presented a tutorial on LLM-driven Instruction Following at EMNLP 2023.<br /><br />------<br /><br />Create your own Add to Calendar links with addevent.com/r/a <br>
UID:1faf899a4ec74c429db5b0faa5c1ed79addeventcom
SUMMARY:Qinyuan Ye - Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition (Geo Asia)
DTSTART;TZID=America/Los_Angeles:20250903T090000
DTEND;TZID=America/Los_Angeles:20250903T100000
DTSTAMP:20260817T104622Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/yhv-tiir-ava?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR