BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Large language models (LLMs) are remarkably efficient across a wide range of natural language processing tasks and well beyond them. However\, a comprehensive theoretical analysis of the LLMs' generalization capabilities remains elusive. In our paper\, we approach this task by drawing an equivalence between autoregressive transformer-based language models and Markov chains defined on a finite state space. This allows us to study the multi-step inference mechanism of LLMs from first principles. We relate the obtained results to the pathological behavior observed with LLMs with high temperatures\, such as repetitions and incoherent replies. Finally\, we leverage the proposed formalization to derive pre-training and in-context learning generalization bounds for LLMs under realistic data and model assumptions. Experiments with the most recent Llama and Gemma herds of models show that our theory correctly captures their behavior in practice.\n\nAmbroise Odonnat is a PhD student at Huawei Noah's Ark Lab and Inria\, advised by Romain Tavenard\, Laetitia Chapel\, and Ievgen Redko. His goal is to improve the core understanding of Transformers by conducting theoretical analysis and large-scale experiments on large language models\, transformer fine-tuning\, and out-of-distribution generalization.\n\n------\n\nCreate your own Add to Calendar links with addevent.com/r/a \n
X-ALT-DESC;FMTTYPE=text/html:Large language models (LLMs) are remarkably efficient across a wide range of natural language processing tasks and well beyond them. However, a comprehensive theoretical analysis of the LLMs' generalization capabilities remains elusive. In our paper, we approach this task by drawing an equivalence between autoregressive transformer-based language models and Markov chains defined on a finite state space. This allows us to study the multi-step inference mechanism of LLMs from first principles. We relate the obtained results to the pathological behavior observed with LLMs with high temperatures, such as repetitions and incoherent replies. Finally, we leverage the proposed formalization to derive pre-training and in-context learning generalization bounds for LLMs under realistic data and model assumptions. Experiments with the most recent Llama and Gemma herds of models show that our theory correctly captures their behavior in practice.<br><br>Ambroise Odonnat is a PhD student at Huawei Noah's Ark Lab and Inria, advised by Romain Tavenard, Laetitia Chapel, and Ievgen Redko. His goal is to improve the core understanding of Transformers by conducting theoretical analysis and large-scale experiments on large language models, transformer fine-tuning, and out-of-distribution generalization.<br /><br />------<br /><br />Create your own Add to Calendar links with addevent.com/r/a <br>
UID:e11b50ab2c304b118e186fd2e66da695addeventcom
SUMMARY:Ambroise Odonnat - Large Language Models as Markov Chains (ML theory)
DTSTART;TZID=America/Los_Angeles:20250619T090000
DTEND;TZID=America/Los_Angeles:20250619T100000
DTSTAMP:20260916T143607Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/nvg-ptgt-ucf?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR