BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:In this talk\, we introduce KPOP\, a novel preconditioned optimizer developed at Exo Labs that performs adaptive optimization in the Kronecker-factored eigenbasis (KFE). By leveraging efficient Kronecker-factored curvature approximations (KFAC)\, KPOP brings the benefits of second-order information to modern large-scale training while maintaining scalability and simplicity. Despite a higher memory footprint than AdamW\, KPOP consistently outperforms it\, both in training iterations as well as wall-clock time\, and achieves performance close to state-of-the-art optimizers on the competitive NanoGPT benchmark on 8x NVIDIA H100s set-ups. We also present TopKPOP\, a memory- and bandwidth-efficient variant that retains only the top eigenvectors of the KFE. This enables fine-grained control over resource usage and makes the optimizer well-suited for adapting the training algorithm to specific hardware constraints of the training environment. Combined with a hybrid CPU-GPU implementation that fully exploits Apple’s unified memory architecture\, these optimizers enable effective training of large language models on consumer hardware. We are the first to successfully demonstrate large-scale distributed training of LLMs on clusters of Apple devices\, spanning 2 to 16 nodes.\n\n------\n\nCreate your own Add to Calendar links with addevent.com/r/a \n
X-ALT-DESC;FMTTYPE=text/html:In this talk, we introduce KPOP, a novel preconditioned optimizer developed at Exo Labs that performs adaptive optimization in the Kronecker-factored eigenbasis (KFE). By leveraging efficient Kronecker-factored curvature approximations (KFAC), KPOP brings the benefits of second-order information to modern large-scale training while maintaining scalability and simplicity. Despite a higher memory footprint than AdamW, KPOP consistently outperforms it, both in training iterations as well as wall-clock time, and achieves performance close to state-of-the-art optimizers on the competitive NanoGPT benchmark on 8x NVIDIA H100s set-ups. We also present TopKPOP, a memory- and bandwidth-efficient variant that retains only the top eigenvectors of the KFE. This enables fine-grained control over resource usage and makes the optimizer well-suited for adapting the training algorithm to specific hardware constraints of the training environment. Combined with a hybrid CPU-GPU implementation that fully exploits Apple’s unified memory architecture, these optimizers enable effective training of large language models on consumer hardware. We are the first to successfully demonstrate large-scale distributed training of LLMs on clusters of Apple devices, spanning 2 to 16 nodes.<br /><br />------<br /><br />Create your own Add to Calendar links with addevent.com/r/a <br>
UID:382ad2e9bd514509ac0894ab554bedb9addeventcom
SUMMARY:Tycho van der Ouderaa and Matt Beton - KPOP: Kronecker Preconditioned adaptive OPtimization (Eff)
DTSTART;TZID=America/Los_Angeles:20250814T100000
DTEND;TZID=America/Los_Angeles:20250814T110000
DTSTAMP:20260823T081658Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/wdk-yipf-zjd?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR