BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Toronto
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
TZNAME:EST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
TZNAME:EDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Classical statistics teaches us that overparameterization causes overfitting\, which prevents good generalization. However\, highly overparameterized neural network architectures generalize surprisingly well. This is because the training of these models tends towards low rank or sparse solutions\, without requiring explicit constraints. This preference is known as implicit regularization\, and it can be found in a variety of contexts\, including attention layers\, LoRA\, matrix sensing\, and diagonal linear networks. As a result\, implicit regularization helps explain how overfitting is avoided and generalization is improved in neural networks.\n\nIn this talk\, I will show how weight decay controls implicit regularization beyond its explicit role of constraining the model capacity. For instance\, it moves the implicit regularizer from $L_2$ to $L_1$\, which leads to more sparsity in the model. This demonstrates how weight decay not only serves as a model constraint\, but also has an implicit effect. By turning off weight decay during training\, only the implicit effect remains\, resulting in better generalization overall. Besides better generalization\, I use these insights to induce sparsity in deep neural networks. Sparsity aims to reduce model size and inference time by removing as many weights as possible. This results in a new method: PILoT (Parameteric Implicit Lottery Ticket)\, a sparsification method based on overparameterization and weight decay that uses the transition of the implicit regularization from $L_2$ to $L_1$ to gradually sparsify\, achieving high sparsity with a smaller performance drop.\n\nTom Jacobs is a PhD student at the CISPA Helmholtz Institute in Saarbrucken\, Germany\, under supervision of Rebekka Burkholz. He currently works on understanding training dynamics in deep learning and designing methods for efficiency. Before that he has graduated with a master in applied mathematics specializing in mathematical analysis and probability theory from the Technical University of Eindhoven\, in the Netherlands.\n\n------\n\nPowered by addevent.com \nShare your next event with us!\n
X-ALT-DESC;FMTTYPE=text/html:Classical statistics teaches us that overparameterization causes overfitting, which prevents good generalization. However, highly overparameterized neural network architectures generalize surprisingly well. This is because the training of these models tends towards low rank or sparse solutions, without requiring explicit constraints. This preference is known as implicit regularization, and it can be found in a variety of contexts, including attention layers, LoRA, matrix sensing, and diagonal linear networks. As a result, implicit regularization helps explain how overfitting is avoided and generalization is improved in neural networks.<br><br>In this talk, I will show how weight decay controls implicit regularization beyond its explicit role of constraining the model capacity. For instance, it moves the implicit regularizer from $L_2$ to $L_1$, which leads to more sparsity in the model. This demonstrates how weight decay not only serves as a model constraint, but also has an implicit effect. By turning off weight decay during training, only the implicit effect remains, resulting in better generalization overall. Besides better generalization, I use these insights to induce sparsity in deep neural networks. Sparsity aims to reduce model size and inference time by removing as many weights as possible. This results in a new method: PILoT (Parameteric Implicit Lottery Ticket), a sparsification method based on overparameterization and weight decay that uses the transition of the implicit regularization from $L_2$ to $L_1$ to gradually sparsify, achieving high sparsity with a smaller performance drop.<br><br>Tom Jacobs is a PhD student at the CISPA Helmholtz Institute in Saarbrucken, Germany, under supervision of Rebekka Burkholz. He currently works on understanding training dynamics in deep learning and designing methods for efficiency. Before that he has graduated with a master in applied mathematics specializing in mathematical analysis and probability theory from the Technical University of Eindhoven, in the Netherlands.<br /><br />------<br /><br />Powered by addevent.com <br>Share your next event with us!<br>
UID:5cba3b21d23b424d82a192fbec8d5f78addeventcom
SUMMARY:Tom Jacobs - Weight Decay Controls Implicit Regularization: Insights on Generalization and Sparsity (Theory)
DTSTART;TZID=America/Toronto:20250814T100000
DTEND;TZID=America/Toronto:20250814T110000
DTSTAMP:20260804T225753Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/nvg-ptgt-ucf?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR