BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Training MoE (Mixture of Experts) models is significantly more challenging than training dense models. A key difficulty lies in the widely-used TopK router\, which is non-differentiable and leads to challenges such as unstable optimization\, difficulty in end-to-end training\, and suboptimal solutions. To address this\, we propose DenseMixer\, which introduces a single forward pass through the non-activated experts (with controllable computational overhead) to provide the router with more accurate gradient signals\, thereby improving post-training performance. We observe consistent improvements across three types of models — OLMoE\, Qwen1.5-MoE\, and Qwen3-MoE — with sizes 7B/14B/30B\, on 10+ datasets. For example\, on the GPQA-Diamond benchmark\, Qwen3-30B-MoE achieves a 3.7 percentage point improvement over standard training when using just 1k SFT examples. In terms of efficiency\, DenseMixer only adds one extra forward pass for non-activated experts\, increasing theoretical FLOPs by 46% (1.46x the original). However\, the actual increase in training time is lower than the theoretical estimate. When the model size or data volume is moderate\, the additional time cost remains acceptable\, achieving the goal of improving training quality with controllable computational overhead.\n\nFeng Yao is currently a second-year PhD student in the CSE department at the University of California\, San Diego (UCSD)\, advised by Prof. Jingbo Shang and Prof. Vish Krishnan. Previously\, he received his master's degree from Tsinghua University\, advised by Prof. Zhiyuan Liu and Prof. Weixing Shen.\n\n------\n\nPowered by addevent.com \nShare your next event with us!\n
X-ALT-DESC;FMTTYPE=text/html:Training MoE (Mixture of Experts) models is significantly more challenging than training dense models. A key difficulty lies in the widely-used TopK router, which is non-differentiable and leads to challenges such as unstable optimization, difficulty in end-to-end training, and suboptimal solutions. To address this, we propose DenseMixer, which introduces a single forward pass through the non-activated experts (with controllable computational overhead) to provide the router with more accurate gradient signals, thereby improving post-training performance. We observe consistent improvements across three types of models — OLMoE, Qwen1.5-MoE, and Qwen3-MoE — with sizes 7B/14B/30B, on 10+ datasets. For example, on the GPQA-Diamond benchmark, Qwen3-30B-MoE achieves a 3.7 percentage point improvement over standard training when using just 1k SFT examples. In terms of efficiency, DenseMixer only adds one extra forward pass for non-activated experts, increasing theoretical FLOPs by 46% (1.46x the original). However, the actual increase in training time is lower than the theoretical estimate. When the model size or data volume is moderate, the additional time cost remains acceptable, achieving the goal of improving training quality with controllable computational overhead.<br><br>Feng Yao is currently a second-year PhD student in the CSE department at the University of California, San Diego (UCSD), advised by Prof. Jingbo Shang and Prof. Vish Krishnan. Previously, he received his master's degree from Tsinghua University, advised by Prof. Zhiyuan Liu and Prof. Weixing Shen.<br /><br />------<br /><br />Powered by addevent.com <br>Share your next event with us!<br>
UID:20f79186c6a64c419b2d33692618adcbaddeventcom
SUMMARY:Feng Yao - DenseMixer: Improving MoE Post-Training with Precise Router Gradient (Asia)
DTSTART;TZID=America/Los_Angeles:20250723T090000
DTEND;TZID=America/Los_Angeles:20250723T100000
DTSTAMP:20260730T200536Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/yhv-tiir-ava?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR