BEGIN:VCALENDAR
PRODID:-//AddEvent Inc//AddEvent.com v1.7//EN
VERSION:2.0
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
BEGIN:STANDARD
DTSTART:20261101T010000
RRULE:FREQ=YEARLY;BYDAY=1SU;BYMONTH=11
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RRULE:FREQ=YEARLY;BYDAY=2SU;BYMONTH=3
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
DESCRIPTION:Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes\, a challenge that existing Large Vision–Language Models (LVLMs) struggle to meet. Although these models can describe visual content\, their lack of explicit grounding leads to object hallucinations and unreliable depth reasoning\, limiting their usefulness for accessibility guidance. Sultan et al. introduce WalkGPT\, a pixelgrounded LVLM for the new task of Grounded Navigation Guide\, unifying language reasoning and segmentation within a single architecture for depth-aware accessibility guidance. Given a pedestrian-view image and a navigation query\, WalkGPT generates a conversational response with segmentation masks that delineate accessible and harmful features\, along with relative depth estimation. The model incorporates a Multi-Scale Query Projector (MSQP) that shapes the final image tokens by aggregating them along text tokens across spatial hierarchies\, and a Calibrated Text Projector (CTP)\, guided by a proposed Region Alignment Loss\, that maps language embeddings into segmentation-aware representations. These components enable fine-grained grounding and depth inference without user-provided cues or anchor points\, allowing the model to generate complete and realistic navigation guidance. They also introduce PAVE\, a large-scale benchmark of 41k pedestrian-view images paired with accessibility-aware questions and depth-grounded answers. Experiments show that WalkGPT achieves strong grounded reasoning and segmentation performance.\n\nRafi Ibn Sultan is a Ph.D. candidate in Computer Science at Wayne State University and a researcher in the Trustworthy AI Lab. His work focuses on computer vision and vision–language models\, studying how multimodal systems can better understand spatial relationships in visual scenes. He develops models that combine segmentation\, depth\, and language to enable more grounded reasoning about the physical world\, with applications ranging from medical image analysis to pedestrian navigation and accessibility technologies. His broader research goal is to build vision–language systems that move beyond surface-level description toward deeper spatial understanding and practical real-world use.\n\n------\n\nCreate your own Add to Calendar links with addevent.com/r/a \n
X-ALT-DESC;FMTTYPE=text/html:Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision–Language Models (LVLMs) struggle to meet. Although these models can describe visual content, their lack of explicit grounding leads to object hallucinations and unreliable depth reasoning, limiting their usefulness for accessibility guidance. Sultan et al. introduce WalkGPT, a pixelgrounded LVLM for the new task of Grounded Navigation Guide, unifying language reasoning and segmentation within a single architecture for depth-aware accessibility guidance. Given a pedestrian-view image and a navigation query, WalkGPT generates a conversational response with segmentation masks that delineate accessible and harmful features, along with relative depth estimation. The model incorporates a Multi-Scale Query Projector (MSQP) that shapes the final image tokens by aggregating them along text tokens across spatial hierarchies, and a Calibrated Text Projector (CTP), guided by a proposed Region Alignment Loss, that maps language embeddings into segmentation-aware representations. These components enable fine-grained grounding and depth inference without user-provided cues or anchor points, allowing the model to generate complete and realistic navigation guidance. They also introduce PAVE, a large-scale benchmark of 41k pedestrian-view images paired with accessibility-aware questions and depth-grounded answers. Experiments show that WalkGPT achieves strong grounded reasoning and segmentation performance.<br><br>Rafi Ibn Sultan is a Ph.D. candidate in Computer Science at Wayne State University and a researcher in the Trustworthy AI Lab. His work focuses on computer vision and vision–language models, studying how multimodal systems can better understand spatial relationships in visual scenes. He develops models that combine segmentation, depth, and language to enable more grounded reasoning about the physical world, with applications ranging from medical image analysis to pedestrian navigation and accessibility technologies. His broader research goal is to build vision–language systems that move beyond surface-level description toward deeper spatial understanding and practical real-world use.<br /><br />------<br /><br />Create your own Add to Calendar links with addevent.com/r/a <br>
UID:9a81281713fd43f8b7dd497d379ce37faddeventcom
SUMMARY:Rafi Ibn Sultan - WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation 
DTSTART;TZID=America/Los_Angeles:20260526T080000
DTEND;TZID=America/Los_Angeles:20260526T090000
DTSTAMP:20261010T102900Z
TRANSP:OPAQUE
STATUS:CONFIRMED
SEQUENCE:0
LOCATION:https://meet.google.com/pqd-higu-fwa?hs=122&authuser=0
X-MICROSOFT-CDO-BUSYSTATUS:BUSY
BEGIN:VALARM
TRIGGER:-PT30M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR