Running frontier-like open models locally means no cloud dependency, no per-token costs after hardware, and all data staying on-device. At SIGGRAPH, NVIDIA announced the NVIDIA Agent Toolkit stack on DGX Station — NemoClaw, Nemotron 3 Ultra, Omniverse libraries, and OpenShell secure runtime in a single deskside supercomputer. This session unpacks what that stack means for developers and walks through a working example: deploying Nemotron 3 Ultra on DGX Station.
We'll get hands-on — launching Nemotron 3 Ultra (550B, NVFP4) on a GB300 DGX Station with vLLM, including MoE expert CPU offloading, speculative decoding, and tool-call routing to an OpenAI-compatible API. And cover the SIGGRAPH announcements and what the full Agent Toolkit stack enables.
What you'll learn:
- What the NVIDIA Agent Toolkit stack on DGX Station includes and why it matters for local AI
- How to serve Nemotron 3 Ultra on DGX Station using vLLM with NVFP4 quantization and CPU offloading
- How to configure speculative decoding, prefix caching, and tool-call routing for agentic workloads
- How to verify and query a locally running Nemotron 3 Ultra endpoint
Running open models like Nemotron locally on DGX Station or DGX Spark? Bring your setup questions — we'll answer them live.