Build Your Own Local Routing Layer (SwitchYard)
Share

Post Content

 

 [[{“value”:”Thanks to ‪@NVIDIADeveloper‬ for DGX Spark. Check it out here: https://nvda.ws/3XIkwsh

In this video I explain why a routing layer is becoming essential for agentic systems and show how to run your own router locally so you control cost, speed, specialization, and privacy. I compare proprietary and open-source options (OpenRouter, Devin Fusion, RouteLLM) and then focus on NVIDIA Switchyard (built on RouteLLM) and how it supports multiple routing strategies: random, LLM classifier, stage routing, and escalation—plus their cost tradeoffs. I also cover NVIDIA Nemotron 3.5 Lightning (30B latent MoE) and why its high throughput matters, including a speed test versus Kimi K3. Finally, I demo local-first alert analysis workflows where Nemotron generates private digests and Switchyard escalates to Kimi K3 only when needed, and explain why routing should be done per task/session rather than per turn.

LINKS:
Blog: https://nvda.ws/3RAV34N
Nemotron 3.5 Lightning: https://nvda.ws/3RPClGE
NeMo Switchyard: https://nvda.ws/4wO4WLE
NVFP4 DFlash: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash
NVFP4 DSpark: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
NVFP4: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4

My voice to text App: whryte.com
Website: https://engineerprompt.ai/
RAG Beyond Basics Course:
https://prompt-s-site.thinkific.com/courses/rag
Signup for Newsletter, localgpt:
https://tally.so/r/3y9bb0

00:00 Why Routers Matter
01:29 Cost Speed Privacy
03:09 Routing Options Today
04:01 Nemotron Lightning Intro
05:19 Speed Test Results
06:34 Switchyard Setup Basics
07:27 Four Routing Strategies
08:50 Signals And Cost Tradeoffs
09:28 Cutting Costs With Hybrid
10:33 Network Alerts Demo
12:22 Per Alert Escalation
12:56 Best Practices And Wrap”}]] Read More Prompt Engineering 

#Promptengineering #AI

By ali

Leave a Reply