Cost-Aware Semantic Gateway for LLM Routing
Sending every question to GPT-4o is expensive, and most questions don't need it. So I put a small classifier in front that guesses how hard each one is. Easy ones stay on Llama 3.2, running locally on vLLM, and only the hard ones go to GPT-4o.
The spend and share numbers come from an 81-query benchmark. Routing accuracy was measured on labeled prompts.