Cost-Aware Semantic Gateway for LLM Routing
A lightweight classifier estimates how complex each request is, then routes it to a local or hosted model. Simpler requests stay local; more complex ones go to GPT-4o.
In an 81-query benchmark, 79% of requests stayed local and cloud-model spend was 68.1% lower than an all-GPT-4o baseline. Routing accuracy on labeled prompts was 93.4%, which measures query classification, not answer quality.