"""Dispatch / router / orchestrator benchmark for LAIC. Motivation ---------- The idea (2026-08-17): use a *cheap, always-on local* model as the front door for all communication, and have IT decide which tier actually does each task — keep trivial things local, escalate genuinely hard/large-context/high-stakes work to a bigger local model or the cloud. This bench measures how good a given model is at BEING that dispatcher. It is deliberately shaped like the coding suite (env-driven OpenAI endpoint, self-contained dataset, validate/run/report, JSONL out, category breakdown) so it runs against the same LM Studio / vLLM endpoints the coding suite uses: python bench.py router run python bench.py router report # rebuild summary from results/router_