josie / llm-bench

1 commit
1 branch
0 tags

llm-bench

A local LLM coding-model benchmark suite. Point it at any OpenAI-compatible endpoint, pick an instrument, and run. Pure stdlib Python — no packages needed.

Quick start

cp .env.example .env          # set ENDPOINT + MODEL for your model
python bench.py coding-suite validate   # prove tasks are well-formed (no model needed)
python bench.py coding-suite run        # run all 24 tasks against the model

Results land in results/ as JSONL + JSON. Work dirs in work/ (both gitignored).

Instruments

python bench.py <instrument> <command> [args...]

| Instrument | Command(s) | What it measures | |---|---|---| | coding-suite | validate / run / report | 24 single-turn bug-fix + tool-calling tasks, real toolchain (Go/Rust/Swift/TS) | | agentic-fix | validate / run | Multi-turn tool loop — find a cross-file bug in a small Go codebase | | agentic-build | validate / run | Build a Go+SQLite web app end-to-end (4 product gates: build/test/serve/persist) | | agentic-multitask | validate <task> / run <task> / cohort <t1> <t2> ... | Multi-task gauntlet: Go+HTMX, Rust CLI, Next.js+Prisma | | router | validate / run / report | Dispatch/router eval (26 ROUTE + 8 ORCH, deterministic scoring, privacy axis) | | needle | run [--depths 1000,16000 --needles 3 --repeats 2] | Needle-in-a-haystack retrieval + speed per depth | | loop-battery | run | CoT runaway probes (5 prompts × N runs, repetition scoring) | | no-think | run | Latency win from suppressing reasoning (think vs no-think A/B) |

Examples

# Coding suite (single-turn, ~5-15 min)
python bench.py coding-suite run

# Agentic build (multi-turn, ~10-40 min)
python bench.py agentic-build run

# Full gauntlet — all three real-app tasks, sequential
python bench.py agentic-multitask cohort go_htmx rust_cli ts_next

# Router/dispatch eval
python bench.py router run

# Needle retrieval at specific depths
python bench.py needle run --depths 1000,16000,64000

# Override config per-run without editing .env
ENDPOINT=http://10.0.0.40:8000/v1/chat/completions \
MODEL=qwen3.8-27b \
MAX_TOKENS=16000 \
python bench.py agentic-build run

Configuration

Copy .env.example to .env and edit. Env vars win over .env, so you can override per-run (see the last example above).

| Var | Default | What | |---|---|---| | ENDPOINT | http://127.0.0.1:1234/v1/chat/completions | OpenAI-compat chat URL | | MODEL | local | model id | | API_KEY | (blank) | bearer token (if endpoint needs auth) | | MAX_TOKENS | 8000 | per-turn generation cap | | TEMPERATURE | 0 | sampling temp (0 = greedy, recommended) | | MAX_TURNS | 40 | agentic loop cap | | HTTP_TIMEOUT | 900 | per-request wall timeout (s) | | NOTHINK | (blank) | 1 = suppress reasoning via empty think prefill |

Compatible endpoints

Anything that speaks OpenAI /v1/chat/completions:

  • LM Studio — http://127.0.0.1:1234/v1/chat/completions
  • Ollama (local) — http://127.0.0.1:11434/v1/chat/completions
  • vLLM / llama-server — http://127.0.0.1:8000/v1/chat/completions
  • Ollama Cloud — https://api.ollama-cloud.com/v1/chat/completions
  • OpenAI — https://api.openai.com/v1/chat/completions

Tool-calling tasks use the OpenAI tools param. The XML/JSON tool-calling tasks use freeform text, so they work even on backends without native function-calling.

Requirements

  • Python 3.10+
  • Language toolchains for the tasks you want to run:

| Toolchain | Needed for | |---|---| | Go 1.22+ | agentic-fix, agentic-build, go_htmx, go bug-fixes | | Rust 1.70+ | rust_cli, rust bug-fixes | | bun 1.0+ | ts bug-fixes | | Node.js 18+ + npm | ts_next | | Swift / Xcode | swift bug-fixes (macOS only; skipped gracefully elsewhere) |

No Python packages needed — everything uses the stdlib.

Adding your own tasks

Coding suite — add a dict to BUGFIX_T1 / BUGFIX_T2 / TOOLCALL in bench/tasks/coding_suite.py. Bug-fix tasks need: id, lang, symptom, buggy, fix, test. validate proves buggy fails + fix passes.

Multitask gauntlet — drop a .py file in bench/tasks/multitask/ exposing TASK_NAME, PLAN, STARTER, REFERENCE, validate(work_dir), gates(work_dir). Auto-discovered. See go_htmx.py for the shape.

New instrument — add a module under bench/tasks/ with a main() function, register it in INSTRUMENTS in bench.py. Import from bench.core (client, runners, tools, reporting).

Project layout

llm-bench/
├── bench.py                      # CLI entry point
├── .env.example                  # config template
├── bench/
│   ├── core/                     # client, runners, tools, reporting
│   └── tasks/                    # instruments (one per file)
│       ├── coding_suite.py
│       ├── agentic_fix.py
│       ├── agentic_build.py
│       ├── agentic_multitask.py
│       ├── router.py
│       ├── needle.py
│       ├── loop_battery.py
│       ├── no_think.py
│       └── multitask/            # gauntlet task modules (auto-discovered)
│           ├── go_htmx.py
│           ├── rust_cli.py
│           └── ts_next.py
├── results/                      # generated (gitignored)
└── work/                         # generated (gitignored)

License

MIT. See LICENSE.