#!/usr/bin/env python3 """ needle.py — repeatable needle-in-a-haystack battery for quant/lever A/Bs. Replaces the ad-hoc one-off probes (EMERALD-7742 @66.5k etc.) with a committed instrument. Measures retrieval quality at depth AND speed telemetry per depth, so a candidate quant can be compared against the incumbent on one run. Method (mirrors halogen-flash's published battery + our 2026-09-03 probes): - Filler text is neutral prose paragraphs (seeded, deterministic per depth). - A needle is a single sentence with a random code token: "The maintenance code for the {place} is {CODE}-{NNNN}." Spliced into filler at a target depth, the document continues, then asks for the code. Exact string match on the retrieved code. Temp 0 (greedy) — quant differences only, no sampling noise. - Speed: prefill tok/s (prompt_eval from usage), decode tok/s (eval), wall. - Multiple needles per depth, different insertion positions -> a distribution, not a single anecdote. Also: one NEAR-MISS needle per depth (a second code exists elsewhere in the doc; the model must return the RIGHT one) to catch confabulation. Usage: python bench.py needle run # full battery python bench.py needle run --depths 1000,16000 # quick subset python bench.py needle run --needles 3 --repeats 2 # smoke config Env (via .env / bench.core.client): ENDPOINT, MODEL, API_KEY, MAX_TOKENS, TEMPERATURE, HTTP_TIMEOUT. Plus: LABEL tag for the JSON/report (default MODEL) N_PREDICT max answer tokens (default 512; reasoning models think first) Output: results/needle-