2026-08-27 - Asked to drop an engine, and the engine was fine

local-cpu scored zero and looked like dead weight. It is the second-best engine on this machine - 4 of 12 novel words against 2 for the four-billion-parameter model on the accelerator above it. The zero was the scorer demanding a JSON array while the proposal loop reads three fallbacks deep. A benchmark stricter than its consumer measures the benchmark (D335).

measure now takes the caller's scoring function. The strict one is a default, not the rule.

Quoting the reply is what found it. The failure line now carries what was said, and it said <think> </think> Here is a list of **12 new variant markers**... - engine working, model on topic, and a reasoning block the managed path suppresses with --reasoning off that the in-process path cannot. None of that was visible behind "not a list of words".

Three measurements, three wrong answers, all mine: by speed, by volume, by strictness. The third nearly deleted something that works.