Cost vs accuracy: pick your spot per task type
Basic Tool Calling (file reader) Simple (prompt only) Tools + Python
The prize is the upper-left: high accuracy at low cost per task. Cost is $ per query (one task attempt incl. all tool round-trips), log scale — not price per token. The right combo depends on the task type and your accuracy bar. Subscription / agentic-CLI combos (GPT-5.5 in Codex, etc.) carry no per-token cost and are charted separately. Free-tier lanes are imputed at the paid-variant price. Preliminary, June 2026; objective-keyed scoring; per-task samples are still small — expanding the bank is what sharpens these clusters.