Terminal state first
Correct tool syntax and a convincing final answer are not enough. The task succeeds only when the database, ticket, order, file, or other business record reaches the required state.
Do not pass an agent because it says “done.” Compare expected terminal state, real backend state, and side effects with deterministic checks to see whether the task actually completed without missing, duplicate, or forbidden actions.
Correct tool syntax and a convincing final answer are not enough. The task succeeds only when the database, ticket, order, file, or other business record reaches the required state.
Payments, refunds, emails, publishes, and deletes must not be missing, duplicated, or forbidden. Exactly-once behavior cannot be inferred from model prose.
One successful run is not reliability. Repeated trials expose intermittent false success, duplicate tools, side-effect drift, and unstable terminal outcomes.
Microsoft / Hugging Face ThinkingBox / ThinkingBox-Bench is an important public reference: v1.0 contains 507 stateful business workflows, each repeated 20 times; 477 tasks are graded from state alone and 30 add narrow response rubrics. The public write-up also reports that, in a 121,680-trial common-set analysis, 67.24% of failed attempts still terminated cleanly, invoked a state-changing tool, and reported no final tool error—the exact false-success gap that transcript-only grading misses. XBSTACK does not copy ThinkingBox-Bench tasks, golden state, or trajectories; this browser tool and the public CLI benchmark apply the broad outcome-first principle to your own agent records.
It does not grade whether the answer sounds successful. It deterministically compares expected state, actual state, and required, duplicate, or forbidden side effects. Semantic rubrics should be reserved for requirements that cannot be represented as state.
An agent can say “done” while the database, ticket, payment, email, file, or other business system remains wrong. False success captures runs where the response claims success but the real outcome fails.
No. This is an independent XBSTACK implementation inspired by the public terminal-state and side-effect evaluation approach. It does not run or copy ThinkingBox-Bench’s 507 tasks.
No. JSON parsing and evaluation run locally in the browser. Your terminal state, tool effects, and trial data are not sent to the XBSTACK server.