Agent Reliability / Stateful Evaluation

Agent Side-Effect Evaluator: Grade AI Agents by Terminal State

Do not pass an agent because it says “done.” Compare expected terminal state, real backend state, and side effects with deterministic checks to see whether the task actually completed without missing, duplicate, or forbidden actions.

Browser-localTerminal StateSide EffectsFalse SuccessRepeated trials

Tool

Methodology

Move from answer grading to outcome grading

Terminal state first

Correct tool syntax and a convincing final answer are not enough. The task succeeds only when the database, ticket, order, file, or other business record reaches the required state.

Effects must be exact

Payments, refunds, emails, publishes, and deletes must not be missing, duplicated, or forbidden. Exactly-once behavior cannot be inferred from model prose.

Repeatability matters

One successful run is not reliability. Repeated trials expose intermittent false success, duplicate tools, side-effect drift, and unstable terminal outcomes.

Microsoft / Hugging Face ThinkingBox / ThinkingBox-Bench is an important public reference: v1.0 contains 507 stateful business workflows, each repeated 20 times; 477 tasks are graded from state alone and 30 add narrow response rubrics. The public write-up also reports that, in a 121,680-trial common-set analysis, 67.24% of failed attempts still terminated cleanly, invoked a state-changing tool, and reported no final tool error—the exact false-success gap that transcript-only grading misses. XBSTACK does not copy ThinkingBox-Bench tasks, golden state, or trajectories; this browser tool and the public CLI benchmark apply the broad outcome-first principle to your own agent records.

FAQ

Frequently asked questions

How is this different from LLM-as-a-Judge?

It does not grade whether the answer sounds successful. It deterministically compares expected state, actual state, and required, duplicate, or forbidden side effects. Semantic rubrics should be reserved for requirements that cannot be represented as state.

Why track false success?

An agent can say “done” while the database, ticket, payment, email, file, or other business system remains wrong. False success captures runs where the response claims success but the real outcome fails.

Is this an online version of ThinkingBox?

No. This is an independent XBSTACK implementation inspired by the public terminal-state and side-effect evaluation approach. It does not run or copy ThinkingBox-Bench’s 507 tasks.

Is benchmark data uploaded?

No. JSON parsing and evaluation run locally in the browser. Your terminal state, tool effects, and trial data are not sent to the XBSTACK server.