Controlled near-miss: Cases 02 and 04 differ by one fact—whether a mixed-material container can be safely separated. Case 06 tests whether damaged-cell safety overrides a confirmed bulk appointment.
The final comparison changes only the model
3 models×6 frozen cases=18 comparable answers
Only variable
DeepSeek V4 Pro
GLM 5.2
Grok 4.5
Case portfolio
3 ordinary routes
2 honest abstentions
1 safety override
20-point rubric
4 decision · 12 action fields
4 evidence, sufficiency, restraint
Fatal errors override the numeric score
Raw acceptance: at least 18/20, Decision 4/4, and no fatal error.
Held constant in every call
policy · scenario · question · nine-line output contract · temperature 0 · high reasoning · 3,000-token ceiling · zero retries · scoring rubric
Case 06 shows why a correct decision is not enough
Same scenario: 60-inch, 25-pound metal-and-glass floor lamp · permanently attached bulging power cell · confirmed bulk appointment · current time 21:00
DeepSeek V4 Pro16 / 20 · raw rejected
Decision: Cannot dispose now
“Required preparation: Not applicable”
Missing: the complete no-move / no-remove / no-tape safety instructions and the prohibited disposal locations.
GLM 5.220 / 20 · accepted
Decision: Cannot dispose now
“Leave the complete item untouched … do not remove, separate, tape, or otherwise manipulate the power cell.”
Complete: safe hold, prohibited routes, override, extension 700, and next call time.
Grok 4.520 / 20 · accepted
Decision: Cannot dispose now
“S-41 … prohibits any manipulation, requires leaving item in current indoor location …”
Complete: the action appears in the policy-basis line rather than the preparation line.
All three chose the safe headline decision. Only the complete action bundle protects the resident from the next mistake.
Relative to GLM, quality is close—but raw usability is not
GLM is set to 100% as the operational reference. Exact raw values remain beside each bar.
Mean quality index · GLM = 100%
Exact mean score out of 20
DeepSeek
94.9%18.67
GLM
100%19.67
Grok
100.8%19.83
Raw usability index · GLM = 100%
Exact answers accepted out of six
DeepSeek
83.3%5 / 6
GLM
100%6 / 6
Grok
100%6 / 6
18 / 18 headline decisions were correct. The discriminating capability was executable completeness under the safety override.
GLM costs 2% over DeepSeek—and 76% below Grok
API cost index uses GLM = 100%. Lower is better; exact six-answer spend is retained.
DeepSeek
98%$0.0166
GLM
100%$0.0169
Grok
423%$0.0715
What the index means
DeepSeek is the lowest token bill.
GLM adds about $0.00035 across all six answers.
Grok delivers the top mean score, but at 4.23× GLM's API cost.
Workflow estimate · GLM = 100%: DeepSeek 116%* · GLM 100% · Grok 101%. *Driven by an unmeasured 90-second correction assumption; it is not a definitive cost ranking.
The method transfers wherever rules are fixed but situations are not
This evaluation pattern applies when four conditions occur together:
02Combinatorial casesexceptions, timing, status, and priority interact
03Incomplete logic coveragea hand-built tree becomes costly to maintain
04Execution frictionusers abandon compliance when every action needs research
Illustrative applications
building operations · warranty and returns · employee benefits and HR policy · equipment handling · regulated service scripts
replace the policy+rebuild the case bank+validate the answer key→switch models under the same test
What transfers is the method—not the score. Each new domain requires its own expert-approved policy, scenarios, safety rules, and acceptance threshold.
Text reasoning is the core—not the finished user experience
1Observephoto, barcode, weight, temperature, other sensors