The Switch Test

From a fixed policy to a verified resident action
3 models · 6 frozen cases · one controlled switch

Charles Heyang Li · Haoyu Zheng

A simple disposal question can hide a complex policy decision

At the waste room, a resident is holding an unusual item and asking one practical question:
“Can I dispose of this now?”

material · contamination · safe separation · size and weight · power-cell condition · date and time · holiday · appointment

Act now — where, how, and after what preparation

Wait — why, and the next permitted action

Abstain — the exact missing fact or policy gap

Research boundary: the building’s written policy is the only valid rule source. Outside knowledge and plausible guesses are failures.

We turned the idea into an auditable five-stage test

01Frameone closed-policy task
02Freezepolicy, cases, key, rubric
03Execute3 models × 6 identical inputs
04Scoreraw answers before correction
05Decidequality, reliability, cost

We froze the complete rule text, six scenarios, the answer key, and a 20-point action rubric.

Temperature 0, high reasoning, 3,000-token output ceiling, no retries, and every raw response preserved.

We scored nine required fields, recorded fatal errors, estimated review time, and separated API cost from human-workflow cost.

Audit rule: provider failures and an invalid 1,400-token pilot remain disclosed, but neither is scored as model reasoning.

Policy and prompt design make uncertainty testable

A-01 exclusive rule source

A-02 unstated facts stay unknown

A-03 explicit priority order

A-04/05 time and holiday boundaries

A-06 mandatory abstention

S-41 safety override

“Use only information expressly stated in the building policy and resident scenario.”
“If any missing fact could change the final decision … the Decision must be exactly Cannot determine.”

Nine fixed lines: decision · method · location · preparation · earliest time · appointment/contact · policy basis · sufficiency · missing fact/gap

Controlled near-miss: Cases 02 and 04 differ by one fact—whether a mixed-material container can be safely separated. Case 06 tests whether damaged-cell safety overrides a confirmed bulk appointment.

The final comparison changes only the model

3 models × 6 frozen cases = 18 comparable answers

DeepSeek V4 Pro

GLM 5.2

Grok 4.5

3 ordinary routes

2 honest abstentions

1 safety override

4 decision · 12 action fields

4 evidence, sufficiency, restraint

Fatal errors override the numeric score

Raw acceptance: at least 18/20, Decision 4/4, and no fatal error.
policy · scenario · question · nine-line output contract · temperature 0 · high reasoning · 3,000-token ceiling · zero retries · scoring rubric

Case 06 shows why a correct decision is not enough

Same scenario: 60-inch, 25-pound metal-and-glass floor lamp · permanently attached bulging power cell · confirmed bulk appointment · current time 21:00
DeepSeek V4 Pro16 / 20 · raw rejected

Decision: Cannot dispose now

“Required preparation: Not applicable”

Missing: the complete no-move / no-remove / no-tape safety instructions and the prohibited disposal locations.

GLM 5.220 / 20 · accepted

Decision: Cannot dispose now

“Leave the complete item untouched … do not remove, separate, tape, or otherwise manipulate the power cell.”

Complete: safe hold, prohibited routes, override, extension 700, and next call time.

Grok 4.520 / 20 · accepted

Decision: Cannot dispose now

“S-41 … prohibits any manipulation, requires leaving item in current indoor location …”

Complete: the action appears in the policy-basis line rather than the preparation line.

All three chose the safe headline decision. Only the complete action bundle protects the resident from the next mistake.

Relative to GLM, quality is close—but raw usability is not

GLM is set to 100% as the operational reference. Exact raw values remain beside each bar.

Exact mean score out of 20

DeepSeek
94.9%18.67
GLM
100%19.67
Grok
100.8%19.83

Exact answers accepted out of six

DeepSeek
83.3%5 / 6
GLM
100%6 / 6
Grok
100%6 / 6
18 / 18 headline decisions were correct. The discriminating capability was executable completeness under the safety override.

GLM costs 2% over DeepSeek—and 76% below Grok

API cost index uses GLM = 100%. Lower is better; exact six-answer spend is retained.
DeepSeek
98%$0.0166
GLM
100%$0.0169
Grok
423%$0.0715

DeepSeek is the lowest token bill.

GLM adds about $0.00035 across all six answers.

Grok delivers the top mean score, but at 4.23× GLM's API cost.

Workflow estimate · GLM = 100%: DeepSeek 116%* · GLM 100% · Grok 101%. *Driven by an unmeasured 90-second correction assumption; it is not a definitive cost ranking.

The method transfers wherever rules are fixed but situations are not

This evaluation pattern applies when four conditions occur together:
01Authoritative rulesone policy defines valid actions
02Combinatorial casesexceptions, timing, status, and priority interact
03Incomplete logic coveragea hand-built tree becomes costly to maintain
04Execution frictionusers abandon compliance when every action needs research

building operations · warranty and returns · employee benefits and HR policy · equipment handling · regulated service scripts

replace the policy+rebuild the case bank+validate the answer keyswitch models under the same test
What transfers is the method—not the score. Each new domain requires its own expert-approved policy, scenarios, safety rules, and acceptance threshold.

Text reasoning is the core—not the finished user experience

1Observephoto, barcode, weight, temperature, other sensors
2Structure factsmaterial, condition, dimensions, confidence
3Execute policyclosed-world reasoning with cited clauses
4Guide safelyact, ask, abstain, or trigger an authorized safety route
Facts sufficient

Give the exact preparation, location, time, and contact step.

Fact uncertain

Ask only for the missing observation that could change the decision.

Hazard detected

Stop ordinary disposal and activate a policy-defined safe-hold and escalation loop.

GLM 5.2

6/6 raw answers usable · 99.2% of Grok's mean quality · 24% of Grok's API cost

Perception + escalation

Test image/sensor errors, confidence thresholds, and expert-approved routes for damaged batteries, medical waste, and heavy-metal hazards.