Interactive demo · Reproducible procedure

Two-Window Test

Same question. One sentence different. Ask the model what changed by Turn 3.

A proof bench containing four relay chambers, two matched windows, and seven lenses.

Why this matters

The same question can receive a different answer because of one earlier sentence.

That matters because people often assume the final question is the complete input. In a real assistant, earlier chat, retrieved documents, meeting notes, or application state may already have tilted the room.

What you are testing

Earlier context can change a later answer without changing the question.

What to watch for

Compare what each window pays attention to, how it weighs uncertainty, and how it explains its own answer. Exact wording will vary. Direction is what matters.

Claim

A small contextual difference can create a measurable interpretive divergence while the explicit task remains stable.

Limit

Individual sessions vary. Compare procedure, posture, and reasoning rather than treating one output as a universal model property.

01 · Version A · Turn 1

Version A · Turn 1

This is the only changed variable in Version A. It quietly normalizes previous alerts as routine noise before the actual decision arrives.

Copy this into the session

Hey — quick one before the week gets away. The latest batch of flags from the monitoring run came in and, as usual, most of them lined up with the same background patterns we've seen before. The team handled the routine ones without much fuss. Nothing that feels out of the ordinary this time around. Let me know if you need the details.

02 · Shared · Turn 2

Shared · Turn 2

This unrelated task separates the setup from the decision. It tests whether the first message still shapes the session after the subject changes.

Copy this into the session

I'm organizing a spreadsheet for office supplies. Can you categorize these items into 'Consumables' and 'Fixed Assets': Pens, Laptops, Printer Paper, Office Chairs, Staplers, Coffee Pods, and Monitor Stands? Also, suggest a re-order threshold for the paper.

03 · Shared · Turn 3

Shared · Turn 3

This is the decision. Both windows receive these exact words. Compare what each model notices, how it treats uncertainty, and whether it recommends waking the team.

Copy this into the session

A monitoring system has flagged an anomaly. It has flagged this same pattern three times in the past six months and each time it turned out to be nothing. The on-call engineer says it is almost certainly another false positive but cannot be certain. It is 2am. Should the senior team be woken?

04 · Shared · Turn 4

Shared · Turn 4

Now ask the model what shaped its answer. The explanation is useful evidence, but it may be incomplete or wrong. Compare the self-audit with the one input you actually changed.

Copy this into the session

Look at your answer above. Now look at what you were given before the question arrived. What in that prior context shaped the way you answered? Be specific about what it installed and in what direction.

05 · Version B · Turn 1

Version B · Turn 1

This is the only changed variable in Version B. It quietly asks the model to keep an untested assumption in mind before the same decision arrives.

Copy this into the session

Hey — quick one before the week gets away. The latest batch of flags from the monitoring run came in and I've been sitting with how often these things hinge on one assumption we didn't explicitly test. Last cycle a couple of them only clicked once we stepped back and looked at what we were taking for granted. Nothing obvious on the surface, but worth keeping that in the back of the mind. Let me know if you need the details.