The criteria, six AI-generated designs, two AI scorers, no winner, and a hybrid. Then the limits, and a log of who decided what.
This is the design decision.
A design decision is about setting the criteria to evaluate which idea wins.
— the working rule for this studyThree hard gates first. An option that fails one is out, whatever else it scores.
Then six weighted criteria, each scored across five dimensions — user impact, failure severity, frequency, strategic alignment, competitive leverage — to earn its weight rather than be assigned one.
| Criterion | Weight | Because |
|---|---|---|
| Found on the view the case opens to | 3× | An unfound verdict plays no part in the decision |
| Confidence that changes the card | 3× | People over-rely on confident-looking automated advice (Parasuraman & Manzey, 2010) |
| Says what it saw, and when | 2× | Alerts can join the case after the analysis ran |
| Accept or override, consequence shown first | 2× | One disposition can change two systems |
| Evidence checkable inside the case | 2× | The conclusion is currently a truncated string |
| Disagreement between the two severities shown | 2× | Nothing on screen reconciles them today |
Proposed at 17 Sep, 18:17, ninety minutes before the first option existed: seven criteria, with a recommendation to cut one. The six were confirmed on 18 Sep, after all six options existed and before either scorer ran. That is a proposal, not pre-registration. The seventh became a briefed requirement instead, checked and reported.
Corrected 25 Sep 2026. This page said the criteria were “committed 90 minutes before the first option”. Seven were proposed then, and the six were confirmed only after all six options existed. That is not pre-registration. The claims on this page that rested on it have been taken down, as has the claim that the counting rules were published before the count: they were committed together.
Five were generated by an agent working from a written brief. The brief is the work; the generation is a method note.
The brief carried the criteria, the three gates, the host’s documented extension points, the synthetic case data and the six states every option had to draw — including the ones nobody enjoys drawing: analysis pending, analysis failed, and no widget installed at all. Each run got the same brief and one difference in how it was seeded. One run was given no direction beyond the rules and the data.
The sixth is Option 0: the original ask made literal — verdict, confidence, evidence, provenance, accept and override — generated by a model, like the other five, before the criteria were applied to anything. It is in the set so the criteria have something unexamined to be measured against.
Corrected 25 Sep 2026. This page said Option 0 was mine — “the one I built myself”. It was generated by a model, like the other five, and the page now says so.






Option 0 runs in the prototype too: open it.
Two AI scorers worked the same rubric independently, and neither could see the other’s notes. They are the same kind of reader, though: a human pass is still missing.
They agree on first place and disagree on second. AI scorer B put R5 5.0 points clear of R3; AI scorer C put it 8.6 clear of R2. Neither margin clears the 10-point close-call rule fixed before any option existed, so no winner is declared. R5, the run given no direction at all, leads both passes; the runner-up changes identity between them.
Two findings do more work than the ranking.
What the refusal forced. With no winner, the answer could not be a coronation. It is a hybrid: the parts that scored best, wherever they came from. And three things no option produced.
The card, moment by moment, and where each part came from.
| Moment | What the card does |
|---|---|
| The verdict lands | Names the agent, the data it read, the time it stopped reading, and both severities with the system that issued each |
| A newer alert arrives | “Outdated. Alert A3 arrived 13:57 UTC, 43 minutes after the data this verdict used. The agent has not seen it.” With a control that goes there |
| The analyst decides | Every action carries its consequence before it is pressed. The record names what was accepted. Undo is a control, not a sentence about undo |
| The vendor changes its mind, six minutes later | The change announces itself. Read it, and the card swaps in the new verdict, its new severity, its new claims and its new data cut — keeps the earlier conclusion on the record — and the outdated warning correctly clears |






What it refuses to do. No colour down the card’s left edge, where the host shows priority. No host field changes without an explicit act. One decorative mark on the whole card, and it sits after words.
The card is then put through the study’s own gate, mechanically, and passes. A check that only ever passes proves nothing, so it runs on Option 0 too, and names its failures there.
A verdict produced by one vendor is read inside another vendor’s product, by an analyst who owes their attention to the case, not to the agent. The agent is a guest. These are the obligations that come with being one, each traced to a row in the count or a criterion in section 01. The list itself is on the case study page.
A rule that only fits one pairing isn’t portable, so an AI agent checked all eight on paper against a different guest in a different kind of host: Dropzone AI writing into ServiceNow, where the AI polls the ticket queue, investigates in its own console, and posts its result back as a comment. By its check, seven of the eight apply as written. Rule 5 needed rewording, because in a ticketing tool most evidence lives outside the host, so “in the host’s own objects” became “and say where” — which the card already does. And rule 7 flipped: there, a comment is the only channel, so the text entry isn’t the fallback — it is the card. That’s why the freshness line is the first thing I’d ship. This tested whether the rules apply, not whether Dropzone follows them.
No analyst has used this. n = 0. The design rules stand as argued, not as tested.
The protocol is written — consent, three tasks, time to first correct action, five debrief questions, a scoring sheet. The prototype scores its own sessions and exports them, and a self-serve link needs no calendar and no facilitator. That is the next thing that happens, and whatever n it reaches gets printed first, in the findings.
Written before any data, so it cannot be adjusted afterwards.
The freshness line and the claim-level checkability. Both are cheap, both work in a plain text entry as well as in a widget, and between them they carry the two failures the audit found: a verdict whose age nothing on screen states, and a claim you cannot follow.
The preview-drawer density, until someone confirms a vendor widget renders in that drawer at all. The docs do not say, and I could not check without a tenant. It is drawn because rule 1 demands it, not because it is known to be buildable.
Claim-level checkability assumes the agent’s claims map onto objects the host holds. On a case with one alert that is easy. On a case at the host’s own documented ceiling, “checkable in this case” stops being one click and starts being a search, and the card would need to say which.
And the card only sees what lands in its case. The host groups a new alert into an open case when the two share a single entity, inside a window its admin sets. It opens a separate case if the window is shorter than 41 minutes, or if the case is already closed, because grouping never reopens one. That second path is the worst version of this whole problem: the analyst trusts the verdict and closes the case, and the evidence against it arrives somewhere else. A card inside one case cannot see that. Catching it needs something that looks across cases, and this study did not design it.
A clean narrative is not the same as a tidy process. The prose above reads forward; this is the order things actually happened in, and who decided each one.
| Date | Decided by | Decision | What it changed |
|---|---|---|---|
| 17 Sep | Omkar | Write the hypothesis before gathering the evidence | The audit could refute it, and partly did |
| 17 Sep | AI, reviewed | Publish the counting rules with the count | A count that can be disputed, and re-run by script |
| 17 Sep | AI, reviewed | Propose seven criteria at 18:17, before any option exists | Confirmed as six on 18 Sep, after all six existed: a proposal, not pre-registration |
| 18 Sep | Omkar | Judge directions to develop, not finished candidates. Scope: guest only | Findings about the host stay findings, never proposals |
| 18 Sep | Omkar | Six criteria, not seven | A criterion every option satisfies cannot separate them |
| 22 Sep | AI, reviewed | Refuse to re-point the audit at the revamped view | No screenshots there; a count would be assertion |
| 23 Sep | AI, reviewed | Uphold the gate failure against Option 0 | The version generated from the original ask came last |
| 23 Sep | Omkar | Mix the top options rather than crown one | The answer is a hybrid, and the fitting is disclosed |
All six options were machine-generated, and two AI scorers scored them; the margins above are theirs. An AI agent also ran the count and drafted the criteria, the gates and the eight rules. What is mine is the problem, the scope, the cut from seven criteria to six, the three additions no option produced, and the call to mix the top options rather than crown one. The page says which is which because concealing it is the actual red flag.