← Omkar Khadamkar Verdict card
Verdict card · the method

How it was made

The criteria, six AI-generated designs, two AI scorers, no winner, and a hybrid. Then the limits, and a log of who decided what.

Scope the vendor’s card, not the consoleAnalyst sessions 0
WHAT HAPPENED WHERE IT TURNED 1AI counted it 72 marks that are not text 51 cannot be read without leaving the console 2Set the criteria Six criteria, weighted Proposed before the first option; confirmed after all six existed 3Six answers Five to one brief, plus one, literal The one built from the brief came last in both passes 4Scored twice Two AI scorers, blind to each other No winner: both margins fell inside the gate 5Wrote the rule Eight obligations a guest owes The host shipped a new view. All eight held
Five stages, and the turn in each. Every box links to the section that carries it.
Contents
  1. 01What good had to mean
  2. 02Six designs
  3. 03Scored, and no winner
  4. 04The card, part by part
  5. 05The rules, tested
  6. 06Limits
  7. 07Who decided what
01What good had to mean

This is the design decision.

A design decision is about setting the criteria to evaluate which idea wins.

— the working rule for this study

Three hard gates first. An option that fails one is out, whatever else it scores.

Then six weighted criteria, each scored across five dimensions — user impact, failure severity, frequency, strategic alignment, competitive leverage — to earn its weight rather than be assigned one.

CriterionWeightBecause
Found on the view the case opens to3×An unfound verdict plays no part in the decision
Confidence that changes the card3×People over-rely on confident-looking automated advice (Parasuraman & Manzey, 2010)
Says what it saw, and when2×Alerts can join the case after the analysis ran
Accept or override, consequence shown first2×One disposition can change two systems
Evidence checkable inside the case2×The conclusion is currently a truncated string
Disagreement between the two severities shown2×Nothing on screen reconciles them today

Proposed at 17 Sep, 18:17, ninety minutes before the first option existed: seven criteria, with a recommendation to cut one. The six were confirmed on 18 Sep, after all six options existed and before either scorer ran. That is a proposal, not pre-registration. The seventh became a briefed requirement instead, checked and reported.

Corrected 25 Sep 2026. This page said the criteria were “committed 90 minutes before the first option”. Seven were proposed then, and the six were confirmed only after all six options existed. That is not pre-registration. The claims on this page that rested on it have been taken down, as has the claim that the counting rules were published before the count: they were committed together.

02Six designs

Five were generated by an agent working from a written brief. The brief is the work; the generation is a method note.

The brief carried the criteria, the three gates, the host’s documented extension points, the synthetic case data and the six states every option had to draw — including the ones nobody enjoys drawing: analysis pending, analysis failed, and no widget installed at all. Each run got the same brief and one difference in how it was seeded. One run was given no direction beyond the rules and the data.

The sixth is Option 0: the original ask made literal — verdict, confidence, evidence, provenance, accept and override — generated by a model, like the other five, before the criteria were applied to anything. It is in the set so the criteria have something unexamined to be measured against.

Corrected 25 Sep 2026. This page said Option 0 was mine — “the one I built myself”. It was generated by a model, like the other five, and the page now says so.

R5's card at low confidence: the verdict, a limited-confidence banner naming the missing evidence line, a priority comparison block, and four actions labelled with bracketed words.
R5 · no seedGiven nothing but the rules and the data. Its action icons are bracketed words, and at low confidence it re-aims the action at the gap.First in both passes
R3's escalation confirmation: key changes listed, a sync-status list, a required analyst reason, and a split of what can and cannot be undone.
R3 · claim by claimThe only option where the required reason is a mechanism, not a label — the confirm button stays disabled until it is written.
R2's card leading with when the analysis ran and a bordered block comparing case priority with agent severity.
R2 · as of whenLeads with freshness, and writes each action’s consequence next to the action.
R4 drawn as a paused playbook step offering three choices, each captioned 'No case changes'.
R4 · ask firstThe most literal use of the host’s extension points: a paused playbook step with three choices, nothing pre-selected.
R1's fallback states: analysis pending, analysis failed with a retry, and the plain-text Activity entry for a tenant with no widget installed.
R1 · one line firstIts fallback states — pending, failed, and the plain-text entry for a tenant with no widget installed. Every run had to draw these.
Option 0: a verdict card with a red bar down its left edge marking the verdict, just below the case's own yellow priority bar.
Option 0 · the brief made literalModel-generated from the literal brief, before the criteria were applied. Note the red bar down the card’s left edge. It marks the verdict, but the host already uses a left-edge bar for priority, and there red means Critical.Last in both passes, and it fails two gates

Option 0 runs in the prototype too: open it.

03Scored, and no winner

Two AI scorers worked the same rubric independently, and neither could see the other’s notes. They are the same kind of reader, though: a human pass is still missing.

They agree on first place and disagree on second. AI scorer B put R5 5.0 points clear of R3; AI scorer C put it 8.6 clear of R2. Neither margin clears the 10-point close-call rule fixed before any option existed, so no winner is declared. R5, the run given no direction at all, leads both passes; the runner-up changes identity between them.

Two findings do more work than the ranking.

What the refusal forced. With no winner, the answer could not be a coronation. It is a hybrid: the parts that scored best, wherever they came from. And three things no option produced.

04The card, part by part

The card, moment by moment, and where each part came from.

MomentWhat the card does
The verdict landsNames the agent, the data it read, the time it stopped reading, and both severities with the system that issued each
A newer alert arrives“Outdated. Alert A3 arrived 13:57 UTC, 43 minutes after the data this verdict used. The agent has not seen it.” With a control that goes there
The analyst decidesEvery action carries its consequence before it is pressed. The record names what was accepted. Undo is a control, not a sentence about undo
The vendor changes its mind, six minutes laterThe change announces itself. Read it, and the card swaps in the new verdict, its new severity, its new claims and its new data cut — keeps the earlier conclusion on the record — and the outdated warning correctly clears
A confident verdict goes out of date, the analyst acts, and the vendor changes its mind Five moments: the agent stops reading at 13:14 and concludes Benign with high confidence at 13:16; an alert arrives at 13:57, forty-three minutes after the data it used; the analyst accepts Benign at 14:04; the vendor revises the verdict to Malicious at 14:10, after the decision. 13:16 Benign High confidence data to 13:14 13:57 Alert A3 arrives public access enabled 43 min after its data cut 14:04 The analyst accepts the record names what was accepted 14:10 Malicious the vendor changes its mind WHAT THE CARD DOES AT EACH MOMENT names the data cut → “Outdated. The agent has not seen it.” → keeps undo a control → the change announces itself, and the card swaps in the new verdict, its new claims and its new data cut The alert arrived 43 minutes after the data the verdict used, and 41 minutes after the analysis itself. Both numbers are true, and both are printed. Building it found two defects no drawing would have: a record reading “you accepted this verdict” is silently rewritten when the verdict changes, and a staleness warning keeps accusing an analysis that has since seen the alert. Both are fixed.
In the prototype: open CS-4133, accept the verdict, and wait twenty seconds. In the case's own clock the revision lands six minutes after the decision; in real time it arrives while you are still looking at the card, because a state that is already there when you look is not the state being designed for.
The card showing a Benign verdict with an amber band reading Outdated, naming the alert that arrived 43 minutes after the data the verdict used.
OutdatedFreshness as a stated condition, not two timestamps to compare.
The card at low confidence: no action recommended, the missing claim named, and a link to the alert that contradicts it.
Low confidenceThe recommendation is withdrawn, the missing claim is named, and the alert that contradicts it is one click away.
The card after a decision: a line reading You accepted Benign at 14:04 UTC, an Undo this button, and a note about what undo cannot take back.
DecidedThe record names what was accepted, so a later revision cannot quietly rewrite it. Undo is a button, with one line on the part that cannot be undone.
After the analyst accepted Benign, a bordered block announces that the vendor changed the verdict to Malicious.
The vendor changes its mindAfter the decision, not before it. Read the new verdict, or record that you stood by yours.
The revised card reading Malicious, with a line saying it was revised from Benign after the analyst's decision, the new conclusion, the earlier conclusion kept below it, and new claims.
The revised verdict, readNew verdict, new severity, new claims, new data cut — the earlier conclusion kept on the record, and the outdated warning correctly cleared.
The same card compressed into the host's narrow preview drawer, ending in a control that opens the case.
The preview drawerThe second width, which the host shipped on 18 September. Nothing is decided here.

What it refuses to do. No colour down the card’s left edge, where the host shows priority. No host field changes without an explicit act. One decorative mark on the whole card, and it sits after words.

Every part of the final card, with the option it came from Nine parts of the card. Six come from the scored options — R5, R2, R4, R3, Option 0. Three were decided by the author and produced by no option: low confidence withdraws the recommendation, undo is a control, and the verdict revised after the analyst acted. V1 · THE HYBRID CAME FROM Verdict and confidence, with both severities named and attributed When it was analysed, and whether anything is newer The staleness sentence, in plain words Every claim, with whether you can check it in this case Evidence links as controls, not prose The consequence written under the action, before it is pressed Low confidence withdraws the recommendation Undo is a control The verdict revised after you acted R5 · no seed R2 · as of when R4 · ask first R3 · claim by claim Option 0 · its one strength R2 · as of when Omkar Omkar Omkar The three in red were produced by no option in the set. They are the unmet-by-all list, and they are the parts of this card a machine did not propose.
Six options were machine-made and are labelled as such; the three additions are not. What the card deliberately does not carry is a verdict colour down its left edge. The host uses that device for priority, and it is one of the two gate failures that ruled Option 0 out.

The card is then put through the study’s own gate, mechanically, and passes. A check that only ever passes proves nothing, so it runs on Option 0 too, and names its failures there.

05The rules, tested

A verdict produced by one vendor is read inside another vendor’s product, by an analyst who owes their attention to the case, not to the agent. The agent is a guest. These are the obligations that come with being one, each traced to a row in the count or a criterion in section 01. The list itself is on the case study page.

Eight obligations a third-party finding owes the console it lands in Eight numbered rules: arrive where the work is at whatever width you are given; say what you saw and when; declare what you could not see; let the host's claim stand beside yours; make every claim checkable, and say where; change nothing without an explicit human act; degrade to text; explain your own marks. THE AGENT IS A GUEST. THESE COME WITH BEING ONE. 1Arrive where the work already is, at whatever width the host gives you. 2Say what you saw, and when. 3Declare what you could not see. 4Let the host's claim stand beside yours. 5Make every claim checkable, and say where. 6Change nothing without an explicit human act — and say what the act will do first. 7Degrade to text. 8Explain your own marks. got sharper got sharper
Each rule cites the audit row it came from, the criterion it serves, and what the five design runs did with it. This is the portable part: the Wiz-inside-Google-SecOps example dates, the obligations do not. The two marked rules are the ones the host's 18 September release sharpened.

A rule that only fits one pairing isn’t portable, so an AI agent checked all eight on paper against a different guest in a different kind of host: Dropzone AI writing into ServiceNow, where the AI polls the ticket queue, investigates in its own console, and posts its result back as a comment. By its check, seven of the eight apply as written. Rule 5 needed rewording, because in a ticketing tool most evidence lives outside the host, so “in the host’s own objects” became “and say where” — which the card already does. And rule 7 flipped: there, a comment is the only channel, so the text entry isn’t the fallback — it is the card. That’s why the freshness line is the first thing I’d ship. This tested whether the rules apply, not whether Dropzone follows them.

06Limits

No analyst has used this. n = 0. The design rules stand as argued, not as tested.

The protocol is written — consent, three tasks, time to first correct action, five debrief questions, a scoring sheet. The prototype scores its own sessions and exports them, and a self-serve link needs no calendar and no facilitator. That is the next thing that happens, and whatever n it reaches gets printed first, in the findings.

What would make this wrong

Written before any data, so it cannot be adjusted afterwards.

What I’d ship first

The freshness line and the claim-level checkability. Both are cheap, both work in a plain text entry as well as in a widget, and between them they carry the two failures the audit found: a verdict whose age nothing on screen states, and a claim you cannot follow.

What I’d cut

The preview-drawer density, until someone confirms a vendor widget renders in that drawer at all. The docs do not say, and I could not check without a tenant. It is drawn because rule 1 demands it, not because it is known to be buildable.

What breaks at scale

Claim-level checkability assumes the agent’s claims map onto objects the host holds. On a case with one alert that is easy. On a case at the host’s own documented ceiling, “checkable in this case” stops being one click and starts being a search, and the card would need to say which.

And the card only sees what lands in its case. The host groups a new alert into an open case when the two share a single entity, inside a window its admin sets. It opens a separate case if the window is shorter than 41 minutes, or if the case is already closed, because grouping never reopens one. That second path is the worst version of this whole problem: the analyst trusts the verdict and closes the case, and the evidence against it arrives somewhere else. A card inside one case cannot see that. Catching it needs something that looks across cases, and this study did not design it.

07Who decided what

A clean narrative is not the same as a tidy process. The prose above reads forward; this is the order things actually happened in, and who decided each one.

DateDecided byDecisionWhat it changed
17 SepOmkarWrite the hypothesis before gathering the evidenceThe audit could refute it, and partly did
17 SepAI, reviewedPublish the counting rules with the countA count that can be disputed, and re-run by script
17 SepAI, reviewedPropose seven criteria at 18:17, before any option existsConfirmed as six on 18 Sep, after all six existed: a proposal, not pre-registration
18 SepOmkarJudge directions to develop, not finished candidates. Scope: guest onlyFindings about the host stay findings, never proposals
18 SepOmkarSix criteria, not sevenA criterion every option satisfies cannot separate them
22 SepAI, reviewedRefuse to re-point the audit at the revamped viewNo screenshots there; a count would be assertion
23 SepAI, reviewedUphold the gate failure against Option 0The version generated from the original ask came last
23 SepOmkarMix the top options rather than crown oneThe answer is a hybrid, and the fitting is disclosed

All six options were machine-generated, and two AI scorers scored them; the margins above are theirs. An AI agent also ran the count and drafted the criteria, the gates and the eight rules. What is mine is the problem, the scope, the cut from seven criteria to six, the three additions no option produced, and the call to mix the top options rather than crown one. The page says which is which because concealing it is the actual red flag.