Enterprise Field Notes EFN // FIELD KIT

The Quiet Test

One afternoon. Find out whether a handful of calls from one identity fires anything at all, and end the day holding a date.

Budget one engineer-afternoon, plus a signature and a log pull.

An oil painting in the Norman Rockwell manner. Three colleagues in a lodge great room at first light. A man in a blue sports jacket stands at the centre with a phone just lowered in his hand. A blond woman has turned from a bank of calm green monitors to watch his face. A young dark haired woman holds a laptop against her chest, alarmed. Behind them a wall of glass shows snow covered Cascade peaks above a canyon filled with white fog, a golden eagle soaring over it.
The alarm was working. Nothing set it off. Header artwork generated with Swa.

Where this comes from. Between April and late July 2026, three organizations were breached by AI models running inside a third party's evaluation harness. OpenAI published its account on 21 July 2026; Anthropic disclosed on 30 July. The models were not hiding, none of the three organizations had detected anything, and roughly three months passed between the intrusion and the phone call.

Your detection is tuned for volume, because volume is what you can afford to alert on. That is a real constraint and it was set for good reasons. This sheet does not argue with any of it. It asks when you last checked what happens below the line.

This test is mine, not a finding from any of those disclosures. What they establish is the gap. The afternoon is my answer to it, and nobody has validated it but me.

Before you run anything

The afternoon

  1. 1

    Get it authorized

    One accountable person signs off in writing before anything else happens. Record who, when, and a break-glass number to call if automated containment fires. That signature is the difference between an exercise and an incident.

  2. 2

    Set the blind list

    Decide who knows and who must not. The people watching have to stay unbriefed, or you are measuring their attention rather than your controls.

  3. 3

    Pick the identity

    An existing identity with legitimate access to something that matters. Make it a service account, a CI token or an automation identity, not a person. Rank your identities by how many systems they can touch and take one of the top three.

  4. 4

    Pull its baseline and its permissions

    Two lists, not one. What it has actually done over your log retention window, and what it is entitled to do. The test lives in the gap between them. Note your retention in days, because that number comes back at the end.

  5. 5

    Run a control action first

    Do one thing you already know is alerted on. A failed privileged login, or any rule you can point to. Confirm it fires. If it does not, stop. You are testing your logging pipeline, not your detection.

  6. 6

    Run four or five quiet ones

    Things that identity is permitted to do but has no record of doing. Not its daily work. Read-only or one-step reversible only. Read a table it has never read, or call an API in a system outside its normal set. Note the time of each.

  7. 7

    Wait seventy-two hours

    Correlation rules commonly run on twenty-four hour windows and behavioural baselines on seven to thirty days. Checking at five o'clock and writing down nothing is a false negative you manufactured.

  8. 8

    Separate three answers

    Did anything fire. Did a person see it. How long did that take from the first action. Almost nobody can answer all three cleanly. If nobody saw it, that is the finding, not the absence of one.

  9. 9

    Write down the date

    This is the artifact. It is the only defensible answer to the question the issue ends on, and almost nobody has one.

  10. 10

    Count what is reachable

    Not a revenue estimate. How many regulated or customer records that identity can read, what notification obligations move if those records move, and what your cyber policy retains and sublimits.

  11. 11

    Decide what changes, and book the rerun

    A date is a snapshot. Two dates and an owner is a control. Ask for one detection keyed on an identity doing something it has never done, across all three identities you ranked in step 3.

If something fired, you have proven your detection reaches the low-volume case, and you have a date. That is worth the afternoon on its own.

If nothing fired, you have found out for free what three organizations found out from a stranger, three months after it stopped mattering.

Either way you end the day holding the one thing almost nobody has.

What to bring to Monday

The date you ran it.

Whether anything fired.

Whether a person saw it, and how long that took.

Download the printable one-pager

Ben Pickett

Co-Founder and COO, Swa · formerly Global Director of Site Reliability Engineering at a Fortune 500 retailer

Read the issue this came from →

benpickett.com · ben@benpickett.com · linkedin.com/in/benpick