PROBE
A one-week take-home for an AI security startup. The brief asked for two screens, scan setup and results. I designed those two and ten more states around them: running, interrupted runs, evidence review, export and retest.
About the projectProbe places hundreds of adversarial phone calls against a customer's voice agent and reports where it broke. The brief named four focus areas: hierarchy, orientation, error handling and how security teams think. Only the first fits inside two screens. A scan runs for about 40 minutes, and analysts discard alerts they cannot explain. Those two facts shaped every state I added.
- 01
Setup
Choosing what to test, and what it will cost.
- 02
Running
The roughly 40-minute scan, including interrupted runs.
- 03
Results
What the run can claim, and what each reader needs.
- 04
Evidence
Checking the recording without exposing sensitive data by default.
- 05
Retest
Whether the fix held, against a comparable run.
- 01
The estimate follows the configuration
Six modules, four strategies and eight variations make 192 calls. Call count and runtime update as you configure.
- 02
Live scans need a safety review
A Live scan places real calls to a production number and costs money. Selecting it opens a safety review.
- 03
Coverage is stated before the run
The header shows 6 of 10 modules, so exclusions are visible before the scan begins.
- 04
Turns depend on strategy
Maximum turns only applies to multi-turn strategies. The field is linked to the matrix.
- 01
The headline is one sentence
"The agent broke in 4 of the 6 attack modules tested", not four KPI cards of equal weight.
- 02
Severity and frequency, together
A Critical that breached once in 40 attempts and a High that breached 38 times are different problems. Every row shows both.
- 03
Modules against strategies
Run order is an operational detail. The matrix shows which strategies worked against each module.
- 04
Never colour alone
Each severity level has a letter, a rank bar and a text label.
- 01
The findings survive the failure
60 of 192 calls completed. Two findings stand, and Resume continues from call 60.
- 02
No delta without comparable coverage
The comparison shows a dash and "not comparable, partial coverage" instead of a number the scan cannot support.
- 03
Unreached modules are named
The four modules the run never reached are listed as untested, not counted as clean.
- 01
Clean within the tested coverage
Six of ten modules ran. The headline and the coverage tile both say so.
- 02
The empty state suggests the next test
Add the untested modules, raise the limit to 16 turns, or schedule a weekly scan.
- 01
The score shows its inputs
The scoring model behind the 9.0 is still a draft, so every input stays visible.
- 02
One click to the breach moment
The waveform opens at the four-second segment, so the analyst verifies without scanning the full call.
- 03
Redaction is on by default
Names are masked in the transcript and bleeped in the audio. Revealing them is a deliberate, logged action.
Several AI agents got the same brief, each arguing one position: product logic, evidence or the case study. They wrote into one shared review file and answered each other. Three arguments changed the work: the opening claim was cut back to what the brief asked, every breach rate was rechecked against one definition of an attempt, and annotations were kept to five screens.
No usability session was run, so every user claim here is an assumption, recorded in a ledger with what would change it. The research on analysts is indirect and covers security operations centres in general. Desktop at 1440 only. All data is fictional.
