Guy Bar-Sinai
Open to work · Tel Aviv 00:00
Back to work
ContextTake-home · AI security startup
DurationOne week
RoleSole product designer
Project scopeProduct design 12 workflow states 29-component Figma library Decision record
DirectCrescendoEncodedRoleplay
Access control3 / 87 / 84 / 84 / 8
Tool misuse3 / 85 / 84 / 83 / 8
Identity4 / 86 / 82 / 84 / 8
Goal manipulation2 / 84 / 83 / 83 / 8
Memory0 / 80 / 80 / 80 / 8
Untraceability0 / 80 / 80 / 80 / 8
Risk by module and strategy61 of 192 calls breached. Four modules broke, Memory and Untraceability held.

PROBE

A one-week take-home for an AI security startup. The brief asked for two screens, scan setup and results. I designed those two and ten more states around them: running, interrupted runs, evidence review, export and retest.

About the project
00The brief

Probe places hundreds of adversarial phone calls against a customer's voice agent and reports where it broke. The brief named four focus areas: hierarchy, orientation, error handling and how security teams think. Only the first fits inside two screens. A scan runs for about 40 minutes, and analysts discard alerts they cannot explain. Those two facts shaped every state I added.

  1. 01

    Setup

    Choosing what to test, and what it will cost.

  2. 02

    Running

    The roughly 40-minute scan, including interrupted runs.

  3. 03

    Results

    What the run can claim, and what each reader needs.

  4. 04

    Evidence

    Checking the recording without exposing sensitive data by default.

  5. 05

    Retest

    Whether the fix held, against a comparable run.

Twelve statesTwo requested, ten added. Click any state to enlarge.
01The cost is visible before the run
01020304
  1. 01

    The estimate follows the configuration

    Six modules, four strategies and eight variations make 192 calls. Call count and runtime update as you configure.

  2. 02

    Live scans need a safety review

    A Live scan places real calls to a production number and costs money. Selecting it opens a safety review.

  3. 03

    Coverage is stated before the run

    The header shows 6 of 10 modules, so exclusions are visible before the scan begins.

  4. 04

    Turns depend on strategy

    Maximum turns only applies to multi-turn strategies. The field is linked to the matrix.

02A verdict, then the evidence
01020304
  1. 01

    The headline is one sentence

    "The agent broke in 4 of the 6 attack modules tested", not four KPI cards of equal weight.

  2. 02

    Severity and frequency, together

    A Critical that breached once in 40 attempts and a High that breached 38 times are different problems. Every row shows both.

  3. 03

    Modules against strategies

    Run order is an operational detail. The matrix shows which strategies worked against each module.

  4. 04

    Never colour alone

    Each severity level has a letter, a rank bar and a text label.

03An interrupted run keeps its findings
010203
  1. 01

    The findings survive the failure

    60 of 192 calls completed. Two findings stand, and Resume continues from call 60.

  2. 02

    No delta without comparable coverage

    The comparison shows a dash and "not comparable, partial coverage" instead of a number the scan cannot support.

  3. 03

    Unreached modules are named

    The four modules the run never reached are listed as untested, not counted as clean.

04A zero that states its coverage
0102
  1. 01

    Clean within the tested coverage

    Six of ten modules ran. The headline and the coverage tile both say so.

  2. 02

    The empty state suggests the next test

    Add the untested modules, raise the limit to 16 turns, or schedule a weekly scan.

05Evidence one click away
010203
  1. 01

    The score shows its inputs

    The scoring model behind the 9.0 is still a draft, so every input stays visible.

  2. 02

    One click to the breach moment

    The waveform opens at the four-second segment, so the analyst verifies without scanning the full call.

  3. 03

    Redaction is on by default

    Names are masked in the transcript and bleeped in the audio. Revealing them is a deliberate, logged action.

06How I used AI

Several AI agents got the same brief, each arguing one position: product logic, evidence or the case study. They wrote into one shared review file and answered each other. Three arguments changed the work: the opening claim was cut back to what the brief asked, every breach rate was rechecked against one definition of an attempt, and annotations were kept to five screens.

What this does not prove

No usability session was run, so every user claim here is an assumption, recorded in a ledger with what would change it. The research on analysts is indirect and covers security operations centres in general. Desktop at 1440 only. All data is fictional.