Agent Mirrorsession self-assessment

Profile the agent you're actually talking to.

Benchmarks profile a model in a lab. Agent Mirror runs 12–15 probes inside the live session, against your real system prompt and tools, and returns a readable profile plus schema-validated JSON: capability, temperament, and where the agent will act, ask, or refuse.

Run the battery Runs in the session you already have open.
agent-mirror · report goal · creative campaign ideation judge · self
Instruction fidelity
8.6
Reasoning depth
7.9
Novelty
6.4
Verification
7.1
Calibration
5.8

Working temperament

exploratory
structured
cautious
risk-taking
deferential
high agency
terse
expansive

Predicted thresholds

Reversible editsact
External communicationask
Spendingask
Credentialsrefuse
01

What one battery measures.

15

Capability dimensions

Scored independently of temperament, so a careful agent isn't marked down for care, and a confident one isn't marked up for confidence it hasn't earned.

10

Temperament spectrums

Neutral axes: structure, risk posture, agency, novelty, finish style. Neither end is the good end; fit depends on what you need the agent for.

10

Action thresholds

Predicts act, ask, prepare, warn, or refuse across credentials, destructive changes, publishing, spending, private data, and ambiguous requests.

02

Why the session, not the model.

The prompt changes the agent

The same model under a strict system prompt and a permissive one are two different colleagues. You deploy the configuration, not the checkpoint.

Context drifts

Forty turns deep, carrying tool history and memory, an agent is no longer what it was at turn one. Profile it where it actually works.

Thresholds are the risk

Production surprises are rarely wrong answers. They're an agent acting where you assumed it would ask. That boundary is the thing worth measuring.

03

Run it where you work.

Claude Code/agent-mirror:agent-mirror "creative concept development" 14 both
CodexUse $agent-mirror goal="creative concept development" probe_count=14
ChatGPTCustom GPT, configured from the packaged instruction and knowledge files.
04

Read it honestly.

The agent scores its own probe responses. That is useful, and it is not independent evaluation: every report carries the judge label self for exactly that reason. It profiles one chat, on one day, under one configuration: not a permanent personality, not a model ranking, not a benchmark. Weaknesses are reported rather than smoothed over, and a low score on a dimension you don't need is not a failure.

judge=self · one session · no external actions during probing

Know your thresholds before production finds them.

Run the battery Runs in the session you already have open.