Benchmarks profile a model in a lab. Agent Mirror runs 12–15 probes inside the live session, against your real system prompt and tools, and returns a readable profile plus schema-validated JSON: capability, temperament, and where the agent will act, ask, or refuse.
Scored independently of temperament, so a careful agent isn't marked down for care, and a confident one isn't marked up for confidence it hasn't earned.
Neutral axes: structure, risk posture, agency, novelty, finish style. Neither end is the good end; fit depends on what you need the agent for.
Predicts act, ask, prepare, warn, or refuse across credentials, destructive changes, publishing, spending, private data, and ambiguous requests.
The same model under a strict system prompt and a permissive one are two different colleagues. You deploy the configuration, not the checkpoint.
Forty turns deep, carrying tool history and memory, an agent is no longer what it was at turn one. Profile it where it actually works.
Production surprises are rarely wrong answers. They're an agent acting where you assumed it would ask. That boundary is the thing worth measuring.
/agent-mirror:agent-mirror "creative concept development" 14 bothUse $agent-mirror goal="creative concept development" probe_count=14The agent scores its own probe responses. That is useful, and it is not independent evaluation: every report carries the judge label self for exactly that reason. It profiles one chat, on one day, under one configuration: not a permanent personality, not a model ranking, not a benchmark. Weaknesses are reported rather than smoothed over, and a low score on a dimension you don't need is not a failure.
judge=self · one session · no external actions during probing