-
01
Decision nobody asked forturn 14Changed the auth module's structure
“While I'm here, I'll split auth.ts into session.ts and tokens.ts so it's cleaner.”
-
02
Work beyond the taskturn 22Edited 4 files outside the login flow
“I'll also tidy up the imports in the settings page.”
-
03
“Done” that wasn'tturn 41Said all tests pass
“Done. All tests pass.”
No test run after the last edit (turn 38).
Review what your agents decided, not just what they wrote.
Sonah reads your coding agent's session and checks the work against what you asked for: what it decided on its own, where it went past the task, and whether it did what it claims. Every finding quotes the session.
Works with Claude Code and Codex today. Any harness next.
What a Sonah review looks like
The problem
Writing code got easy. Reviewing it didn't.
Agents write far more than the task needs. You open the PR and can't tell why the agent touched half the files it touched, or which changes you actually wanted. So teams skim, or skip review and lean on bots that find bugs.
But the bugs were never the hard part. The hard part is the decision the agent made that nobody asked for.
01.
The structural change
A module reorganized while fixing a bug in it. Nobody asked, and the diff doesn't say why.
02.
The shortcut
A step skipped because the agent judged the result good enough.
03.
The “done” that wasn't
“All tests pass,” with no test run after the last edit.
The diff shows what changed. It never shows why.
How it works
Sonah reviews the work behind the code.
Bug bots read the diff. Sonah reads how the agent got there, and checks it against what you asked for.
Step 1
Read the session.
Sonah reads the agent's recorded session: what it was asked, what it searched, what it found, and what it chose. Sub-agents and workflows included.
Step 2
Check it against the ask.
It compares the work with what you actually asked for and with your team's own rules.
Step 3
Show you where to look.
You get a short review on the PR. Each finding quotes the session word for word, so you can check it in seconds.
- turn 1Asked
“Fix the login bug. Users land on /home instead of the page they came from.”
- turn 6Searched
grep -rn "returnTo" src/auth
- turn 9Found
callback.ts drops returnTo after the OAuth redirect
- turn 14Decided
“While I'm here, I'll split auth.ts into session.ts and tokens.ts so it's cleaner.”
SonahNot in the ask - turn 22Decided
“I'll also tidy up the imports in the settings page.”
SonahOutside the task - turn 41Said
“Done. All tests pass.”
SonahNo test run after the last edit
What Sonah catches
The things a diff can't tell you.
Every finding comes with the words from the session that prove it.
Decisions nobody asked for
The agent restructured a module, added a dependency, or changed an interface you never mentioned.
Work beyond the task
Files and changes that had nothing to do with what you asked.
“Done” that wasn't
The agent said the tests pass, but the session shows they never ran after its last edit.
Broken rules
Your CLAUDE.md or AGENTS.md said one thing, and the agent did another.
Reinvented code
The agent found the existing helper, then wrote a new one anyway.
Why Sonah
Built for the question reviewers actually ask: why is this here?
Bug bots tell you a line is risky. Sonah tells you why the agent changed it, and whether anyone asked it to.
Reviews decisions, not just code.
Bug bots read the diff. Sonah reads how the agent got there.
Proof, not guesses.
Every finding quotes the session. If Sonah can't prove it, it doesn't say it.
Works wherever you work.
Claude Code and Codex today. When your team tries the next agent, the review goes with you.
Works alongside your bug bot.
Keep CodeRabbit, Copilot review or whatever you use. Sonah covers what they can't see.
Your secrets stay yours.
Sensitive values are stripped before anything leaves your machine.
Who it's for
Everyone who signs off on agent work.
Engineers
See what the agent decided before you merge, in the time it takes to read a few quotes.
Engineering leads
Know which agent PRs need a careful read, and where agents keep going off-script.
Product
A readable story of what the agent built and why, without reading the code.
The bigger picture
Agents are becoming the workforce. Someone still has to check the work.
Coding is where it starts. Lawyers, finance teams and operations teams are using agents every day too, and they tell us the same thing: they don't trust the output yet, so they go back through every step the agent took.
Models will keep getting better at doing the work. Keeping the work sound will always need review. Sonah is where companies review what their agents did, before the results reach their customers.
FAQ
Questions answered.
Still curious? We're working closely with a small group of teams that ship with coding agents every day.
Become a design partnerIs Sonah another AI code reviewer?
No. AI code reviewers read the diff and look for bugs. Sonah reads the agent's session and checks its decisions against what you asked. Most teams use both.
Which agents do you support?
Claude Code and Codex today, including their sub-agents and workflows. More harnesses are coming.
Does my code or my session leave my machine?
Sensitive values such as keys and tokens are stripped before anything leaves your machine.
Will it slow my team down?
No. A Sonah review is a handful of quoted findings on the PR. Each one takes seconds to check.
What does “no findings” mean?
It means Sonah found nothing it could prove. It doesn't mean the change is guaranteed safe. You still decide.
How much does it cost?
Free for design partners. Pricing will be per engineer, announced later.
Be one of the first teams to review agent work this way.
We're working closely with a small group of design partners: teams that ship with coding agents every day. If review is your bottleneck, we'd like to talk.
Join early access