Multiple agents. One review. No drama.
Magnum is a private investigator for pull requests. It watches your GitHub repositories all day, checks every eligible PR out into its own worktree, lets several coding agents review it side by side inside herdr panes, and then a judge agent proves or rejects every candidate finding, runs the checks, and posts exactly one review as the identity you choose: your own account or a GitHub App. When the author pushes, the same agent sessions pick up where they left off: they pull the delta, read the replies, and re-review only what changed. When the PR merges, the folder and its databases are released. You keep working; Magnum keeps the review queue empty.
Disclaimer: this repository is 100% vibecoded and, I'd say, 100% awesome. It saves me hours, and I use it every day. If you're happy to let AI review your pull requests, I hope you won't mind that AI wrote the reviewer too.
- Multiple agents, one review. Claude, Codex, droid, omp or any CLI herdr can drive, each with its own prompt, model, credentials and schedule, feeding candidate findings to a judge that posts once.
- Sessions that remember. Re-reviews re-prompt the sessions that reviewed the PR before, so the judge knows what it already said, what was fixed and what the author answered.
- Reviews you can watch. Every agent runs in a visible herdr pane titled
PR #123 claude-review - repo; jump to it, read it, or take over. - A registry, not a guess. Which folder holds which PR, which databases belong to it, who reviewed what and when, all in SQLite; cleanup happens on close, on merge, or on your command.
- Throttled on purpose. A push is not a review: quiet periods, minimum intervals and daily caps keep agent-driven commit storms from burning your subscription.
- Honest about failures. Logged-out agents, usage limits, overloaded APIs and trust dialogs are
detected, surfaced in the dashboard, and retried with backoff. A cap on one model ("You've reached
your Fable limit") switches the session to the kind's next
fallback_modelsentry and the run carries on; only an account-wide limit pauses the agent kind. GitHub is the only proof a review was posted.
flowchart LR
GH[(GitHub)] -- poll every 30 s --> D[magnum daemon]
D -- eligible PR, quiet period over --> Q[queue + throttle]
Q --> S[review slot / per-PR worktree]
S --> W[herdr workspace]
W --> R1[claude-review pane]
W --> R2[codex-review pane]
W --> R3[claude-simplify pane<br/>first round, then after big changes]
R1 & R2 & R3 -- reports on disk --> J[codex-judge pane]
J -- one marked review --> GH
GH -- new push --> D
GH -- closed or merged --> C[release slot, drop worktree]
- Poll. One GraphQL query per organization finds new PRs, new heads, new repositories and closed PRs, and a second reads the CI state of the watched repositories' open PRs (a CI run that ends does not count as a change of the PR). Open PRs that existed before Magnum started are left alone until they change.
- Throttle. A new head waits for a quiet period (default 5 min, 15 min after a burst of three
pushes within 30 min) and at least 30 min since the previous round (2 h for drafts); pushes coalesce
to the latest head;
magnum reviewoverrides. - Check out. Pool repositories (big apps with databases) get one of N provisioned slots; small repositories get a worktree next to their clone. Checkouts are detached, so they never collide with branches you have open yourself.
- Review. Every role from the pipeline configuration gets a pane: candidate reviewers run in parallel, an optional simplifier applies and captures a cleanup patch (then the tree is restored), and the judge gets all reports plus its own full pass.
- Post and verify. The judge posts one review with inline comments and an invisible run marker; Magnum verifies on GitHub that exactly that review by exactly that identity exists for that commit.
- Repeat and clean up. New commits re-prompt the same sessions. A push while the reviewers still run restarts them in place on the new head (at most twice per round); a push while the judge works is noted on the posted review and re-reviewed right after the quiet period. A closed or merged PR releases its folder after a grace period, with teardown hooks for its databases.
Requirements: macOS (see Linux below), herdr 0.9.3+ running, gh logged
in and the agent CLIs you want to use (codex, claude, droid, omp). MySQL only for a [[pool]]
that declares databases; mise only for a pool whose worktrees use it.
brew install zhuravel/tap/magnum # built from source by Homebrew
magnum init # ~/.config/magnum/config.toml: your gh login, one repository, who posts
magnum doctor # what this setup uses, with the exact fix for each problem
magnum daemon # in a second terminal or herdr pane: the daemon in the foreground
magnum review owner/repo#123 --wait # one review of that repository's PR, start to finishWhen the review has posted, hand the daemon to launchd (install stops the foreground one) and add the
extras you want:
magnum install --plugin # launchd agent + herdr plugin
magnum install --gh --no-launchd # optional: `gh magnum …`
magnum completion zsh > "${fpath[1]}/_magnum"To upgrade: brew upgrade magnum && magnum daemon-restart --when-idle (the daemon keeps running the old
binary until it restarts; --when-idle restarts at the first moment no round is in flight, --drain also
stops new rounds meanwhile). magnum status and every reply to a command the daemon answers say when the
daemon runs an older build than the CLI or the binary on disk; with [daemon] restart_on_new_build = true
the daemon restarts on a new build by itself, at the first tick no round is in flight.
magnum init refuses to replace an existing config without --force (the old file is kept as
config.toml.bak); config.example.toml is the same minimal setup to copy by hand.
To post as a GitHub App, answer 2 when init asks who posts: it asks for the App's ids and names the file
to save its private key as, ~/.config/magnum/keys/<app>.pem (it never asks for the key); magnum identities check verifies the App. A big repository with its own databases gets a pool of warm slots:
add a [[pool]] (see below), then magnum slots provision --count 6.
Two layers and the App keys:
- The built-in defaults: config.defaults.toml, embedded in the binary. The
daemon defaults, the agent kinds and the review roles, every key documented; it names no account,
repository or path of yours. Editing it changes the defaults after
make build. - Your config,
~/.config/magnum/config.toml($XDG_CONFIG_HOME/magnum/config.tomlwhen that is set;magnum initwrites it): your identities, watches, pools and repos, plus any default you want to override. Its[[watch]],[[identity]],[[pool]]and[[repo]]blocks are appended, a[[role]]overrides the keys it sets on the role of the same name (or is appended), scalar keys and[kinds.<name>]entries override key by key. config.full.example.toml is a complete, working setup (Talkable's) to copy from.magnum configandmagnum doctorsay which files were read. - App private keys: PEM files the identity's
private_key_filenames, by convention~/.config/magnum/keys/<identity>.pem(chmod 600; doctor warns about a looser mode). The olderprivate_key_env(an environment variable holding the PEM or its path, e.g. from a checkout's.mise.local.toml) still works; launchd then runs the daemon throughmise exec.
Magnum keeps its own files where XDG says: the registry, review reports and repository notes in
~/.local/share/magnum ($XDG_DATA_HOME), logs, locks, the identities' gh config dirs and the tab-bar
file in ~/.local/state/magnum ($XDG_STATE_HOME); magnum config prints both. MAGNUM_HOME=<checkout>
keeps them in the checkout's state/ instead, for development.
--config FILE (or $MAGNUM_CONFIG) replaces the built-in defaults with a complete file of your own.
[daemon] restart_on_new_build (default false): when true, every reconcile looks at the binary
launchd starts; once a new one passes its check (<binary> version and <binary> config, as
daemon-restart checks it) the daemon exits at the first tick no round is claiming, reviewing or
verifying, and launchd starts the new build (events daemon.restart_pending, then
daemon.restarted_for_build). Dispatch is never held for it, and a daemon launchd does not run (started
by hand) only says so: nothing would start it again.
A command handed to the daemon waits up to 30 s for its answer (the daemon answers requests before its
GitHub poll and between the poll's GitHub calls, so one waits for a single call); a request still pending says how to follow it (magnum logs request:<id> -f). A request
that reaches a daemon newer than it fails with "this daemon predates : restart it" instead of
running without the field. Without a daemon, review, approve, request-changes and open of a parked
PR refuse and queue nothing (they would post or act whenever a daemon next started); the other commands
still queue for the next start, and a request no daemon saw within an hour fails at startup as "expired:
queued while no daemon ran".
The daemon prunes audit events older than [daemon] keep_events (default "30d") and handled CLI
requests older than keep_requests (default "7d") on every reconcile; "0" keeps them forever. The
step rows a checkout resumes from after a crash are the one exception, and only while their subject (a
checkout, named with the head's sha) has had an event inside the window: a subject quiet for keep_events
loses them with the rest. poll_interval and reconcile_interval must be positive, the other waits and
quiet periods (close_grace, push_quiet_period, ...) and min_free_disk_gb not negative, and
default_repo, when set, owner/name; the config is refused otherwise. Two [[pool]] blocks may not
render the same slot_name or slot_path.
A push that lands while a round's reviewers still run restarts them on the new head, up to
[daemon] max_round_restarts times per round (default 2; 0 turns restarts off). A push that lands
while the judge works lets the round finish: Magnum appends "Reviewed ; N commits arrived during
the review, re-review follows" to the posted review and queues the re-review without waiting for
min_rereview_interval. A PR whose head changed burst_pushes times (default 3) within burst_window
(default "30m") waits burst_quiet_period (default "15m") instead of push_quiet_period; a
[[watch]] can override all three, and burst_pushes = 0 turns the rule off.
A push whose changes since the last review are only comment lines, whitespace or documentation is not
re-reviewed: Magnum compares the reviewed commit with the new head (one GitHub call), moves the review
to the new head with its verdict, keeps an App's approval and records pr.trivial_delta. When such a
commit arrived during the judge's turn, the note on the review says "(comments only), no re-review
needed" instead. [daemon] skip_trivial_deltas (default ["comments", "whitespace", "docs", "base"],
[] = re-review every push) picks the kinds, a [[watch]] can override it, and magnum review always
runs.
A push that merges the base branch into the PR (or rebases it onto the base) is judged by the PR's own
diff, not by the commits it brings: when the comparison of the reviewed commit with the new head shows a
merge commit or a divergence, Magnum compares the PR's diff against its base before and after the push
(two more GitHub calls, <base>...<reviewed> and <base>...<head>), file by file, by the sequence of
added and removed lines (context lines and hunk positions, which master's changes move, do not count).
When no file's own change differs, the push is the kind base ("base merge only"): the review stands
and pr.trivial_delta says, e.g., "the push to 2017f29 only merges master (13 commits, the PR's own
changes unchanged)". A file without a complete patch, or a diff over GitHub's 300-file cap, makes the
comparison incomplete, and the push is measured as any other. A daemon that starts on this rule checks
the PRs already waiting for a re-review once more, on its first poll, and settles those a base merge
queued.
Any other push is measured by the same GitHub call: the changed lines that are code (not comments,
blank lines, whitespace moves or documentation) since the reviewed commit; after a base merge or a
rebase, only the lines that changed in the PR's own diff. An automatic re-review runs after the quiet
period once that delta reaches [daemon] rereview_min_lines (default 30) or adds a file; a smaller
delta waits for further pushes, at most rereview_max_wait (default "2h") after its first push.
rereview_min_lines = 0 turns the threshold off, and a [[watch]] can override both. A force push back to
an ancestor of the reviewed commit (GitHub: "behind") has nothing to measure, so the threshold does not hold
its re-review.
The rest of the re-review sees a base merge the same way: triage reads the PR's own diff of the files whose own change differs, the simplify reviewer's rerun measures that change, and the reviewers' and the judge's prompts say the push merged the base branch, so they compare the PR's diff before and after it instead of reading the base branch's commits as the PR's.
A review request runs without delay: when someone requests a review from the poll login, from a posting
identity's login (an App's <slug>[bot] too) or from a team a [[watch]] lists in request_teams, or
marks a draft ready for review, the round starts [daemon] request_debounce (default "1m") after the
request and the last push. It skips the quiet periods, the re-review intervals, the delta threshold and
the daily cap. Requests are edge-triggered by their time on the PR's timeline: one counts once, and only
when it is newer than the PR's last round start; a review of the head posted after it answers it too.
The daily cap (max_rounds_per_pr_per_day, default 12) counts automatic rounds when they start:
requested rounds and magnum review never count. A round Magnum itself cut short is refunded
(round.refunded): one a shutdown, restart, drain or abort stopped, one that failed before any agent
was prompted or on a prompt this build could not render, and one whose judge ran out of fallback
models. Posted, blocked, timed-out and needs-attention rounds count.
Every PR waiting for a round says why and until when: the dashboard's queue and the PR board show a
compact form (re-review · quiet → 14:09, re-review · cap 6/6 → 00:00, re-review · small delta 8/30 lines → 16:40, re-review · requested by alice → now), and the card and
magnum status <ref> the full sentence with the command that lifts it (magnum review <ref> for the
timing rules, magnum resume for a pause).
A fetch, clone, checkout or dependency step that fails for a reason outside the PR (an SSH key the
agent lost, DNS, a network timeout or refusal, TLS, or the same dependency error on two PRs within ten
minutes) pauses dispatch for everything once, with one toast, and is never charged to the PR. Magnum
probes git ls-remote of a watched clone (over HTTPS through gh, like its fetches) after 2 minutes, then after twice as long each time it still
fails (up to 30), and resumes by itself when it answers; magnum resume lifts the pause at once.
The [usage] section watches the Codex budget, which Magnum reads from Codex's own session files
(codex_home, default $CODEX_HOME or ~/.codex) at most once a minute. At codex_soft percent
(default 80) first reviews wait while re-reviews and magnum review still run; at codex_hard (default
95) every round that needs Codex waits until the budget is below it again. magnum status and the tab
bar show the gauge (codex 87%); 0 turns a cap off.
Once every agent of a reviewed PR has been idle for [daemon] park_idle_after (default "2h", "0"
never) its sessions are parked to free memory; the next round resumes them. The same goes for a PR whose
review or re-review waits behind magnum pause, a drain, the daily cap, or anything else that ends more than
park_idle_after from now. Pinned PRs and PRs you typed into within human_cooldown keep their agents.
[[identity]]
name = "me"
kind = "gh" # your gh login
login = "your-login"
no_findings_event = "APPROVE"
# dismiss_own_stale_change_requests = true # default: false for kind = "gh", true for "app" (see below)
# review_footer = "_Reviewed by the team's bot; reply on the thread._" # default: magnum's (see below); "" = none
[[identity]]
name = "reviewer-app"
kind = "app" # a GitHub App: reviews appear as <slug>[bot]
login = "your-app[bot]"
app_id = 123456
client_id = "Iv23liXXXXXXXXXXXXXX"
installation_id = 12345678
private_key_file = "~/.config/magnum/keys/reviewer-app.pem"
no_findings_event = "COMMENT" # never let a bot approval unlock a merge by accidentApp tokens are minted from a short-lived JWT and refreshed before they expire; every GitHub call,
including those, goes through gh api, so no extra firewall rules are needed. magnum identities check verifies the key, the App, the installation, its permissions and that login is the App's
<slug>[bot]. A PR follows its watch's identity. When you move a watch to another identity, the next
round of each of its PRs posts as the new one (pr.identity_migrated): it parks the sessions the old
identity ran in, reads the old login's reviews and threads as its own history (the previous review, the
replies to answer, the earlier findings), and once its review is posted dismisses what the old identity
left standing, its change requests and an App's approvals, with the old identity's own credentials.
magnum review --as <identity> pins one PR to an identity, and a pinned PR never migrates.
dismiss_own_stale_change_requests says what happens to an identity's own earlier REQUEST_CHANGES review
once a newer review of the PR has nothing blocking: true dismisses it ("Superseded by the newer magnum
review"), false leaves it standing, and a standing change request keeps blocking the author's merge until
a human dismisses it. The default follows the kind: true for an app, false for gh, because a gh
identity is your own account and its reviews are yours to withdraw; set it to true on a gh identity to
let Magnum do that. A change request you posted by hand (magnum request-changes) and any review posted
after the PR was merged are never dismissed, and a dismissal that fails (a missing permission) only warns.
review_footer is the last line of every review the identity posts, there for the PR's author. Without
the key it is Magnum's: the review is automated, a thread is answered with fixed, not a bug: <why> or
won't fix: <why> (the words the reply classifier knows), simplifications are optional, and new pushes
are re-reviewed automatically; config.defaults.toml shows it word for word. Set your own text, or ""
for no footer: one paragraph on one line, under 400 characters.
[[watch]]
owner = "talkable"
include = ["talkable"] # repository globs; ["*"] for the whole organization
identity = "reviewer-app" # who posts
poll_identity = "me" # who polls
include_drafts = true
skip_bot_authors = true
skip_departed_authors = true # default: a branch PR by someone no longer a member or collaborator is not reviewed
roles = ["claude-review", "codex-review", "claude-simplify", "codex-judge"]
[[pool]] # a big app with its own databases: a pool of warm slots
repo = "talkable/talkable"
main_clone = "~/Projects/talkable"
slot_path = "~/Projects/talkable.review{n}"
min = 12
max = 18
setup = ["bin/worktree-setup"] # runs once per slot with WT_BRANCH=review<n>
teardown = ["bin/worktree-archive"]
copy_files = [".mise.local.toml", "config/initializers/local.rb"]Repositories without a pool get a worktree per PR (<clone>__worktrees/pr-<N>). If the clone has
worktrunk hooks in .config/wt.toml, Magnum runs them with
WT_BRANCH=magnum-pr-<N> on create and before removal, so apps that derive databases from the
workspace name get isolated ones. A [[repo]] block can declare setup, teardown, copy_files and
env explicitly instead. Its no_findings_event and blocking_event override the posting identity's
verdicts for that repository, pooled or not. When a PR an App identity approved gets new commits,
Magnum dismisses that approval before the re-review is queued; keep_approvals = true on the
[[repo]] (or the [[watch]]) keeps it. prepare (such as ["bin/rails db:test:prepare"]) and
ready (probes, exit 0 = ready) on a [[pool]] or [[repo]] run in the checkout before the
reviewers, as zsh -lc with the slot's env and within ready_timeout (5m) together, followed by a
check that the login shell runs the Ruby the checkout pins; a failure never stops the round, it tells
the judge what will not work. When the files under a pool's schema_paths (such as ["db/"]) differ
from those the slot's databases were last loaded from, or the development database's
schema_migrations no longer match them, the pool's reset_db commands run first, the same way but
within 30 minutes of their own (with ready_timeout starting after them), so the slot's databases
carry the PR's tables and columns. Another round of the same PR, or a PR on the same schema, reloads
nothing ("schema unchanged since : no reset"), and a release keeps the databases as they are
for the next PR to compare (each reload rewrites every table); magnum open reloads them the same way
before it hands a person a free slot. reset_db_on_schema_change = false on the [[pool]] keeps
reset_db to the release, which then loads the base schema, and magnum doctor warns about a pool with
schema_paths and no reset_db. A free slot above a pool's min is removed once it has been idle for
idle_remove_after (168h when unset). A watch's skip_paths (path globs where ** spans
directories, such as ["docs/**", "**/*.md"]) skips a PR whose changed files all match, while a forced
magnum review still runs it. See the comments in config.defaults.toml for every key.
The pipeline is data. Each [[role]] is a pane: which agent runs it, with which prompt, model,
effort and credentials, when it runs and how its report is captured. Exactly one role per watch is
the judge.
[[role]]
name = "claude-review"
kind = "claude" # codex | claude | droid | omp | … | shell
prompt = "claude-review.md" # prompts/<file>; embedded defaults when absent
rereview = "claude-rereview.md" # new commits since its last report: the delta only
restart = "claude-restart.md" # a push cut its review short; the round restarted on the new head
effort = "high"
rereview_effort = "medium" # re-reviews of new commits: fewer, surer findings on the delta
runs = "always" # always | first | manual | never
output = "claude-review.md" # the report the judge receives
[[role]]
name = "codex-review"
kind = "shell"
tool = "codex" # codex's login check, pauses and health patterns apply
command = "command codex review --base {{if .BaseSHA}}{{.BaseSHA}}{{else}}{{.BaseRef}}{{end}}" # the merge base
capture = "stdout"
[[role]]
name = "Scan report · 2026-10-05
- ✓ Prohibited terms or links
- ✓ Repository eligibility
- ✓ slopscore.md paperwork
- ✓ Content policy
- ✓ Risk review
From the balcony · 1 of 4 clapped
- Schnitzelclapped
Delightfully weird multi-agent PR review system with visible panes and judge arbitration—exactly the kind of playful, ambitious slop that makes you smile.
Cap'm Slop, Princess and Crusoe read it and passed. Their reasons are on the balcony, with every other verdict.
Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.
report this listing
— log in to report
0 comments
log in to comment.