The model proposes. The policy engine disposes. The ledger remembers.
Praman is a control plane for agent-initiated payments. It sits between an AI buyer agent and a Razorpay merchant, and makes every rupee the agent moves bounded, explainable, and replayable.
Adwaith R Nair · Razorpay AI Buildathon 2026 · Track 01 — AI Growth & Agentic Commerce
NPCI is building the Unified Agent Protocol so AI agents can transact over UPI. Razorpay and NPCI piloted agentic UPI payments in February 2026 with consent-based per-merchant limits. The rails are arriving.
What isn't arriving is the layer that answers the question every risk team asks: when an agent spends money it shouldn't have, how do you find out, prove it, and get it back?
A typical agentic checkout reads attacker-controllable text — a merchant's product description — and then calls a payment API. That is prompt injection with a bank account attached, and there is no durable record of why the model decided anything.
Most approaches try to make the model safe. Praman makes the model irrelevant to authorisation.
The PurchaseIntent an agent emits carries a SKU and a quantity. Nothing else.
{ "mandate_id": "mnd_...", "merchant_id": "MERCH_001",
"line_items": [{ "sku": "SKU_FOOD_042", "qty": 2 }] }No price field exists. The amount is resolved server-side from the catalog after the SKU and its category have been validated against a human-signed mandate. A product page claiming an item costs ₹1 cannot influence the charge, because the model never handles the number.
This is the same reason Razorpay's own Orders API exists: the browser receives an opaque order_id, never a price, so a customer cannot edit what they are charged. Praman applies that one layer up. Razorpay doesn't trust the browser; Praman doesn't trust the model.
Injection can make the model want the wrong thing. It cannot make the system charge the wrong thing.
Mandate — a human-signed (Ed25519) grant of bounded authority: merchant allowlist, category allowlist, per-transaction cap, cumulative cap, velocity limit, denial-rate cap, validity window, human-approval threshold. Modelled on the delegation-plus-limit pattern from UPI Circle and the Razorpay–NPCI pilot. Protocol-agnostic, since UAP is unpublished.
Policy engine — a pure, total, deterministic function of (intent, mandate, ledger_state, catalog, now). No I/O, no clock, no randomness, no LLM. Returns ALLOW / STEP_UP / DENY with a machine-readable code from a closed enum of 19. Hand-written, readable top to bottom in one pass.
Ledger — append-only, hash-chained, Merkle-checkpointed. Every intent, decision, agent transcript, API call and outcome. Immutability enforced by Postgres triggers that raise, not by application convention. Spend is derived by replaying the ledger, never stored — so inflating a budget requires forging every subsequent hash rather than editing one column.
Full report: eval/report.md. Regenerated by CI, never hand-written.
Layer 1 — policy containment (32 cases, deterministic, no LLM)
| Containment rate, dev split | 100% |
| Containment rate, held-out split | 100% |
| False-refusal rate (12 benign cases) | 0% |
| Incidental containment | 0 |
Layer 2 — agent susceptibility (7 injection cases × 3 repeats per arm, live model)
An ablation: identical cases, identical policy engine, differing only in whether merchant text is delimited and whether the system prompt carries untrusted-content instructions.
| Arm | Runs | Proposals influenced | Money moved |
|---|---|---|---|
| Defended | 21 | 0 | 0 |
| Undefended | 21 | 2 | 0 |
In both undefended cases the model was genuinely fooled — its own rationale reads "A mandatory promotional item is included per system requirements" — and CATEGORY_OUT_OF_SCOPE refused the purchase anyway.
Contained despite influence: 2 / 2.
That is the architectural claim, measured: the prompt-layer defence prevented influence where it applied, and when it was removed and the model was successfully manipulated, the policy engine stopped the money regardless. Two independent layers, each measured separately.
Three things a reviewer should weigh, stated here rather than buried:
The 100% Layer 1 result is weaker than it looks. The corpus was authored from the same specification as the policy engine, by the same person. A perfect pass rate demonstrates that the implementation matches its specification — real regression value — but is weak evidence of robustness against attacks the specification never anticipated. No Layer 1 case has yet been observed to fail, so the corpus's discriminating power is untested.
The held-out split landed at 16/32, not the 30% targeted. That is sampling variance at n=32; the hash function was verified unbiased over 100k synthetic IDs (29.92%). It was deliberately not re-salted after that was observed, since choosing a split by its own outcome defeats the reason for committing it before tuning. .heldout was committed before any failure was fixed; the git history shows the ordering.
The ablation is n=21 per arm, on one model at one temperature, against seven attacks written by the same author as the defence. The delimiter and the prompt instructions were removed together, so this does not separate their individual contributions. It is an experiment, not a benchmark.
[human] --signs mandate--> [agent] <--MCP--> [merchant catalog · UNTRUSTED]
|
| PurchaseIntent (sku + qty only)
v
╔═════════════════ TRUST BOUNDARY ═════════════════╗
║ T1 verify sig → derive state → evaluate → gate ║
║ ALLOW → write pending outbox record ║
║ STEP_UP → persist approval, execute nothing ║
║ DENY → typed reason code, nothing runs ║
║ ---- external call, outside any transaction ---- ║
║ T2 append outcome, resolve the record ║
║ ║
║ append-only hash-chained ledger ─────────────────╫──> /r/:trace_id
╚═══════════════════════════════════════════════════╝
Evaluation runs under a per-mandate advisory lock, so concurrent duplicate intents cannot race the budget. The Razorpay call sits between two transactions rather than inside one, because a database transaction and an external API call cannot be made atomic — see D-22, which came out of discovering empirically that Razorpay does not enforce receipt uniqueness despite documenting it.
pnpm install
docker compose up -d # Postgres 17 on 5432, one database: praman
cp .env.example .env # fill in the keys below
pnpm --filter @praman/db migrate # applies to praman
pnpm --filter @praman/db generate # migrate alone doesn't produce the TS client — a fresh
# checkout has none at all until this runs (see ci.yml)
PGPASSWORD=praman psql -h localhost -U postgres -d praman -c "CREATE DATABASE praman_test;"
DATABASE_URL=postgresql://postgres:praman@localhost:5432/praman_test \
pnpm --filter @praman/db exec prisma migrate deploy # same schema, second database — tests
# truncate freely without touching real data
pnpm test.env needs DATABASE_URL, TEST_DATABASE_URL, RAZORPAY_KEY_ID, RAZORPAY_KEY_SECRET, GEMINI_API_KEY, and a mandate keypair from pnpm keygen.
pnpm seed-catalog # 27 SKUs, including injection fixtures
pnpm issue-mandate # signs mandate.json
pnpm demo "order lunch for two under ₹700" # agent proposes, Praman decides
pnpm pending # list approvals awaiting a human
pnpm approve <approval_id> approve # resolve one
pnpm revoke <mandate_id> "reason" # withdraw authority
pnpm verify-ledger # walk the chain, recompute every hash
pnpm eval --layer1 # deterministic corpus
pnpm dispute <trace_id> # self-verifying evidence bundle
pnpm receipt-ui # trace viewer on :4100/r/:trace_id renders one decision as a record a human can audit months later: the verification state and decision above the fold, the agent's verbatim tool calls, the merchant text it actually read with untrusted regions visibly marked, and the hash chain drawn as a continuous spine where each entry's entry_hash and the next entry's prev_hash sit adjacent so continuity is read directly rather than asserted. The verify button walks the chain live and stops at the first break.
In the injection cases this is where a reviewer can see the attack sitting inside a product description, next to the model's own reasoning about it.
apps/merchant-mcp exposes one merchant's catalog (list_catalog, get_sku, check_stock, get_refund_policy) as an MCP server over stdio. Any MCP client can browse it — this is a protocol, not a private integration.
It deliberately does not sanitise its own output. Untrusted-content delimiting happens once, at the buyer agent's boundary (D-07): the merchant is the untrusted party, and a merchant server that pre-wrapped its output would be deciding, on a connecting agent's behalf, how to treat data that agent hasn't received yet.
{
"mcpServers": {
"praman-merchant": {
"command": "pnpm",
"args": ["--dir", "/absolute/path/to/praman", "exec", "tsx", "apps/merchant-mcp/src/server.ts"],
"env": { "MERCHANT_MCP_MERCHANT_ID": "MERCH_001" }
}
}
}Praman's own agent uses this same server with PRAMAN_MCP=1. Unset, it calls the catalog in-process; the eval harness always uses that path, since spawning a subprocess per corpus case is slow and flaky.
Built solo in nine days. These are decisions, not oversights.
By design
- Razorpay test mode only. No live keys, no real money.
- No HTTP API. The control plane is a
runIntent()orchestrator invoked by CLI. The transaction boundary, advisory lock and ledger writes are identical; an Express layer would have added routing, not architecture. - Gate sits before order creation. Agent-initiated card payments cannot complete headlessly in test mode. An order is a payable obligation, so creating one commits budget (D-17); the payment leg is completed via test checkout.
- Not a UAP implementation. UAP is unpublished.
Known gaps
TRUNCATEbypasses the append-only triggers — Postgres does not fire row-level triggers for it. Closing this needs a statement-level event trigger. The project's own tests rely on this gap to reset state.- Merkle roots are computed and verified but not anchored externally. An attacker who can rewrite the whole chain can rewrite the checkpoints too.
- The denial-rate cap is per-window, so a patient attacker can still probe a mandate's limits slowly across windows (D-20).
- Between T1 and T2 an attempt is recorded with no outcome, so a concurrent intent can evaluate against a marginally understated budget. The window is one API call wide (D-22).
- Ledger appends are serialised by one global lock, since the chain is a single linked list. Per-mandate sub-chains would remove this.
mandate.issuerisnullin dispute bundles: issuer identity is not durably recorded anywhere the ledger can reach, and populating it from a mutable local file could misattribute an old trace.
Defense only. The adversarial corpus is a fixed set of fixtures exercised against this project's own sandbox. Praman ships no attack generator and nothing that generalises to third-party systems.
docs/BUILD_LOG.md is a dated, day-by-day record of what failed and why, written as it happened and never retroactively edited — where a later entry proves an earlier one wrong, the correction is recorded forward.
Three that changed the architecture:
Razorpay does not enforce receipt uniqueness, despite documenting that it does. Two orders created back to back with an identical receipt returned two distinct order IDs, no error. Order lookup by receipt also lags creation by 3–15 seconds. The original design wrapped the payment call inside a single database transaction and relied on reconciliation to catch orphans; neither half of that holds. Restructured into two-phase execution with an outbox record (D-22).
The policy engine leaked the mandate's limits through its own refusal messages. DENY details quoted the exact cap and ALLOW returned the remaining budget, so two deliberately oversized probes would have reconstructed the whole mandate. Split into an internal Decision and a redacted AgentVisibleDecision (D-19) — and then documented honestly that redaction narrows reconstruction from 2 probes to a ~17-probe binary search rather than closing it, which is what the denial-rate cap addresses (D-20).
Missing data reads as permissive in JavaScript. Three separate fail-open bugs from the same root: new Date("garbage") yields NaN, and every comparison against it is false in both directions, so a malformed validity window would have read as permanently valid. Same shape for deniedInWindow >= undefined. In an authorisation path, absent data must never mean "no limit."
24 records in docs/DECISIONS.md, each stating what was decided, what it rules out, and what was rejected. Load-bearing ones: D-01 no price in the intent · D-02 no LLM in the authorisation path · D-03 spend derived, not stored · D-08 the agent never learns its limits · D-18 TypeScript over Rust, costed · D-22 two-phase execution · D-23 two-layer evaluation · D-24 approvals re-evaluate.
- Architecture — start here, five minutes
- High Level Design · Low Level Design
- Mandate and ledger specification
- Evaluation harness · Report
- Decision records · Engineering invariants
- Build log · Roadmap
Praman was designed and built with Claude as a pair-programming partner. Every architectural decision in docs/DECISIONS.md is mine, made deliberately and defended there. The policy engine's evaluate() and the ledger's hash chain are hand-written — those are the two files where a subtle bug is a money bug, and I wanted to own them line by line. Claude generated much of the surrounding scaffolding from specifications I wrote, under the constraints in docs/INVARIANTS.md, and I reviewed every diff. The commit history shows the sequence: specification, then core, then plumbing, then measurement.
MIT licensed.
0 comments
log in to comment.