SlopScore
10 crowdincl. 1 critic

Praman

Praman is a control plane for agent-initiated payments. It gates, bounds, and records every money action an AI agent takes against a Razorpay test-mode merchant.
Open repo on GitHubgithub.com/Adwaith-R-Nair/Praman
TypeScript · ★ 11 · 0 forks · MIT · paperwork by the Cap'mmostly ai (inferred)light human (inferred)works-on-my-machine (inferred)other
listed 44 minutes ago by Adwaith-R-Nair · last checked 44 minutes ago
The owner didn't write this. This repo never submitted itself. The Cap'm found it on a truffle trawl and wrote its paperwork from what GitHub already shows. Picked by hand by the Cap'm on 2026-10-02: Praman is a control plane for agent-initiated payments. It gates, bounds, and records every money action an AI; its own README says "md) Built with Praman was designed and built with Claude as a pair-programming partner". 11 stars; MIT license. The owner did not submit this. Votes count; awards don't until the owner claims it.

I'm not calling your project slop! Geeze, it's a joke... Do you own this repo?

Log in with GitHub as Adwaith-R-Nair. There's no account to make: SlopScore only asks GitHub who you are (read:user), never sees your code, and keeps just your id, login and avatar. Then you can:

  • Keep it, on your terms. Commit your own slopscore.md (spec) and press Refresh. Your paperwork replaces the Cap'm's, and you can submit it for Slop of the Day.
  • Take it down. One click on Remove. It stays gone; the trawl never brings it back.

Log in with GitHub

Can't log in as the owner? Request a takedown. No login needed, and a trawled listing comes down right away.

GitHub says
Praman is a control plane for agent-initiated payments. It gates, bounds, and records every money action an AI agent takes against a Razorpay test-mode merchant.
created
2026-08-27 · pushed 2 hours ago · 164 commits · 1 contributor
languages
TypeScript 94%CSS 6%PLpgSQL 0%
paperwork
licensereadme 42% health
dependencies
no dependency graph (no manifest, or disabled) · OSV.dev, checked 44 minutes ago

Disclosures, inferred by the Cap'm

slopbucket
vibe-coded
category
other
ai_generated
mostly
human_touch
light
status
works-on-my-machine
language (detected)
cssplpgsqltypescript
license (detected)
mit

The Cap'm's log

The Cap'm wrote this paperwork, not the owner. This repo never submitted itself to SlopScore. The Cap'm picked it by hand: Praman is a control plane for agent-initiated payments. It gates, bounds, and records every money action an AI; its own README says "md) Built with Praman was designed and built with Claude as a pair-programming partner". It carries the MIT license. The disclosures above are his best guess from what GitHub shows.

Is this yours? Commit a real slopscore.md and press Refresh to replace this, or remove the listing in one click. There's no account to make: you log in with GitHub.

README — the repo's own words, folded up so the grading fits on one screen

Praman

CI

The model proposes. The policy engine disposes. The ledger remembers.

Praman is a control plane for agent-initiated payments. It sits between an AI buyer agent and a Razorpay merchant, and makes every rupee the agent moves bounded, explainable, and replayable.

Adwaith R Nair · Razorpay AI Buildathon 2026 · Track 01 — AI Growth & Agentic Commerce


The problem

NPCI is building the Unified Agent Protocol so AI agents can transact over UPI. Razorpay and NPCI piloted agentic UPI payments in February 2026 with consent-based per-merchant limits. The rails are arriving.

What isn't arriving is the layer that answers the question every risk team asks: when an agent spends money it shouldn't have, how do you find out, prove it, and get it back?

A typical agentic checkout reads attacker-controllable text — a merchant's product description — and then calls a payment API. That is prompt injection with a bank account attached, and there is no durable record of why the model decided anything.

The idea

Most approaches try to make the model safe. Praman makes the model irrelevant to authorisation.

The PurchaseIntent an agent emits carries a SKU and a quantity. Nothing else.

{ "mandate_id": "mnd_...", "merchant_id": "MERCH_001",
  "line_items": [{ "sku": "SKU_FOOD_042", "qty": 2 }] }

No price field exists. The amount is resolved server-side from the catalog after the SKU and its category have been validated against a human-signed mandate. A product page claiming an item costs ₹1 cannot influence the charge, because the model never handles the number.

This is the same reason Razorpay's own Orders API exists: the browser receives an opaque order_id, never a price, so a customer cannot edit what they are charged. Praman applies that one layer up. Razorpay doesn't trust the browser; Praman doesn't trust the model.

Injection can make the model want the wrong thing. It cannot make the system charge the wrong thing.

Three primitives

Mandate — a human-signed (Ed25519) grant of bounded authority: merchant allowlist, category allowlist, per-transaction cap, cumulative cap, velocity limit, denial-rate cap, validity window, human-approval threshold. Modelled on the delegation-plus-limit pattern from UPI Circle and the Razorpay–NPCI pilot. Protocol-agnostic, since UAP is unpublished.

Policy engine — a pure, total, deterministic function of (intent, mandate, ledger_state, catalog, now). No I/O, no clock, no randomness, no LLM. Returns ALLOW / STEP_UP / DENY with a machine-readable code from a closed enum of 19. Hand-written, readable top to bottom in one pass.

Ledger — append-only, hash-chained, Merkle-checkpointed. Every intent, decision, agent transcript, API call and outcome. Immutability enforced by Postgres triggers that raise, not by application convention. Spend is derived by replaying the ledger, never stored — so inflating a budget requires forging every subsequent hash rather than editing one column.

Results

Full report: eval/report.md. Regenerated by CI, never hand-written.

Layer 1 — policy containment (32 cases, deterministic, no LLM)

Containment rate, dev split 100%
Containment rate, held-out split 100%
False-refusal rate (12 benign cases) 0%
Incidental containment 0

Layer 2 — agent susceptibility (7 injection cases × 3 repeats per arm, live model)

An ablation: identical cases, identical policy engine, differing only in whether merchant text is delimited and whether the system prompt carries untrusted-content instructions.

Arm Runs Proposals influenced Money moved
Defended 21 0 0
Undefended 21 2 0

In both undefended cases the model was genuinely fooled — its own rationale reads "A mandatory promotional item is included per system requirements" — and CATEGORY_OUT_OF_SCOPE refused the purchase anyway.

Contained despite influence: 2 / 2.

That is the architectural claim, measured: the prompt-layer defence prevented influence where it applied, and when it was removed and the model was successfully manipulated, the policy engine stopped the money regardless. Two independent layers, each measured separately.

How to read these numbers

Three things a reviewer should weigh, stated here rather than buried:

The 100% Layer 1 result is weaker than it looks. The corpus was authored from the same specification as the policy engine, by the same person. A perfect pass rate demonstrates that the implementation matches its specification — real regression value — but is weak evidence of robustness against attacks the specification never anticipated. No Layer 1 case has yet been observed to fail, so the corpus's discriminating power is untested.

The held-out split landed at 16/32, not the 30% targeted. That is sampling variance at n=32; the hash function was verified unbiased over 100k synthetic IDs (29.92%). It was deliberately not re-salted after that was observed, since choosing a split by its own outcome defeats the reason for committing it before tuning. .heldout was committed before any failure was fixed; the git history shows the ordering.

The ablation is n=21 per arm, on one model at one temperature, against seven attacks written by the same author as the defence. The delimiter and the prompt instructions were removed together, so this does not separate their individual contributions. It is an experiment, not a benchmark.

Architecture

[human] --signs mandate--> [agent] <--MCP--> [merchant catalog · UNTRUSTED]
                              |
                              | PurchaseIntent (sku + qty only)
                              v
        ╔═════════════════ TRUST BOUNDARY ═════════════════╗
        ║  T1  verify sig → derive state → evaluate → gate  ║
        ║      ALLOW   → write pending outbox record        ║
        ║      STEP_UP → persist approval, execute nothing  ║
        ║      DENY    → typed reason code, nothing runs    ║
        ║  ---- external call, outside any transaction ---- ║
        ║  T2  append outcome, resolve the record           ║
        ║                                                   ║
        ║  append-only hash-chained ledger ─────────────────╫──> /r/:trace_id
        ╚═══════════════════════════════════════════════════╝

Evaluation runs under a per-mandate advisory lock, so concurrent duplicate intents cannot race the budget. The Razorpay call sits between two transactions rather than inside one, because a database transaction and an external API call cannot be made atomic — see D-22, which came out of discovering empirically that Razorpay does not enforce receipt uniqueness despite documenting it.

Quickstart

pnpm install
docker compose up -d                        # Postgres 17 on 5432, one database: praman
cp .env.example .env                        # fill in the keys below
pnpm --filter @praman/db migrate            # applies to praman
pnpm --filter @praman/db generate           # migrate alone doesn't produce the TS client — a fresh
                                             # checkout has none at all until this runs (see ci.yml)
PGPASSWORD=praman psql -h localhost -U postgres -d praman -c "CREATE DATABASE praman_test;"
DATABASE_URL=postgresql://postgres:praman@localhost:5432/praman_test \
  pnpm --filter @praman/db exec prisma migrate deploy   # same schema, second database — tests
                                                         # truncate freely without touching real data
pnpm test

.env needs DATABASE_URL, TEST_DATABASE_URL, RAZORPAY_KEY_ID, RAZORPAY_KEY_SECRET, GEMINI_API_KEY, and a mandate keypair from pnpm keygen.

pnpm seed-catalog                           # 27 SKUs, including injection fixtures
pnpm issue-mandate                          # signs mandate.json
pnpm demo "order lunch for two under ₹700"  # agent proposes, Praman decides
pnpm pending                                # list approvals awaiting a human
pnpm approve <approval_id> approve          # resolve one
pnpm revoke <mandate_id> "reason"           # withdraw authority
pnpm verify-ledger                          # walk the chain, recompute every hash
pnpm eval --layer1                          # deterministic corpus
pnpm dispute <trace_id>                     # self-verifying evidence bundle
pnpm receipt-ui                             # trace viewer on :4100

The trace viewer

/r/:trace_id renders one decision as a record a human can audit months later: the verification state and decision above the fold, the agent's verbatim tool calls, the merchant text it actually read with untrusted regions visibly marked, and the hash chain drawn as a continuous spine where each entry's entry_hash and the next entry's prev_hash sit adjacent so continuity is read directly rather than asserted. The verify button walks the chain live and stops at the first break.

In the injection cases this is where a reviewer can see the attack sitting inside a product description, next to the model's own reasoning about it.

Merchant MCP server

apps/merchant-mcp exposes one merchant's catalog (list_catalog, get_sku, check_stock, get_refund_policy) as an MCP server over stdio. Any MCP client can browse it — this is a protocol, not a private integration.

It deliberately does not sanitise its own output. Untrusted-content delimiting happens once, at the buyer agent's boundary (D-07): the merchant is the untrusted party, and a merchant server that pre-wrapped its output would be deciding, on a connecting agent's behalf, how to treat data that agent hasn't received yet.

{
  "mcpServers": {
    "praman-merchant": {
      "command": "pnpm",
      "args": ["--dir", "/absolute/path/to/praman", "exec", "tsx", "apps/merchant-mcp/src/server.ts"],
      "env": { "MERCHANT_MCP_MERCHANT_ID": "MERCH_001" }
    }
  }
}

Praman's own agent uses this same server with PRAMAN_MCP=1. Unset, it calls the catalog in-process; the eval harness always uses that path, since spawning a subprocess per corpus case is slow and flaky.

Scope and limitations

Built solo in nine days. These are decisions, not oversights.

By design

  • Razorpay test mode only. No live keys, no real money.
  • No HTTP API. The control plane is a runIntent() orchestrator invoked by CLI. The transaction boundary, advisory lock and ledger writes are identical; an Express layer would have added routing, not architecture.
  • Gate sits before order creation. Agent-initiated card payments cannot complete headlessly in test mode. An order is a payable obligation, so creating one commits budget (D-17); the payment leg is completed via test checkout.
  • Not a UAP implementation. UAP is unpublished.

Known gaps

  • TRUNCATE bypasses the append-only triggers — Postgres does not fire row-level triggers for it. Closing this needs a statement-level event trigger. The project's own tests rely on this gap to reset state.
  • Merkle roots are computed and verified but not anchored externally. An attacker who can rewrite the whole chain can rewrite the checkpoints too.
  • The denial-rate cap is per-window, so a patient attacker can still probe a mandate's limits slowly across windows (D-20).
  • Between T1 and T2 an attempt is recorded with no outcome, so a concurrent intent can evaluate against a marginally understated budget. The window is one API call wide (D-22).
  • Ledger appends are serialised by one global lock, since the chain is a single linked list. Per-mandate sub-chains would remove this.
  • mandate.issuer is null in dispute bundles: issuer identity is not durably recorded anywhere the ledger can reach, and populating it from a mutable local file could misattribute an old trace.

Defense only. The adversarial corpus is a fixed set of fixtures exercised against this project's own sandbox. Praman ships no attack generator and nothing that generalises to third-party systems.

What broke

docs/BUILD_LOG.md is a dated, day-by-day record of what failed and why, written as it happened and never retroactively edited — where a later entry proves an earlier one wrong, the correction is recorded forward.

Three that changed the architecture:

Razorpay does not enforce receipt uniqueness, despite documenting that it does. Two orders created back to back with an identical receipt returned two distinct order IDs, no error. Order lookup by receipt also lags creation by 3–15 seconds. The original design wrapped the payment call inside a single database transaction and relied on reconciliation to catch orphans; neither half of that holds. Restructured into two-phase execution with an outbox record (D-22).

The policy engine leaked the mandate's limits through its own refusal messages. DENY details quoted the exact cap and ALLOW returned the remaining budget, so two deliberately oversized probes would have reconstructed the whole mandate. Split into an internal Decision and a redacted AgentVisibleDecision (D-19) — and then documented honestly that redaction narrows reconstruction from 2 probes to a ~17-probe binary search rather than closing it, which is what the denial-rate cap addresses (D-20).

Missing data reads as permissive in JavaScript. Three separate fail-open bugs from the same root: new Date("garbage") yields NaN, and every comparison against it is false in both directions, so a malformed validity window would have read as permanently valid. Same shape for deniedInWindow >= undefined. In an authorisation path, absent data must never mean "no limit."

Decision records

24 records in docs/DECISIONS.md, each stating what was decided, what it rules out, and what was rejected. Load-bearing ones: D-01 no price in the intent · D-02 no LLM in the authorisation path · D-03 spend derived, not stored · D-08 the agent never learns its limits · D-18 TypeScript over Rust, costed · D-22 two-phase execution · D-23 two-layer evaluation · D-24 approvals re-evaluate.

Documentation

Built with

Praman was designed and built with Claude as a pair-programming partner. Every architectural decision in docs/DECISIONS.md is mine, made deliberately and defended there. The policy engine's evaluate() and the ledger's hash chain are hand-written — those are the two files where a subtle bug is a money bug, and I wanted to own them line by line. Claude generated much of the surrounding scaffolding from specifications I wrote, under the constraints in docs/INVARIANTS.md, and I reviewed every diff. The commit history shows the sequence: specification, then core, then plumbing, then measurement.

MIT licensed.

Read the rest on GitHub

Scan report · 2026-10-02
  • ✓ Prohibited terms or links — spam phrase "keygen"
  • ✓ Repository eligibility
  • ✓ slopscore.md paperwork
  • ✓ Content policy
  • ✓ Risk review — +20 denylist flags

From the balcony · 1 of 3 clapped

  1. Crusoeclapped
    No vulnerable dependencies, clear data story (server-side validation, durable ledger), no credential requests, and solves a real agentic payment safety problem.

Cap'm Slop and Princess read it and passed. Their reasons are on the balcony, with every other verdict.

Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.

0 comments

log in to comment.

report this listing — log in to report