Memory Court is an auditable autonomous-agent demo built for OpenAI Build Week. GPT-5.6 investigates a memory case, chooses one structured action per turn, and proposes bounded cognitive-state interventions. sonuv-guard independently adjudicates every intervention as COMMIT, REPAIR, REJECT, or FORGET before any state is applied.
The interface makes the boundary explicit: GPT-5.6 supplies the investigation and proposal; Guard controls structured state execution. Guard does not review or certify the model's natural-language reasoning.
- Application: https://memory-court-build-week.vercel.app
- API health: https://memory-court-api-production.up.railway.app/api/health
- Source: https://github.com/billgaohub/memory-court-build-week
The public build keeps the server-side GPT-5.6 Responses API path intact. Its current key has no billable quota, so the default judgeable artifact is a competition-period trace generated by GPT-5.6 Sol inside Codex, executed against the real Guard adapter, and displayed as REPLAY MODE. It is explicitly not presented as an API-live run.
- Open the application and select The Silent Lifeboat, the competition-period case.
- Inspect the gold REPLAY MODE trace. Its provenance banner states that GPT-5.6 Sol generated the actions inside Codex—not through the OpenAI API—and the checked-in source action file is reproducibly adjudicated by sonuv-guard.
- If the server later has usable API quota, choose Run live audit. Otherwise the failed live request automatically returns to the same labeled replay.
- Choose Export audit JSON to inspect the complete evidence artifact.
No account is required. The public demo is free to test and rate-limits each client to five new live sessions per ten minutes.
Vercel: React + TypeScript + Vite
|
v
Railway: FastAPI -> OpenAI Responses API (GPT-5.6)
|
v
schema validation -> sonuv-guard -> committed state + audit event
The live loop is deliberately bounded to 8 model calls, 8 events, 3 intervention proposals, and 600 output tokens per call. Each OpenAI request has a 30-second timeout; timeout and rate-limit responses retry once. Sessions live in one API process for one hour. This is a transparent hackathon-grade design, not a distributed production control plane.
Requirements: Python 3.11+, Node.js 22+, and an optional OpenAI API key for live mode.
git clone https://github.com/billgaohub/memory-court-build-week.git
cd memory-court-build-week
python3 -m venv .venv
. .venv/bin/activate
pip install -e 'backend[test]'
export OPENAI_API_KEY='your server-side key'
export ALLOWED_ORIGINS='http://localhost:5173'
uvicorn memory_court.app:app --app-dir backend --host 127.0.0.1 --port 8000In a second terminal:
cd frontend
npm ci
VITE_API_BASE_URL=http://localhost:8000 npm run devWithout OPENAI_API_KEY, /api/health reports live_available=false and the frontend presents the labeled replay.
bash scripts/verify.shThe gate runs 36 backend tests, including exact reproduction of the Codex-generated Guard trace, frontend interaction tests, TypeScript checking, a production build, vendored-source hashes, and—inside the original source workspace—the upstream 121-test sonuv-guard regression suite.
The competition-period Silent Lifeboat replay was generated by GPT-5.6 Sol inside Codex in primary task 019f725e-6f43-78c2-8587-4ad6a3725d9f. The exact six source envelopes are checked in at replay/silent_lifeboat.codex-actions.json. A backend test executes them through the same AgentSession and sonuv-guard adapter and requires the published six-event replay, including REPAIR followed by COMMIT, to match exactly.
Separately, the retained live backend uses the async OpenAI Python SDK Responses API with gpt-5.6, Pydantic structured outputs, low reasoning effort, low verbosity, store=false, and a 600-token ceiling. Each response must be exactly one of:
inspect_memorypropose_interventionfinalize
Only propose_intervention reaches Guard. Invalid fields, booleans masquerading as integers, out-of-range values, and empty patches are rejected before adjudication.
The recorded Codex trace and API-live path are never conflated: replay uses mode=replay, api_live=false, and the model label gpt-5.6-sol via Codex (recorded).
Codex was used to review the original game plan, isolate a standalone architecture, create the approved design specification, implement the backend and frontend test-first, package the pre-existing Guard runtime with provenance hashes, prepare deployment, and run the completion audit. The primary build task ID and commit evidence are recorded in CODEX_EVIDENCE.md.
The autonomous loop, FastAPI service, React audit console, Silent Lifeboat case, replay system, deployment packaging, tests, and submission materials were created during the competition period. A precise snapshot of pre-existing sonuv-guard commit 62157c5 and the pre-existing Last Birthday case are disclosed in PREEXISTING_VS_NEW.md.
- Railway builds
backend/Dockerfilefrom the repository root. ConfigureOPENAI_API_KEY,OPENAI_MODEL=gpt-5.6,TRUST_PROXY=true, andALLOWED_ORIGINS. - Vercel uses
frontendas Root Directory,npm run buildas Build Command,distas Output Directory, and the Railway origin asVITE_API_BASE_URL. - After Vercel assigns the production domain, set that exact origin in Railway's
ALLOWED_ORIGINSand redeploy.
See SECURITY.md for the threat boundary and SUBMISSION.md for competition copy.
MIT. See LICENSE.
0 comments
log in to comment.