You read something, it makes sense, and you feel like you understand it—until you try to explain it out loud. Then the missing step, vague phrase, or confidently wrong fact suddenly appears.
Quack is the listener that catches what you gloss over. It is a restrained voice tutor for Feynman-style learning: explain a topic aloud, hear nothing while your reasoning is sound, and receive one sharp question when a real gap appears. At the end, Quack turns the complete conversation into a focused recap of what you understood, what you repaired, and what still needs review.
Quack runs locally. There's no hosted demo because it needs your own OpenAI API key to make live voice and reasoning calls.
Watch the demo: https://youtu.be/t_CSa1QR9Tc
Run it yourself:
Clone the repo.
Backend: cd backend, create and activate a venv (python -m venv .venv then source .venv/bin/activate), install with python -m pip install -e '.[dev]'. Create backend/.env containing OPENAI_API_KEY=your_key_here. Start it with uvicorn app.main:app --reload.
Frontend (new terminal): cd frontend, then pnpm install, then pnpm dev.
Open http://localhost:5173 and allow microphone access.
Try saying something wrong on purpose — e.g. explain basketball and mention "Lionel Messi plays basketball" — and Quack will interrupt.
Topic, pasted context, and PDF/TXT/Markdown sources are converted into a structured GPT-5.6 evaluator brief. The browser then opens a live GPT-Realtime-2.1 WebRTC session with speech-to-speech audio, streaming transcription, and a parallel mid-speech error detector.
Quack does not rely only on a prompt that says “stay quiet.” When an unprompted learner explanation is correct, the Realtime model calls a wait_for_user tool. The browser records that structured decision, returns its tool result, and deliberately does not request another spoken response. Silence is therefore an application state, not hopeful wording that the model may drift away from. The one narrow exception is an open question loop: after Quack asks a direct diagnostic question, the browser disables tool use until Quack verbally resolves the answer. A correct answer receives exactly “Right.” or “Correct.”; a still-wrong answer receives one new sharp question and keeps the loop open.
A general voice assistant can be prompted to speak less, but it is still optimized to respond. It may acknowledge a pause, encourage the user, or ask a follow-up question. Quack gives the model a concrete non-speaking action and makes that action part of the protocol.
The primary voice session uses low-eagerness Semantic VAD so pauses, lists, and unfinished thoughts remain protected. In parallel, a second WebRTC session commits small audio windows for low-delay transcription. As soon as that channel produces a complete, falsifiable clause, a conservative GPT-5.6 classifier checks it against the prepared ground truth. A confirmed contradiction disables the pending VAD response, locks the turn against duplicates, and forces one immediate response.create interruption.
Prompting a normal voice assistant cannot create true mid-speech interruption on its own: the model is ordinarily invoked only after VAD decides that the user’s turn has ended. Quack’s parallel channel evaluates speech while the learner is still talking and explicitly triggers inference when an error is found.
The entire project was scaffolded and built through Codex—from the product specification and architecture to the FastAPI/React implementation, prompts, tests, live debugging, and visual design.
Key architecture decisions developed through Codex include:
- Dual WebRTC sessions: one patient speech-to-speech session for the conversation and one transcription-only session optimized for mid-speech claim detection.
- Tool-enforced silence: the
wait_for_userfunction makes “say nothing” an explicit control path instead of a fragile prompt preference. The browser makes that tool unavailable while Quack has an outstanding question, ensuring direct answers receive either a one-word verdict or one new sharp question. - Conservative, fail-closed checking: only complete claims with evidence in the prepared brief may trigger an interrupt. Ambiguity, fragments, timeouts, and classifier failures fall back to the patient voice path.
GPT-5.6 runs at three distinct moments:
- Before the session:
gpt-5.6converts the topic and supplied sources into a ground-truth evaluator brief containing key facts, required reasoning steps, and common misconceptions. - During the explanation: the GPT-5.6 family’s low-latency
gpt-5.6-lunamodel performs narrow contradiction checks on completed live clauses against that brief. - After the session:
gpt-5.6receives the evaluator brief and full ordered transcript to produce the evidence-grounded recap.
These examples make the interrupt easy to reproduce. Prepare the topic, begin explaining naturally, then include the suggested incorrect claim in the middle of a longer sentence.
| Topic | Suggested wrong statement |
|---|---|
| Basketball rules | “Basketball is played on a grass field, and each team tries to move the ball…” |
| Photosynthesis | “Plants get most of their mass from the soil, which is why roots are so important…” |
| Cell biology | “The nucleus is where the cell produces usable energy, and then that energy…” |
| Earth’s seasons | “Summer happens because Earth moves closer to the Sun, while winter happens farther away…” |
frontend/ React + Vite + TypeScript single-page experience
backend/ FastAPI preparation and secure Realtime credential service
outputs/ Product and engineering specification
cd frontend
pnpm install
pnpm devOpen http://localhost:5173.
Use a current Chromium, Safari, or Firefox browser and allow microphone access when the prepared-session screen starts the live experience. localhost is treated as a secure browser context for microphone development.
cd backend
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[dev]'
uvicorn app.main:app --reloadThe API runs at http://localhost:8000; readiness is available at GET /health.
Create .env in backend/ before starting the API:
OPENAI_API_KEY=your_key_here
QUACK_ALLOWED_ORIGINS=http://localhost:5173,http://127.0.0.1:5173
QUACK_PREPARATION_MODEL=gpt-5.6
QUACK_ERROR_DETECTION_MODEL=gpt-5.6-luna
QUACK_REALTIME_MODEL=gpt-realtime-2.1
QUACK_REALTIME_TRANSCRIPTION_MODEL=gpt-realtime-whisper
QUACK_REALTIME_VOICE=marincd frontend && pnpm test && pnpm build
cd backend && pytestCopy .env.example to backend/.env and set the real API key there. Never expose OPENAI_API_KEY to the frontend. The browser receives only a one-minute Realtime client credential minted for the prepared session.
- The backend uses GPT-5.6 to prepare a compact ground-truth evaluator brief.
POST /api/realtime-tokenmints the short-lived GPT-Realtime-2.1 voice credential, whilePOST /api/transcription-tokenmints a separate low-delay transcription credential.- The browser sends the same microphone track to both WebRTC sessions. The voice session keeps patient Semantic VAD; the transcription-only session commits short audio windows while speech continues.
- Complete declarative claims are sent to
POST /api/error-check, where GPT-5.6 Luna performs a narrow, no-reasoning contradiction check against the prepared brief. - A confirmed material error disables automatic VAD response generation, commits the current speech, mutes the microphone, and sends a forced
response.createinterruption immediately. A turn-scoped lock cancels any late duplicate response, and patient VAD is restored only when the learner begins the next turn. - Correct speech, fragments, lists, and thinking pauses stay on the ordinary low-eagerness Semantic VAD path.
The standard API key never leaves FastAPI. Browser audio tracks, remote audio, the data channel, and the peer connection are closed when the session ends, reconnects, or unmounts.
- Session preparation uses a live GPT-5.6 Responses API call and requires an API key.
- The in-memory prepared-session store is intended for the current single-process demo deployment; production scaling needs shared session storage.
- The fast detector is intentionally conservative and fails closed to the normal patient Realtime path if its classification request times out.
- Session state is intentionally ephemeral.
See the complete build specification for the architecture, prompt design, test matrix, and Build Week delivery plan.
Quack AI is available under the MIT License.
0 comments
log in to comment.