A lonely developer's friend and assistant. An AI voice-conversation IoT device I built myself, wired into my own systems, that I can tell to do anything.
Tap a button, talk to the device → audio travels through the cloud → an AI agent (Claude) thinks → the device answers back in speech, in a voice I designed. It doesn't just reply — it actually controls my home and my computer.
Built from a single cable up, on an ESP32-S3 (Freenove FNK0082) + a brain (a Mac at home, or any small Linux box/VPS) + cloud MQTT.
daijin_demo_readme.mp4
One take on the voice side: I talk to daijin at home, the Claude brain on my Mac decides, and two nodes on their own networks act — a cat-treat dispenser in the next room and a device waiting in a park.
daijin is not "a better Siri." It's a different kind of thing.
- The brain is an agent that reasons and acts, not a command matcher. Siri pattern-matches fixed commands. daijin's brain is Claude — it holds a real conversation and acts on my scripts, MQTT devices, and data through shell access on my Mac.
- I own it and I can rewire it. Personality, prompts, and the set of allowed actions are mine to define. Closed assistants (Siri, Alexa, even ChatGPT voice) can't give you an agent that works inside your own infrastructure.
- It's a separate being, not an app on my phone. Something you place on a desk or carry around. A relationship, not a tool.
→ daijin's differentiation is ownership + agency: a fully-owned, open-ended agent device embedded in the systems of my life.
flowchart LR
subgraph device["ESP32-S3 device (ears + mouth)"]
MIC["INMP441 mic<br/>(tap-toggle record)"]
SPK["Speaker + I2S amp<br/>(jitter ring buffer)"]
end
BROKER{{"HiveMQ Cloud<br/>MQTT / TLS<br/>(both sides rendezvous outbound)"}}
subgraph brain["Brain: Mac or Linux server (outbound only)"]
STT["Groq Whisper API / whisper.cpp<br/>(STT, Korean)"]
AGENT["Claude<br/>(claude -p, streaming)"]
TTS["Edge TTS (default) / ElevenLabs / macOS<br/>(sentence by sentence)"]
STT --> AGENT --> TTS
end
MIC -- "PCM chunks<br/>daijin/audio/in" --> BROKER
BROKER --> STT
TTS -- "ADPCM chunks<br/>daijin/audio/out" --> BROKER
BROKER --> SPK
AGENT -. "controls" .-> HOME["Home LEDs / devices"]
- They meet at a cloud broker → works on any network: home WiFi or a phone hotspot on LTE. The brain only goes outbound (no inbound exposure needed), so it can be a Mac at home or a headless Linux server — the device does not care which. See Run the brain on a Linux server.
- Fully streaming reply pipeline: Claude's answer is split into sentences as it generates → each sentence is synthesized and published immediately → the device starts speaking after the first sentence, not the full answer. Honest numbers: a full round trip (record → STT → agent → TTS → playback) typically takes 6–22 s depending on how much the agent thinks and does — streaming hides part of that, but this is a conversation, not real-time.
- Audio protocol (docs/chunking-protocol.md): PCM 16 kHz/16-bit mono, chunked over MQTT with a 4-byte header
[clipId][seq:2][flags]. Both directions are IMA ADPCM-compressed 4:1 — the ~19KB/s bandwidth ceiling applies to the ESP32's send buffer too, so raw 32KB/s PCM can't keep up in either direction. Downstream plays through a jitter ring buffer; upstream is drained by a dedicated capture task so a blocking TLS publish never drops samples. - The brain =
claude -p(it's an agent, so it controls the home's LEDs and devices by voice) with a pinned session for cross-conversation memory. Security trade-off, stated plainly: the agent runs with broad tool access (shell, file read/write, web search/fetch) so it can genuinely act on the Mac — which means the MQTT broker credential is the security boundary: anyone who can publish to the audio topic can drive the agent. Mitigations in place: TLS to the broker with certificate verification on every side (the firmware keeps verifying even when NTP is blocked), credentials kept only in gitignored files, the brain makes outbound connections only, and the system prompt requires spoken confirmation before destructive actions (a soft guard, not a hard one). Give the device and the nodes their own broker logins limited to their topics, so a node left outside cannot reach the agent: docs/broker-credentials.md.
The free broker lives in the EU: ~300 ms RTT from Korea. The ESP32's lwIP TCP receive window is ~5.7 KB and fixed. Max throughput = window ÷ RTT ≈ 19 KB/s — physically below the 32 KB/s that raw PCM playback needs. No amount of code fixes that; compressing to 8 KB/s does.
daijin separates the "ears + mouth" (device) from the "brain" (agent) over MQTT topics. So the brain is swappable, and daijin itself can be used as the voice I/O device for any agent.
- Swap the brain: anything in place of
claude -p— OpenAI, Gemini, a local LLM (EXAONE), the Claude Agent SDK, etc. STT/TTS and the device stay the same. - Generic voice I/O: subscribe/publish just the
daijin/text/in(what I say) anddaijin/text/out(the reply) topics, and any agent can use daijin as its ears and mouth (daijin handles audio, STT, TTS). - Multi-agent routing: dispatch by transcribed intent to a home-control agent / coding agent / chit-chat agent.
- MCP integration: expose daijin's actions (speaking, device control) as MCP tools, or have the brain wire up multiple MCP servers to extend its abilities modularly.
The topic contract (
daijin/audio/*,daijin/text/*) already acts as the interface, so swapping the brain or attaching an external agent is just one adapter layer. (Current brain = Claude; the extensions are roadmap.)
- Phase 2 shipped — the device itself listens and speaks: tap BOOT to record, tap again to send; the reply streams back and plays with zero underruns, at home or on a phone hotspot outside.
- Voice: free Edge TTS is the default engine (neural multilingual voice); a custom voice designed on ElevenLabs (streamed as PCM, flash model) is opt-in via
TTS_ENGINE=eleven, with macOS TTS as the offline-safe final fallback. - Device UX: ready chirp through the speaker, dim-green idle LED, status colors for record/upload/play (LTE connects can take 30+ seconds — sound beats a blinking LED).
- Earlier milestones: HTTP LAN control → MQTT LAN → HiveMQ Cloud portable control (verified from a phone on LTE) → cloud voice round-trip with a fake device.
- Board: ESP32-S3-WROOM (Freenove FNK0082 kit) on the GPIO extension board + breadboard
- Mic: INMP441 I2S MEMS (WS=1, SCK=2, SD=42, L/R→GND) — must be read as 32-bit I2S frames, top 16 bits used
- Speaker: 8Ω 2W through the kit's Audio Converter & Amplifier module (PCM5102 DAC: BCK=14, LCK=12, DIN=13, SCK unconnected)
- Power: USB or a plain power bank — it's a take-it-with-you device
sketches/ ESP32 firmware (arduino-cli)
01~06 learning sketches (LED/WiFi/servo/sensor)
10-remote-led HTTP remote LED (early)
11-mqtt-led LAN MQTT
12*-funnel Tailscale Funnel attempt (unstable for portable use — kept as a record)
13-cloud-led HiveMQ Cloud (portable ✅)
14-daijin the daijin device: mic → cloud → speaker ✅
15-home-node home fleet node: one firmware per NODE_NAME — retained heartbeat, MQTT LWT, servo/relay/DHT/PIR switches
test-mic mic verification (record → HTTP → whisper)
test-speaker speaker verification (I2S tone melody)
mic-diag auto-sweeps every pin/slot combo to find I2S wiring empirically
voice/ daijin
daijin_mqtt.py brain: MQTT client (STT → Claude → streaming TTS → ADPCM chunks) — runs on macOS or Linux
daijin_common.py shared by the brain, the fleet watcher and the tools: secrets, local hooks, broker client
nodes.json the home nodes in one place: id, number, spoken name, LED colour, firmware flags
nodectl.py node/device control on one broker connection; dev.sh, fleet.sh, status.sh, say.sh wrap it
broker.sh broker choice and CA bundle for console.sh (Mac/Linux)
daijin_server.py legacy: old HTTP brain (superseded by daijin_mqtt.py, kept as record)
talk.sh legacy: Mac-only voice loop (Phase 0, superseded by daijin_mqtt.py)
dev.sh command a home node and verify this command's ack
fleet.sh fleet status: ONLINE / STALE(heartbeat gap) / OFFLINE(broker LWT)
say.sh proactive speech: daijin/say (verbatim) · daijin/ask (agent composes)
install-agents.sh macOS: register brain + fleet watcher as LaunchAgents
install-services.sh Linux: register them as systemd user services (systemd/*.service)
test_device.py device stand-in: sends a question as ADPCM chunks, decodes the reply (`--say "..."` needs no recording)
mic_test_server.py HTTP endpoint for mic verification
tests/ unit tests (ADPCM pair, clip reassembly, session handling, TTS fallback, node table, node commands) — run `python3 -m unittest discover -s voice/tests`; stdlib only, no network or hardware
docs/ roadmap · research · audio protocol · broker credentials
ESP32-S3 · Arduino (arduino-cli) · MQTT (HiveMQ Cloud, TLS/Let's Encrypt) · Groq Whisper API or whisper.cpp (STT, Korean) · Claude CLI (agent brain, stream-json) · Edge TTS (default) / ElevenLabs / macOS TTS · IMA ADPCM · launchd / systemd
Everything here runs on a brain machine (a Mac, or a Linux server — see the next section) plus one or more ESP32-S3 boards. Clone anywhere; paths are derived from the repo location.
Mac prerequisites
- Claude Code CLI, logged in (
claudeon PATH or in~/.local/bin) brew install whisper-cpp mosquitto arduino-cli— whisper for STT, mosquitto for themosquitto_pub/mosquitto_subhelpers, arduino-cli for flashingpip3 install paho-mqtt edge-tts- whisper model: download
ggml-large-v3-turbo-q5_0.binintovoice/models/ - A cloud MQTT broker with TLS. The free HiveMQ Cloud tier is what this repo was built against.
Secrets (gitignored)
cp secrets.local.example secrets.local.txt # broker, TTS engine, model
cp sketches/secrets.h.example sketches/14-daijin/secrets.h
cp sketches/secrets.h.example sketches/15-home-node/secrets.h
Fill in the broker credentials and your WiFi networks. ca_cert.h (Let's Encrypt root) is committed; nothing else needs a certificate.
Firmware
arduino-cli lib install "PubSubClient" "WiFiManager" "ESP32Servo" "Adafruit SSD1306" "Stepper" "DHT sensor library"
cd sketches
arduino-cli compile --fqbn esp32:esp32:esp32s3 --board-options CDCOnBoot=cdc 14-daijin && \
arduino-cli upload --fqbn esp32:esp32:esp32s3 --board-options CDCOnBoot=cdc -p /dev/cu.usbmodemXXXX 14-daijin
bash flash-node.sh dev1 /dev/cu.usbserial-XXXX # home node: servo + buzzer + relay
bash flash-node.sh dev2 /dev/cu.usbserial-XXXX # home node: stepper + buzzer
Wiring for the device (mic, amp, OLED) is in Hardware; node pins are at the top of 15-home-node.ino.
Brain
python3 -u voice/daijin_mqtt.py # foreground, logs to stdout
bash voice/install-agents.sh # or: register brain + fleet watcher as LaunchAgents (auto-start, auto-restart)
Then tap BOOT on the device and talk. bash voice/console.sh shows every transcript, reply, and node command live; bash voice/fleet.sh lists node status; bash voice/dev.sh red motor drives a node by hand (nodes are called red and blue; they light up in their colour).
Optional: local broker bridge. Running mosquitto on the brain machine and bridging it to the cloud broker cuts each node command from ~3 s to well under a second. The helper scripts use localhost:1883 automatically when it is up. Bridge config is outside the repo; the shape is a standard connection block with daijin/dev/+/cmd out and daijin/dev/+/state|status|event in.
The brain is the only piece that ran on the Mac, and nothing in it needs macOS. Put it on an always-on Linux box — a €5 VPS, a Raspberry Pi, an old laptop — and the Mac can sleep or leave the house. What changes:
| On the Mac | On Linux |
|---|---|
| whisper.cpp on the GPU (~2 s per clip) | Groq Whisper API (STT_ENGINE=groq): the same whisper-large-v3-turbo model, same vocabulary prompt, free tier covers home use. Local whisper.cpp still works as the fallback if you download the model, but on a 2-vCPU VPS it takes tens of seconds per clip. |
claude logged in via browser |
claude setup-token on your Mac → paste into CLAUDE_CODE_OAUTH_TOKEN in secrets.local.txt (1-year subscription token; Pro/Max). |
afconvert, say |
ffmpeg. No offline TTS fallback on Linux; Edge TTS is the default anyway. |
| LaunchAgent | systemd user service (install-services.sh) |
| The agent's shell is your Mac | The agent's shell is the server. Home nodes are still driven over MQTT, so dev.sh/fleet.sh work unchanged; anything that needed the Mac itself (its files, its apps) does not. |
Conversation memory (the pinned Claude session) lives on the machine that runs the brain, so a new machine starts with a blank memory.
# Debian/Ubuntu
sudo apt install ffmpeg mosquitto-clients python3-pip git
pip3 install paho-mqtt edge-tts # or: pip3 install --break-system-packages …, or use a venv
curl -fsSL https://claude.ai/install.sh | bash # Claude Code CLI → ~/.local/bin/claude
git clone https://github.com/nonasking/daijin.git && cd daijin
cp secrets.local.example secrets.local.txt # broker creds + GROQ_API_KEY + CLAUDE_CODE_OAUTH_TOKEN (from `claude setup-token` on the Mac)
python3 -u voice/daijin_mqtt.py # foreground first: tap BOOT, talk, watch the log
bash voice/install-services.sh # then register as systemd user services (auto-start, auto-restart, survives logout via linger)
journalctl --user -u daijin-brain -f # the services log to journald
python3 voice/test_device.py --say "지금 몇 시야?" # no device at hand: a stand-in sends the question and saves the spoken replyRun only one brain at a time against a broker: two subscribers on daijin/audio/in would both answer. Stop the Mac's LaunchAgent (bash voice/install-agents.sh remove) once the server is up.
A server in the EU sits next to the HiveMQ broker, so node command acks drop from ~3 s to well under a second even without the local bridge. Moving later to a home server is the same install; the local mosquitto bridge simply becomes available again.
Apache 2.0. See LICENSE and NOTICE.
- ESP32 is 2.4GHz WiFi only, and the USB connection must be a data cable; ESP32-S3 serial needs
CDCOnBoot=cdc. - The INMP441 speaks 24-bit-in-32-bit I2S frames — reading 16-bit frames yields full-scale white noise (
k×512+1patterns). - Press-fit header pins are not connections. Solder them — and check for solder bridges (a VDD–GND bridge shorts the whole board: "plug GND anywhere and the lights die").
- Breadboard power rails are dead plastic unless something actually feeds them.
- When wiring is uncertain, don't stare at photos — sweep: firmware that tries every pin permutation and reports signal levels found our miswiring in one pass (
mic-diag). - MQTT QoS1 gates every message on a broker round-trip — for real-time audio use QoS0 and let TCP stream.
WiFi.setSleep(false)— or the ESP32's receive throughput silently collapses to ~12 KB/s.- Bandwidth-delay product is real: far broker × tiny TCP window = a hard throughput ceiling. Compress, or move the broker closer.
- For latency, stream sentences, not essays: the user hears the first sentence while the rest is still being thought.
- Name nodes by sound, not by number. Measured with whisper large-v3-turbo on the same 144 synthetic utterances (pink noise at ∞/20/10/5 dB): IMA ADPCM 4:1 costs about 1 CER point over raw PCM, mostly smeared consonants; with a vocabulary prompt about half that. The node names mattered more than the codec: 일 번 / 이 번 ("one"/"two") differ by one final consonant and were confused 4/16 times even from raw PCM, while red / blue were 0/16. Nodes are now called by colour and light up in it.
- ESP32-S3 TLS needs the right time to verify certificates, and LTE carriers often block NTP. Falling back to no verification hands the broker password to anyone on the same network; fall back to the
Dateheader of a plain HTTP response instead (tls_time.h). - For a portable backend, a cloud broker beats a home Mac + tunnel (unstable).
Built with Claude Code.
0 comments
log in to comment.