![]()
An agentic harness you can watch.
The agent's browser, terminal, desktop and file edits stream into the conversation in real time — each command, page and edit shown at the point it happened, and still there to scroll back to afterwards. Every session records itself, and reading one back is reading the thread. Approval in chat is off: nothing waits for you, and every command is still shown as it runs.
Built on one idea: the event log is the product. Everything the agent does emits a typed event to an append-only log. The live view is a subscriber to that log. A recording is that log read back. There is no separate recording pipeline to bolt on, and no way for the live view and the recording to disagree — they are the same code path.
agent ──emit──> event log (jsonl + blobs) ──┬──> websocket ──> live UI
│ └──> replay ─────> same UI
└── asciinema export, grep, tail -f
Autora is an agent with your machine's reach, and it is built to use it. This is the part to read before installing it anywhere.
- It has no login of its own. Run from the command line it binds
127.0.0.1and refuses everything else: whoever can reach the port is the operator. Installed on Umbrel it sits behind the dashboard's own login. Do not put it on the open internet, and do not forward a port to it. - Approval in chat is off. There is no per-call gate: commands run when the agent decides to run them, and you watch them happen rather than being asked first. That is the whole point of the app, and it is also the risk.
- The container is deliberately privileged. The Umbrel compose mounts the host
filesystem read-write at
/hostand setsAUTORA_WORKDIR=/host, because working on the server -- including deleting things -- is what it is for. The comments inblofstedt-autora/docker-compose.ymlsay so next to the mount. Run it on a machine you own, on a network you trust, in a VM if you can, and treat that machine as expendable. - Your provider keys stay on the server. They are called from the server and
never sent to the browser; a key saved in Settings is written to
.autora/settings.json, readable only by the account Autora runs as.
git clone https://github.com/blofstedt/Autora && cd Autora
npm install
export GEMINI_API_KEY=... # optional -- keys can also be added in Settings
npm run devOpen http://localhost:3000. Type a task. Watch it work.
It listens on this machine only. Opening it to the network is a choice, and
Before you run it has the terms: AUTORA_HOST=0.0.0.0 npm run dev when you want it from a phone or another computer, on a network you
trust. Requests and live connections started by
other websites are refused either way.
For production, npm run build then npm start. The container is the same
two steps: docker build -t autora . && docker run -p 8817:8817 -v autora-data:/data autora.
There is no Windows build and no installer: Autora is a Linux container, and on Windows that means Docker Desktop, which runs it on WSL2. Install Docker Desktop with the WSL2 backend, then the same commands, from PowerShell or WSL:
git clone https://github.com/blofstedt/Autora; cd Autora
docker build -t autora .
docker run -p 8817:8817 -v autora-data:/data autoraThen open http://localhost:3000 for npm run dev, or http://localhost:8817
for the container. Four things are worth knowing before you start:
- The container is deliberately privileged — it watches the agent's terminal, browser and desktop, so it is not sandboxed and Docker Desktop's defaults will not confine it. Before you run it has the terms, and they are the same on Windows as anywhere else.
- Paths are the container's, not Windows'. A tool call's
/host/datais inside the Docker VM.C:\Users\you\workis mounted with-v /c/Users/you/work:/host/work(PowerShell:-v C:\Users\you\work:/host/work), and that mapping is what the agent sees. npm installfor development needs the WSL side, not a Windows Node: the code runs on Linux, and a Windows install oftsx,esbuildandnode-ptyproduces modules the container cannot load. Do development in WSL (wslfrom PowerShell), where the repo behaves exactly as it does on Linux.- Line endings and file watching.
git config core.autocrlf falsein the clone: shell scripts with CRLF fail withbad interpreter, and the file watcher makes WSL2 spin on/mnt/c. Keep the repo on the Linux side (~/Autora) rather than under/mnt/c, which is slow for exactly this kind of work.
The desktop relay runs on the machine with the screen, so on Windows that is a Windows process talking out to the Autora container — see Controlling a desktop. Nothing about it needs the container to be on the same side of the WSL boundary.
Replies come from whichever provider you connect, called from the server -- the key never reaches the browser. On a fresh install the chat opens on a card that asks for one: pick a vendor, paste its key, and Autora checks it before saving. For the model list and every other option, open Settings in the sidebar:
| Provider | Get a key | Notes |
|---|---|---|
| OpenAI | https://platform.openai.com/api-keys | GPT-5, GPT-4.1, GPT-4o, o3/o4-mini |
| Google Gemini | https://aistudio.google.com/apikey | Free tier covers Flash |
| Anthropic Claude | https://console.anthropic.com/settings/keys | Sonnet, Haiku, Opus |
| DeepSeek | https://platform.deepseek.com/api_keys | Cheapest of the hosted options |
| OpenRouter | https://openrouter.ai/keys | One key, hundreds of models; list and prices fetched live |
| Local server | — | Anything speaking the OpenAI API: vLLM, Ollama, llama.cpp |
Each provider keeps its own key, model and endpoint, so switching between them costs nothing. Which provider answers picks one, or leave it on Automatic and the first connected provider is used, in the order shown in the panel.
Every model in the picker is quoted with its price per million tokens, and those are the same numbers the billing card counts with. Check key asks the vendor whether the key works before you rely on it, and fills the model list from the same call; Refresh models re-asks later, which is how the OpenRouter catalogue stays current.
Keys pasted into the panel are saved on the server in .autora/settings.json,
created readable only by the account Autora runs as. Keys set in the
environment (GEMINI_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY,
DEEPSEEK_API_KEY, OPENROUTER_API_KEY) are still honoured and are never
written to that file -- they are the fallback when the app has no key of its
own, and the panel says which of the two is in force.
Without any key the console still runs -- every panel, the event stream, memory, approvals and the schedule all work -- but nothing answers, and the thread says which key is missing rather than improvising a reply.
docs/GEMINI.md covers the Gemini specifics: choosing a model, the thinking budget, and what each error in the thread means.
Everything Autora says out loud -- the v key, replies read aloud, spoken
approvals -- uses one voice for the whole install. Left alone it is the
browser's own (speechSynthesis), which is a different voice on every
device and on most of them the worst one the platform has.
With a Deepgram key it speaks through Deepgram's hosted Aura voices instead:
nothing to run, one request a sentence, and the same voice on the phone, the
laptop and the desktop. Paste the key into Settings -> Secrets as
DEEPGRAM_API_KEY (or set it in the environment) and that is the whole setup.
Settings -> Voice lists the voices the service offers -- the list comes from
the service, so it is whatever it actually has -- plays a sample, remembers the
one picked on every device, and says which service is speaking.
There is one service and one key: Deepgram's hosted voices. A local voice
server is no longer looked for, and a key is the whole of the setup. Which
voice speaks can be changed in Settings -> Voice, and the choice is
remembered on every device. When the agent is asked to say something out loud
it uses its speak tool, which plays the words on the open page straight
away rather than making an audio file to hand over.
| Variable | Purpose |
|---|---|
DEEPGRAM_API_KEY |
Speak through Deepgram's hosted Aura voices. A key saved in Settings -> Secrets wins over this |
AUTORA_DEEPGRAM_VOICE |
The Deepgram voice to start with (default aura-2-thalia-en); a voice picked in Settings wins |
The page never talks to Deepgram itself. POST /api/speech takes a fragment
of text and returns an audio file, which keeps the key on the server and means
a phone on the tailnet needs no route of its own. POST /api/speech/stream
takes a whole reply and returns raw 24 kHz PCM as it is rendered, which is what
a narrated answer uses: Deepgram draws the voice fresh for every request -- the
same sentence sent twice comes back at a different pitch and pace -- so a
request per sentence was a voice that changed from sentence to sentence. The
first samples arrive in about half a second and run three or four times faster
than they are spoken, so one rendering for the turn still starts straight away. The audio arrives from the
origin the page already trusts. GET /api/speech says whether there is a key,
which voice it would use, and which voices the service has -- the list comes
from Deepgram's own model catalogue, so it is whatever it actually offers.
If there is none -- no key, wrong address, container down -- the panel says so and the browser's own voice carries on being used. Nothing goes quiet, and nothing is required to be set up.
The mark at the left of the composer, or v, is talk mode: the strip you type
into becomes the conversation and the thread above it keeps running. It needs a
microphone, which browsers only open on a secure page -- https, or
localhost.
There is one thing to do, and it is to say the name. Say "Autora" and whatever follows it -- "Autora, what's the weather" -- is the request; say the name on its own and whatever you say next is. It goes out when you pause: a sentence that has finished goes about a second later, one that trails off mid-thought waits longer for you to come back to it. The bar says "Say 'Autora' to add to the conversation…" while it is listening, and shows the words as they form once it is yours.
The microphone is open the whole time talk mode is, and everything that is not the name, at the start of a phrase, is dropped where it stands. That is what makes talking over an answer work: say the name while Autora is speaking, or while it is off working, and it stops and the words are yours. It is also how an open microphone avoids hearing the answer and sending it back as a question -- nothing is listening for anything but its own name.
There is no switch for any of it, and nothing on the bar to press except the X that takes you back to typing. The camera lives in Settings -> Voice; on, the picture runs the width of the bar, and the agent is shown about a frame a second of it.
The month's spend stays over the live bar in talk mode: a spoken turn costs what a typed one does.
A model decides; tools are what it decides with. Settings -> Tools lists the four groups, says whether each can actually be used right now, and sets how tightly each is held:
| Group | What it is | Asks first, by default |
|---|---|---|
| Terminal | A real shell on the machine Autora runs on, via bash -lc. Output streams into the thread as it arrives. |
Every command |
| Web browser | The Chromium the agent drives: open, read, click, fill, scroll, screenshot. | Anything that changes the page |
| Computer control | Somebody's actual desktop, over the relay. | Anything but looking |
| Memory | The workspace graph, which outlives the session. | Never |
The same list generates three things that used to be written down separately: the tool schemas the model receives, the sentence the agent is told about what it can do, and this panel. They cannot disagree, which is the point -- an agent that has been told it has a terminal it does not have will narrate command output rather than admit it, and one told nothing will deny having a shell it is holding.
A group can be off (you turned it off, and the agent is told so) or unusable (on, but the thing it needs is absent -- no Chromium on the server, no relay dialled in). Those are different sentences, and the agent gets the one that is true, because "I cannot browse" and "Chromium is not installed on the server" send you to two very different places.
Nothing waits for approval: every call runs straight away and is shown in the
thread as it happens. Stop kills the whole process group, so sleep 300 inside a
command dies with the command rather than outliving it.
The terminal is a shell, not a terminal emulator: there is no TTY, so vim,
top and anything that pages will hang rather than work. Use non-interactive
flags. sudo works wherever the account Autora runs as can use it -- in the
Umbrel container that is usually root, where it is unnecessary rather than
unavailable.
Integrations in the sidebar connects Model Context Protocol servers; their
tools reach the agent as mcp__<server>__<tool> from its next step.
You rarely need to go there first. When a task lives on a service with an API -- GitHub, Slack, Notion, a Postgres database, Google Maps, library docs -- and no connected server covers it, the agent offers one in the conversation instead of reaching for the browser: a card saying what it is, why it beats clicking through the website, and exactly what will run. Nothing is installed until you tap Set it up, and Not now is remembered for the rest of the session. Asking "what MCP servers could help with X?" gets an answer from the same list.
- Keys a server needs are typed on the card and saved straight to
Settings › API Keys › Secrets. The server's config only refers to them as
${secret:NAME}, filled in when it starts, so a key never enters the chat, the session log or the model's context. - Beyond the built-in list, the agent can offer any MCP server published on
npm (run with
npx -y), a remote one at a URL, or one it writes itself: a few tools, each a short piece of JavaScript, generated into a real stdio server under the settings directory (mcp-servers/). - A launcher that downloads a binary for the platform is given the build
for this machine before it starts, fetched once into the data directory
rather than a cache a container update wipes. webclaw's launcher has no musl
entry, so on an Alpine container it used to fetch a glibc build it could not
run and the only symptom was
MCP error -32000: Connection closed.
Servers the agent set up are marked set up by Autora on the Integrations page, where they can be edited, switched off or removed like any other.
The page also has a Suggested for you card, for the times you go looking rather than being offered one in a conversation. It is deliberately quiet and easy to satisfy: a server is suggested either because this install already holds the key it wants — paste a GitHub token anywhere and the GitHub server is a fact about the install rather than a guess — or because a word from its list appears in your standing instructions or in the titles of your recent chats. Nothing is installed by looking at the card, the line under each one says which of the two it was, and no suggestions at all is a normal state of affairs: a list padded to look helpful would be worse than an empty one.
Settings carries a billing card: this month, today, and all time, with a thirty-day chart and a breakdown by provider and model down to individual turns. It is counted locally -- every turn records the tokens the vendor reported and prices them from the table the model picker quoted -- so it is an accurate running estimate, not an invoice. It cannot see usage from outside this app, nor a vendor's own discounts, cache credits or minimums, and a turn whose vendor reported no token count is marked as estimated rather than silently averaged in.
An optional monthly budget draws a meter and warns as it fills. It does not block: a console that stops answering mid-sentence because of a number typed weeks ago is a worse surprise than the bill it was meant to prevent.
# any OpenAI-compatible server: vLLM, Ollama, llama.cpp, LM Studio
AUTORA_LLM_BASE_URL=http://localhost:8000/v1 npm run devThen pick Local server in Settings and type the model's name
(qwen3-coder, say). The address can be set there too, which wins over the
variable.
Autora speaks plain HTTP, which is fine over localhost and fine for
reading the thread from another device. Two things need more than that:
| Wants a secure page | Why |
|---|---|
| Install to the home screen | Browsers only install a PWA from an https origin |
| Dictation and live voice chat | Microphone capture is gated on a secure context |
A private network is not enough — http://box.tailnet.ts.net:8817 is an
insecure origin as far as the browser is concerned, and no amount of site
settings changes that.
On a tailnet, this is one command and it is the best answer there is. Tailscale issues a real certificate, so there is no warning to click through, and whatever was guarding the plain port — Umbrel's login, say — still guards it. On the machine running Autora:
tailscale serve --bg --https=8443 8817 # https://<machine>.<tailnet>.ts.net:8443 -> :8817
tailscale serve status # confirm, and see the URL to open
tailscale serve --https=8443 off # undoNot on 443. Plain tailscale serve --bg 8817 serves on 443, and on an
Umbrel that takes the port umbrelOS 2.0 needs for its own HTTPS dashboard:
umbreld then fails at boot with EADDRINUSE ... :::443 and restarts in a loop,
and the whole box is unreachable. If that has happened, free the port with
tailscale serve --https=443 off and start Umbrel again
(sudo systemctl start umbrel).
Open the https:// URL it prints, port and all, and the microphone,
the install prompt and everything else appear. It needs MagicDNS and HTTPS
Certificates enabled for the tailnet, both in the admin console under DNS;
tailscale serve will tell you if they are off. With a proxy in front you can
also drop --host 0.0.0.0 and go back to the default 127.0.0.1 bind, so
nothing has to listen on the tailnet interface at all.
Anything else that terminates TLS works the same way: Caddy or nginx with a
certificate from tailscale cert, or whatever reverse proxy you already run.
With nothing in front of it, AUTORA_TLS=1 is the fallback. Autora puts an
https listener on the next port up with a certificate it generates and keeps:
AUTORA_TLS=1 AUTORA_HOST=0.0.0.0 npm start # http on 3000, https on 3001Nothing moves — the plain port serves what it always did, and the http page offers a link across.
That certificate is signed by a small authority Autora generates and keeps, and
the authority is offered for download at /autora-ca.crt (Settings → Trust
this server has it, with the steps for Android and iPhone). Install it on a
device and the warning stops, the microphone works, and the page installs to
the home screen as a real app — three complaints that are one fact underneath:
nothing on that device had vouched for the certificate. Leaves are issued per
hostname during the handshake, from the name the browser asks for, so reaching
the box by tailnet name, .local or bare IP all match.
Install it only on devices you own: an authority you trust can vouch for any site to that device, and this one is worth exactly what the server holding its key is. Skipping it costs nothing but the warning — clicking through still opens the microphone, it just will not install to a home screen.
The listener itself has whatever authentication Autora has, which is none. That
is fine on a network you control and a poor thing to expose anywhere else,
which is why it is off by default in the Umbrel app (AUTORA_TLS=1 to turn it
on) and why a real certificate in front is better still.
Until then the microphone and Live are still in the composer, greyed, and
tapping either says what is in the way and offers the secure page if one is
running — rather than vanishing and leaving you to conclude voice was never
built.
| Env | |
|---|---|
AUTORA_TLS=1 |
Also serve https, with a generated certificate |
AUTORA_TLS_PORT |
Where (default: AUTORA_PORT + 1) |
AUTORA_TLS_CERT / AUTORA_TLS_KEY |
Use a real certificate instead |
AUTORA_TLS_NAMES |
Extra names or addresses for a browser that connects by bare IP (which names no host, so gets the certificate listing this machine's own addresses) |
The Umbrel app publishes 8818 for this and ships with AUTORA_TLS: "0". Put a
proxy in front if you can; set it to 1 if you cannot.
Once the page is on https — the tailnet URL above, or a certificate from
/autora-ca.crt — the browser will install it as an app, and it is worth
doing: it opens without the browser's chrome, it gets its own icon and
switcher entry, and the session list, the thread and the composer use the
whole screen.
| Android (Chrome, Edge) | Menu (⋮) → Add to Home screen → Install. Or the banner Autora shows when it sees an install prompt available. |
| iPhone / iPad (Safari) | Share → Add to Home Screen. iOS ignores the manifest's display mode, so installing is the only way to lose the browser bars. |
| Desktop (Chrome, Edge) | The install icon in the address bar, or menu → Install Autora. |
What it installs is public/manifest.webmanifest — standalone display,
portrait-friendly, icons at 192, 512 and a maskable 512 for Android's adaptive
shapes — plus the service worker in public/sw.js, which serves the app's
shell from cache so it opens instantly and does not show a browser error page
when the server is briefly away.
An installed copy is a copy: it keeps working while the server is down, which
also means it does not notice an update until it asks for one. Autora checks
/api/system and says so in a bar at the top when the version on screen is
behind the one being served — including a Reload button, and the service
worker calls skipWaiting so the reload actually lands on the new build
rather than the cached one.
Autora can drive a real computer's screen — click, type, scroll — through a small relay that runs on that computer. The relay connects out to Autora, so nothing has to be opened up on the machine with the screen, and it is a single file with no Autora imports, so that machine does not need Autora installed.
Get it from the server you want it to reach, which fills the address in for you:
# on the machine you want the agent to control
curl -O http://YOUR-SERVER:8817/relay.py # Windows: curl.exe -O ...
pip install websockets mss pyautogui pillow # Windows: py -m pip install ...
python relay.py # Windows: py relay.pySettings → Desktop relay in the UI has those same three commands with your server's own address already in them, says whether a relay is connected, and lists what to check when one will not connect.
python relay.py 192.168.1.5:8817 points it at a different address; any form
works — host:port, http://host:port, ws://host. Over https with Autora's
own certificate, the relay needs --insecure or the certificate from
/autora-ca.crt installed on that machine.
One thing catches people out: outside the container Autora listens on
127.0.0.1 unless told otherwise, so start it with AUTORA_HOST=0.0.0.0 if
the relay is on another machine.
Frames reach the conversation only while a session is using the desktop; a connected but idle relay is not an hour of screenshots in your transcript. The live preview in Settings tells you it is alive in the meantime.
macOS needs Accessibility permission for the terminal you run it from (System Settings → Privacy & Security → Accessibility) — capture works without it, clicking and typing do not. Linux needs an X display, so install XWayland under Wayland. Windows needs nothing special.
curl localhost:3000/api/sessions # list recordings
curl localhost:3000/api/sessions/<id>/events # one session's events, as JSONEach session's full log is also on disk as sessions/<id>/events.jsonl under
the state directory, one event per line.
One column: the conversation, with the work inside it.
| In the thread | Shows |
|---|---|
| A command | The shell call where the agent ran it, its own outScan report · 2026-09-30
From the balcony · 3 of 4 clapped
Cap'm Slop read it and passed. Their reasons are on the balcony, with every other verdict. Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are. report this listing— log in to report |





0 comments
log in to comment.