An open-source AI agronomy agent that runs locally on your own computer or server, connects to messaging apps like Telegram and WhatsApp, and computes aquaponics system designs from a deterministic, source-cited engineering core instead of guessing them.
A self-hostable agent specialized for agriculture: a domain application in the spirit of Hermes / OpenClaw, rather than another agent framework. Its first deep domain is aquaponics. Describe your water, space and species in one sentence and it returns a buildable design, with a bill of materials, an operating envelope, a source for every number, and an explicit list of what it does not model. It will also search fish × crop mixes for the ratio that grows the most food from the least water.
Built by a hands-on aquaponics operator to cut the pain he lived: years of reading papers and losing fish to figure out what the math could have told him up front.
The sizing method behind it is a granted Taiwan utility model patent (TW M661364). The code is MIT, runs on open weights, and needs no proprietary API.
pip install agronaut
agronaut size --fish tilapia --crop lettuce --area 12 --temp 27 --water 3000That prints a sized system: tank and system volume, fish count, feed rate, biofilter media, pump duty, a bill of materials, a source for every number, and an explicit list of what it does not model. Nothing in that command touches a network.
Try agronaut list for the species and crops it knows, or agronaut optimize --area 10 --temp 28 --water 5000 --objective food to search fish × crop ratios.
agronaut setupIt asks which model and which channel you want, checks each key against the live service as
you paste it, reads your Telegram id off a message you send your own bot, and writes
~/.config/agronaut/.env itself. Nothing to hand-write.
The model is your choice: Claude or NVIDIA with your own key, or a local model through
Ollama with no key at all. Run agronaut setup again later and it shows what you have and
asks what to change, so switching from Claude to a local model (or back) or adding a
channel touches only that part. Saved keys are kept, so switching back needs nothing
re-typed. To jump straight to one part:
agronaut setup model # switch between Claude, a local model and NVIDIA
agronaut setup telegram # connect, keep, or allow another Telegram account
agronaut setup whatsapp # add or update WhatsApp; saved values are kept on EnterPrefer to configure it by hand?
Chat needs a language model. Ollama is the shortest path, and Agronaut already defaults to it:
ollama pull qwen3.5:4b # ~3.4 GB, once. Any tool-calling model works.
agronaut # chat in your terminalThen just say what you have: "I have a 20 m² greenhouse in Bobo-Dioulasso, water sits around 28 °C, I want tilapia and lettuce."
Prefer a browser? agronaut web serves the Streamlit app on
localhost:8501.
Photos too, if you want them: ollama pull llama3.2-vision and set VLM_PROVIDER=ollama.
Config lives in ~/.config/agronaut/.env for an installed copy, or ./.env in a checkout.
Tight on memory?
qwen3.5:2b(2.7 GB) is smaller: pull it and setLLM_MODEL=qwen3.5:2b. Would rather not run a model at all? A free hosted key works instead:LLM_PROVIDER=nvidiawithNVIDIA_API_KEYfrom build.nvidia.com.
# .env in the project root
TELEGRAM_BOT_TOKEN=... # from @BotFather
AGRONAUT_ALLOWED_IDS=... # your Telegram user id, so it is yours aloneagronaut botMessage your bot. /log ammonia 0.5 nitrate 40 temp 27 puts a reading into your live twin
and /forecast tells you what the week ahead does to it — both with no model in the path,
so they work even when the LLM is slow or unreachable.
| You want | You need |
|---|---|
size, size-hydro, optimize, list, the web calculator |
Nothing. Pure aqua_model: deterministic, offline, cited. |
| Chat, the Telegram bot, photo understanding | A model — Ollama locally, or a hosted key |
/log, /forecast, /advise, /approve |
A model for setup, then nothing: the twin commands never call one |
On a phone, Telegram is the recommended channel; WhatsApp works too but takes more setup (guide). Trouble, or want Docker or a hosted demo? See Install and run: all the options.
A chatbot retrieves what a paper said. Agronaut computes the answer for your specific system. The trustworthy part is a deterministic engineering model — the LLM only collects facts, routes to the right tool, and explains results in plain language.
YOU ──▶ agent layer (LLM: collect facts, route, explain)
│ proposes values
▼
validation gate ── rejects bad/uncertain input ──┐
│ typed, validated │
▼ │
aqua_model (TRUST ZONE — pure, tested, cited) │
coefficients ▸ mass balance ▸ sizing ▸ optimizer ◀───┘
│
▼
a sized system + bill of materials + operating envelope
+ cited coefficients + an explicit "what's NOT modeled" list
The math is verifiable on its own — you can audit every coefficient (with its source) without trusting the model. Calibration ≠ validation: the engine ships with seed defaults from published sources, meant to be calibrated against a real running system.
Three modes in the app (sidebar Mode switch):
- Assistant (chat) — troubleshoot a running system (low DO, yellow leaves, pump sizing…).
- Design Calculator — fixed inputs → a fully sized system: tank/system volume, fish count, feed/day, pump turnover, biofilter, makeup water, bill of materials, operating envelope, maintenance checklist, and a downloadable funder-ready report.
- Optimize Ratio — search fish × crop-mix combinations for the best ratio under your binding constraint (e.g. a fixed water budget), maximizing food, protein, or water-use efficiency, and showing the gain over a naive even split.
The design and optimizer modes are fully deterministic and need no LLM at all.
Photograph a yellowing leaf, a sick fish, or green water — on Telegram, WhatsApp, or the web chat — and you get a cited differential, not a guess:
- A vision model describes what it sees. It only observes.
- A deterministic guard strips any measurement or prescription out of that description, so a
fabricated
pH 6.4or an inventedadd 5 mL of saltcan never enter the conversation as though you had said it. A named condition is kept but flagged unverified. - A fixed, cited table (
aqua_model/triage.py) maps the visible symptoms to a ranked list of candidate causes — each naming the knowledge document it came from, and each with the checks that would tell it apart from its neighbours.
It will not hand you a single confident diagnosis, because a photograph cannot support one: iron deficiency and pH lockout look identical in an image, so you get both plus the check that separates them. Ordering follows the knowledge base's own rules — pH before iron, water quality before any fish pathogen. Nothing it says states a dose.
Every design renders as a self-contained 3D page — greenhouse, tanks, filtration, beds, graded plumbing with the flow animated in the direction the water actually goes. One HTML file, no server and no CDN, so it opens from a double-click on a laptop that has never been online.
With a twin bound to it, the same drawing stops being a picture:
# a design, plus the season it would have at a real site, on a slider
python scripts/render_3d.py --crop basil --site taichung_2025 --days 365 -o first_year.htmlDrag the scrubber and the fish grow, the water turns amber and then red as ammonia and
nitrite cross the bands aqua_model/advisory.py acts on, and the crop is drawn as vigorously
as it is actually growing. You watch the nitrite spike of week two arrive instead of reading
about it afterwards. On Telegram or WhatsApp, show_my_system_3d does the same for the system
you run: your fish count, your water, advanced through the weather that actually happened
since you last spoke to the bot, then forward through the forecast.
The badge always says which of the three you are looking at (as designed, today, or a forecast), because confusing them would be the worst thing this view could do. "Today" is the same state the bot calls "Now", so the picture and the conversation never describe different water. The panel says in as many words that the geometry is a proposed arrangement, never a survey of your site.
Speak instead of typing, on Telegram or WhatsApp. The transcript runs through a normal turn, so memory, tools and cited knowledge all apply.
Agronaut runs a consultation, not a one-shot Q&A. It identifies your goal (design a system, optimize a ratio, or troubleshoot a problem), asks for the few essentials that goal needs, then gives a first-cut recommendation tied to your system — and remembers it (a typed System Profile + episodic notes) across sessions.
You can also set the mode explicitly with /design, /optimize, or /troubleshoot —
the bot then jumps straight to gathering what that goal needs. All commands appear in
Telegram's / menu.
Agronaut also learns from outcomes: after suggesting a fix it can check back later ("did the water change fix the ammonia?"), and whatever worked is remembered and shapes its future advice.
Lessons can also become shared knowledge: a generalized, PII-stripped version of a verified
fix is nominated, the owner approves it in a local review CLI (python -m agronaut_agent.review),
and approved insights then help other operators — labeled as community experience, never as
verified science.
And it calibrates to reality: when you report real measured outcomes (harvest weight, FCR, crop yield), Agronaut tunes your future sizings toward your system — bounded to the published empirical ranges, so a measurement can only move a coefficient within what the literature allows, and every calibrated number is labeled.
The deterministic sizing model now covers five fish (tilapia, clarias, channel catfish, trout, common carp) and 30+ crops — leafy greens (lettuce, kale, chard, spinach, pak choi, arugula, watercress…), culinary herbs (basil, mint, cilantro, parsley, dill…), and fruiting crops (tomato, cucumber, pepper, strawberry, eggplant, zucchini…) — each with cited, calibratable seed coefficients placed within FAO 589's published feeding-rate band for its category.
Every result lists the coefficients it used (value + range + source: FAO 589, UVI/Rakocy, literature) and an explicit list of what it does not model (pH/alkalinity, micronutrients, salinity, solids, pests, cohort logic, per-crop ET). A confidently-wrong design can't masquerade as complete.
The same rule governs the advice layer. Citation is enforced in code, not asked for in a prompt: every retrieved passage is labelled with its source before the model ever sees it. And retrieval is allowed to say no — a question the corpus cannot answer returns "no matching passages" rather than the three closest paragraphs wearing source labels. Ask Agronaut the capital of Canada and it will decline, not cite an aquaponics paper at you.
Parametric, not machine-learned — buildable today from published equations:
- Feeding-rate ratio (FRR) sizes the system: grams of feed per m² of plant area/day.
- Nitrogen balance is an independent consistency check (feed → fish-retained → excreted → plants + solids + water-exchange + denitrification), flagging disagreement with FRR rather than silently reconciling — this guards against over-sizing the grow beds.
- Water balance (evapotranspiration + evaporation + sludge − rainfall) drives the water-budget feasibility check.
- Optimizer is bounded enumeration over a small species×crop palette (no heavyweight solver), with the even-split baseline inside the search space so it can never do worse.
Sizing is computed. Troubleshooting advice is retrieved, from a corpus of 22 hand-written operator guides plus openly licensed publications — currently 3941 chunks, led by Goddek et al. (2019) and FAO 589.
Retrieval is measured, not assumed. docs/dpg/retrieval_eval/golden_set.json holds queries in
real operator voice ("my tilapia are gasping at the surface", not "dissolved oxygen") plus
off-topic controls that must be refused:
python -m scripts.retrieval_eval # recall@k, precision@k, MRR, MAP@k + floor separation
python -m scripts.retrieval_sweep --all # re-pick floor / per-source cap / hybrid β
python -m scripts.corpus_report # what each declared source actually contributesNine techniques were implemented and measured. Four ship, four lose, one is available but
unused — and the losses are recorded in docs/dpg/retrieval_eval/techniques.json with the
conditions that would reverse them, which is how hybrid search went from rejected to shipped when
the corpus grew:
| ships | why | |
|---|---|---|
| Relevance floor | on (1.50) | refuses 8/10 off-topic queries, silences 0/33 real ones, keeps 0.117 headroom |
| Hybrid BM25 + RRF | on (β=0.90) | lost at 362 chunks, won at 1354, re-confirmed at 3941 |
| Per-source cap | on (1) | two books hold 97% of the corpus; at cap=2 they take 2 of 3 slots |
| PDF cleaning | on | drops contents pages; running header removed from 111 chunks → 4 |
| Metadata filtering | available, off | the third leg of hybrid search. Filters source_type, kb_tag, chapter, page, url_category on both pools before fusion. A capability, not a ranking change — no golden-set number moves, and none is claimed |
| Header chunking · context prefix · PDF chapter labels · cross-encoder rerank | off | each measured worse on this corpus |
Current: hit 0.879 · recall 0.833 · MAP 0.604 · 8/10 off-topic refused · 0/33 real silenced.
techniques.json was written 2026-08-25. The next day, commit 70b2d00 added a second book and
took the corpus from 1354 to 3935 chunks. Nothing was re-measured, and every constant silently
became wrong for the corpus that actually shipped:
| at 1354 (recorded) | at 3941, old constants | at 3941, re-tuned | |
|---|---|---|---|
| hit_rate | 0.939 | 0.818 | 0.879 |
| recall@k | 0.894 | 0.727 | 0.833 |
| MAP@k | 0.697 | 0.548 | 0.604 |
| off-topic refused | 8/10 | 4/10 | 8/10 |
Two constants moved, one did not. The floor tightened 1.65 → 1.50, because the distance bands separated as the corpus grew (on-topic worst 1.383, closest off-topic 1.411, where at 1354 chunks they overlapped and no floor could work). The cap tightened 2 → 1, because cap=2 was calibrated against one oversized source and there are now two. β stayed at 0.90 — it describes the relationship between two ranking signals, which is a property of the query language, not of how much text sits behind it.
Worth being precise about which change did what, because they pull opposite ways. The floor costs retrieval quality: at cap=1, staying at 1.65 would score hit 0.909 and MAP 0.624 against 1.50's 0.879 and 0.604. That is bought deliberately, to double off-topic refusal from 4/10 to 8/10. The cap is what pays for it: at floor 1.50, cap=1 gives 0.879/0.833/0.604 against cap=2's 0.788/0.697/0.538.
The floor was not tightened to 1.40, though that refuses all 10 controls: it clears the worst real query by 0.017, and this project had already rejected a 0.032 margin as too thin. 33 golden queries say nothing about the 34th; headroom is the only thing that does.
python -m scripts.retrieval_sweep --all # re-pick all three, with the evidence tableThat command exists because the drift was not carelessness. Re-measuring three constants was an afternoon of ad-hoc scripting, so it did not happen. It also earned its keep immediately: while this work was in review the corpus moved again (a 22nd knowledge file, 3935 → 3941 chunks) and re-running was one command rather than an afternoon. Run it after any corpus or embedding-model change.
"Run it after any corpus change" was still a rule someone had to remember, so CI now enforces it.
The baseline records a fingerprint of the corpus it measured (urls.txt plus the guides in
knowledge/), and test_corpus_fingerprint.py fails when the shipped corpus has moved away
from it: always when a source is added or removed, and once the guides have drifted more than
10% in size, so one contributed guide does not block its author. It cannot see a change to the
chunking code or to a page behind a URL.
Three of the four failures share one mechanism: they add topic words to chunks in a corpus where every document already shares a vocabulary domain, which dilutes rather than disambiguates. What worked was structural — refusing irrelevant passages, refusing error pages, refusing to let one source fill the whole answer.
Corpus licensing is mixed and deliberately explicit. The code is MIT; FAO 589 is
non-commercial-only. See docs/dpg/CORPUS.md — commercial users should drop
that entry from urls.txt and rebuild. Vet any source before adding it:
python -m scripts.corpus_report --candidate "<url>" --label "<expected topic>"It checks four things, because a source can fail in four ways: unreachable, empty, wrong subject (a guessed publication ID once resolved to "Sharks for the Aquarium" — 28k characters that pass every check except being about aquaponics), or not openly licensed.
Retrieval quality is measured offline against a golden set. Production behaviour is a different question, and needs a different instrument.
Every turn is one trace. All the events a turn produces — the message, each model call, each tool call, the retrieval, the turn summary — carry the same random per-turn id, so the log reads as a path rather than as counters:
agronaut traces # recent turns: which tools ran, what retrieval returned, where the ms went
agronaut analytics # p50/p95/max latency for turn / model / retrieval, tokens, thumbs up-downThe trace holds shape, never content. No prompt, no reply, no passage text, no query is
recorded, and that is enforced by an allowlist that drops unknown fields rather than by callers
remembering not to pass them (agronaut_agent/tests/test_turn_tracing.py asserts it). The trace
id is minted fresh per turn and is never derived from the user, so it groups a turn without
following anyone between turns.
What gets measured, and why those things. Turn latency and model latency separately, because
the course is blunt that the transformer is the bottleneck and this project previously timed only
retrieval — the fast, cheap stage. Token counts in and out, omitted entirely rather than recorded
as 0 when a provider reports no usage, so a quiet provider cannot drag every cost aggregate
toward zero. And a failed turn is still written, because dropping the turns that broke is how a
p95 comes to look healthier than the service is.
retrieval_eval scores whether the right documents were found. It cannot score whether the reply
used them, and a system can hit recall 0.894 while inventing every number in its answer.
AGRONAUT_FAITHFULNESS_EVAL=1 python -m scripts.faithfulness_evalThree metrics of three deliberately different kinds:
| judged by | what it catches | |
|---|---|---|
faithfulness |
an LLM, per atomic claim | claims the retrieved context does not support — the grounding measure |
response_relevancy |
an LLM + embeddings | an answer that is true but does not address the question |
citation_accuracy |
code, no model | [source: ...] labels that were never retrieved — a fabricated citation |
The judge is treated as a witness, not an oracle: rubrics are binary with named labels, an
unparseable verdict counts as unjudged rather than being folded into either side, and
n_unjudged is printed beside every score. It calls the network, so it is opt-in and never runs
in CI; the scoring arithmetic is pure and unit-tested without a model.
The judge is a different model from the one that writes the answer (AGRONAUT_JUDGE_PROVIDER /
AGRONAUT_JUDGE_MODEL), and it must copy the sentence of the context that backs a claim before
it may call the claim supported; code then checks that the quote really is in the context.
Sentences about the sources themselves ("the context does not specify…", "consult other
resources") are not claims and are left out.
Measured (2026-10-03): 33 golden-set answers written by Claude Sonnet 5, 479 claims.
| faithfulness | agreement with a person (kappa, 29 claims) | |
|---|---|---|
| gpt-oss-20b, quote-first prompt (default) | 0.84 | 0.24 blind, 0.39 after review |
| gpt-oss-20b, earlier prompt (two runs) | 0.90, 0.90 | 0.34, 0.30 |
| Claude Sonnet 5 as judge | 0.88 | 0.03 |
Citation accuracy is 1.00 (no fabricated sources) and response relevancy 0.44. Read the
faithfulness figure as a range, not a verdict: the judges agree with themselves (kappa 0.76
between two runs) much better than with the one person who has labelled claims so far, which is
"fair" agreement on a small sample. Claude was dropped as a judge because it agreed with the
person no better than chance. Report and labels: docs/dpg/faithfulness_eval/.
python -m scripts.label_claims 2026-09-30_baseline.json # label claims blind
python -m scripts.label_claims 2026-09-30_baseline.json --review # second look at disputes
python -m scripts.faithfulness_eval --agreement 2026-09-30_baseline.json
AGRONAUT_FAITHFULNESS_EVAL=1 python -m scripts.faithfulness_eval --rejudge 2026-09-30_baseline.jsonMost turns never retrieve anything: the model calls the sizing engine and explain
0 comments
log in to comment.