SlopScore
10 crowdincl. 2 critics

trochilus

Local MoE inference in C, token-exact against transformers on any CPU: it configures itself from the machine it runs on and never swaps.
Open repo on GitHubgithub.com/namespaceMarcello/trochilus
C · ★ 1 · 0 forks · Apache-2.0 · paperwork by the Cap'mmostly ai (inferred)light human (inferred)works-on-my-machine (inferred)other
listed 1 hour ago by namespaceMarcello · last checked just now
The owner didn't write this. This repo never submitted itself. The Cap'm found it on a truffle trawl and wrote its paperwork from what GitHub already shows. Picked by hand by the Cap'm on 2026-09-28: Local MoE inference in C, token-exact against transformers on any CPU: it configures itself from the machine i; its own README says "How it is built Trochilus is written by Claude Code (Anthropic's Opus) as the coding agent, with Marcello Costagliola leading — the directio". 1 stars; Apache-2.0 license. The owner did not submit this. Votes count; awards don't until the owner claims it.

I'm not calling your project slop! Geeze, it's a joke... Do you own this repo?

Log in with GitHub as namespaceMarcello. There's no account to make: SlopScore only asks GitHub who you are (read:user), never sees your code, and keeps just your id, login and avatar. Then you can:

  • Keep it, on your terms. Commit your own slopscore.md (spec) and press Refresh. Your paperwork replaces the Cap'm's, and you can submit it for Slop of the Day.
  • Take it down. One click on Remove. It stays gone; the trawl never brings it back.

Log in with GitHub

Can't log in as the owner? Request a takedown. No login needed, and a trawled listing comes down right away.

GitHub says
Local MoE inference in C, token-exact against transformers on any CPU: it configures itself from the machine it runs on and never swaps.
topics
avx2avx512ccpu-inferencedeterministicggufinference-enginelarge-language-modelsllmllm-inferencelocal-inferencelocal-llmmixture-of-expertsmoeolmoesimdzero-dependency
created
2026-09-17 · pushed 3 hours ago · 61 commits · 1 contributor
languages
C 73%Python 16%Shell 10%Makefile 1%PowerShell 0%JavaScript 0%
paperwork
licensereadme 42% health
dependencies
no dependency graph (no manifest, or disabled) · OSV.dev, checked 1 hour ago

The Cap'm's log

The Cap'm wrote this paperwork, not the owner. This repo never submitted itself to SlopScore. The Cap'm picked it by hand: Local MoE inference in C, token-exact against transformers on any CPU: it configures itself from the machine i; its own README says "How it is built Trochilus is written by Claude Code (Anthropic's Opus) as the coding agent, with Marcello Costagliola leading — the directio". It carries the Apache-2.0 license. The disclosures above are his best guess from what GitHub shows.

Is this yours? Commit a real slopscore.md and press Refresh to replace this, or remove the listing in one click. There's no account to make: you log in with GitHub.

README — the repo's own words, folded up so the grading fits on one screen

Trochilus

Trochilus

Exact Mixture-of-Experts inference in C, on the computer you already own.

Trochilus runs Mixture-of-Experts language models with no dependencies: libc and the operating system's threads, nothing else. One binary picks its kernels at runtime for whatever CPU it lands on, reads the experts from disk when they do not fit in RAM, and puts part of the work on an NVIDIA GPU when one is there — without a CUDA toolkit. And it gives the same tokens as transformers, checked in every build.

Speed may change; the result may not. Threads, SIMD tier, batch size, expert budget, speculative decoding, Windows or Linux, CPU or GPU: each one changes how fast the answer comes, never a bit of it. Every one of them is a test in the gate.

What it does today

  • Runs OLMoE-1B-7B and any GGUF v3 file of its family, straight from Hugging Face: F32, F16, Q8_0, Q4_K and Q6_K weights (so the Q4_K_M files people download).
  • CPU kernels for scalar, AVX2, AVX-512 and AVX-512 with VNNI and VBMI, chosen at runtime, each bit-identical to the scalar definition, with a test that the tier you think ran is the one that ran.
  • The decode's attention on an NVIDIA GPU, loaded at runtime (hand-written PTX through the driver) with the CPU's bits: the logits of 600 positions and the tokens after a 4000-token prompt are the CPU's byte for byte. Without a GPU nothing changes.
  • Experts from disk under a RAM budget: a slot store with an O(1) LRU, unbuffered reads, the prompt read layer by layer, the next layer's experts read while this one computes.
  • Speculative decoding (--spec): drafts taken from the text already in the context, several tokens verified in one pass that loads each weight row from memory once for all of them (a pass of three tokens costs 1.7 times a pass of one).
  • run, chat (the model's own chat template, byte for byte apply_chat_template), generate, and serve, which keeps the model loaded between commands.
  • A byte-level BPE tokenizer built from the GGUF metadata, token for token Hugging Face tokenizers, at 22.0 MB/s.

Results

OLMoE-1B-7B, the same GGUF for both engines, both in the same Linux container on one laptop (Ryzen 9 7940HX, 16 cores, 31 GB), a free machine, the night of 2026-09-26 to 27. Tokens per second, the median of 10 runs: two series of 5 for each engine, alternated with the other's.

llama.cpp Trochilus Trochilus / llama.cpp
Q4_K, prompt of 512 tokens, 4 threads 220.6 220.0 1.00×
Q4_K, generation after that prompt, 4 threads 50.9 52.4 1.03×
Q8_0, prompt of 512 tokens, 16 threads 370.7 492.9 1.33×
Q8_0, prompt of 2048 tokens, 16 threads 341.3 485.2 1.42×
Q8_0, generation at 512 tokens of context, 8 threads 33.1 33.1 1.00×
Q8_0, generation at 2048 tokens of context, 8 threads 27.2 26.3 0.96×

The difference is in what is computed: llama.cpp rounds the activations to 8 bits before each matrix product and keeps its attention cache in 16 bits; Trochilus computes what transformers computes, the activations in float (for Q4_K, exact integers) and the cache in 32 bits. Every number, the conditions, and the optimizations that did not pay are in docs/MEASUREMENTS.md.

Under a RAM budget the model is not read twice: with a partial expert budget a 2048-token prompt runs 1.89x faster reading 6 273 MiB instead of 22 880, and 1.22x more with the disk reading the next layer while the cores compute this one.

How it keeps up, bit for bit

Every tier must give the scalar definition's bits, so the speed has to come from how the work is arranged, not from rounding. Four ideas do it, each checked in the gate:

Sixteen rows in sixteen lanes. The scalar definition adds a row's products in sixteen interleaved chains and joins them with a fixed tree. Rather than spread one row across a register's sixteen lanes, which must then be joined at the end of every row, Trochilus gives each lane its own weight row: every chain sums in the scalar order, the input is broadcast to sixteen rows at once, and nothing is decoded, scaled or joined inside the loop. The prompt's matrix-product tile runs at 164.7 GFLOP/s on one core, 98.6% of what the CPU can do in float without fused multiply-adds (which would round once where the scalar definition rounds twice), and every output is the scalar definition's, bit for bit.

Q4_K in exact integers, each input converted once. A Q4_K matrix product takes each block of 256 inputs as 32-bit fixed point and sums exact integers, then rounds once: the input's 32 bits are the only approximation, where a float dot product rounds at every step. The order of the sums no longer matters, so any thread count or SIMD width gives the same bits, as fast as the float kernels this replaced. Each input row is converted once and shared: q, k and v read the same converted rows, gate and up read a token's row through a map instead of eight gathered copies, and the SwiGLU converts the down projection's input while it is still in the core's cache.

A byte that can be deduced is not read. In OLMoE's GGUF files the router matrices are 32-bit floats whose low 16 bits are all zero: the model was converted from bf16. The load checks every value, keeps only the top halves, and each row widens them back with a shift: the 32-bit row's bits, from half the bytes, every token. A matrix with one value that does not fit stays in 32 bits.

A correctly rounded exp, proved on every float. Our expf is a 64-entry table and a polynomial, with the eight hard cases computed at 200 bits; the gate checks it on all 4 278 190 082 float arguments that are not NaN, against the exact value, in every SIMD tier. (glibc's rounds 0.004% of them otherwise; ggml's 3.36%, by up to 2 units in the last place.)

What makes it different

Exactness is an invariant, not a hope. Windows and Linux give the same logits byte for byte on the real model; tiny models built for the purpose are compared with transformers in every gate, and so is the real model cut to two layers. Speculative decoding cannot change the answer: every verified row is bit for bit the computation of a single-token pass.

Nothing to configure. At startup the engine measures the cores, the instructions, the free RAM and the disk, and places the work itself; while it runs it times how many threads each kind of pass wants (the decode, and each size of verify pass on its own). It refuses to load if it would leave the machine with less than 2 GB or 10% of its RAM.

Get started

git clone https://github.com/namespaceMarcello/trochilus.git
cd trochilus
make                                   # build/trochilus: no dependencies
build/trochilus cpu                    # what the engine sees of this machine
build/trochilus run  -m OLMoE-1B-7B-0125-Instruct-Q4_K_M.gguf -f prompt.txt -n 200
build/trochilus chat -m OLMoE-1B-7B-0125-Instruct-Q4_K_M.gguf

C11 with intrinsics and a Makefile: gcc, clang or MinGW-w64, on Linux and Windows. make check is the gate (lint, a 0-warning build, the tests, ASan, TSan, the oracles). The Python tools under tools/ are for conversion and for the oracles only; the engine never needs them.

Status

Pre-alpha, under active development, measured on one machine.

Milestone Content State
M0 the exact engine: GGUF, CPU kernels for every tier, the OLMoE graph, tokenizer, chat done
M1 experts from disk under a RAM budget, with no options to set the store, the budget and the overlapped reads done; the first prompt's cost next
M2 smaller weights: Q4_K, Q6_K, then Q2_K and IQ2 Q4_K and Q6_K done, Q4_K_M runs
M3 CUDA: the attention, the dense weights, the experts in VRAM the decode's attention done
M4 a model of hundreds of gigabytes on the same laptop —
M5 KV checkpoints on disk, an HTTP server, a draft model speculation from the context, serve
M6 Vulkan and Metal, so the GPU is not only NVIDIA —

Not there yet: one model family and one chat template; the GPU does only the decode's attention, on NVIDIA only; no 2-bit formats; no NEON kernels (ARM takes the scalar path); greedy decoding only, and no HTTP API; a model larger than RAM has not been run yet.

How it is built

Trochilus is written by Claude Code (Anthropic's Opus) as the coding agent, with Marcello Costagliola leading — the direction, the questions, the reviews and every decision; every commit says so. One piece of the engine at a time: first read how colibri, ds4 and llama.cpp solve it, measure theirs against ours, then build and measure again. A comparison alternates its two sides run by run and carries an A/A control, the prediction is written before the run, and every mistake becomes an automatic check (a test, a lint rule, a mutation that must turn red). The engineering log is in docs/:

Document Content
ARCHITECTURE.md principles, layers, the correctness ladder, milestones
STATUS.md where the project stands, the decisions, the next step
MEASUREMENTS.md every measurement, including the optimizations that were rejected
LESSONS.md every mistake and discovery, with the check that now prevents it
ORIGINS.md where each idea and file comes from; every piece against the three references
COMMANDS.md benchmarks, measurements, mutations, reports

Acknowledgements

Trochilus is written from scratch, on what four other engines taught. They are pinned at a commit in ref/ and read as primary sources, and every idea taken from one of them is recorded in docs/ORIGINS.md with the file and function it came from.

Project Some of what we learned from it
colibri — Apache-2.0 threads on physical cores; weights read with pread; the pretokenizer regex replayed over codepoints; oracles on tiny generated models; drafts from the prompt's own text; one expert in one slot
ds4 — MIT a persistent thread pool instead of OpenMP; the vocabulary from GGUF metadata; (token, expert) pairs sorted by a counting sort; one weight row against several tokens in registers; the GGUF type table
llama.cpp / ggml — MIT prompts in passes of 512 tokens; the K-quant block layouts; pinning a thread to a processor; undoing a rejected draft; the engine we race
ik_llama.cpp — MIT its K-quant CPU kernels, read before we wrote ours

transformers and Hugging Face tokenizers define what a correct result is here, and the first model is OLMoE-1B-7B, from AI2. Only two files carry code from elsewhere — the GGUF type table in src/format/gguf.c and tools/make_tiny_olmoe.py — and each names its origin in its header.

License

Apache-2.0 — see LICENSE, and NOTICE for third-party material.

Read the rest on GitHub

Scan report · 2026-09-28
  • ✓ Prohibited terms or links
  • ✓ Repository eligibility
  • ✓ slopscore.md paperwork
  • ✓ Content policy
  • ✓ Risk review

From the balcony · 2 of 4 clapped

  1. Princessclapped
    Clear working implementation with Apache-2.0 license, declared status 'works-on-my-machine', token-exact verification in CI, and concrete technical details about CPU/GPU support and model compatibilit
  2. Crusoeclapped
    No vulnerable dependencies, local-only inference with no telemetry or credential requirements, and clear technical story about data handling.

Cap'm Slop and Schnitzel read it and passed. Their reasons are on the balcony, with every other verdict.

Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.

0 comments

log in to comment.

report this listing — log in to report