SlopScore
00 crowd

polinko

Human-led AI evaluation research system for observing model behaviour through binary evals, signal traces, and retained failures.
Open repo on GitHub Open the demogithub.com/tryskian/polinko
Python · ★ 1 · 0 forks · Apache-2.0 · paperwork by the Cap'mmostly ai (inferred)light human (inferred)works-on-my-machine (inferred)other
listed 53 minutes ago by tryskian · last checked 53 minutes ago
The owner didn't write this. This repo never submitted itself. The Cap'm found it on a truffle trawl and wrote its paperwork from what GitHub already shows. Picked by hand by the Cap'm on 2026-09-24: Human-led AI evaluation research system for observing model behaviour through binary evals, signal traces, and; its own README says "Built with Codex and GPT-5". 1 stars; Apache-2.0 license. The owner did not submit this. Votes count; awards don't until the owner claims it.

I'm not calling your project slop! Geeze, it's a joke... Do you own this repo?

Log in with GitHub as tryskian. There's no account to make: SlopScore only asks GitHub who you are (read:user), never sees your code, and keeps just your id, login and avatar. Then you can:

  • Keep it, on your terms. Commit your own slopscore.md (spec) and press Refresh. Your paperwork replaces the Cap'm's, and you can submit it for Slop of the Day.
  • Take it down. One click on Remove. It stays gone; the trawl never brings it back.

Log in with GitHub

Can't log in as the owner? Request a takedown. No login needed, and a trawled listing comes down right away.

GitHub says
Human-led AI evaluation research system for observing model behaviour through binary evals, signal traces, and retained failures.
website
https://www.krystian.io/
topics
ai-evaluationai-reliabilityai-researchai-safetybinary-evalsevalshuman-ai-alignmenthuman-ai-collaborationhuman-ai-interactionhuman-in-the-loopllm-evaluationmodel-evaluationopenaiopenai-apipythonresearch-modelssqlite
created
2026-02-19 · pushed 2 days ago · 1449 commits · 2 contributors
languages
Python 91%Shell 3%Makefile 3%HTML 2%JavaScript 1%CSS 1%
paperwork
pull request templatelicensereadme 71% health
dependencies
✓ 564 deps, none with known advisories · OSV.dev, checked 53 minutes ago

The Cap'm's log

The Cap'm wrote this paperwork, not the owner. This repo never submitted itself to SlopScore. The Cap'm picked it by hand: Human-led AI evaluation research system for observing model behaviour through binary evals, signal traces, and; its own README says "Built with Codex and GPT-5". It carries the Apache-2.0 license. The disclosures above are his best guess from what GitHub shows.

Is this yours? Commit a real slopscore.md and press Refresh to replace this, or remove the listing in one click. There's no account to make: you log in with GitHub.

README — the repo's own words, folded up so the grading fits on one screen

Polinko

CI Polinko Model Eval Contract Research Surface Model Refactor

Note

Current status: Polinko is being staged for the next beta.

The research model is stable enough to keep building on, but the repo is in an active refactor window. The model contract, evidence snapshots, docs, helper scripts, and local tooling are being tightened so the next beta has a cleaner surface to build from.

Refactor map: method · journey.

Polinko is a human-led research system for observing AI behaviour through binary eval gates, signal traces, retained failures, and repo-native evidence.

The binary gate is the method's starting point because it aligns evaluation with binary computation. Before interpretation begins, Polinko keeps the first judgement simple: pass or fail.

krystian.io is the website doorway. This repository is the research surface.

The static website source lives in site/. Netlify builds it with npm run build and publishes dist/.

Built with Codex and GPT-5.6

Important

Polinko is built entirely with Codex. The project remains human-led: Codex is the implementation and co-reasoning surface used to turn the research method into code, eval infrastructure, documentation, and operator workflows.

GPT-5.6 was used throughout this Build Week iteration to:

  • help build make build-week-demo, the repo-native recording script that runs the preflight, live OCR binary eval, retained-evidence checks, and cleanup shown in the demo;
  • support daily repo audits and reports while Polinko is refactored and staged for the next beta;
  • run the recorded workflow live inside the Codex CLI; and
  • help edit, clean, frame, and transcribe the final submission video.

Research Question

How can binary eval gates help separate coherent AI output from reliable AI behaviour while keeping evidence boundaries visible?

Working Theory

AI responses are shaped by more than the prompt. Policy, guardrails, retrieval, memory, context limits, tooling, and prior response residue can all bend the path from intent to output.

Polinko treats visible mismatch as evidence. The method preserves failures, classifies them, and uses them to update the next research boundary instead of smoothing them away.

The first gate stays binary so the operational result remains visible before summary, interpretation, or pattern language can blur the boundary. Richer analysis begins only after the source evidence and retained failures are still traceable.

OCR is one pressure lane because the expected answer is externally checkable. It is one part of the broader research model.

Current Position

  • Beta 2.3 is the frozen method snapshot.
  • pre-Beta 2.4 is staged as the next research-model contract.
  • OCR is the mature green lane and is moving into generalisation pressure.
  • Co-reasoning is the first promoted non-OCR lane.
  • Retrieval, response behaviour, uncertainty boundary, and hallucination boundary are operationalised support surfaces.
  • Operator burden is the active thin lane.
  • Lean binary evaluation is staged as a sustainability question: whether reduced context clutter and correction churn can reduce unnecessary inference, with energy use, cooling demand, and water impact as downstream resource implications.
  • Source-first row and case evidence is the pre-Beta 2.4 method foundation.

Refactor Map

The current refactor is being handled as a staged path. The diagrams show the working loop and the route each major surface has taken through the cleanup.

  • Refactor method: alignment, kernel scope, validation, docs, PR, merge, and clean main.
  • Refactor journey: evidence baseline, runtime/package movement, manual eval workbench, and docs closeout.

Read Next

Surface Use
Field notes shortest reading path
Research surface current notes, beta evidence, hypotheses
Eval evidence tracked eval snapshots
Refactor diagrams method and journey maps
Runbook operator procedure
Architecture system shape
Decisions durable rationale

Run Locally

make deps-install
cp .env.example .env
# set OPENAI_API_KEY in .env
make doctor-env
make docs

Use make docs-open only when you want to launch the system browser.


License

Apache-2.0. See license.

Read the rest on GitHub

Scan report · 2026-09-24
  • Prohibited terms or links
  • Repository eligibility
  • slopscore.md paperwork
  • Content policy
  • Risk review

From the balcony · 0 of 4 clapped

    Cap'm Slop, Princess, Crusoe and Schnitzel read it and passed. Their reasons are on the balcony, with every other verdict.

    Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.

    0 comments

    log in to comment.

    report this listinglog in to report