vibecodedetector.fanficnow.com
Paste a URL or some source code. Get a percentage, the evidence behind it, and an honest account of why you should not trust it too much.
Paste anything into the Auto field and it works out which of the four
readings it wants — an address, a GitHub repository, source, or a git log —
and says which way it went. Pick a mode by hand if you disagree.
No account, and nothing stored unless the operator has deliberately configured a database for the optional admin panel (see Privacy).
- Live page — fetches a URL and up to four of its own stylesheets and scripts,
then reads it the way you would with View Source open: builder fingerprints
first, then structure, then the look of the thing. When the page builds itself
in the browser it reads the bundle rather than the empty shell it serves — a
classNameis a string literal, and a minifier has no reason to touch the inside of a string — and when the bundle points at a source map it reads the original source, comments and all. The response headers are read too, for what somebody configured rather than for where it happens to be hosted. - Whole site — tick the box and it follows links from that page and from
its sitemap, reads as many as it can manage in about twenty seconds (up to
fifty), and compares them against each other. This is not the page check run
fifty times: a signal only counts site-wide when a quarter of the pages carry
it, and whether the pages resemble each other is itself evidence no single
page can give you. The sitemap earns its place twice: it lists pages the
navigation never links to, and its
lastmodcolumn is a dated record of a site being worked on that nobody thinks to fake. Honoursrobots.txtand finishes inside a shared-hosting request — on a slow site that means fewer pages, and the report says how many it managed. - Pasted code — reads a source file in any language for the tells that survive
in text: comment habits, error handling, naming, dependency incoherence, the
security profile. Stylesheets get their own reading — declaration order, rules
crushed onto one line, a label on every block and a reason for nothing — and
a page's own CSS files and
<style>blocks go through it too. - GitHub repository — give it
owner/nameand it reads three things at once: the commit history through the same checks the git tab uses, the file tree, and a few of the largest source files in full. The tree is the part nothing else can see — an assistant's own configuration file committed with the code, a pile ofIMPLEMENTATION_SUMMARY.md-shaped session reports, a.envin version control, thirty source files and no tests. Public repositories only, and it reads a sample: the newest and oldest hundred commits and three files, which the report says out loud rather than implying it read the codebase. - Git history — paste the output of
git logand it reads how the code arrived: one enormous opening commit, hundreds of lines in minutes, a trail of one-line fixes behind it. This is the strongest evidence the tool has, and the hardest to fake after the fact. Use this when the repository is not public; when it is, the repository mode does the same reading without the pasting.
Either way you get a score out of 100, a verdict, a confidence level, and every signal that fired with the excerpt that triggered it — shown in the code it was found in. Each excerpt carries the lines above and below the match, the file it came from, and how many times that pattern was found in the subject; the count is not decoration, it feeds the score. You can then download a signed one-page PDF certificate, which anyone can check at /verify.
It does not prove anything, and the interface says so before it says anything else.
Peer-reviewed benchmarks put off-the-shelf detection of AI-generated source at around chance. Pan et al., Assessing AI Detectors in Identifying AI-Generated Code (ICSE-SEET 2024), measured five detectors at roughly 0.5 accuracy across 5,069 human-written Python solutions. Linters and formatters normalise away the statistical signal these tools depend on, masking is cheap, and every stylistic tell weakens as models improve and as human developers adopt the same frameworks.
So this tool is built to be read, not quoted. The number is the least interesting thing on the results page.
Do not use it to accuse anyone of anything. False-positive harm in academic and hiring contexts is real and falls hardest on people who did nothing wrong.
Every signal carries a weight in log-odds. Scoring starts from a prior of −1.0 — the assumption that a given subject is not generated — sums the weights of what was found, and pushes the total through a logistic curve.
Evidence is ranked, because evidence is not equal:
| Tier | Example | Weight |
|---|---|---|
| Platform fingerprint | cdn.gpteng.co, a lovable-tagger marker, a builder's generator meta tag |
4.5 — decisive |
| Repository history | a big-bang first commit, 600 lines in four minutes, a run of "fix typo" | 0.6–1.4 |
| Repository contents | a committed CLAUDE.md, a pile of session summaries, a .env in git, nothing tested |
0.5–1.5 |
| Site-wide | every page one template with the words swapped; or pages from visibly different eras | 0.7–1.2 |
| Structural | a scaffold nobody renamed, uniform comment density, the same problem solved several ways | 0.35–1.1 |
| Code style | what-not-why comments, swallowed exceptions, tests that assert nothing | 0.4–1.4 |
| Content & security | generic testimonials, a key shipped to the browser, auth decided in localStorage | 0.35–0.9 |
| Aesthetic | indigo gradients, neon cyan and magenta, a gradient for a background, Inter, three identical cards | 0.25–0.5, capped as a group |
| Human authorship | ticket references, exasperated comments, commented-out code, mixed indentation | subtracts |
How often a tell fired is part of the reading. One what-not-why comment is a coincidence; thirteen of them across a file is a habit, and habits are what this is reading. Every signal carries an occurrence count, and that count earns it a multiplier of up to 1.5× on a log scale — bounded, because the difference between one occurrence and ten is real and the difference between forty and eighty is the length of the file. Fingerprints are exempt: finding a builder's badge three times identifies it exactly as well as finding it once.
Eight rules the scoring will not break:
- Aesthetic evidence is capped. A subject with nothing but aesthetic tells cannot exceed 55%, however purple it is.
- Every category has a ceiling. Signals within a category are not independent — code that swallows its exceptions usually over-wraps them too — so tripping eight weak style tells must never outweigh one hard fingerprint. Fingerprints are the only category with no ceiling.
- Inference stops short of identification. Without a fingerprint no reading passes 92%, however many families of evidence agree. The category ceilings stop one family running away; this stops several near their ceilings from adding up to the number a builder naming itself would produce.
- No reading reaches 0% or 100%. The scale is clamped to 3–97.
- Thin input is pulled toward the middle and reports insufficient confidence, rather than being quietly guessed at. Thin means nothing to read — a page that serves 300 bytes of markup and half a megabyte of readable bundle is not thin.
- Confidence never exceeds moderate without a platform fingerprint — pattern reading without repository history has not earned more than that.
- Human signals are first-class and weighted on the same scale as the rest.
- Repetition is bounded and never manufactures a category. The multiplier is capped at 1.5×, applies inside the category ceilings rather than around them, and does not apply to fingerprints at all.
All 130 signals, with their weights and reasoning, are on the site itself at
/catalogue and in
docs/SIGNALS.md — both rendered from lib/Catalog.php, so
neither can drift from the code.
Repository history. One enormous opening commit followed by a trail of "fix typo"
commits is far harder to fake retroactively than anything in a served page or a
pasted file — which is why two of the tabs exist. It cannot be reached from a URL.
If the project is public on GitHub, the repository tab fetches it for you; if it is
not, you need the repository in hand and a pasted git log. Either way, start there.
Even then it reads the shape of the work, not who did it: a developer who commits carefully while an agent writes the code produces a history that looks entirely human, because in every respect git records, it is. Where trust is not adversarial, ask the developer.
Roughly half and half, and it is worth being precise about which half.
The AI half wrote the code: the detection engine, the website, the PDF writer, the logo, the test suite. The human half decided things: the research the signal catalogue is built on, which tells are worth trusting and what each one is worth, the calibration, the design direction, and the bug reports that fixed what the machine got wrong.
So: a half-vibecoded app for detecting vibecoded apps. That is on the front page of the site, not buried here, because it is the most useful thing this project has to say about its own reliability.
Point the detector at this site and it comes back in the AI-leaning band. The exact figure is deliberately not written down here or on the site: it has read 55%, then 73%, then lower again after the redesign, and a number typed into a paragraph is stale the week after it is typed. Run it yourself — that is what the front page is for, and a claim you can check in ten seconds beats one you have to take on trust.
It moves in both directions and neither is an adjustment made in the project's favour: up when the detector got better at reading JavaScript and started catching its own, down when the front page stopped being four screens of prose under a text box. No signal that fires on this site has ever been exempted.
- Nothing to fingerprint. Agentic editors write into an ordinary repository. No badge, no builder subdomain, no injected runtime. Signs run in one direction only, and this repo is the direction they do not run in.
- The tells were avoided on purpose. No what-comments, no docblock on every trivial function, no indigo gradient, no Inter, no frosted glass, no neon, no three-card grid — and, since the redesign, no gradient anywhere on the site at all. That is masking, and masking is cheap. It took no particular effort, and it is the reason the aesthetic family scores this site at zero: not restraint, just knowing what the list says.
What fires is small and fair: formal error messages, heavy em-dash use, vocabulary in the front-end script that names nothing in particular. Every one is genuinely present. The reading is not flattering and it has not been adjusted, because a detector that quietly exempts the site it runs on is worth nothing at all.
None of it is special-cased away, and a detector that exempted itself would be worth less than one that takes the hit.
git clone https://github.com/goldo9824/vibecode-detector.git
cd vibecode-detector
php -S localhost:8000Then open http://localhost:8000. There is nothing to install first.
php tests/run.php # the whole suite, ~780 assertions
php tools/gen-signals-doc.php # regenerate docs/SIGNALS.md
php tools/build-assets.php # regenerate the SVG files from lib/Brand.php
php tools/build-social.php # regenerate the 1280x640 social preview cardBuilt for LWS shared hosting: upload the folder over FTP and it runs. No Composer, no npm, no build step, no database, no cron.
The PDF certificates are generated by a hand-written PDF 1.4 writer
(lib/Pdf.php) using the standard-14 fonts, precisely so that there is nothing to
install on the host.
Full instructions, including the two 403 checks to run afterwards, are in docs/DEPLOY-LWS.md.
api/website.php gives a caller with an API key programmatic access to the
Live page / Whole site check, with a much higher rate limit than the
anonymous UI. There is no key by default; the operator sets one by hand in
data/api-keys.txt, which never goes in the repo — or, with a database
configured, creates named, individually revocable keys through admin/
instead. See docs/API.md for setup, and hand
llms.txt to anyone you give a key to — it's written for an
AI agent to read and call the endpoint correctly on its own.
An optional, password-gated dashboard at /admin/ for creating and revoking
named API keys and seeing usage — total analyses by mode and per day, the
most-analysed websites, traffic, and the same broken down per key. Every website
that has ever been analysed has its own searchable, sortable list forty rows to
a page, and its own page of charts.
Two of its five pages answer questions nothing else can:
- GitHub — every repository that has been searched, every time GitHub refused, and how many repositories an hour gets through before it does. GitHub allows 60 API requests an hour per address and a repository read spends up to eight of them, so that ceiling is what decides how many people can use the repository tab at once.
- Reports — what people say when a reading looks wrong. Every result on the front page carries a Does this reading look wrong? block: too high, too low, or about right. The panel shows where on the scale the disagreement sits, and how often a reading is disputed per hundred analyses.
Inert with no database configured; see docs/ADMIN.md.
How the public pages describe themselves to a search engine — titles, canonicals, Open Graph, structured data and the sitemap — is in docs/SEARCH.md.
index.php the analyser, and nothing else
method.php how the number is arrived at, where it is wrong, who wrote this
signs.php the visual field guide, every specimen rendered live
catalogue.php every signal, its weight and its reasoning, on the site
verify.php certificate verification
llms.txt instructions for an AI agent calling api/website.php
robots.txt disallows /admin/, points at the sitemap
sitemap.php served as /sitemap.xml, built for whatever domain it runs on
api/ analyze.php, website.php, certificate.php, feedback.php
admin/ optional password-gated key management and usage dashboard
lib/
Catalog.php every signal, its weight and its reasoning — the source of truth
Subject.php works out whether a paste is an address, a repo, source or a log
Report.php scoring, verdict bands, guard rails
SiteAnalyzer.php live-page analysis
CodeAnalyzer.php source analysis
GitAnalyzer.php repository-history analysis
RepoAnalyzer.php public GitHub repositories: history, tree and source together
GitHub.php the few read-only GitHub endpoints, and the request budget
Crawler.php polite same-origin crawl, robots.txt and a time budget
SiteSurvey.php multi-page aggregation and cross-page comparison
Fetcher.php HTTP with SSRF protection
Db.php optional MySQL connection, used only by admin/
ApiKeys.php API key CRUD, backed by Db.php
UsageLog.php usage recording and stats, backed by Db.php
VisitLog.php page views, counted without identifying anyone
GitHubLog.php what the GitHub allowance is spent on, and when it runs out
Feedback.php readings people reported as wrong, and what they said
AdminAuth.php admin session, login and CSRF
AdminUi.php how the panel words dates, modes and page numbers
Num.php counters past a thousand: 1.2k, 17k, 1.5M, and small ones in words
Seo.php title, canonical, Open Graph and structured data
Pager.php page-number arithmetic for a long list
Chart.php charts, drawn as SVG on the server
Pdf.php a small PDF 1.4 writer
Brand.php the mark, as geometry
Certificate.php certificate layout
assets/ css, js, svg
specimens/ one file per visual sign, rendered inside signs.php
tests/ fixtures and the runner
tools/ doc and asset generators
No account, no cookies, no third-party analytics, nothing loaded from anyone else's server. Pasted code and git history are read once in memory and discarded, never written to disk — that's true with or without anything below. Certificates are signed rather than stored, which is why verification is a signature check and not a lookup — there is nothing to look up.
The one opt-in exception: an operator can configure a database for admin/
(see docs/ADMIN.md) to manage named API keys and to see how the
site is used. With no database configured, that code path is inert and the site
keeps nothing at all — which is the default for anyone who forks this and does not
set one up. With one configured, four things are recorded:
- Analyses — the mode, and for URL and repository checks the address or repository name analysed. Never the pasted content, never the fetched page, never the repository's source, never anything about who is asking.
- Visits — one row per page view: the path, a timestamp, the referring site's host, and a coarse client class. No address, no cookie, no session, and never the query string, which is where the URL somebody asked about would be.
- GitHub requests — one row per request this site makes to GitHub's API: the repository, which endpoint, the status, and what GitHub said was left of the hourly allowance. Nothing about who asked for it.
- Reported readings — when somebody says a reading looks wrong: the reading itself, which way they say it is wrong, and their note. Nothing about who they are, and no field to leave an email in — there is nowhere for an answer to go.
Counting people rather than page views needs something per visitor, and the
something is a token: an HMAC of address and user agent under this installation's
secret and today's date. It cannot be reversed into an address, and because the
date is in the salt it cannot recognise the same person tomorrow. Counting who came
today is the whole of what it can do. Rows are deleted after 90 days, and
'log_visits' => false in the database config turns the whole thing off while
leaving the API-key management working.
New signals are welcome, especially ones that are hard to mask. A signal has to justify its weight and come with a fixture; see CONTRIBUTING.md.
If the tool got something wrong, that is the most useful bug there is:
The detection reference this is built on is summarised in docs/REFERENCE.md, with sources.
The indigo problem has a named origin: on 7 August 2025 Tailwind co-creator Adam
Wathan apologised for making every button in Tailwind UI bg-indigo-500 five years
earlier, "leading to every AI generated UI on earth also being indigo".
Built by Landfall studio.
MIT. See LICENSE.
A Landfall studio product
0 comments
log in to comment.