Phase 1 was a rule-based phishing detector: keyword, URL, IP, domain, attachment and sender heuristics, using only the standard library. Phase 2 adds a machine-learning classifier on top of it. It trains the classifier on public datasets and measures it honestly, including on real phishing from sources the model never saw during training.
The shipped model was tested as-is, with no re-tuning, on three hold-out sets. None of their sources appear in training:
| Set | Phishing | Legitimate |
|---|---|---|
| A | 454 real phishing from 2025 (Nazario). Only 4 are near-copies of training emails. | 1,478 from 8 Apache mailing lists, 2025 |
| B | the same 454 | 917 from Ubuntu mailing lists, 2025 |
| C | 214 English phishing from Phishing Pot, a honeypot collection (2026) | 767 from OSGeo mailing lists, 2025 |
Brackets show 95% confidence intervals.
| Set | Recall | False alarms | F1 | ROC-AUC | |
|---|---|---|---|---|---|
| Phase 1: rule engine | A | 30.0% (25.9–34.3) | 14.1% (12.4–15.9) | 0.34 | 0.64 |
| B | 30.0% | 14.0% (11.9–16.4) | 0.38 | 0.61 | |
| C | 27.1% (21.6–33.4) | 31.3% (28.1–34.7) | 0.23 | 0.46 | |
| First ML version (v1) | A | 97.6% | 8.6% | 0.87 | 0.98 |
| B | 97.6% | 24.9% | 0.79 | 0.94 | |
| Current model | A | 99.1% (97.8–99.7) | 1.5% (1.0–2.2) | 0.97 | 0.998 |
| B | 99.1% (97.8–99.7) | 1.4% (0.8–2.4) | 0.98 | 0.998 | |
| C | 93.0% (88.8–95.7) | 0.4% (0.1–1.1) | 0.96 | 0.998 |
- Phishing caught: 30% (rules) → 99% (ML) on 2025 Nazario phishing, and 93% on Phishing Pot.
- False alarms: 8.6–24.9% (v1) → 0.4–1.5%.
- Stable across retraining: over 5 training seeds, recall on 2025 Nazario phishing stays within 98.9–99.1% and false alarms on set A within 1.6–2.3% (see below).
- What it still gets wrong: the remaining false alarms are mostly one-line "Unsubscribe" emails that users send to a list by mistake, plus a few project announcements. The missed phishing is mostly short, link-free lures and phishing disguised as a product newsletter. Some Phishing Pot samples are plain marketing spam, so its labels are partly noisy.
- But not on commercial mail: on the author's own promotional emails, 88% of English promotions are flagged as phishing (see below). Every legitimate test set above is a mailing list, so these false-alarm rates hold for that kind of mail only.
- The lockbox test confirms it (see below): 96.9% recall at 0.5% false alarms on emails and organizations never used during development.
These are the CLI's own decisions (see How it works). Details by source and example errors: reports/HOLDOUT_2025.md.
By the end of development every hold-out had been looked at at least once. So one more set was put aside. Its rules (protocol) were written down before it was used. Then the finished model was evaluated on it exactly once, and nothing changed afterwards.
- Phishing: 457 English emails from 1,000 random older Phishing Pot samples. Hold-out C used only the newest 800, so none of these were ever loaded during development.
- Legitimate: 562 emails from the 2026 mailing lists of three organizations used nowhere else: ISC (
bind-users), Samba and Mailman (mailman-users). - None of these emails is a near-copy of a training email.
| Recall | False alarms | F1 | ROC-AUC | |
|---|---|---|---|---|
| Phase 1: rule engine | 42.2% (37.8–46.8) | 7.8% (5.9–10.3) | 0.56 | 0.80 |
| Current model | 96.9% (94.9–98.2) | 0.5% (0.2–1.6) | 0.98 | 0.999 |
The errors match the patterns seen on the hold-outs:
- 3 false alarms in 562 emails.
- The missed phishing is mostly crypto-wallet "desktop app" impersonations (CoinTracker, Rabby, Lido) and very short link-free messages.
- Label noise: one "missed phishing" sample is actually an ordinary mailing-list reply.
Full report: reports/LOCKBOX.md.
What these numbers mean in a real inbox. The test sets are 22–33% phishing; a real inbox, after the provider's filters, is usually well under 1%. Recall and the false-alarm rate don't depend on that share, but precision (the share of alerts that are real) does. The table shows the precision implied by the measured recall and false-alarm rate, at the same threshold, for different phishing shares:
| Set | Precision on the test set | 0.1% phishing | 1% phishing | 5% phishing |
|---|---|---|---|---|
| A | 95.3% | 6.2% (4.3–9.3) | 40.2% (31.2–50.9) | 77.8% (70.3–84.4) |
| B | 97.2% | 6.5% (4.1–11.1) | 41.4% (29.9–55.8) | 78.6% (69.0–86.8) |
| C | 98.5% | 19.2% (8.2–45.5) | 70.6% (47.3–89.4) | 92.6% (82.4–97.8) |
| Lockbox | 99.3% | 15.4% (6.4–38.9) | 64.7% (40.8–86.5) | 90.5% (78.2–97.1) |
- At 1% phishing, only about 2 alerts in 5 are real on sets A and B. Precision of 95–99% on the test sets says almost nothing about a real inbox.
- The intervals are wide because the false-alarm rates rest on only 3–22 errors.
- PR-AUC drops the same way: 0.994–0.998 on the test mix, but 0.86–0.93 once the legitimate emails are re-weighted to 0.1% phishing. Details:
reports/HOLDOUT_2025.md. - The lockbox row comes from its saved confusion counts. The lockbox model was not run again.
- The calibrated probability has the same limit: it assumes 37% phishing, as in training. Re-reading it at the real share is tested below.
This text-only model would be one layer of a mail filter, not the whole filter.
The production threshold (0.519) maximizes F1 on the training mix, which is 37% phishing. A real inbox has much less phishing, and a missed phishing email costs more than a false alarm. The textbook fix has two steps:
- Re-read the calibrated probability at the real phishing share (prior shift).
- Flag an email when passing it would cost more than flagging it.
This fix has no free parameter. But it assumes that only the share changes, and that phishing and legitimate mail each look the same as in training. The hold-outs come from other sources, so they test exactly that assumption. The decision rule was committed first (protocol).
Costs are in false-alarm units: a missed phishing email costs r false alarms. The table shows expected cost per 1,000 emails at the pre-registered primary setting, 1% phishing and r = 100:
| Threshold | A | B | C | |
|---|---|---|---|---|
| Production | 0.519 | 23.5 | 22.8 | 74.0 |
| Prior shift (the candidate) | 0.370 | 29.6 | 38.0 | 54.5 |
| Prior shift + analyst review band | 0.128–0.994 | 16.5 | 22.2 | 25.3 |
| Best threshold on the hold-out itself (optimistic) | 0.74 / 0.72 / 0.12 | 20.4 | 18.6 | 41.9 |
- The prior shift does not qualify, so production stays.
- At the primary setting it lowers the threshold to 0.370 and raises the cost significantly on A (+6.0 per 1,000) and B (+15.1).
- On A and B every extra alert is a false alarm: 9 and 14 more legitimate emails, and not one more phishing email.
- Only C gains: 5 more phishing emails caught against 3 more false alarms.
- It works when it raises the threshold. This happens when false alarms dominate the cost: 0.1% phishing with
r≤ 100, or 1% phishing withr≤ 10.- In those settings it lowers the cost on A and B every time, by 32–98%.
- On A and B it lands within 1.5 per 1,000 of each set's own best threshold.
- At 0.1% phishing and
r= 100, for example, the threshold rises to 0.856 and the cost falls from 15.8 to 10.7 (A) and from 15.0 to 7.2 (B). - C is the exception at that setting: its cost rises by 20% (10.9 → 13.1).
- It fails when it lowers the threshold into the middle of the score range. That range does not carry over between sources.
- The hold-outs disagree on the best threshold: 0.74 and 0.72 on A and B, 0.12 on C.
- Between 0.1 and 0.6, the re-read probability is too high on A and B, where mostly legitimate mailing-list email sits, and too low on C, where mostly phishing sits.
- How low the threshold can safely go depends on the inbox, and a formula cannot know that.
- A review band helps on all three sets, but nothing ships. In the band, an analyst looks at every email scoring between 0.128 and 0.994; a review is assumed to cost a quarter of a false alarm.
- Cost falls on every set, no legitimate hold-out email is blocked, and 2.6–7.1% of all mail goes to review.
- Not blocking any legitimate email rests on a thin margin: the highest-scoring legitimate email in A scores 0.9931.
- The pre-registered rule tested the band only on top of the prior shift, which did not qualify, so the band does not ship either.
- Changing the threshold does not fix commercial mail.
- At 0.1% phishing and
r= 100, 82.8% of English promotions are still flagged (production: 87.6%). - Only the most extreme threshold (0.998) brings this down to 23%.
- 19.5% of English promotions score above 0.9994. The problem is in the model, not in the threshold.
- At 0.1% phishing and
What it means: keep 0.519 as the default. An inbox that knows its phishing share can safely raise the threshold when false alarms are the main cost. Lowering it for high-stakes settings needs labeled mail from that inbox.
Full tables (all 12 cost × share settings, calibration on each hold-out, risk–coverage): reports/DECISION_POLICY.md.
Every legitimate test email above comes from open-source mailing lists. The first test on commercial mail shows that the low false-alarm rate does not carry over.
Data. The set is 674 promotional emails (newsletters and campaigns) from the author's own inbox.
- 45 sender domains, at most 30 emails per sender; 169 of the emails are English.
- Private: the emails are not published, so this result cannot be reproduced from the repository.
- Masking: before scoring, the author's name and addresses, every recipient address (in plain, URL-encoded and base64 form), and phone and card numbers were removed.
- Report: it holds counts only, with no subjects, senders or bodies.
The rules were committed before the run (protocol).
| Emails | Flagged as phishing (every one is a false alarm) | |
|---|---|---|
| English promotions | 169 | 87.6% (81.8–91.7) |
| Other languages (mostly Turkish) | 505 | 94.7% (92.3–96.3) |
| For comparison: hold-out A's mailing lists | 1,478 | 1.5% (1.0–2.2) |
- Why (exploratory, found after the result): the text features put promotional mail on the phishing side.
- The mean word + character contribution is +2.4 for English promotions, against −7.4 for the 2024 mailing lists in training and +5.5 for Kaggle "phishing".
- The words that push hardest are marketing language: "your", "email", "unsubscribe", "this email".
- The cause is mainly the missing marketing on the legitimate side: no legitimate training email is marketing. The first guess, that Kaggle's spam-heavy "phishing" class (finding 2) taught it, explains little: removing it barely helps (see below), while adding legitimate promotions does.
- Not a language effect: English promotions are flagged almost as often as Turkish ones.
- The URL features don't fix it: 91.8% instead of 92.9% (McNemar p = 0.17).
- HTML signals, measured on the same mail: none passes the pre-registered rule (at least 5% of training phishing and at least 3× the rate on promotions).
- Hidden text, for example, is in 75% of promotions (newsletter preview lines) and 7% of phishing.
- Only free-hosting links pass: 19% of phishing, 0% of promotions.
- Limits:
- One inbox, 45 senders, mostly Turkish brands.
- The comparison phishing is older (2019–2024).
- "Legitimate" relies on the Promotions category plus a junk filter (iCloud routed the email to the inbox and it passes DMARC).
What it means: as it stands, the model behaves more like a bulk-mail detector than a phishing detector.
Full tables: reports/COMMERCIAL.md.
Three candidates against production, with the decision rule committed first (protocol). Only P1 and P2 use public data and could replace production. P3 is a diagnostic: its promotions are scored by models that never saw their sender (5 folds by sender domain).
| English promotions | All promotions | A false alarms | 2025 recall | C recall | |
|---|---|---|---|---|---|
| Production | 87.6% | 92.9% | 1.5% | 99.1% | 93.0% |
| P1: no Kaggle | 84.6% | 76.1% | 2.6% | 99.3% | 66.4% |
| P2: only Kaggle's legitimate emails | 82.2% | 76.7% | 3.6% | 98.9% | 75.2% |
| P3: + private promotions in training | 37.3% | 15.0% | 0.9% | 97.1% | 85.0% |
- Removing Kaggle's spam does not fix it.
- English promotions barely move: 84.6% (p = 0.23) and 82.2% (p = 0.01).
- False alarms on the hold-outs rise significantly (A: 1.5% → 2.6–3.6%).
- Phishing Pot recall collapses to 66–75%.
- Neither candidate qualifies, so production stays as it is.
- Seeing legitimate marketing does most of the work. With the private promotions in training, English promotions from unseen senders drop from 87.6% to 37.3%.
- It costs recall: 9 more of the 454 phishing emails from 2025 are missed (p = 0.004). Phishing and marketing look alike.
- 37% is still far too high. Most of the training promotions were Turkish; the non-English rate falls to 7.5%, partly because the model can learn "Turkish = legitimate". The fix needs English legitimate marketing mail at scale, which no public dataset provides.
Full tables: reports/COMMERCIAL_FIX.md.
Public legitimate marketing mail does not exist: the one candidate (marketeam/Marketing-Emails) turned out to be internal team correspondence with no links. So the second attempt changed the model instead. The generic words that carry tone ("your" has the largest coefficient, +5.4) were removed, or the character n-grams, or both (protocol).
| English promotions | All promotions | A false alarms | C recall | LLM-written phishing recall | |
|---|---|---|---|---|---|
| Production | 87.6% | 92.9% | 1.5% | 93.0% | 97.1% |
| M1: no English stop words | 88.8% | 89.9% | 1.8% | 91.1% | 97.2% |
| M2: no character n-grams | 89.9% | 89.6% | 2.2% | 89.7% | 98.4% |
| M3: both | 88.2% | 82.8% | 2.8% | 88.8% | 98.0% |
- No variant qualifies. English promotions do not move, and removing character n-grams raises false alarms on the mailing lists.
- The tone is carried by content words, not by function words or spelling patterns: "unsubscribe", "offer", "account", "click". A model that sees no legitimate marketing has no reason to treat those as legitimate.
- This confirms the earlier diagnostic: the fix has to come from training data, not from the model.
Full tables: reports/TONE.md.
An exploratory diagnosis, run after the decision-policy study, found that the model still partly recognizes where an email came from:
- Strongest legitimate features:
enron(−7.6), the quote marker>,vince,2002andwrote. - Strongest phishing features:
your,2005,2004andspamassassin sightings(the name of a mailing list). - Why: every training source holds a single class, so anything that identifies a source also identifies the label.
- The shortcut is exploitable. Thread-hijacking attacks hide phishing inside a real conversation. Appending a fake three-line quoted reply in that style drops recall from 99.1% to 72.9% (Nazario 2025) and from 93.0% to 68.2% (Phishing Pot).
A pre-registered fix removed three families of shortcuts the same way in every source (protocol):
- Reply structure: quote lines, "… wrote:" lines, reply and forward headers.
- Dates and years.
- Text format: Kaggle's legitimate mail is lowercased, with spaced punctuation.
| Production | Cleaned (S) | |
|---|---|---|
| Nazario 2025 recall, with a fake quoted reply (T1) | 72.9% | 98.7% |
| Hold-out C recall, with a fake quoted reply (T1) | 68.2% | 92.1% |
| Nazario 2025 recall, with an Outlook-style reply and benign text (T4) | 78.2% | 93.8% |
| False alarms, A | 1.5% | 3.9% |
| False alarms, B | 1.4% | 3.1% |
| English promotions flagged | 87.6% | 89.3% |
-
S does not qualify, so production stays. It closes the exploit, but false alarms on A and B rise significantly (p < 0.001 and p = 0.014).
-
The low false-alarm rate on mailing lists partly rested on the shortcut (exploratory, after the result).
- 35 of the 42 new false alarms on A are replies whose quote was removed.
- What is left is short (median 45 words, against 104 for all of A's legitimate mail), addressed to "you", and often has a link.
- To this model, "legitimate" largely meant "quotes earlier mail".
-
Removing shortcuts makes the model find new ones. After cleaning, these are among its strongest features:
713, Houston's area code: 99% of the emails that contain it are Enron's;enronandvince;spamassassin sightings.
-
Benign padding without quote markers still works on both models. A plain friendly sentence (T5) costs 4–7 points of recall.
-
What it means: the cause is the data, not the text processing. As long as each source holds one class, any trait of a source is a trait of the label. The fix needs two things:
- legitimate mail that looks like what phishing imitates;
- ideally, sources that contain both classes.
That is the next step: a research inbox used only for this project.
Full tables, ablations and the source audit: reports/SHORTCUTS.md.
These headers record whether a mail really comes from the domain it shows. They are not model features: the mailing-list archives used as legitimate training mail strip them (0% present), so any header feature would learn the source. The measurement compares training phishing (Nazario 2019–2024, 1,571 emails) with the author's private promotions (676 emails, without the DMARC part of the junk filter). Phishing Pot and the lockbox are kept back for a later blind test (protocol).
| Signal | Phishing | Promotions |
|---|---|---|
DMARC recorded, not pass |
58.2% | 0.6% |
DKIM not aligned with the From domain |
52.4% | 4.6% |
| No DKIM signature | 40.6% | 0.0% |
Reply-To on another domain |
7.1% | 0.0% |
| Sender on a freemail domain | 2.5% | 4.4% |
- Strong and complementary to the text. Eight of the nine signals pass the pre-registered rule. Freemail is the exception.
- The receiver-independent checks are the trustworthy ones. DKIM alignment and
Reply-Toare computed from the headers themselves. The recorded SPF/DKIM/DMARC results were written by different servers (iCloud for the promotions, the collector or a forwarder for Nazario). "No results at all" (26.6% of phishing, 0% of promotions) is a collection artifact, not a phishing trait. - They cannot catch everything (exploratory, after the result). 95.0% of promotions come from a verified sender (DMARC pass and aligned DKIM), and so does 27.0% of phishing: phishers authenticate their own lookalike domains.
- What this suggests: a second stage for verified senders could remove most commercial false alarms. Suppressing alerts for them outright would cost up to a quarter of phishing recall. The design needs its own protocol and a blind test on Phishing Pot.
Full tables: reports/HEADERS.md.
The pre-registered design (protocol) works like this. If the sender is verified and not suspicious (no brand on a foreign domain, no punycode, no free-hosting link, no Reply-To elsewhere), the model needs a higher score t_v to flag the email. t_v comes from training out-of-fold scores.
- The rule picked
t_v= the production threshold. More than 5% of verified training phishing already scores below it, so the stage changes nothing and does not qualify. Nothing ships. - No other bar would work either. The sensitivity table below is descriptive and does not decide:
t_v |
English promotions flagged | Nazario 2025 recall | Hold-out C recall |
|---|---|---|---|
| production (0.519) | 87.6% | 99.1% | 93.0% |
| 0.99 | 43.8% | 90.3% | 87.4% |
| 0.999 | 29.6% | 84.1% | 85.0% |
| never alert for verified senders | 8.9% | 55.1% | 75.7% |
- Verified-sender phishing is common in recent mail. 58.6% of Nazario 2025 phishing comes from a verified sender, against 27.0% in Nazario 2019–2024 (same collector) and 27.6% in Phishing Pot. Authentication says who sent a mail, not whether it is honest.
- So the fix has to come from the text model: legitimate English marketing mail in training.
Full tables: reports/SENDER_STAGE.md.
Two test-only sets, with the decision rule committed first (protocol):
- Zenodo (Gutierrez et al. 2026): 4,986 phishing emails written by GPT-4.1, DeepSeek 3.2 and Llama 3.3 70B in five themes.
- Greco et al. 2024: 1,000 phishing emails written by ChatGPT and WormGPT.
None is a near-copy of training phishing.
| Emails | Recall | |
|---|---|---|
| Zenodo, all | 4,986 | 97.1% (96.6–97.5) |
| GPT-4.1 | 1,665 | 98.7% (98.0–99.1) |
| DeepSeek 3.2 | 1,665 | 97.5% (96.7–98.2) |
| Llama 3.3 70B | 1,656 | 95.0% (93.9–96.0) |
| HR theme | 996 | 85.7% (83.4–87.8) |
| Banking, parcel, IT support, tax themes | 3,990 | 99.7–100% |
| Greco (ChatGPT, WormGPT) | 1,000 | 99.5% (98.8–99.8) |
| For comparison: hold-out C, human-written | 214 | 93.0% (88.8–95.7) |
- Under the pre-registered rule, LLM-written phishing is not harder: recall is higher than on human-written hold-out C (Fisher p = 0.003).
- Weak spots:
- HR lures: benefits enrolment deadlines and staff reorganisation memos, the same kind of lure the model missed in finding 2.
- Llama 3.3.
- Emails without a link: 94.5%, against 98.8% with one.
- But high recall here proves little. In an exploratory audit, 40 random emails from Greco's "legitimate" LLM set were labelled by one person, after the result and with the model's scores visible (labels). 15 of the 16 clearly benign ones are flagged: COVID notices, event invitations, tips.
- As with the promotions, polished corporate-sounding text is flagged almost regardless of intent.
- So recall on LLM-written phishing means little without a false-alarm rate on LLM-written legitimate mail of the same style.
- Caveats:
- The emails were generated on request, in themes that overlap the training phishing, so this tests style more than new lures.
- 13% of the Zenodo emails carry a bracketed placeholder where a link would be.
Full tables: reports/LLM_PHISHING.md.
The first ML version flagged 25% of Apache release announcements and half of Ubuntu security notices as phishing.
1. Diagnosis. The problem was a single feature, not the text. On the false alarms, the text features actually voted "legitimate" (−2.4 on average). The rule features voted "phishing" (+5.7), and almost all of that came from one raw count, url_count. A linear model extrapolates without limit: an announcement with 30 links scored far beyond anything seen in training. Real phishing has a median of 1 link.
2. Bounded features. Raw totals were replaced with ratios, maximums and yes/no flags: "share of links that look suspicious", "worst link's score", "contains an IP link". Counts now use a log scale, and every scaled feature is clipped to ±3.
3. Modern legitimate training data. 2,205 emails from 2024 Python, Fedora and GNU mailing lists were added: release announcements, update notices and Q&A. Apache lists were left out on purpose, because Apache release announcements are also posted to a list in hold-out A.
| Version | Hold-out B false alarms (blind) | Hold-out A false alarms | Recall (2025 Nazario) |
|---|---|---|---|
| v1: raw-count features | 25.7% | 7.0% | 85.7% |
| v2: bounded features | 15.2% | 10.0% | 97.4% |
| v2 + threshold re-tuned with modern data | 15.2% | 10.0% | 97.4% |
| v3: bounded features + modern training data | 0.9% | 1.2% | 96.5% |
These are the current re-runs of the experiment, after all the fixes below. Every split keeps near-duplicates together, and Kaggle no longer contains SpamAssassin copies.
- The gain from modern data is real, not noise. v3 vs v2 on the same emails (McNemar test): 132 emails that only v3 gets right against 5 that only v2 gets right, p < 0.001.
- Bounded features alone are not enough. They cut false alarms on B (p < 0.001) but not on A (p = 0.62).
- Re-tuning the threshold does not help at all. At the max-F1 point it picks the same threshold. The model has to see this kind of email in training.
- A validation false-alarm rate does not carry over to a new source. v1 had at most 1% false alarms on validation and 29.7% on hold-out B.
- No regression on the original data. In-domain test F1 is 0.982 for v1 and 0.981 for v3.
Full tables: reports/FALSE_ALARMS.md.
Cutting false alarms first cost recall: 97.6% → 94.9% on 2025 phishing. Of the 23 phishing emails missed at that point:
- 57% had no link.
- 43% were very short.
- 15 used a payment pretext: invoice, ACH remittance, RFQ, order confirmation.
Real phishing of this kind was scarce in training.
Two fixes were compared on hold-out C, built for this step from two sources used nowhere else. The production version is the one with the best validation F1, a rule fixed before any set-C result was seen.
| Version | Change | C recall | C false alarms | 2025 Nazario recall | Validation F1 |
|---|---|---|---|---|---|
| v3 | – | 91.1% | 0.5% | 97.6% | 0.9830 |
| v3+lure | + lure features: payment / attachment pretext, generic greeting, urgency, credential request, short body | 87.4% | 0.3% | 97.4% | 0.9830 |
| v4 | + 705 real phishing emails frScan report · 2026-10-09
From the balcony · 2 of 4 clapped
Cap'm Slop and Schnitzel read it and passed. Their reasons are on the balcony, with every other verdict. Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are. report this listing— log in to report |
0 comments
log in to comment.