Multi-source threat actor intelligence for everyone — with memory.
THEORY is an open-source alternative to enterprise threat intelligence platforms. It generates analyst-grade dossiers on threat actors by aggregating data from 20 free and keyless-by-default sources — MITRE ATT&CK, MISP Galaxy, CIRCL MISP, Malpedia, AlienVault OTX, SigmaHQ, YARA-Rules, ThreatFox, MalwareBazaar, URLhaus, GreyNoise, AbuseIPDB, Shodan InternetDB, urlscan.io, VulDB, NVD, CISA, CISA KEV, and vendor research blogs — then synthesizes everything using an LLM into a clean executive overview and actor-specific intelligence summaries.
As of v2.0, THEORY isn't stateless. Every run feeds a persistent correlation graph that accumulates across every query you've ever made, so you can ask "have I ever seen this IP?" or "is this actor connected to this CVE?" without re-running collection, track what's changed for an actor over time, and even watch one continuously.
Built for threat intelligence analysts, detection engineers, security researchers, and students who believe good intelligence shouldn't require a six-figure subscription.
For any supported threat actor, THEORY generates:
- LLM-written synopsis — 4-6 sentence executive overview synthesized from all available data, at the top of every dossier
- TTP table — every known technique with tactic, confidence score, and detection guidance
- Detection opportunities — Sigma rules mapped to actor TTPs and YARA rules matched to malware families, plus draft Sigma rule skeletons for every gap
- Malware inventory — all associated families with full descriptions, sample hashes (MalwareBazaar), and YARA detection coverage
- IOC table — deduplicated, defanged indicators from OTX, ThreatFox, MalwareBazaar, URLhaus, and CIRCL MISP with confidence scores and malware family attribution
- IP and domain enrichment — GreyNoise noise/RIOT context, AbuseIPDB reputation scores, Shodan InternetDB open-port/CVE data, and urlscan.io scan history annotate every public IP and domain/URL indicator so analysts can distinguish real infrastructure from internet background radiation
- Vulnerability intelligence — CVEs extracted from actor TTPs and cross-referenced against CISA KEV, NVD, and VulDB, with CVSS scores, exploit availability, and remediation status
- Recent intelligence — LLM-synthesized summaries of recent vendor research articles, with source attribution and links
- Campaigns — full campaign descriptions with dates and ATT&CK links
- Targeted sectors and CISA advisories
- IR playbooks — analyst-ready checklists with IOC blocks, detection checklists, hunt hypotheses, and containment guidance
- ATT&CK Navigator layers — confidence-colored heatmaps importable directly into MITRE Navigator
- HTML dossiers — self-contained, shareable intelligence reports that open in any browser
- Detection coverage gap reports — compare actor TTPs against your local detection rules
And, persisted across every run rather than per-dossier:
- A correlation graph you can query directly — by indicator, technique, CVE, or attack type — with or without re-running collection
- Change tracking — what's new or different for an actor since the last time you checked, on demand or continuously
- Grounded Q&A — ask a question in plain English and get an answer sourced only from what THEORY has actually recorded
Output formats: terminal dossier, markdown, JSON, STIX 2.1 (for MISP/OpenCTI/Sentinel), IOC CSV, HTML, ATT&CK Navigator, IR playbook (markdown or Jira), and Sigma rule skeletons.
# 1. Clone the repository
git clone https://github.com/threatcraft-co/theory
cd theory
# 2. Create a virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# 3. Install THEORY and dependencies
pip install -e .
# 4. Download the ATT&CK bundle (required for MITRE source)
theory --update-bundles
# 5. Configure your API keys
cp .env.example .env
# Edit .env and add your OTX_API_KEY (free at otx.alienvault.com)
# Everything else — including 2 of the 20 sources — needs no key at all
# 6. Run your first dossier
theory --actor APT28That's it. Your first dossier renders in the terminal and saves to output/dossiers/apt28.md — and the run is now in your local correlation graph at output/graph/graph.json, ready to query.
# Ask the graph directly, any time after that first run
theory --ioc 1.1.1.1
theory ask "what do we know about APT28's malware?"20 sources, 15 of which need no API key or signup at all.
| Key | Source | Auth Required | Cache |
|---|---|---|---|
mitre |
MITRE ATT&CK — techniques, malware, campaigns (local bundle) | None | 7 days |
cisa |
CISA advisories + KEV catalog | None | Per request |
cisa_kev |
CISA KEV — 1,600+ confirmed-exploited CVEs, ransomware flags | None | 24 hours |
malpedia |
Malpedia malware family database | None | Per request |
misp_galaxy |
MISP Galaxy — 1,000+ actors, aliases, attribution, target sectors | None | 7 days |
circl_misp |
CIRCL OSINT MISP feed — event-level indicators + campaign context | None | 24h / 7 days |
otx |
AlienVault OTX pulses + IOCs | OTX_API_KEY |
Per request |
nvd |
NIST NVD — CVSS scores, CWE classification, references | None (optional key raises rate limit) | 30 days/CVE |
sigma |
SigmaHQ detection rules mapped to ATT&CK (local clone) | None (GITHUB_TOKEN optional) |
7 days |
yara |
YARA file detection rules matched to malware families (local clone) | None | 7 days |
threatfox |
ThreatFox IOCs by malware family | None | 24 hours |
malware_bazaar |
MalwareBazaar sample hashes by malware family | ABUSECH_API_KEY |
24 hours |
urlhaus |
URLhaus malware distribution URLs by family | ABUSECH_API_KEY |
24 hours |
greynoise |
GreyNoise IP noise/RIOT context | GREYNOISE_API_KEY |
7 days |
abuseipdb |
AbuseIPDB IP reputation scores from community reports | ABUSEIPDB_API_KEY |
3 days |
shodan_internetdb |
Shodan InternetDB — open ports, hostnames, CPEs, known CVEs on an IP | None | 24 hours |
urlscan_io |
urlscan.io public scan history for domains/URLs | None (optional key raises rate limit) | 24 hours |
vuldb |
VulDB actor-CVE correlation and exploit intelligence | VULDB_API_KEY |
7 days |
vendor |
Vendor intelligence synthesis — 50 research blogs, LLM-synthesized | LLM API key | 7 days |
personal |
Your own local research indicators — gitignored, never leaves your machine | None | none (reads live) |
shodan_internetdb is not the paid Shodan search API — it's Shodan's own free, keyless InternetDB lookup service, a deliberately different thing. The paid Shodan and Censys APIs remain unimplemented (see pyproject.toml's internet-scan extra); VirusTotal was considered and skipped in favor of urlscan.io, which covers similar ground without a signup.
# Live status of all sources
theory --list-sourcestheory --actor APT28
theory --actor "Fancy Bear" # alias resolution — same output
theory --actor "Forest Blizzard" # same actor, different name# Default (mitre + cisa + cisa_kev + malpedia + misp_galaxy, no auth needed)
theory --actor APT28
# Add community IOCs
theory --actor APT28 --sources mitre,cisa,malpedia,otx
# Full enrichment including detection rules (Sigma + YARA)
theory --actor APT28 --sources mitre,cisa,malpedia,otx,sigma,yara,threatfox
# Complete abuse.ch trifecta (network + file + delivery IOCs)
theory --actor APT28 --sources mitre,malpedia,threatfox,malware_bazaar,urlhaus
# IP + domain enrichment (GreyNoise, AbuseIPDB, Shodan InternetDB, urlscan.io)
theory --actor APT28 --sources mitre,otx,threatfox,greynoise,abuseipdb,shodan_internetdb,urlscan_io
# Vulnerability intelligence (CISA KEV + NVD + VulDB)
theory --actor APT28 --sources mitre,cisa_kev,nvd,vuldb
# Your own research alongside public sources
theory --actor APT28 --sources personal,mitre,cisa
# Everything, with vendor intelligence synthesis (requires LLM key in .env)
theory --actor APT28 --sources mitre,cisa,cisa_kev,malpedia,misp_galaxy,circl_misp,otx,nvd,sigma,yara,threatfox,malware_bazaar,urlhaus,greynoise,abuseipdb,shodan_internetdb,urlscan_io,vuldb,vendor# Terminal + markdown file (default)
theory --actor APT28
# Raw JSON profile
theory --actor APT28 --output json
# STIX 2.1 bundle (import into MISP, OpenCTI, Sentinel)
theory --actor APT28 --output stix
# IOC-only CSV (for SIEM lookup tables)
theory --actor APT28 --sources mitre,otx,threatfox --output csv
# Self-contained HTML dossier (shareable, opens in any browser)
theory --actor APT28 --sources mitre,malpedia,otx --output html
# ATT&CK Navigator layer (import at mitre-attack.github.io/attack-navigator)
theory --actor APT28 --sources mitre,malpedia,otx --output navigator
# IR playbook with detection checklist and IOC blocks
theory --actor APT28 --sources mitre,sigma --output playbook
# IR playbook in Jira wiki markup
theory --actor APT28 --sources mitre,sigma --output playbook --playbook-format jira
# Non-technical executive summary (BLUF format, requires LLM key)
theory --actor APT28 --output exec
# Executive summary with sector context
theory --actor "Lazarus Group" --output exec --sector finance
# Draft Sigma rule skeletons for every detection gap
theory --actor APT28 --sources mitre,sigma --output sigma-skeleton
# All formats at once
theory --actor APT28 --output all
# Print only — don't write files, don't touch the correlation graph
theory --actor APT28 --no-save --no-graphEvery --actor run feeds a persistent correlation graph at output/graph/graph.json — a typed, bidirectional graph of actors, indicators, techniques, malware, CVEs, and campaigns that accumulates across every run you've ever made. You can query it directly, two ways:
# Standalone — what does THEORY know about this, across every past run?
theory --ioc 1.1.1.1
theory --technique T1566
theory --attack-type ransomware
theory --cve CVE-2023-23397
# Combined with --actor — one question: is THIS actor connected to THIS
# entity, in anything THEORY has ever recorded? (Not two reports stapled
# together — checks for a direct link first, then a shared "bridge" node.)
theory --actor APT28 --ioc 1.1.1.1
theory --actor APT28 --technique T1566
theory --actor APT28 --attack-type espionage
theory --actor APT28 --cve CVE-2023-23397--attack-type is deliberately grounded in two fields THEORY already collects reliably — actor motivations and malware type — rather than a separate, hand-built attack-pattern taxonomy. Disable graph writes for a one-off run with --no-graph.
theory ask "what do we know about 1.1.1.1?"
theory ask "is APT28 connected to T1566?"
theory ask "what are my own notes on Lazarus Group?"Answers come only from what THEORY has actually recorded — the persistent graph and your personal research notes — never from the model's own training data passed off as current intel. Works with any configured LLM provider (Claude, OpenAI, Ollama) via a provider-agnostic tool-calling protocol.
# Compare the last two JSON runs for an actor (free — one step of history
# is kept automatically whenever you run --output json or all)
theory --actor APT28 --output json
# ... time passes, new intel appears, you run again ...
theory --actor APT28 --output json
theory diff --actor APT28
# Compare any two profile files directly
theory diff --from old.json --to new.json
# Re-run on a timer and report only what changed (Ctrl+C to stop)
theory --actor APT28 --watch
theory --actor APT28 --watch --watch-interval 900 # every 15 minutes# One-time setup — creates a gitignored redirect + starter file
theory --init-personal
# Add your own indicators to ~/.theory/personal_indicators.yaml, then:
theory --actor APT28 --sources personal,mitre,cisaTwo-layer, gitignored indirection: config/local_sources.yaml → a private file outside the repo entirely (default ~/.theory/personal_indicators.yaml, overridable with --personal-path). Even an accidental commit of the redirect leaks only a path, never your research.
# Compare actor TTPs against your local detection rules
theory --actor APT28 --sources mitre,sigma --detection-path ~/my-sigma-rules
# Output: coverage %, covered techniques, and gaps sorted by confidence
# Turn every gap into a starting-point Sigma rule (logsource/selection left as TODO)
theory --actor APT28 --sources mitre,sigma --output sigma-skeletontheory --list-actors # 1,019 supported actors with 2,455 aliases
theory --list-sources # all 20 sources with auth and cache info# Refresh ATT&CK bundle, Sigma rules, YARA rules, MISP Galaxy cluster, and CISA KEV catalog
theory --update-bundlestheory --actor APT28 --sources mitre,cisa --verbosepip install -e ".[serve]"
theory serve # localhost:8088, opens browserSame pipeline as the CLI, nothing leaves your machine. See server/README.md. (The v2.0 graph queries, theory ask, diff, and --watch are CLI-only for now.)
THEORY knows 1,019 actors by all their names (2,455 aliases total — expanded from an initial hand-curated ~35 via a one-time MISP Galaxy import). Any alias resolves to the same canonical dossier:
theory --actor "Cozy Bear" # → APT29
theory --actor "Midnight Blizzard" # → APT29
theory --actor "Nobelium" # → APT29
theory --actor "NOBELIUM" # → APT29 (case-insensitive)The output file is always named by the canonical actor — --actor "Fancy Bear" produces apt28.md, not fancy_bear.md. Querying an actor THEORY doesn't have an alias entry for still works — it just falls back to the raw name you typed instead of resolving the full alias set, which costs some recall on sources other than misp_galaxy itself.
theory --list-actors # see all 1,019 actors and their aliasesEvery dossier opens with an Intelligence Overview — a 4-6 sentence executive synopsis written by Claude (or your configured LLM) using the full aggregated profile as context.
The synopsis:
- Uses the name you queried, not aliases
- Covers origin, motivations, target sectors, signature TTPs, notable malware, and recent activity
- Works with or without
--sources vendor— synthesizes from structured MITRE data alone if needed - Appears at the top of both the terminal output and the markdown file
LLM provider resolution order: Claude → OpenAI → Ollama. Set THEORY_LLM_PROVIDER in .env to override, or leave blank to auto-detect. Ollama runs fully offline. The same provider resolution backs theory ask and the LLM-generated playbook/exec-summary content.
When you add vendor to your sources, THEORY fetches recent articles from 50 threat research blogs (Mandiant, Google TAG, Unit 42, Secureworks, Recorded Future, CrowdStrike, Kaspersky GReAT, Check Point Research, Sophos, Proofpoint, and more) and uses an LLM to synthesize what each article reveals about your actor specifically. Sites without an RSS feed can be ingested via a sitemap-based feed type instead — see "Adding custom feeds" below.
# Set your preferred provider and API key in .env
THEORY_LLM_PROVIDER=claude
ANTHROPIC_API_KEY=your_key_here
# Run with synthesis
theory --actor "Lazarus Group" --sources mitre,malpedia,otx,vendorThe dossier includes a Recent Intelligence section with actor-specific summaries, source attribution, and direct links to original articles.
THEORY uses a local clone of the SigmaHQ repository — no rate limits, no API, instant results.
# First run clones the repo (~150MB, ~1-2 minutes, one time only)
theory --actor APT28 --sources mitre,sigma --no-save
# Every subsequent run is instant
theory --actor APT28 --sources mitre,sigma --no-saveDetection rules are linked directly to actor TTPs in the dossier. Techniques with no matching rule show up in the detection-coverage gap analysis below, and can be turned into draft rule skeletons with --output sigma-skeleton. See docs/SIGMA_RATE_LIMITS.md for full architecture details.
Where Sigma covers log-based and network detection, YARA covers file-based and memory detection. THEORY uses a local clone of the Yara-Rules/rules repository — no rate limits, no API, instant results.
Scan report · 2026-10-03
- ✓ Prohibited terms or links
- ✓ Repository eligibility
- ✓ slopscore.md paperwork
- ✓ Content policy
- ✓ Risk review
From the balcony · 0 of 2 clapped
Princess and Crusoe read it and passed. Their reasons are on the balcony, with every other verdict.
Critics are accounts on this site with no GitHub account behind them. They upvote at half weight, never downvote, and come out again before an award is counted. Who they are.

0 comments
log in to comment.