Post

Building barb — A Phishing URL Analyzer Built with Claude Code

Building barb — A Phishing URL Analyzer Built with Claude Code

barb is a command-line tool for heuristic phishing URL analysis. It runs entirely offline, requires no API keys, and delivers structured verdicts directly in the terminal. This post covers how it was built, how the workflow has evolved since Building vex — An IOC Enrichment Tool Built with Claude Code, and what a more deliberate approach to AI collaboration looks like in practice.

FieldDetails
TypePython CLI Project
Repositorygithub.com/duathron/barb
PyPIpypi.org/project/barb-phish
Version1.1.0
StackPython, Typer, Rich, Pydantic v2
Built withClaude Code

The Problem

Phishing URLs land in inboxes, get extracted from email gateways, show up in SIEM alerts. Every analyst who has had to triage a list of suspicious links knows the drill: open a browser tab, paste the URL into VirusTotal or URLScan.io, wait, read, move to the next one. For five URLs that’s fine. For fifty, it becomes the job instead of part of it.

The existing options all involve trade-offs. VirusTotal URL scan requires an API key and a network request that reaches out to a cloud service for every URL you check. URLScan.io submits the URL to an external system — which is not always appropriate in a sensitive investigation. PhishTank is community-driven and works only for known campaigns. There’s a gap for something that runs locally, offline, instantly, and without registering anything anywhere.

That was the starting point for barb.


From vex to barb — How the Workflow Changed

When I built Building vex — An IOC Enrichment Tool Built with Claude Code, the MeetUps were already part of the process, but they were primarily reactive — agents reviewed code and decisions as they came up. With barb, the MeetUp structure was more deliberate from the beginning: before a single line of code was written, two full MeetUps defined what the tool would be, how it would be named, and what every key design decision meant.

The first MeetUp debated which project to build next. Five agents — AI Specialist, Architect, SOC Analyst, DFIR, and Marketing — each submitted a brief. The options ranged from a vex reporting layer to a phishing analyzer to an alert triage summarizer. The vote was unanimous for the phishing analyzer: it’s independently usable, testable with real public URLs, directly addresses one of the most common attack vectors named in job descriptions, and is explainable in a single sentence.

The second MeetUp — seven agents this time — ran through the architecture, scope, and every design decision with formal votes. Not all of those votes were unanimous. The Code Security Agent dissented on the default for send_url when using the LLM explanation feature, preferring opt-in over opt-out. The AI Specialist dissented on deferring Ollama support. Both dissents are documented in the MeetUp log. The decisions that prevailed have documented rationale; the minority positions have documented reasoning too.

A third MeetUp — just Marketing and UX Design — focused entirely on naming. That produced barb: the sharp backward-pointing part of a fishhook, the mechanism that catches and holds. Short, no CLI or PyPI conflicts, phonetically pairs with vex, zero collision with developer vocabulary. The PyPI suffix pattern follows the same logic as vex-ioc — the package is barb-phish, the CLI command is barb.

What changed between vex and barb is the formalisation. The MeetUp structure went from a useful pattern to an actual methodology: clear participant lists, briefing phases, structured discussion, recorded votes with rationale, and explicit documentation of dissents and deferred items. The result is a codebase with a traceable decision history — not just what was built, but why every significant choice was made and what the alternatives were.


What barb Does

Eight Heuristic Analyzers

barb runs eight independent checks on every URL. Each analyzer examines a specific dimension of phishing infrastructure and returns a list of signals with severity levels.

AnalyzerWhat it detectsExample
EntropyUnusually high randomness in domain or pathx7k2m9p.evil.com
HomoglyphUnicode characters that visually mimic ASCIIpаypal.com (Cyrillic ‘а’)
TLDHigh-risk top-level domainspaypal-login.tk
SubdomainExcessive depth or squatting patternssecure.paypal.com.evil.com
BrandBrand name appearing in a non-brand domainpaypal-secure.evil.com
ShortenerKnown URL shortener servicesbit.ly/abc123
EncodingPercent-encoding or punycode abuse%70%61%79pal.com
IP URLIP address used instead of domainhttp://192.168.1.1/login

All data — homoglyph mappings, brand lists, known shorteners, suspicious TLDs — is bundled as static JSON files. Nothing is fetched at runtime. The tool makes no HTTP requests to analyzed URLs.

Five-Tier Verdict

Signals are aggregated into a weighted risk score and mapped to one of five verdict levels: SAFE, LOW_RISK, SUSPICIOUS, HIGH_RISK, or PHISHING. The thresholds are configurable in ~/.barb/config.yaml, as are the per-analyzer weights. A URL with zero signals scores 0.0 and gets SAFE; a URL that combines brand impersonation, a suspicious TLD, and a homoglyph character can exceed the PHISHING threshold (13+) easily.

The five-tier system was a MeetUp decision — UX Design and Marketing initially favored three tiers for simplicity, but the SOC Analyst and Architect argued that granular verdicts matter for automation: --threshold filtering, exit code mapping to SOAR playbooks, scripting decisions based on risk level. The visual distinction — green, blue, yellow, orange, red — made the five tiers legible enough that UX Design came around.

1
barb analyze https://pаypal.com --explain
1
2
3
4
5
6
7
8
╭──────────────────────── barb ────────────────────────╮
│ URL       hxxps[://]pаypal[.]com                     │
│ Verdict   🔴 PHISHING                                │
│ Score     15.0                                        │
╰──────────────────────────────────────────────────────╯
 Severity   Analyzer     Finding
 CRITICAL   homoglyph    Cyrillic 'а' (U+0430) mimics Latin 'a'
 HIGH       brand        'paypal' in non-PayPal domain

Automation-Ready by Design

Exit codes map to verdict levels: 0 for SAFE or LOW_RISK, 1 for SUSPICIOUS or HIGH_RISK, 2 for PHISHING, 3 for errors. That makes barb directly usable in shell scripts:

1
2
barb analyze -f daily_urls.txt -o json | jq '.[] | select(.verdict == "PHISHING")'
cat suspicious_links.txt | barb analyze --threshold 8 -o csv > triage_report.csv

The --threshold flag filters batch output to only show URLs that meet or exceed the specified risk score — so a list of 500 URLs produces a focused list of the ones worth investigating.

The --explain Flag

When --explain is passed, barb generates a natural-language summary of its findings. By default, this uses a template-based system that constructs an explanation from the signal data — no API key, no network request, no LLM required. For analysts who want more context, Anthropic Claude or OpenAI can be configured as providers via pip install barb-phish[llm]. The LLM receives the defanged URL and the signal breakdown, not the original URL.

The decision to make the template-based system the default — not the LLM — was unanimous. The tool has to work for anyone, with or without an API key.


Technical Architecture

The architecture inherits directly from vex: Typer CLI, Pydantic v2 data models, typing.Protocol-based analyzer system, Rich output, config priority hierarchy. The decision to reuse these patterns rather than experiment with something different was deliberate — a proven foundation lets the focus stay on the domain-specific logic.

1
2
3
4
5
6
7
8
9
10
11
12
barb/
├── main.py               # CLI entrypoint: analyze, config, version
├── config.py             # Pydantic v2 config with priority hierarchy
├── models.py             # AnalysisResult, Signal, RiskVerdict, ParsedURL
├── url_parser.py         # URL decomposition via urllib.parse
├── scoring.py            # Weighted signal aggregation → verdict
├── defang.py             # Defanging/refanging (copied from vex)
├── batch.py              # Parallel batch processing
├── analyzers/            # Eight heuristic analyzer modules
├── data/                 # Bundled JSON data files
├── explain/              # Template + LLM explanation system
└── output/               # Rich + console output, JSON + CSV export

One difference from vex is visible here: url_parser.py and scoring.py exist as standalone modules. In vex, similar logic was embedded in the enrichers. Separating URL decomposition from analysis and scoring from verdict assignment makes each part testable in isolation — something the Architect pushed for from the start, and the test suite reflects: 58 unit tests at v1.1.0 (40 at v1.0.0, 18 new for the OSINT enrichers), all passing.

CI was also part of the plan from the beginning, unlike vex where it came later. pytest and ruff lint run on every pull request.


Beta Testing and the One Bug

The QM Agent ran a structured test matrix before the v1.0.0 release — four levels: install and import, functional smoke tests, edge cases, and output verification. Thirty-one of thirty-two tests passed on the first run. The one that failed surfaced a real issue: a URL longer than 2048 characters triggered a ValueError inside parse_url() that propagated all the way to the user as a raw Python traceback.

The fix was wrapping the analyze path in a try/except ValueError that outputs a clean error message and exits with code 3. The beta test was re-run after the fix. All 32 tests passed. The QM verdict was SHIP.

That process — define the test matrix before the release, find the bug, fix it, re-verify — was the same pattern that caught the pyproject.toml issues in vex. The difference is that with barb it was planned upfront rather than discovered during a post-release install.


Lessons Learned

The MeetUp format is more useful when it comes first. With vex, agents reviewed decisions that had already been made. With barb, they shaped decisions before implementation began. The architecture is cleaner for it — the single analyze subcommand, the five-tier verdict system, the template-first explainer — each of those came out of the MeetUp phase, not a refactor cycle.

Documenting dissent is as useful as documenting decisions. The Code Security Agent’s objection to send_url: true as default is in the log. Anyone reading the project documentation now knows that this setting exists, why the default was chosen, and what the privacy-conservative alternative is. That context wouldn’t survive a future refactor without the log.

Zero network dependencies in the core is a real differentiator. The offline-first constraint felt like a limitation when it was first proposed. In practice it’s the tool’s strongest feature: it can run on an airgapped analysis machine, in a CI pipeline, or in an environment where outbound traffic to analysis services isn’t permitted. The --explain LLM flag is optional precisely so the core use case never depends on it.

CI from day one changes the feedback loop. With vex, the first real install test was manual. With barb, every commit runs the test suite. The bug that was found in the QM session was a genuine edge case — not something that would have shown up in a normal usage run. Having tests in place made the fix low-risk: one change, re-run, green.


What’s New in v1.1.0

v1.1.0 adds OSINT enrichment via a new --osint flag — opt-in, so the offline-first default stays intact. Two enrichers run when the flag is passed: DNS resolution via Python’s stdlib socket.getaddrinfo(), which flags loopback addresses, known sinkholes, and NXDOMAIN responses; and RDAP-based domain age lookup using the IANA bootstrap registry (RFC 7480–7484) without any external packages.

The domain age signals are tiered: domains registered within the last 30 days trigger a HIGH signal, under 90 days a MEDIUM, and privacy-protected registrants a LOW. A bootstrap cache at ~/.barb/rdap_bootstrap.json has a 7-day TTL to avoid redundant registry lookups.

The MeetUp for this release (2026-04-01) included OSINT Agent, SOC Analyst, Architect, Code Security, and UX Design. The vote was unanimous on --osint as opt-in with DNS and RDAP. One option that came up and was rejected: scraping nslookup.io or dnschecker.org for DNS data. The Code Security Agent blocked it — scraping third-party sites to enrich untrusted URLs creates a privacy and reliability problem. stdlib and RDAP were the right call.

What’s Next

The remaining v1.1 backlog items — Ollama support, STIX 2.1 export, SQLite cache for repeat analysis — are documented with rationale. The integration with vex via --from-barb is already live in vex v1.2.0, so the pipeline works now: barb handles the offline heuristic URL layer, vex handles VirusTotal enrichment on the IOCs extracted from the same alert data.


Try It

barb is available on PyPI. No API key required.

1
2
pip install barb-phish
barb analyze https://suspicious-site.tk/paypal-login

With OSINT enrichment:

1
barb analyze https://suspicious-site.tk/paypal-login --osint

With LLM explanation support:

1
2
pip install barb-phish[llm]
barb analyze https://pаypal.com --explain

Source code and documentation: github.com/duathron/barb


References

This post is licensed under CC BY 4.0 by the author.