fetch-url-as-markdown

$npx mdskill add CodeAlive-AI/ai-driven-development/fetch-url-as-markdown

Fetch a web page and return clean Markdown via trafilatura or Exa MCP

  • The user asks to fetch and display content from a given URL in Markdown format.
  • This skill depends on local installation of trafilatura for processing the fetched content, and Exa MCP as a fallback for complex pages.
  • It decides whether to use trafilatura or Exa MCP based on the exit code of the trafilatura command. If not installed, it installs trafilatura.
  • The skill delivers the cleaned Markdown content directly to the user.

SKILL.md

.github/skills/fetch-url-as-markdownView on GitHub ↗
---
name: fetch-url-as-markdown
description: Fetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages. Use when the user asks to read, fetch, scrape, summarize, or quote a URL — prefer this over the built-in WebFetch tool. Don't use for binary files (PDFs, images, archives) or for fetching API/JSON endpoints.
---

# URL to Markdown

Fetch any web URL and get clean, readable Markdown — main content only, no
navigation/footer/ads. Local + free by default; smart fallback to Exa MCP
when the page can't be extracted locally.

## Workflow (the only thing the agent needs to remember)

1. **Try trafilatura first**:

   ```bash
   python3 ~/.claude/skills/fetch-url-as-markdown/scripts/fetch_url.py "<URL>"
   ```

2. **If exit code is 1 or 2 → fall back to Exa MCP** with the same URL:

   ```
   mcp__exa__web_search_advanced_exa(
       query="<URL>",
       includeDomains=["<host of URL>"],
       numResults=1,
       textMaxCharacters=50000,
       type="auto"
   )
   ```

   (`mcp__exa__crawling` works too if the server exposes it; the `web_search_advanced_exa`
   call above is the always-available variant — pin the host with `includeDomains` and
   use the URL itself as the query.)

3. Exit code `3` means trafilatura is not installed — install once:

   ```bash
   python3 -m pip install --break-system-packages trafilatura
   ```

## Exit codes (what they mean for the fallback decision)

| Code | Meaning | Action |
|---|---|---|
| 0 | Markdown printed to stdout | done |
| 1 | DownloadError — network/HTTP/timeout/anti-bot block at fetch | fall back to Exa |
| 2 | ExtractionError — empty extract, JS/Cloudflare wall, or stub body (<200 chars) | fall back to Exa |
| 3 | trafilatura missing | install (see above), then retry |
| 4 | UnsupportedContentTypeError — URL is binary (PDF, image, archive) | **don't** fall back to Exa; use the right specialized skill (e.g. `pdf` for PDFs) |

## Defaults baked into the script

- `output_format="markdown"`, `include_formatting=True` — keeps headings/lists/code structure where the source HTML uses real `<h1..h6>` etc.
- `include_links=True`, `include_tables=True`
- `with_metadata=True` → emits a YAML frontmatter (`title`, `author`, `date`, `url`, `hostname`)
- `favor_recall=True`, `deduplicate=True` — readable but trims duplicates
- Real-browser User-Agent + 30s timeout configured in `scripts/settings.cfg`
- Anti-stub guards (built into the script):
  - rejects `Content-Type` other than `text/html|application/xhtml+xml|text/plain|application/xml|text/xml` → exit `4`
  - sniffs raw HTML for Cloudflare / "Please enable JavaScript" / Imperva / DataDome wall markers → exit `2`
  - rejects extracted bodies under 50 chars (configurable via `--min-body N`, `0` to disable) → exit `2`

## Useful flags

```bash
... fetch_url.py "<URL>" --no-links     # strip hyperlinks
... fetch_url.py "<URL>" --no-tables    # strip tables
... fetch_url.py "<URL>" --no-metadata  # omit YAML header
... fetch_url.py "<URL>" --comments     # include user comments (off by default — usually noise)
... fetch_url.py "<URL>" --images       # include image refs (experimental)
... fetch_url.py "<URL>" --precision    # terser output, drops borderline content
```

## When to choose what

| Situation | Tool |
|---|---|
| Article, blog post, docs, README, wiki | trafilatura (default) — local, free |
| JS-heavy SPA, login-walled, Cloudflare | Exa fallback (the script will signal exit 2) |
| Bulk / many URLs | trafilatura — no quota, no API key |
| Already failed twice on a domain | Exa directly |

More from CodeAlive-AI/ai-driven-development

SkillDescription
agentic-readinessAudit and improve repositories for reliable agentic work across Codex and Codex App, Claude Code, and OpenCode. Use when reviewing AGENTS.md or CLAUDE.md quality and discovery, instruction routing in monorepos or meta-repos, agent settings, MCP configuration, skills, subagents, context budgets, or repository organization for coding agents.
agents-consiliumQuery external AI agents (Codex, Gemini, OpenCode, Claude Code headless) in parallel for independent second opinions, code review, bug investigation, and consensus on high-stakes decisions. Agents and models are configurable in config.json. Use for architecture choices, security review, or ambiguous problems where independent perspectives matter. Not for simple questions answerable from docs or the codebase — use web search or repo exploration instead.
bug-fix-protocol8-step disciplined bug-fix protocol that treats every production bug as two failures — the code defect itself and the testing system that allowed it through. Use when fixing a production bug, investigating a regression, writing a post-mortem, or auditing a missed defect. Triggers on "fix this bug", "production bug", "regression test", "post-mortem", "test gap", "why did the tests miss this".
code-that-fits-in-your-headSoftware engineering heuristics from Mark Seemann's book Code That Fits in Your Head (2021). Use when writing new code, reviewing code, refactoring, designing APIs, handling validation and invariants, writing unit tests, debugging defects, performing security review (STRIDE), or setting up a new code base. Covers decomposition (cyclomatic complexity, 80/24 rule, cohesion, fractal architecture), encapsulation (invariants, parse-don't-validate, Postel's law), outside-in TDD (walking skeleton, AAA, triangulation, devil's advocate), API design (affordance, poka-yoke, CQS), git/PR hygiene (50/72 commits, small commits, code review), feature flags, Strangler pattern, bisection debugging, logging with decorators, and STRIDE threat modelling. Not for language-specific syntax, framework tutorials, production incident response, or performance profiling.
fpf-problem-solvingFirst Principles Framework (FPF) — thinking amplifier. Use when user wants to think through a complex problem, architect a system, evaluate alternatives, decompose complexity, classify problems, define quality attributes, plan rigorously, apply an FPF pattern to a first useful result, decide under uncertainty, establish causality, reason about time and trends, describe or synthesize architecture, check mathematical model fit, distinguish relation kinds or occurrences, govern ontic/U-kind admission, publish multi-view artifacts, refresh SoTA packs, trace provenance, or improve pattern quality. Also triggers on: FPF, bounded contexts, SoTA packs, assurance calculus, decision theory, causal reasoning, temporal reasoning, architecture description, modularity, constraint-governed unfolding, narrative rendering, structural adequacy, cultural evolution, quality gates, lexical discipline, FPF Parts A-I. Not for simple task planning, general philosophy, or Agile unrelated to FPF.
hooks-managementManage hooks and automation for coding agents (Claude Code, Codex CLI, OpenCode). Use when users want to add, list, remove, update, or validate hooks. Triggers on requests like "add a hook", "create a hook that...", "list my hooks", "remove the hook", "validate hooks", or any mention of automating agent behavior with shell commands or plugins.
investigating-repository-historyInvestigate GitHub repository history before risky code changes using git blame/log, GitHub PRs, review comments, squash/rebase/cherry-pick/rename heuristics, and cited evidence. Use when asking why code exists, whether a change is safe, what PR introduced behavior, or before editing API, compatibility, security, concurrency, persistence, migration, or performance-sensitive code.
maintaining-macos-healthHands-on playbook for macOS disk cleanup, dev-machine optimization, and proactive health alerting. Use when the Mac is full or slow, when a process persistently burns CPU, when a kernel panic / watchdog timeout / vm-compressor-space-shortage / Jetsam event happened, when the user asks to free disk space, audit storage, set up disk/memory/CPU alerts, or restore the same monitoring on a new Mac. Built around Mole (`mo` CLI) for safety guards plus a custom LaunchAgent-based alerter for active warnings. Covers Apple Silicon laptops with heavy AI/Docker workloads. Not for general macOS support, hardware diagnostics, networking issues, GUI / window-manager bugs, Time Machine recovery, or broken app installs.
maintaining-windows-healthHands-on playbook for Windows 11 disk cleanup, dev-machine optimization, and proactive health alerting. Use when the PC is full or slow, when a BSOD / Kernel-Power 41 / crash dump / commit-memory pressure happened, when the user asks to free disk space, audit storage, set up disk/memory alerts, or restore the same monitoring on a new PC. Built around native Microsoft-supported tooling (Storage Sense, cleanmgr, DISM, pnputil, vssadmin, wevtutil, powercfg) as the safety floor, a drift-protected HTML cleanup UI, and a Task Scheduler + BurntToast alerter. Covers dev machines with heavy AI/Docker/WSL workloads. Not for general Windows support, hardware diagnostics, GPU/driver troubleshooting, antivirus/malware removal, Windows Update repair, networking, or app-specific performance problems unrelated to disk or memory pressure.
mcp-managementSearch, install, configure, update, and remove MCP servers across coding agents (Claude Code, Cursor, VS Code, Claude Desktop, Gemini CLI, Codex, Goose, Zed, and more). Supports multi-agent installation via npx add-mcp, the official MCP registry, and direct config editing.