Read URL

SkillWeb & browsing

Extract clean, complete markdown from any web page, articles, docs, READMEs, blog/social posts, academic papers. Also use as a fallback when curl returns noisy HTML or WebFetch returns truncated, summarized, or refused results.

Use Read URL in Claude, ChatGPT or Ahel Desktop

Free. Sign in, add Read URL and connect your AI. About a minute.

Also: Claude Code · Cursor · Codex

Then ask your AI: use the Read URL skill

Details

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Read URLStart free

What this skill tells your AI

The instructions your AI receives, as published by archibate/dotfiles-claude in skills/read-url/SKILL.md and read by Ahel’s review.

Work down this fallback ladder in order. Each step is only tried when prior steps don't apply or fail.

Fallback ladder

  1. Raw .md / .txt / plain-text URL → curl -sL <url> (already clean, no HTML to strip)
  2. Known site → use the dedicated CLI/API from the routing table below
  3. Docs page → try curl -sL <url>.md. Mintlify and other docs platforms serve clean markdown on the .md route — if the response is text/markdown, you're done; otherwise fall through
  4. Blog / newsletter / multi-post index → try RSS first: curl -sL <url>/feed (also /rss, /feed.xml, /atom.xml, /index.xml). Most static-site generators and CMS platforms expose one; RSS gives you clean <content:encoded> or <summary> bodies without chrome
  5. Generic site (articles, docs, tech blogs, unknown) → npx defuddle parse <url> --markdown — see references/defuddle.md
  6. JS-rendered page (defuddle returns empty / skeleton-only content) → /agent-browser skill
  7. Cloudflare / anti-bot protection (Turnstile, blocked responses, 403/503) → /scrapling skill
  8. Still blocked and genuinely need this page → ask the user to open it and paste the content, or offer the /chrome-cdp skill (requires explicit user approval first). Otherwise, give up and report the failure.

Routing table

Step 2 — URLs matching a known domain:

Domain / PatternPreferred path
github.com / gist.github.comFile via raw.githubusercontent.com; issue/PR via gh issue view / gh pr view; search: api.github.com/search/{code,issues,repositories}?q= (anonymous) — see references/github.md
x.com / twitter.com / t.cocurl -sL https://api.fxtwitter.com/<user>/status/<id> | jq
bilibili.combilibili_api Python library (uv run --with bilibili-api-python): sync(video.Video(bvid="BV...").get_info()) for title/description; comment module for comments
youtube.com / youtu.beyt-dlp --dump-json --skip-download for title/description/metadata; yt-dlp --write-auto-sub --sub-lang en --skip-download for transcript
arxiv.org / ssrn.com/jina-ai skill
mp.weixin.qq.com (微信公众号)/scrapling skill — scrapling extract get <url> works without a browser
www.cnblogs.com (博客园)Plain defuddle works — server-rendered HTML with the article body inline. For a user's post index: curl -sL 'https://www.cnblogs.com/<user>/rss' (Atom feed)
blog.csdn.net (CSDN)/scrapling skill — plain curl returns a JS-skeleton (content is JS-loaded) and defuddle hits 404 anti-bot. For a summary-only index: curl -sL 'https://blog.csdn.net/<user>/rss/list' returns RSS with 摘要 (not full bodies)
zhihu.com / zhuanlan.zhihu.com (知乎)scripts/fetch_zhihu.py <url> — see references/zhihu.md
juejin.cn (掘金)/scrapling skill — Nuxt SPA; escalate to /chrome-cdp if stealthy-fetch returns only shell
segmentfault.com (思否)/scrapling skill — custom HTTP 468 anti-bot; escalate to /chrome-cdp if stealthy-fetch fails
weibo.com (微博)/scrapling skill — JS-rendered status pages; escalate to /chrome-cdp if stealthy-fetch returns only chrome
xiaohongshu.com (小红书)/scrapling skill — aggressive anti-bot; escalate to /chrome-cdp if stealthy-fetch fails
douban.com / movie.douban.com (豆瓣)Desktop returns an anti-bot 载入中… shell. Use the mobile host with an iPhone UA: curl -sL -A 'Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.0 Mobile/15E148 Safari/604.1' 'https://m.douban.com/movie/subject/<id>/' | defuddle parse --markdown — full server-rendered body (评分, 简介, 影评). Fall back to /scrapling if blocked
y.qq.com (QQ 音乐)Hard — stealthy-fetch returns the homepage shell instead of song data. Use /chrome-cdp with the user's logged-in session, or ask them to paste
music.163.com (网易云音乐)Plain defuddle for basic info — <title> has song + artist. For lyrics / comments / playlists use the community-maintained NeteaseCloudMusicApi (self-hosted Node proxy over the internal API)
wallstreetcn.com (华尔街见闻)Plain defuddle works — server-rendered with _articleBody_… class; no auth needed for public articles
www.v2ex.com (V2EX)curl -sL 'https://www.v2ex.com/api/topics/show.json?id=<id>' | jq — returns topic + full content; api/replies/show.json?topic_id=<id> for replies
gitee.comKnown file path: curl -sL 'https://gitee.com/<owner>/<repo>/raw/<ref>/<path>'. Repo metadata: curl -sL 'https://gitee.com/api/v5/repos/<owner>/<repo>' | jq. Shape mirrors GitHub
instagram.cominstaloader CLI
reddit.comHard — .json endpoints are blocked since the 2023 API changes, and scrapling's stealthy-fetch gets a captcha page. Use the official OAuth API (PRAW / snoowrap) with credentials, or /chrome-cdp with the user's logged-in session
stackoverflow.com / *.stackexchange.com / superuser.com / serverfault.com / askubuntu.comStack Exchange API; search: /2.3/search/advanced?intitle=<q> or ?q=<q> — see references/stackexchange.md
*.fandom.com/scrapling skill — Fandom sits behind Cloudflare, plain curl returns the "Just a moment..." challenge regardless of path or User-Agent
Any other MediaWiki site — Wikipedia, Arch Wiki, cppreference, *.wiki.gg, etc.Wikimedia-run wikis use the REST API + prop=extracts; third-party wikis use ?action=raw or api.php?action=parse; search: api.php?action=query&list=search&srsearch=<q> (or action=opensearch) — see references/mediawiki.md
www.rfc-editor.org / any RFCcurl -sL 'https://www.rfc-editor.org/rfc/rfc<N>.txt' — canonical plaintext, no chrome. .html and .json also available (the JSON has metadata like obsoleted-by, authors, status)
peps.python.orgIndividual PEP: curl -sL 'https://peps.python.org/pep-<N>/' (clean HTML). All PEPs indexed: curl -sL 'https://peps.python.org/api/peps.json' | jq — number, title, status, authors, created date
docs.claude.com / docs.anthropic.com (Anthropic & Claude Code docs)Append .md to the URL. Indexes: platform.claude.com/llms.txt (+ llms-full.txt) for API/SDK pages; code.claude.com/llms.txt for Claude Code
news.ycombinator.comcurl -sL 'https://hn.algolia.com/api/v1/items/<id>' | jq — returns story + full comment tree as nested JSON; search: hn.algolia.com/api/v1/search?query=<q> (and /search_by_date)
pypi.orgcurl -sL 'https://pypi.org/pypi/<package>/json' | jq -r '.info.description' for README; .info.summary / .info.version for metadata
npmjs.com / registry.npmjs.orgnpm view <package> readme for README; curl -sL 'https://registry.npmjs.org/<package>' | jq for full metadata; search: registry.npmjs.org/-/v1/search?text=<q>
lobste.rsappend .json to the story URL (e.g. lobste.rs/s/<id>.json), fetch with curl
dev.tocurl -sL 'https://dev.to/api/articles/<id>' | jq -r '.title, .body_markdown' — <id> is the numeric article ID
*.substack.com<subdomain>.substack.com/feed — RSS with full post HTML in <content:encoded>
medium.com / *.medium.comcurl -sL 'https://medium.com/feed/@<user>' — RSS returns the last ~10 posts with full content:encoded HTML. Direct article URLs return a ~4KB paywall shell and need /scrapling if the piece isn't in the user's recent feed
bsky.appcurl -sL 'https://public.api.bsky.app/xrpc/app.bsky.feed.getPostThread?uri=<at-uri>' | jq — no auth needed for public posts; convert bsky.app/profile/<handle>/post/<rkey> to at://<handle>/app.bsky.feed.post/<rkey>
gitlab.comKnown file path (preferred): curl -sL https://gitlab.com/<owner>/<repo>/-/raw/<ref>/<path>. Repo metadata / MR / issue bodies: curl -sL 'https://gitlab.com/api/v4/projects/<owner>%2F<repo>' | jq (URL-encode the slash in the project path); search: api/v4/search?scope=projects&search=<q>
codeberg.org / any Gitea or Forgejo instanceKnown file path: curl -sL https://codeberg.org/<owner>/<repo>/raw/branch/<ref>/<path>. Metadata: curl -sL 'https://codeberg.org/api/v1/repos/<owner>/<repo>' | jq
crates.iocurl -sL 'https://crates.io/api/v1/crates/<crate>' | jq for metadata; .../<version>/readme for README
formulae.brew.sh / any brew formulacurl -sL 'https://formulae.brew.sh/api/formula/<name>.json' | jq — name, desc, versions, deps, caveats
aur.archlinux.orgcurl -sL 'https://aur.archlinux.org/rpc/v5/info/<pkg>' | jq -r '.results[0]' — Name, Version, Description, Maintainer, Depends, URL; search: rpc/v5/search/<pkg>
doi.org / any bare DOIcurl -sL 'https://api.crossref.org/works/<doi>' | jq -r '.message | .title[0], (.author[].family | tostring)' — reliable for title + authors + citation metadata (abstract hit-or-miss). Prefer /jina-ai or the publisher page for full text
Any Discourse forum (discuss.python.org, meta.discourse.org, forum.rust-lang.org, discuss.pytorch.org, etc.)Append .json to the topic URL: curl -sL '<forum>/t/<slug>/<id>.json' | jq — returns topic + all posts in post_stream.posts; search: <forum>/search.json?q=<q>
huggingface.coREADME via <repo>/raw/main/README.md; metadata via /api/models, /api/datasets, /api/papers; search: /api/models?search=<q> (and /api/datasets?search=) — see references/huggingface.md
web.archive.org / any Wayback lookupFind closest snapshot: curl -sL 'https://archive.org/wayback/available?url=<url>&timestamp=<YYYYMMDD>' | jq. Fetch raw archived response: curl -sL 'https://web.archive.org/web/<timestamp>id_/<url>' — the id_ suffix strips Wayback's toolbar injection and returns the original response body
store.steampowered.comcurl -sL 'https://store.steampowered.com/api/appdetails?appids=<appid>&cc=us&l=en' | jq -r '.["<appid>"].data' — name, short_description, release_date, developers, categories, price
speedrun.comcurl -sL 'https://www.speedrun.com/api/v1/games/<slug>' | jq — game metadata; further endpoints at /games/<id>/categories, /runs?game=<id> for leaderboards
Any WordPress site (self-hosted or *.wordpress.com)curl -sL '<site>/wp-json/wp/v2/posts?per_page=10' | jq for recent posts; /wp-json/wp/v2/posts/<id> for a single post (.content.rendered has the HTML body); search: wp-json/wp/v2/posts?search=<q>. Works on any WP install with the REST API enabled — still the default
openlibrary.orgcurl -sL 'https://openlibrary.org/works/OL<id>W.json' | jq for works; /isbn/<isbn>.json for ISBN lookup; /authors/OL<id>A.json for authors. Note: .description is sometimes a string, sometimes {type, value} — handle both
gutenberg.org (Project Gutenberg)curl -sL 'https://www.gutenberg.org/cache/epub/<id>/pg<id>.txt' — full plaintext of out-of-copyright books

Rows tagged search: expose a dedicated search API for when you have a topic, not a URL. Otherwise run WebSearch or /jina-ai skill with a site: filter, then fetch the result URL via this ladder.

Bulk discovery

For whole-site ingestion, probe <site>/llms.txt (URL index) and /llms-full.txt (full corpus). Convention adopted by Mintlify, Cloudflare, Stripe, Next.js, and others. On 404, fetch the index page <site>/ instead.

vs. WebFetch

This skill returns full page text (markdown), parsed locally — no summarization, no information loss. WebFetch routes through a remote small model that may summarize, refuse, or truncate; reach for it only when you want an AI summary, not the content itself.

When to bypass the ladder

  • Need a quick AI summary → built-in WebFetch
  • No specific URL yet, need to search → built-in WebSearch or /jina-ai skill

Signals

GitHub stars
40
Forks
11
Last commit
Sep 2026
Advanced
Item type
skill
Key
read-url
Source
github.com/archibate/dotfiles-claude