research-and-ingest
SkillDocs & knowledgeFind authoritative sources on the web (or via documentation MCP servers like Context7/DeepWiki), download or snapshot them when allowed, normalize them into clean Markdown, and stage them for the library-wiki. Use whenever a new library/spec/API/standard is being added, a wiki page is stale, an ADR needs current evidence, or a behavior best practice needs verification. Always prefer official upstream sources and version-pinned content over training data.
Use research-and-ingest in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add research-and-ingest and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the research-and-ingest skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by llopresto87/cypress in skills/research-and-ingest/SKILL.md and read by ahel’s review.
This skill is invoked by research-scout to fetch external content in a
disciplined way. The output of this skill is two artifacts: raw
snapshots in docs/graph/sources/raw/ (when license allows) and
normalized summaries in docs/graph/sources/normalized/, plus a row in
docs/graph/sources/index.md. The wiki page (created by library-wiki)
draws from these.
When to apply this skill
- A new dependency is being evaluated or added.
- A wiki page's "Last reviewed" date is older than the project's review cadence, or its pin no longer matches the lockfile.
- An ADR is being written and needs current evidence.
- An LLM/VLM feature is being designed and needs the provider's current behavior documentation.
- A spec is being authored and refers to a standard (RFC, schema, protocol) the project hasn't read recently.
Source ranking
When two sources disagree, prefer in this order:
- Official upstream documentation for the exact version in use.
- Official upstream source code (especially public API surface, examples directory, CHANGELOG).
- Official upstream blog posts and migration guides.
- Security advisories from trusted bodies (CVE, CISA, OWASP, official upstream advisories).
- Well-maintained community resources with current dates.
- Recent blog posts from credible authors.
- Anything else, marked clearly with reliability
communityormirror.
Verify a forum answer older than a year against current docs before citing it for a fast-moving library, and test any forum answer before it enters the wiki.
Workflow (per source)
1. Identify
For each source you intend to ingest, note:
- Authority (who maintains it).
- Version coverage (which versions of the library/spec it covers).
- Date (when it was last updated upstream).
- License (whether snapshotting is allowed).
- Slug (the filename you'll use locally).
2. Fetch
Use the host tool's web-fetch capability. Ingest paywalled or
login-walled content only with the user's explicit OK, because its
access terms bind the project. If a documentation MCP server is
configured (Context7, DeepWiki, llms.txt provider, similar), prefer it
for fast, version-aware retrieval.
3. Snapshot (when allowed)
When the license permits, write the raw content to
docs/graph/sources/raw/<slug>-<retrieved-date>.<ext>. Acceptable
extensions: .md, .html, .pdf, .txt, .json. Keep the raw file
whole, to preserve provenance. When the license does not permit a
snapshot, link to the source and store no copy; when the host retrieved
through an MCP summary and holds no page to keep, there is none to
store. Either way, say so in the raw: line of the normalized metadata
block. A snapshot with neither the raw file nor the reason is
UNJUSTIFIED at the coverage gate: the reason recorded is what makes
the omission a decision instead of a habit.
4. Normalize
Produce docs/graph/sources/normalized/<slug>.md:
- Clean Markdown holding the upstream headings and the decision-relevant content only (navigation chrome, ads, tracking pixels and boilerplate footers go).
- At the top, the metadata block:
---
source-title: <title>
source-url: <url>
source-maintainer: <name>
source-version-coverage: <e.g. "1.4 - 1.6">
retrieved: YYYY-MM-DD
license: <SPDX or "see source">
raw: raw/<slug>-<retrieved-date>.<ext> | withheld — <why: the license, a host without fetch, an MCP summary with no page behind it>
reliability: official | community-trusted | community | mirror
---
5. Register
Add a row to docs/graph/sources/index.md:
| <title> | <url> | <maintainer> | YYYY-MM-DD | <version> | <reliability> | <relevance> | <one-line notes> |
6. Draft, then hand back
Draft the wiki page from the normalized sources (skill.library-wiki
says what each section owes) and end the turn with the handback payload
naming the page, the sources index rows, and tester as
recommended_next for the smoke test. The scout is a leaf; the
authoring-class docs-librarian finalizes the page in the close-out,
the phase order being ingest-library.flow's.
Documentation MCP servers (when available)
If the project has any of these configured, prefer them for fetching upstream content:
- Context7 (
@upstash/context7-mcp): current docs for many libraries, addressable by library ID and version. Use the library-ID form for precision. - DeepWiki: open-source repository summaries.
llms.txtproviders: projects that publish a machine-readable docs index.
When you use one, cite the source in the normalized file's metadata with the MCP server name and the date.
The local wiki is still authoritative for the project. The MCP server gets you upstream content faster; the wiki page is your distillation of what this project actually does with that content.
Disagreement handling
If two sources disagree:
- Note the version coverage of each.
- Prefer the more recent official source.
- If a security advisory disagrees with the docs, the advisory wins.
- If the disagreement persists, record both with their versions in the wiki page, and open a question in grill.md §12.
For a non-trivial topic, cross-check against the upstream source code as well, because one source per topic leaves a disagreement nobody can see.
Source reconciliation (lightweight drift check)
Between full research passes, run a cheap periodic reconciliation: diff the currently-resolved dependency versions and manifests/locks against the versions recorded in the library wiki, using only what already resolves locally. Classify each line:
- no mismatch: the wiki pin still matches what resolves; nothing to do.
- refresh before the next API-affecting change: the pin has drifted but no work is about to touch that surface; flag it, and re-ingest when work next touches that surface.
- superseded — treat as historical: the recorded version is gone from the resolved set; mark the wiki content as historical.
This catches silent pin drift that accumulates between full passes, at a
fraction of the cost. It is distinct from a full research pass (which
re-fetches and re-normalizes upstream content) and from
validate-knowledge (which tests whether the wiki prose is navigable
and correct, not whether its pins are still current).
Reference files
docs/graph/templates/library-page.template.md: where this skill's output ultimately lands.docs/graph/protocols/ingest-library.md: the parent protocol.docs/graph/agents/10-research-scout.md: the agent that runs this.docs/graph/skills/library-wiki.md: the wiki-maintenance skill.
Signals
- GitHub stars
- 96
- Forks
- 4
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
research-and-ingest- Source
- github.com/llopresto87/cypress
github.com/llopresto87/cypress