Gutenberg — Public Domain Book Toolkit

SkillSearch

Search, download, and extract public-domain books from Project Gutenberg. Look up books by ID or keyword via gutendex, download plain-text and EPUB editions, strip licensing boilerplate, extract clean text from EPUB for illustrated works, and classify fiction vs non-fiction. Ships a portable CLI script with zero external dependencies. Use when the user says "gutenberg", "public domain", "download a book", "classic literature", "free ebook", "gutenberg.org", or names any public-domain title or author. Do not use this skill for unrelated requests; route to the nearest named specialist.

Available today. Use it from your connected AI after setup.

Connect ahel once, and every AI you use reads what you have installed.

Then ask your AI: use the Gutenberg — Public Domain Book Toolkit skill

What this skill tells your AI

The instructions your AI receives, as published by magnus919/agent-skills in gutenberg/SKILL.md and read by ahel’s review.

Search, download, and extract clean text from Project Gutenberg — 70,000+ free public-domain ebooks. Ships a portable Python CLI with zero external dependencies.

Quick Start

# Search for books
python3 scripts/gutenberg search "Moby Dick"

# Download by Gutenberg ID (plain text)
python3 scripts/gutenberg download 2701 --format txt

# Download EPUB (for illustrated books)
python3 scripts/gutenberg download 2701 --format epub

# Extract clean text (strips PG boilerplate)
python3 scripts/gutenberg extract 2701

# Classify fiction vs non-fiction
python3 scripts/gutenberg classify 2701

# Full pipeline: search → download → extract
python3 scripts/gutenberg pipeline "Alice's Adventures in Wonderland"

How It Works

Project Gutenberg provides 70,000+ free public-domain ebooks in multiple formats. The gutendex API (https://gutendex.com) offers a free, unauthenticated JSON catalog. No API key required — just curl or this CLI.

Data Flow

User provides title/ID/author
       ↓
gutendex API search → pick book by ID
       ↓
Download plain text (preferred) or EPUB (fallback for illustrated books)
       ↓
Strip PG boilerplate → clean text
       ↓
Classify fiction/non-fiction → extract content

CLI Reference

search — Find books by keyword

python3 scripts/gutenberg search "Moby Dick"
python3 scripts/gutenberg search "Dracula" --limit 5
python3 scripts/gutenberg search "Sherlock Holmes" --json
python3 scripts/gutenberg search "Alice" --language en

Returns: ID, title, author (with life dates), language, subjects, download count. Results sorted by download count (most popular first).

metadata — Get full metadata for a book by ID

python3 scripts/gutenberg metadata 2701          # Moby Dick
python3 scripts/gutenberg metadata 11            # Alice's Adventures
python3 scripts/gutenberg metadata 1342          # Pride and Prejudice
python3 scripts/gutenberg metadata 1342 --json   # JSON-only output

Returns: title, author(s), language(s), subjects, bookshelves, summaries, copyright status, download count, and all available format URLs.

download — Download a book by Gutenberg ID

# Plain text (UTF-8, preferred — works for most books)
python3 scripts/gutenberg download 2701 --format txt

# EPUB with images (for illustrated/scientific books)
python3 scripts/gutenberg download 2701 --format epub

# HTML (alternative fallback)
python3 scripts/gutenberg download 2701 --format html

# Specify output directory
python3 scripts/gutenberg download 2701 --format txt --output ./books/

The file is saved to ./gutenberg-<id>.<ext> (or --output path). Large books may take a moment.

extract — Strip PG boilerplate and produce clean text

python3 scripts/gutenberg extract 2701            # from downloaded txt
python3 scripts/gutenberg extract 2701 --input ./gutenberg-2701.txt
python3 scripts/gutenberg extract 2701 --format epub  # extract from EPUB

Output: clean text without the Project Gutenberg license header/footer. For EPUB extraction (illustrated books), extracts text from all XHTML files and merges them into a single cleaned document.

Size detection: if a plain-text download is under 50KB for a known substantial book, warns that the text may be truncated and recommends EPUB mode.

classify — Classify fiction vs non-fiction

python3 scripts/gutenberg classify 2701
python3 scripts/gutenberg classify 2701 --json

Uses the book's subjects and bookshelves to classify:

  • Fiction signals: "Fiction", "novels", "short stories", "poetry", "drama", "fantasy", "horror"
  • Non-fiction signals: "Essays", "History", "Philosophy", "Biography", "Science", "Religion"

Returns: fiction, non-fiction, or ambiguous (with explanation of why).

pipeline — Full fetch pipeline

python3 scripts/gutenberg pipeline "Moby Dick"                           # search first
python3 scripts/gutenberg pipeline 2701                                   # by known ID
python3 scripts/gutenberg pipeline 2701 --clean /tmp/pipeline-output/     # save cleaned text

Runs: search (if title) → metadata → download (txt) → check size → extract (or EPUB fallback) → classify. Prints an executive summary at the end.

Global Flags

FlagEffect
--jsonOutput machine-readable JSON instead of human-readable text
--quietSuppress diagnostic output
--dry-runShow what would be done without executing
--output ./dirSave downloads to a specific directory
--timeout 30Override API timeout (default 15s)

Fiction vs Non-Fiction Handling

When the classified result is fiction, the extracted text comes from an authored imagination. Consider splitting analysis into two tracks:

TrackWhat it coversExample claims
CanonFacts within the fictional world — named entities, quoted lines, story events, world rules"In Stoker's text, Dracula can assume wolf, bat, and mist forms"
CraftReal-world technique — how the author achieves effect, publication history, literary influence"Stoker's epistolary form forces the reader to piece together the narrative like an investigator"
Negative spaceDeliberate omissions — what the author notably leaves unspecified"Dracula is never granted interior voice in the novel"

When classified as non-fiction, claims can be treated as real-world factual assertions about the subject matter.

Known Gotchas

  • Plain text truncation for illustrated books — Books with diagrams, figures, or equations (geometry texts, scientific works, art books) may have plain-text downloads silently cut to 5-10KB (just the PG header). Always check file size. Under 50KB for a known substantial book → switch to EPUB extraction. The pipeline command does this check automatically.
  • Gutendex can be slow or timeout — The API is a free service and can be slow for less popular books. The CLI uses a 15-second default timeout. Use --timeout 30 for slow responses, or navigate directly to https://www.gutenberg.org/ebooks/<id> as a fallback.
  • HTML downloads include navigation markup — HTML downloads contain site navigation and formatting. Prefer plain text or EPUB for clean text extraction.
  • Rare books may 404 on certain format URLs — Not every book has every format. The CLI tries UTF-8 plain text first, falls back to US-ASCII, then to the -0.txt file path, then to EPUB, then to HTML. The download command reports which format was actually retrieved.
  • Rate limiting — Gutendex is unauthenticated but rate-limited. Batch requests with sleep 1 between calls for more than 10 rapid-fire requests.
  • utf-8 vs us-ascii — Gutendex returns both a text/plain; charset=utf-8 and a text/plain; charset=us-ascii URL. Prefer UTF-8; fall back to US-ASCII if the UTF-8 URL returns a 404.
  • Fiction classification ambiguity — Books with both fiction and non-fiction subjects (e.g. "Historical Fiction" + "History") are marked ambiguous. Use --json to inspect the subject list and decide manually.

References

  • scripts/gutenberg — Portable Python CLI. Zero external dependencies (stdlib only). Covers all major Gutenberg workflows: search, download (txt/epub/html), boilerplate stripping, EPUB text extraction, fiction classification, and the full pipeline.
  • references/epub-extraction.md — EPUB text extraction details for illustrated books, with expanded Python walkthrough and format detection tips.
  • Project Gutenberg — 70,000+ free ebooks.
  • Gutendex API — JSON web API for the Project Gutenberg catalog.

Signals

GitHub stars
78
Forks
8
Last commit
Sep 2026
Advanced
Catalog kind
skill
Gateway key
gutenberg
Source
github.com/magnus919/agent-skills