Keep HTML documents under crawl limits
SkillWeb & browsinghtml-size is a skill that lets an AI agent audit HTML page sizes against Google's crawl limits. It measures raw HTML responses, flags pages over 2MB or critically over 5MB, explains common causes of oversized pages, and walks through fixes. It suits anyone auditing page weight for crawl efficiency or investigating why content is missing from Google's index.
Use Keep HTML documents under crawl limits in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Keep HTML documents under crawl limits and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Keep HTML documents under crawl limits skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have an AI agent that can load skills.
What your AI can do with it
- Measure raw HTML response sizes for individual pages
- Flag pages over 2MB and critically over 5MB
- Identify causes such as large inline JSON like __NEXT_DATA__, embedded SVGs, and base64
- Guide fixes including moving assets to external files
- Recommend enabling compression and minifying output
- Support reviewing server-rendered pages that embed large JSON payloads in HTML
Getting started
- Have an AI agent that can load skills.
- Add the html-size skill to the agent's available skills.
- Give the agent a page or set of pages to audit and ask it to check HTML sizes against Google's crawl limits.
- Review the flagged pages and apply the fixes the agent suggests, such as externalizing assets or enabling compression.
What this skill tells your AI
The instructions your AI receives, as published by thedaviddias/front-end-checklist in skills/html-size/SKILL.md and read by ahel’s review.
Googlebot has a documented crawl size limit of approximately 15MB per HTML document. Content beyond this threshold is not parsed or indexed. Excessively large HTML also slows Googlebot crawls, reducing how many of your pages are crawled per budget period.
Quick Reference
- Googlebot stops parsing HTML beyond approximately 15MB; content after that point is not indexed
- Large HTML is usually caused by inline JSON data dumps, excessive inline SVG, or unminified JavaScript in
<script>tags - Target HTML document size under 2MB for optimal crawl efficiency; investigate anything over 5MB
Check
Measure the raw HTML response size (before compression) for each page. Flag pages over 2MB (investigate) and over 5MB (critical). Identify the cause of oversized HTML: (1) Large inline JSON (<script id='__NEXT_DATA__'> or similar). (2) Inline SVG files. (3) Base64-encoded images in HTML. (4) Inline CSS with large amounts of utility classes. (5) Unminified scripts in <script> tags.
Fix
- Measure:
curl -so /dev/null -w '%{size_download}' https://yoursite.com/page | awk '{print $1/1024 " KB"}' - If large inline JSON is the cause (common in Next.js
__NEXT_DATA__):- Reduce data passed to
getServerSideProps/getStaticProps— only pass what the page renders - Use React Server Components (Next.js 13+) to avoid client hydration payloads
- Reduce data passed to
- If inline SVG is the cause: move SVGs to external files and load with
<img>or<use>. - If base64 images are the cause: serve images from a CDN and reference via URL.
- Enable gzip/Brotli compression on the server — Googlebot fetches the compressed response.
- Minify HTML output in production (remove whitespace and comments).
Explain
Google's crawl infrastructure parses only the first ~15MB of an HTML document. Pages that exceed this limit have their tail content silently omitted from Google's index. Beyond the hard limit, large HTML documents consume more crawl budget, meaning fewer of your pages are crawled per day. This particularly affects large e-commerce sites or pages that server-render large datasets into HTML.
Code Review
Check the response Content-Length header or measure the raw HTML byte count. Inspect <script type='application/json'> or <script id='__NEXT_DATA__'> blocks — count their size in bytes. Flag any single block over 500KB. Check for inline SVG elements (look for <svg> in body HTML) that should be external files. Verify HTML is served with gzip or Brotli encoding.
For full implementation details, code examples, and framework-specific guidance,
see references/rule.md.
Rule page: https://frontendchecklist.io/en/rules/seo/html-size
Signals
- GitHub stars
- 74k
- Forks
- 7k
- Last commit
- Oct 2026
Questions
- When should this skill be used?
- Use it when auditing page weight for crawl efficiency, investigating why certain page content is not appearing in Google's index, or reviewing server-rendered pages that embed large JSON payloads in HTML.
- What page sizes does it flag?
- It flags pages over 2MB and marks pages critically over 5MB, based on the raw HTML response size.
- What causes of oversized pages does it identify?
- Common causes include large inline JSON such as __NEXT_DATA__, embedded SVGs, and base64 images.
- What fixes does it suggest?
- It walks through moving assets to external files, enabling compression, and minifying output.
Advanced
- Item type
- skill
- Key
html-size- Source
- github.com/thedaviddias/front-end-checklist
github.com/thedaviddias/front-end-checklist
Related picks
Skill · aaron-he-zhu
The pick for SEOseo-audit
Skill · aeonfun
The pick for SEObrowser-use
Skill · browser-use
More in Web & browsingwebapp-testing
Skill · anthropics
More in Web & browsingplaywright-cli
Skill · microsoft
More in Web & browsingbenchmark
Skill · affaan-m
More in Web & browsing