Optimizing Parquet Storage
SkillDatabases & dataOptimize columnar Parquet storage for analytics, file and row-group sizing, compression codecs (Snappy/ZSTD), partitioning and file layout, column pruning and predicate pushdown, dictionary encoding, and fixing the small-files problem. Use when Parquet reads are slow or costly, files are too small/large, choosing compression or partitioning, or improving scan pruning on a data lake.
Use Optimizing Parquet Storage in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Optimizing Parquet Storage and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Optimizing Parquet Storage skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; Ahel provides instructions and does not run this skill.
No other account needed.
Add Ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by unknown-333/awesome-data-engineering-skills in skills/optimizing-parquet-storage/SKILL.md and read by Ahel’s review.
When to use
- Parquet/lake reads are slow or scan too much data.
- Files are too small (many tiny files) or too large (poor parallelism).
- Choosing compression, row-group size, or partition layout.
- Do NOT use for warehouse-native storage tuning (use the warehouse skills).
Workflow
- [ ] Target ~128MB-1GB files and ~128MB row groups
- [ ] Partition by common filter columns (low/medium cardinality)
- [ ] Sort within files by a filter column to tighten min/max pruning
- [ ] Pick compression: Snappy (speed) or ZSTD (ratio)
- [ ] Compact small files; select only needed columns
- Right-size files and row groups. Aim for ~128MB–1GB files and ~128MB row groups so engines get efficient parallelism and pruning. Tiny files kill performance via per-file overhead.
- Partition on filter columns of low/medium cardinality (date, region). Avoid high-cardinality partitioning (user_id) — it creates millions of tiny files.
- Sort within files by a frequently filtered column so Parquet's per-row-group min/max stats enable predicate pushdown (data skipping).
- Compression: Snappy for hot, latency-sensitive data; ZSTD for better ratio and cheaper storage at similar read speed.
- Read fewer columns — the biggest columnar win is column pruning.
Patterns
Write well-sized, sorted Parquet (Spark):
(df.sort("ordered_at") # tighten row-group min/max for pushdown
.repartition(1, "region") # control file count per partition
.write.partitionBy("region")
.option("compression", "zstd")
.parquet("s3://lake/orders/"))
Predicate pushdown — filtering on the sorted/partition column lets the reader skip row groups and partitions entirely, reading far fewer bytes.
Compact small files — periodically read a partition and rewrite it into a few large files (or use a table format's compaction) to fix the small-files problem from streaming/frequent appends.
Common pitfalls
- Small-files problem — frequent small appends create thousands of tiny files; compact on a schedule.
- High-cardinality partitioning — partitioning by id/timestamp-to-the-second explodes file count; partition coarsely, sort finely.
- No sorting — unsorted files have wide row-group min/max, so pushdown skips nothing.
SELECT *on wide tables — reads every column's bytes; project only needed columns.- Row groups too small/large — too small loses compression/pushdown efficiency, too large hurts parallelism and memory.
- Gzip for analytics — slow to decode; prefer Snappy or ZSTD.
Signals
- GitHub stars
- 21
- Last commit
- Aug 2026
Advanced
- Item type
- skill
- Key
optimizing-parquet-storage- Source
- github.com/unknown-333/awesome-data-engineering-skills
github.com/unknown-333/awesome-data-engineering-skills
Related picks
Skill · a5c-ai
The pick for Sparklark-apps
Skill · larksuite
The pick for Sparkanalytics
Skill · coreyhaines31
More in Databases & datasupabase
Skill · supabase
More in Databases & dataconnect
Skill · composiohq
More in Databases & dataazure-kusto
Skill · microsoft
More in Databases & data