Vector Index Tuning
SkillSearchVector-index-tuning is a skill that guides an AI agent through optimizing vector index performance for latency, recall, and memory. It helps when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.
Use Vector Index Tuning in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Vector Index Tuning and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Vector Index Tuning skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Have an AI agent that can load skills.
What your AI can do with it
- Tune HNSW parameters to adjust the latency-recall tradeoff
- Select quantization strategies to reduce memory use
- Optimize vector index performance for lower latency
- Improve recall of vector search results
- Guide scaling of vector search infrastructure
Getting started
- Have an AI agent that can load skills.
- Add the vector-index-tuning skill to the agent's available skills.
- Ask the agent to tune a vector index, describing the latency, recall, or memory goals.
- Apply the agent's recommended parameter and quantization settings to the vector search setup.
What this skill tells your AI
The instructions your AI receives, as published by wshobson/agents in plugins/llm-application-dev/skills/vector-index-tuning/SKILL.md and read by ahel’s review.
Guide to optimizing vector indexes for production performance.
When to Use This Skill
- Tuning HNSW parameters
- Implementing quantization
- Optimizing memory usage
- Reducing search latency
- Balancing recall vs speed
- Scaling to billions of vectors
Core Concepts
1. Index Type Selection
Data Size Recommended Index
────────────────────────────────────────
< 10K vectors → Flat (exact search)
10K - 1M → HNSW
1M - 100M → HNSW + Quantization
> 100M → IVF + PQ or DiskANN
2. HNSW Parameters
| Parameter | Default | Effect |
|---|---|---|
| M | 16 | Connections per node, ↑ = better recall, more memory |
| efConstruction | 100 | Build quality, ↑ = better index, slower build |
| efSearch | 50 | Search quality, ↑ = better recall, slower search |
3. Quantization Types
Full Precision (FP32): 4 bytes × dimensions
Half Precision (FP16): 2 bytes × dimensions
INT8 Scalar: 1 byte × dimensions
Product Quantization: ~32-64 bytes total
Binary: dimensions/8 bytes
Templates and detailed worked examples
Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.
Best Practices
Do's
- Benchmark with real queries - Synthetic may not represent production
- Monitor recall continuously - Can degrade with data drift
- Start with defaults - Tune only when needed
- Use quantization - Significant memory savings
- Consider tiered storage - Hot/cold data separation
Don'ts
- Don't over-optimize early - Profile first
- Don't ignore build time - Index updates have cost
- Don't forget reindexing - Plan for maintenance
- Don't skip warming - Cold indexes are slow
Signals
- GitHub stars
- 40k
- Forks
- 4k
- Last commit
- Sep 2026
Questions
- When should this skill be used?
- Use it when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.
- What does it optimize for?
- It optimizes vector index performance for latency, recall, and memory.
- Does it change my vector database directly?
- The item does not say. It guides the agent in tuning vector search settings; how changes are applied is not specified.
Advanced
- Item type
- skill
- Key
vector-index-tuning- Source
- github.com/wshobson/agents