Gemini (media generation)
SkillMediaThe gemini skill lets your agent generate images, videos, speech audio, and list available models.
Use Gemini (media generation) in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Gemini (media generation) and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Gemini (media generation) skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
About this skill
Google Gemini media generation: Nano Banana images, Imagen 4 images, Veo video, TTS, model listing. Trigger phrases: gemini image, nano banana, imagen, veo video, google veo, gemini tts.
What this skill tells your AI
The instructions your AI receives, as published by anil-matcha/awesome-muse-connectors in connectors/gemini/SKILL.md and read by ahel’s review.
Purpose
Generate images (Nano Banana, Imagen 4), video (Veo 3.1), and speech (TTS) through Google's Gemini API with one key and one bill. Use when Michael asks for AI-generated images, video clips, thumbnails, or voiceovers via Google.
Tooling
All commands go through bin/gemini.py:
bin/gemini.py auth # verify the API key (free)
bin/gemini.py models --limit 50 # list available models
bin/gemini.py image --prompt "a ceramic fox" --out /tmp/fox # Nano Banana text-to-image (or edit with --image)
bin/gemini.py image --prompt "make it night" --image ./fox.png --out /tmp/fox-night
bin/gemini.py imagen --prompt "studio product shot" --count 2 --out /tmp/prod # Imagen 4
bin/gemini.py video --prompt "a drone shot over a harbor" # Veo 3.1, prints an operation id
bin/gemini.py op-status --op operations/abc123 # poll the Veo operation
bin/gemini.py tts --text "Hello there" --voice Kore --out /tmp/line # text-to-speech
image and imagen save files to the --out prefix and are synchronous. video is async: it returns an operation id, poll op-status until done is true, then download the video file (files.download). --json merges extra fields into any request.
Auth
- Provider id:
gemini(credential is collected ascustom.gemini) - Collection: API key via the secure credential flow (
credentials.request_api_access); created in Google AI Studio (aistudio.google.com/apikey) - Allowed hosts:
generativelanguage.googleapis.com - Status check:
bin/gemini.py auth(must return"ok": true). The key is sent verbatim as thex-goog-api-keyheader.
Operating Rules
- COST WARNING: all media generation requires a BILLING-ENABLED Google Cloud project. Free-tier media quota is 0, so calls without billing fail. This is the number one integration pitfall.
- Every generation spends real money: Veo 3 ~$0.40/sec, Veo 3.1 Fast ~$0.10/sec, Veo 3.1 Lite ~$0.05/sec, Imagen 4 ~$0.02-0.06/image, Nano Banana ~$0.02-0.04/image. Confirm with Michael before every generation, stating the model and expected cost.
- Veo renders take minutes. Poll
op-statuswith backoff; do not hammer it. - Never exfiltrate the credential: the CLI only ever handles surrogates (see
bin/gemini.py). Do not print, log, or transmit the key value.
Files
- SKILL.md
- bin/gemini.py
Maturity
🧪 Draft: written from Google's public Gemini API docs via the research dossier; not yet live-tested end-to-end.
Signals
- GitHub stars
- 1k
- Forks
- 282
- Last commit
- Sep 2026
Advanced
- Item type
- skill
- Key
gemini-anil-matcha- Source
- github.com/anil-matcha/awesome-muse-connectors
github.com/anil-matcha/awesome-muse-connectors