Audio Jingle Skill
SkillFiles & storageaudio-jingle is a skill that lets an AI agent create short audio files such as jingles, background music, voiceovers, and sound effects. It plans the piece first, choosing genre, tempo, script, or texture, then routes the request to a suitable model and saves the result as an MP3 or WAV file in the project folder.
Use Audio Jingle Skill in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add Audio Jingle Skill and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the Audio Jingle Skill skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Make sure the project folder where the audio file should be saved exists and is writable.
What your AI can do with it
- Generate jingles and background music through Suno V5, Udio, or Lyria
- Create voiceovers with MiniMax TTS, FishAudio, or ElevenLabs V3
- Produce sound effects using ElevenLabs SFX or AudioCraft
- Plan genre, tempo, script, or texture before generating audio
- Save the finished audio as one MP3 or WAV file in the project folder
Getting started
- Make sure the project folder where the audio file should be saved exists and is writable.
- Add the skill to your agent's available skills.
- Configure access credentials for the audio services you intend to use.
- Ask the agent for the kind of audio you want, such as a jingle, voiceover, or sound effect.
- Retrieve the generated MP3 or WAV file from the project folder.
What this skill tells your AI
The instructions your AI receives, as published by nexu-io/open-design in design-templates/audio-jingle/SKILL.md and read by ahel’s review.
Three sub-modes. The active project's audioKind decides which one
runs:
audioKind | Models we route to | Plan focus |
|---|---|---|
music | Suno V5 (default), Udio, Lyria 2 | genre + tempo + instrumentation |
speech | MiniMax TTS (default), Fish, ElevenLabs V3 | script + voice + pacing |
sfx | ElevenLabs SFX (default), AudioCraft | texture + impact + duration |
Resource map
audio-jingle/
├── SKILL.md
└── example.html
Workflow
Step 0 — Read the project metadata
audioKind, audioModel, audioDuration (seconds), and (for speech)
voice. Branch by known values and use them verbatim. Missing metadata is not
an instruction to ask: infer a safe default when possible, and emit a
clarifying form only when the missing answer would materially change the
requested output or prevent generation.
Important: voice is provider-specific. For minimax-tts, --voice
must be a valid MiniMax voice_id (for example male-qn-qingse), not
a natural-language description. If you only have a prose voice brief
("warm female narrator", "neutral Mandarin"), keep that in your plan
but omit --voice so the daemon's default voice id applies, or ask the
user to choose a specific id.
Step 1 — Plan
Music
- Genre + reference artists (1-2)
- Tempo (BPM) + key
- Instrumentation (3-5 instruments max)
- Vocals: yes / no / hummed / choir
- Mood arc (intro → chorus → outro)
Speech
- Script (final, not draft — TTS runs verbatim)
- Voice target + pacing
For MiniMax this means a real
voice_id, not prose in--voice - Pronunciation hints for proper nouns / acronyms
SFX
- Texture (impact / whoosh / ambience / foley)
- Duration + envelope (sharp attack vs. gentle swell)
- Layering note (single hit vs. stacked)
State the plan in 2-3 sentences before dispatching.
Step 2 — Compose the prompt
Use the format the upstream model prefers. Bind audioDuration to the
API parameter directly; never put "make it 30 seconds" in prose.
Step 3 — Dispatch via the media contract
Use the unified dispatcher — do not call provider APIs by hand:
"$OD_NODE_BIN" "$OD_BIN" media generate \
--project "$OD_PROJECT_ID" \
--surface audio \
--audio-kind "<music|speech|sfx>" \
--model "<audioModel from metadata>" \
--duration <audioDuration seconds> \
[--voice "<provider voice id (speech only)>"] \
--output "<short-slug>-<duration>s.mp3" \
--prompt "<assembled prompt from Step 2 — for speech, the literal script>"
The command prints one line of JSON: {"file": {"name": "...", ...}}.
The bytes land in the project; the FileViewer renders the audio
transport controls automatically.
Step 4 — Hand off
Reply with: plan summary, the filename returned by the dispatcher, and one sentence on what to try if the user wants a variation (e.g. "swap tempo from 92 to 108 BPM" rather than "make it different").
Hard rules
- TTS runs your script literally. Proof it before dispatching — even one stray comma changes the cadence.
- MiniMax TTS rejects free-form voice prose in
--voice. Use a real MiniMaxvoice_id(for examplemale-qn-qingse) or omit the flag and let the daemon's default voice apply. - Music: under 30s = single section; 30–90s = intro + body; 90s+ = full arc. Don't try to fit a 3-act song into 15 seconds.
- SFX: prefer one well-described layer over a paragraph of "make it cool" — generators reward specific texture words.
- Save the file every turn. The audio viewer shows transport controls the moment the file lands.
Signals
- GitHub stars
- 99k
- Forks
- 11k
- Last commit
- Oct 2026
Questions
- What audio formats does it output?
- It saves the finished audio as one MP3 or WAV file in the project folder.
- Which music models does it use?
- Music requests are routed to Suno V5, Udio, or Lyria.
- Which text-to-speech services does it use?
- Speech is routed to MiniMax TTS, FishAudio, or ElevenLabs V3.
- Which sound effect tools does it use?
- Sound effects are routed to ElevenLabs SFX or AudioCraft.
- Does it plan the audio before generating it?
- Yes. It plans the piece first, covering genre, tempo, script, or texture, then sends a formatted request through a shared dispatcher.
Advanced
- Item type
- skill
- Key
audio-jingle-nexu-io- Source
- github.com/nexu-io/open-design
github.com/nexu-io/open-design