transcribing-audio
SkillSearchUse when a task requires STT/ASR.
Use transcribing-audio in Claude, ChatGPT or Ahel Desktop
Free. Sign in, add transcribing-audio and connect your AI. About a minute.
Also: Claude Code · Cursor · Codex
Then ask your AI: use the transcribing-audio skill
Details
Instructions available. Your AI can read the instructions. Execution depends on the setup they require.
Account requirements not reviewed. Check the skill instructions before use; ahel provides instructions and does not run this skill.
No other account needed.
Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.
What this skill tells your AI
The instructions your AI receives, as published by boldsoftware/shelley in skills/builtin/transcribing-audio/SKILL.md and read by ahel’s review.
Steps
-
By default, save the transcript beside the original audio with the extension replaced by
.transcript.txt(notes/review.m4a→notes/review.transcript.txt). -
Check the input. The gpt-transcribe endpoint accepts mp3, mp4, mpeg, mpga, m4a, wav, webm, flac, and ogg, up to 25 MB. For an existing local file with a supported extension under that limit, use the fast path below immediately. Do not run
ffprobe, inspect duration, list models, probe endpoints, loadreflection-integration, or query available integrations first.Transcode unsupported or larger inputs using ffmpeg. Split and transcribe piecemeal if necessary; for better results, slightly overlap the chunks and then manually stitch together the overlapped outputs. Shelley browser screen recordings have a sibling
<recording-path>.jsonsidecar withpathandduration_ms. Preserve each split chunk's media start offset so chunk-relative timestamps can be rolled up to the original recording timeline. -
Transcribe. Try
https://llm.int.exe.xyz, thenhttps://openai.int.exe.xyzonly when the first returns one of the two recognized routing rejections shown below. Otherwise surface the error and stop. Do not look up integrations first.A JSON response format is required. For
gpt-transcribe, optionalprompt,keywords[], andlanguages[]fields can supply known context, names, and language codes.transcribe() { output=$1 shift for base in https://llm.int.exe.xyz https://openai.int.exe.xyz; do : > "$output" if status=$(curl -sS --fail-with-body -w '%{http_code}' "$base/v1/audio/transcriptions" "$@" -o "$output"); then return fi if { [ "$status" = 400 ] && grep -Fq 'ChatGPT subscriptions do not support transcription; use an LLM integration with managed OpenAI or BYOK' "$output"; } || { [ "$status" = 403 ] && grep -Fq 'integration not found or not attached to this VM (trace: ' "$output"; }; then continue fi cat "$output" >&2 return 1 done cat "$output" >&2 return 1 } transcribe "$tmpdir/response.json" -F model=gpt-transcribe -F response_format=json -F "file=@$upload" jq -er '.text | select(type == "string")' "$tmpdir/response.json" > "$out"When the user asks for word or segment timestamps, run the GPT command above and the Whisper command below as two parallel bash tool calls in one response; each call defines
transcribeand its own variables. The GPT transcript stays canonical; Whisper's verbose JSON supplies timing only. Whisper acceptspromptand singularlanguage, not thegpt-transcribekeyword and language arrays.timestamp_out="${out%.transcript.txt}.timestamps.json" transcribe "$timestamp_out" -F model=whisper-1 -F response_format=verbose_json \ -F 'timestamp_granularities[]=word' -F 'timestamp_granularities[]=segment' \ -F "file=@$upload" jq -e '(.words | type == "array") and (.segments | type == "array")' "$timestamp_out" >/dev/null -
Report the output paths and stop.
Errors
402: LLM credits exhausted; https://exe.dev/user/shelley.- Transcription requires managed OpenAI or OpenAI BYOK; ChatGPT subscriptions return
400on this path.
Signals
- GitHub stars
- 676
- Forks
- 111
- Last commit
- Oct 2026
Advanced
- Item type
- skill
- Key
transcribing-audio- Source
- github.com/boldsoftware/shelley