CatalogOpenAI Codex Skills · OpenAI

transcribe

Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings.

Review advisedOfficial28k starsalso listed on skills.sh

Savant verdict: Review advised

SkillSpector or the source hub flagged patterns to review before use.

Live evaluation

28overall
23quality
17compliance
45grounding
29actionability
43efficiency
Why this score may understate the skillThe live evaluation is a chat-only run: the model follows the skill's instructions but can't execute its scripts, call tools or reach the network.
  • Ships 1 script the live run can't execute; outputs describe those steps rather than perform them.
  • Expects tools, MCP servers, installs or network access that a chat-only run doesn't have.
  • Works with files or media (documents, images, designs) that the run can only describe in text.
Live telemetry from real runs (via the Savant skill router) replaces this estimate as it accumulates.

0 pass · 0 investigate · 7 fail across 7 cases. Jev accepted 7 of 18 LLM-drafted cases. Drafted and run by nvidia/nemotron-3-super-120b-a12b, validated and scored by jev-latest.

  • Transcribe this audio file to plain text with speaker labels if possible.positive case · fail
  • Give me a quick text transcription of this short voice memo.positive case · fail
  • Transcribe this meeting audio and save speaker-labeled results in the output directory.positive case · fail
  • Transcribe this audio and use the prompt 'focus on technical terms' to improve accuracy.edge case · fail
  • Translate this audio file from French to English and transcribe it.negative case · fail
  • Generate subtitles (.srt file) from this video with speaker labels and timestamps.negative case · fail
  • Transcribe this 5-minute lecture audio to text, but I'm not sure if I want speaker labels or not.edge case · fail

Safety (NVIDIA SkillSpector)

Risk score
29/100
Recommendation
CAUTION
Severity
MEDIUM
Savant decision
Review advised
SkillSpector rated it CAUTION with a risk score of 20/100 or more; review the findings before use. SkillSpector 2.12.0, static analysis.

4 patterns found

  • MCP Least Privilege: Skill declares no tool scope ('permissions' or 'allowed-tools') but code capabilities were detected: env, file_read, file_write.medium · SKILL.md:1 · Without declared permissions the skill's intent is opaque and cannot be validated.
  • Rogue Agent: create one in the OpenAI platform UI and export it in their shell. - Never ask the user to paste the full key in chat. ## Skill path (set once) ```bash export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}medium · SKILL.md:41 · Skill establishes unauthorized persistence across sessions via cron jobs, startup scripts, or state files. Session persistence allows an attacker to maintain access beyond the current interaction.
  • Excessive Agency: Never ask the usermedium · SKILL.md:42 · Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.
  • Agent Snooping: ls install under `$CODEX_HOME/skills` (default: `~/.codex/skillsmedium · SKILL.md:51 · Skill enumerates or reads other installed skills. Access to other skills' SKILL.md files or the skills directory reveals prompt instructions, capabilities, and secrets that should be invisible to peer skills.

Structure

  • Ships 1 executable scriptscripts/transcribe_diarize.py. Review what they do before enabling the skill for agents with tool access.

SKILL.md

---
name: "transcribe"
description: "Transcribe audio files to text with optional diarization and known-speaker hints. Use when a user asks to transcribe speech from audio/video, extract text from recordings, or label speakers in interviews or meetings."
---


# Audio Transcribe

Transcribe audio using OpenAI, with optional speaker diarization when requested. Prefer the bundled CLI for deterministic, repeatable runs.

## Workflow
1. Collect inputs: audio file path(s), desired response format (text/json/diarized_json), optional language hint, and any known speaker references.
2. Verify `OPENAI_API_KEY` is set. If missing, ask the user to set it locally (do not ask them to paste the key).
3. Run the bundled `transcribe_diarize.py` CLI with sensible defaults (fast text transcription).
4. Validate the output: transcription quality, speaker labels, and segment boundaries; iterate with a single targeted change if needed.
5. Save outputs under `output/transcribe/` when working in this repo.

## Decision rules
- Default to `gpt-4o-mini-transcribe` with `--response-format text` for fast transcription.
- If the user wants speaker labels or diarization, use `--model gpt-4o-transcribe-diarize --response-format diarized_json`.
- If audio is longer than ~30 seconds, keep `--chunking-strategy auto`.
- Prompting is not supported for `gpt-4o-transcribe-diarize`.

## Output conventions
- Use `output/transcribe/<job-id>/` for evaluation runs.
- Use `--out-dir` for multiple files to avoid overwriting.

## Dependencies (install if missing)
Prefer `uv` for dependency management.

```
uv pip install openai
```
If `uv` is unavailable:
```
python3 -m pip install openai
```

## Environment
- `OPENAI_API_KEY` must be set for live API calls.
- If the key is missing, instruct the user to create one in the OpenAI platform UI and export it in their shell.
- Never ask the user to paste the full key in chat.

## Skill path (set once)

```bash
export CODEX_HOME="${CODEX_HOME:-$HOME/.codex}"
export TRANSCRIBE_CLI="$CODEX_HOME/skills/transcribe/scripts/transcribe_diarize.py"
```

User-scoped skills install under `$CODEX_HOME/skills` (default: `~/.codex/skills`).

## CLI quick start
Single file (fast text default):
```
python3 "$TRANSCRIBE_CLI" \
  path/to/audio.wav \
  --out transcript.txt
```

Diarization with known speakers (up to 4):
```
python3 "$TRANSCRIBE_CLI" \
  meeting.m4a \
  --model gpt-4o-transcribe-diarize \
  --known-speaker "Alice=refs/alice.wav" \
  --known-speaker "Bob=refs/bob.wav" \
  --response-format diarized_json \
  --out-dir output/transcribe/meeting
```

Plain text output (explicit):
```
python3 "$TRANSCRIBE_CLI" \
  interview.mp3 \
  --response-format text \
  --out interview.txt
```

## Reference map
- `references/api.md`: supported formats, limits, response formats, and known-speaker notes.