85453da49f
regression / regression-shards (style-16-prod style-9-prod style-17-prod iframe-render-compat variables-prod mp4-h265-sdr, shard-4) (push) Has been cancelled
regression / regression-shards (style-4-prod style-11-prod style-2-prod animejs-adapter typegpu-adapter parallel-capture-regression, shard-5) (push) Has been cancelled
regression / regression-shards (style-7-prod style-8-prod style-10-prod css-spinner-render-compat webm-transparency mp4-h264-sdr webm-vp9, shard-3) (push) Has been cancelled
regression / regression-shards (sub-composition-video style-18-prod raf-ball-render-compat font-variant-numeric sub-comp-t0 sub-comp-id-selector, shard-7) (push) Has been cancelled
Windows render verification / Detect changes (push) Has been cancelled
Windows render verification / Preflight (lint + format) (push) Has been cancelled
Windows render verification / Render on windows-latest (push) Has been cancelled
Windows render verification / Tests on windows-latest (push) Has been cancelled
CI / Detect changes (push) Has been cancelled
CI / Build (push) Has been cancelled
CI / Lint (push) Has been cancelled
CI / Fallow audit (push) Has been cancelled
CI / Format (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Test (push) Has been cancelled
CI / Producer: integration tests (push) Has been cancelled
CI / Producer: unit tests (push) Has been cancelled
CI / File size check (push) Has been cancelled
CI / Test: skills (push) Has been cancelled
CI / Skills: manifest in sync (push) Has been cancelled
CI / CLI: npx shim (macos-latest) (push) Has been cancelled
CI / CLI: npx shim (ubuntu-latest) (push) Has been cancelled
CI / CLI: npx shim (windows-latest) (push) Has been cancelled
CI / SDK: unit + contract + smoke (push) Has been cancelled
CI / Test: runtime contract (push) Has been cancelled
CI / Studio: load smoke (push) Has been cancelled
CI / Smoke: global install (push) Has been cancelled
CI / CLI smoke (required) (push) Has been cancelled
CI / Semantic PR title (push) Has been cancelled
Player perf / Detect changes (push) Has been cancelled
Player perf / Preflight (lint + format) (push) Has been cancelled
Player perf / player-perf (push) Has been cancelled
Player perf / Perf: drift (push) Has been cancelled
Player perf / Perf: fps (push) Has been cancelled
Player perf / Perf: parity (push) Has been cancelled
Player perf / Perf: scrub (push) Has been cancelled
Player perf / Perf: load (push) Has been cancelled
preview-regression / Detect changes (push) Has been cancelled
preview-regression / Preflight (lint + format) (push) Has been cancelled
preview-regression / Preview parity (push) Has been cancelled
preview-regression / preview-regression (push) Has been cancelled
regression / regression (push) Has been cancelled
regression / Detect changes (push) Has been cancelled
regression / Preflight (lint + format) (push) Has been cancelled
regression / regression-shards (hdr-regression style-5-prod style-3-prod mov-prores, shard-1) (push) Has been cancelled
regression / regression-shards (overlay-montage-prod style-12-prod chat missing-host-comp-id png-sequence portrait-edge-bleed, shard-6) (push) Has been cancelled
regression / regression-shards (style-13-prod style-6-prod vignelli-stacking gsap-letters-render-compat audio-mux-parity, shard-8) (push) Has been cancelled
regression / regression-shards (style-15-prod hdr-hlg-regression style-1-prod many-cuts vfr-screen-recording render-symlinked-assets, shard-2) (push) Has been cancelled
CodeQL / Analyze (actions) (push) Has been cancelled
CodeQL / Analyze (javascript-typescript) (push) Has been cancelled
CodeQL / Analyze (python) (push) Has been cancelled
Docs / Validate docs (push) Has been cancelled
Sync skills to ClawHub / Publish changed skills (push) Has been cancelled
98 lines
5.3 KiB
Markdown
98 lines
5.3 KiB
Markdown
# Transcript Guide
|
|
|
|
For the `transcribe` CLI invocation, the `.en`-translates-non-English rule, and whisper model selection, see [`../transcribe.md`](../transcribe.md). This file covers what to do with the resulting transcript when authoring captions: input formats, mandatory quality checks, cleaning code, external-API fallbacks.
|
|
|
|
## Supported Input Formats
|
|
|
|
The CLI auto-detects and normalizes these formats:
|
|
|
|
| Format | Extension | Source | Word-level? |
|
|
| --------------------- | --------- | --------------------------------------------------------------------------- | ----------------- |
|
|
| whisper.cpp JSON | `.json` | `hyperframes init --video`, `hyperframes transcribe` | Yes |
|
|
| OpenAI Whisper API | `.json` | `openai.audio.transcriptions.create({ timestamp_granularities: ["word"] })` | Yes |
|
|
| SRT subtitles | `.srt` | Video editors, subtitle tools, YouTube | No (phrase-level) |
|
|
| VTT subtitles | `.vtt` | Web players, YouTube, transcription services | No (phrase-level) |
|
|
| Normalized word array | `.json` | Pre-processed by any tool | Yes |
|
|
|
|
**Word-level timestamps produce better captions.** SRT/VTT give phrase-level timing, which works but can't do per-word animation effects.
|
|
|
|
## Transcript Quality Check (Mandatory)
|
|
|
|
After every transcription, **read the transcript and check for quality issues before proceeding.** Bad transcripts produce nonsensical captions. Never skip this step.
|
|
|
|
### What to look for
|
|
|
|
| Signal | Example | Cause |
|
|
| ---------------------------- | -------------------------------------- | ---------------------------------------------------------------------------- |
|
|
| Music note tokens (`♪`, `�`) | `{ "text": "♪" }` or `{ "text": "�" }` | Whisper detected music, not speech |
|
|
| Garbled / nonsense words | "Do a chin", "Get so gay", "huh" | Model misheard lyrics or background noise |
|
|
| Long gaps with no words | 20+ seconds of only `♪` tokens | Instrumental section — expected, but high ratio means speech is being missed |
|
|
| Repeated filler | Many "huh", "uh", "oh" entries | Model is hallucinating on music |
|
|
| Very short word spans | Words with `end - start < 0.05` | Unreliable timestamp alignment |
|
|
|
|
### Automatic retry rules
|
|
|
|
**If more than 20% of entries are `♪`/`�` tokens, or the transcript contains obvious nonsense words, the transcription failed.** Do not proceed with the bad transcript. Instead:
|
|
|
|
1. **Retry with `medium.en`** if the original used `small.en` or smaller:
|
|
```bash
|
|
npx hyperframes transcribe audio.mp3 --model medium.en
|
|
```
|
|
2. **If `medium.en` also fails** (still >20% music tokens or garbled), tell the user the audio is too noisy for local transcription and suggest:
|
|
- Providing lyrics manually as an SRT/VTT file
|
|
- Using an external API (OpenAI or Groq Whisper — see below)
|
|
3. **Always clean the transcript** before building captions — filter out `♪`/`�` tokens and entries where `text` is a single non-word character. Only real words should reach the caption composition.
|
|
|
|
### Cleaning a transcript
|
|
|
|
After transcription (even with a good model), strip non-word entries:
|
|
|
|
```js
|
|
var raw = JSON.parse(transcriptJson);
|
|
var words = raw.filter(function (w) {
|
|
if (!w.text || w.text.trim().length === 0) return false;
|
|
if (/^[♪�\u266a\u266b\u266c\u266d\u266e\u266f]+$/.test(w.text)) return false;
|
|
if (/^(huh|uh|um|ah|oh)$/i.test(w.text) && w.end - w.start < 0.1) return false;
|
|
return true;
|
|
});
|
|
```
|
|
|
|
For model-selection guidance by content type, see [`../transcribe.md`](../transcribe.md) → "Picking a model by content type".
|
|
|
|
## Using External Transcription APIs
|
|
|
|
For the best accuracy, use an external API and import the result:
|
|
|
|
**OpenAI Whisper API** (recommended for quality):
|
|
|
|
```bash
|
|
# Generate with word timestamps, then import
|
|
curl https://api.openai.com/v1/audio/transcriptions \
|
|
-H "Authorization: Bearer $OPENAI_API_KEY" \
|
|
-F file=@audio.mp3 -F model=whisper-1 \
|
|
-F response_format=verbose_json \
|
|
-F "timestamp_granularities[]=word" \
|
|
-o transcript-openai.json
|
|
|
|
npx hyperframes transcribe transcript-openai.json
|
|
```
|
|
|
|
**Groq Whisper API** (fast, free tier available):
|
|
|
|
```bash
|
|
curl https://api.groq.com/openai/v1/audio/transcriptions \
|
|
-H "Authorization: Bearer $GROQ_API_KEY" \
|
|
-F file=@audio.mp3 -F model=whisper-large-v3 \
|
|
-F response_format=verbose_json \
|
|
-F "timestamp_granularities[]=word" \
|
|
-o transcript-groq.json
|
|
|
|
npx hyperframes transcribe transcript-groq.json
|
|
```
|
|
|
|
## If No Transcript Exists
|
|
|
|
1. Check the project root for `transcript.json`, `.srt`, or `.vtt` files.
|
|
2. If none found, run [`../transcribe.md`](../transcribe.md) — pick the starting model from "Picking a model by content type" there.
|
|
3. Run the quality check above. If it fails, retry with a larger model or fall back to manual lyrics / external API.
|