fix(subtitles): tolerate regressing word timestamps

SRT conversion previously rejected otherwise valid word timestamps when a
start preceded the previous word's start. Clamp regressing starts to the
preceding nonempty word's start and extend ends only when necessary.
Preserve transcript order and the original response while continuing to
reject malformed timestamps. Document the adjustment and cover regressions,
cue boundaries, blank words, and invalid input.

Test Plan:
- python3 -B -m unittest -v: all eight tests passed.
- CLI conversion of a synthetic regressing response: passed.
- git diff --cached --check: passed.
- Actual recording verification blocked by unavailable 1Password auth.
This commit is contained in:
2026-09-05 16:27:03 +02:00
parent 4d7fcb1c8f
commit f8e50e4a3e
3 changed files with 55 additions and 5 deletions
+5 -2
View File
@@ -20,8 +20,11 @@ Words are grouped into cues of at most two lines (42 characters per line,
except an unusually long unbroken word) and six seconds. Cues also break at
sentence endings, pauses of at least 0.7 seconds, and speaker changes when
speaker labels are available. Timing comes from the words, rounded to
milliseconds. Subtitle output goes to stdout; missing or invalid word timestamps
produce an error. Without `--srt`, output remains plain text.
milliseconds. If word start times move backwards, they are clamped to the
preceding word's start without reordering the transcript; end times are extended
to the adjusted start only when needed. Subtitle output goes to stdout; missing
or invalid word timestamps produce an error. Without `--srt`, output remains
plain text.
One file per invocation; the whole recording is sent in one request, so
the service's audio length and request size limits apply.