Prd/components/transcription-worker
Lucy Doupalů be9f14ce34 Helpdesk - operator console + patched GrapheneOS Dialer for call handling
A small helpdesk system: an office Pixel running a patched GrapheneOS Dialer
answers technician calls, records both call legs as separate channels, and a
Ruby backend transcribes them through Whisper and files an AI summary against
the caller.

Squashed to a single commit for sharing. No credentials are included; secrets
live outside the repo in /etc/helpdesk/env on the server or a gitignored
.claude/env.local locally. See .claude/env.local.example for the shape.

Start at README.md, then docs/architecture.md.
2026-07-27 18:50:32 +02:00
..
bench Helpdesk - operator console + patched GrapheneOS Dialer for call handling 2026-07-27 18:50:32 +02:00
README.md Helpdesk - operator console + patched GrapheneOS Dialer for call handling 2026-07-27 18:50:32 +02:00
remote_transcriber.rb Helpdesk - operator console + patched GrapheneOS Dialer for call handling 2026-07-27 18:50:32 +02:00
test_remote_transcriber.rb Helpdesk - operator console + patched GrapheneOS Dialer for call handling 2026-07-27 18:50:32 +02:00
test_transcriber.rb Helpdesk - operator console + patched GrapheneOS Dialer for call handling 2026-07-27 18:50:32 +02:00
transcriber.rb Helpdesk - operator console + patched GrapheneOS Dialer for call handling 2026-07-27 18:50:32 +02:00

components/transcription-worker

Turns a call recording into a [Speaker N]: transcript. Two interchangeable implementations with the same transcribe(audio, …) contract; Pipeline.build picks by environment:

  • remote_transcriber.rb - what production uses. Drives the hosted Whisper web service (WHISPER_URL, https://whisper.cajk.org/app) over mTLS with the vmin client cert (WHISPER_MTLS_P12 / WHISPER_MTLS_PASS): submit the file, poll the job, download the speaker-tagged text. Transient HTTP errors are retried with backoff; a 404 on the job stops immediately.
  • transcriber.rb - the local fallback when WHISPER_URL is unset. Runs the WhisperX CLI (transcribe.py, shipped in resources/whisper/; directory via WHISPER_DIR) as a subprocess and maps its exit codes to a state.

Exit codes → states (local path)

exit meaning state retry?
0 with output success transcribed no
0, empty output no speech detected no_speech no
2 / 3 / 5 / 6 bad args or token · no ffmpeg · HF terms not accepted · alignment failed config_error no - fix the config
anything else (1, 137, timeout…) crash, OOM, transient failed yes

The remote path lands on the same states: done with text → transcribed, done empty → no_speech, error/cancelled/interruptedfailed (retryable).

Test

ruby test_transcriber.rb          # 11 assertions, stub scripts, no models
ruby test_remote_transcriber.rb   # 30 assertions, fake HTTP client, no network

Both suites are offline. The HTTP client and the subprocess are injectable, which is what makes that possible.

What the hosted service will accept

Exactly five multipart fields, and we already send all of them: file, model, language, min_speakers, max_speakers. There is no initial_prompt, temperature, beam size or VAD control, so no client-side decoding knob is left to turn - confirmed against the service's own openapi.json and the form its web UI submits. Two models are offered: large-v2, which the service itself labels as the more accurate one for Czech, and large-v3.

Measuring accuracy (bench/)

Changing any of this without measuring is guesswork, and some plausible changes make things worse. bench/ab.rb runs the same recordings through several settings and scores them:

ruby bench/ab.rb [corpus_dir]     # default bench/corpus; needs the same mTLS env as production
ruby bench/test_ab.rb             # 22 assertions for the scoring itself, offline

Put recordings in bench/corpus/ with a matching .txt of what was actually said. Without that reference you still get transcripts and speaker counts, but no error rate - there is nothing to be right or wrong against. Results are cached in bench/out/, because the service is CPU-only and a run takes minutes; delete a file to force that job again.

Both directories are gitignored: call recordings and transcripts are caller PII and must not be committed.

Reading the score

Not every difference from the reference is a transcription error, and counting them all makes the audio look far worse than it is. On a real call the raw rate was 14.6% while only a third of it was actual misrecognition. So each disagreement is classified and the reported figure discounts the ones that are not the model's fault:

kind example counted?
boundary sim kartou -> simkartou no, same letters
variant děkuju -> děkuji no, Czech spoken form written as standard
diacritic balicek -> balíček no
number čtyřicet sedm -> šedesát šest yes, and flagged separately - digits fail differently
real balíček -> badíček yes

Two figures are printed: WER after those discounts, and CER, a character rate with spaces removed that is immune to word-boundary decisions entirely. If they disagree wildly, look at the chunk list before believing either.

Score a dual-channel transcript per speaker, never linearly. A merged transcript is ordered by timestamp while a script is in reading order, so whenever turns group differently the alignment collapses and reports a huge error for text that is correct. On a real call that turned 10% into 44%. score_channels compares each speaker's words against only their own lines. Measured that way, the first dual-channel call scored 9.6% on the operator channel and 28.2% on the caller channel - the operator comes straight off the local microphone, the caller arrives through the cellular codec at about -42 dB. Loudness-normalising the quiet channel was tried and made it slightly worse, so the remaining gap is the far-end audio, not the level.

A caution the harness cannot fix: the reference must be what was actually SAID, not the script you handed the readers. People misread lines, and a deviation scores as a model error. Read the chunk list once and correct the reference against the audio before trusting any number from it.

Two more things. The rate is deliberately not clamped at 100%, because Whisper padding a mumble with invented words is a real failure mode that should look as bad as it is. And where two models disagree at least one is wrong, so comparing variants localises the doubtful words even with no reference at all.