Prd/docs/architecture.md
Lucy Doupalů be9f14ce34 Helpdesk - operator console + patched GrapheneOS Dialer for call handling
A small helpdesk system: an office Pixel running a patched GrapheneOS Dialer
answers technician calls, records both call legs as separate channels, and a
Ruby backend transcribes them through Whisper and files an AI summary against
the caller.

Squashed to a single commit for sharing. No credentials are included; secrets
live outside the repo in /etc/helpdesk/env on the server or a gitignored
.claude/env.local locally. See .claude/env.local.example for the shape.

Start at README.md, then docs/architecture.md.
2026-07-27 18:50:32 +02:00

14 KiB

Architecture

Helpdesk answers a technician's cellular call on an office GrapheneOS Pixel, shows the operator who's calling before they pick up, records and transcribes the call, and files a summary against the caller.

The pieces

  • Backend (components/backend) - one Ruby process: WEBrick and the standard library, no web framework. It serves the operator console, the operator API, and the phone's device API. This is what gets deployed. Entry point is server.rb; the business logic is the Service class in lib/helpdesk/domain.rb; storage sits behind a Store seam.
  • Phone (components/dialer-patch + sepolicy-patch) - a patched GrapheneOS Dialer that records the call and posts webhooks to the backend. Privileged call-audio capture only works in a self-signed build, so this is baked into the OS, not shipped as an app. Built on the dedicated build box (see docs/runbooks/).
  • Transcription (components/transcription-worker) - sends the recording to Whisper. Production uses the hosted Whisper (whisper.cajk.org, over mTLS, selected by WHISPER_URL); without that variable it falls back to a local WhisperX subprocess. A two-channel recording is split and each speaker transcribed separately (below); the result then passes through lib/helpdesk/glossary.rb, which repairs domain words the model cannot know.
  • AI (summariser.rb, context_maintainer.rb, wiki/) - turns transcripts into summaries and keeps a per-caller dossier. Described in its own section below.

The operator PC and the phone are both clients of the backend and never talk to each other directly. The voice call is cellular (phone to carrier) and never touches the server.

The call, start to finish

  1. A call comes in. The phone posts /api/v1/calls/incoming with the caller's number and the number that was dialed.
  2. The backend matches the number to a person and the dialed number to an operator, then pushes a screen-pop to that operator's console.
  3. The operator answers; the phone starts recording and posts /calls/answered.
  4. On hangup the phone posts /calls/ended and uploads the recording (PUT /recordings/<uuid>).
  5. The backend transcribes it, then summarises it, and stores both on the event.

Outgoing calls work too: the operator types a number in the console (or clicks one), the backend queues a dial command, the phone places the call and reports it back through /calls/outgoing. From there the recording and AI pipeline are the same as for an inbound call. The audio always stays on the phone - the PC can't carry call audio, and the server can't inject any (so there is no server-played greeting).

If a webhook goes missing, a reconcile pass on a 60-second timer fixes stuck calls: ring timeouts, a recording that never arrived, an expired quarantine. The exact rules are in docs/device-api.md.

Two auth planes

Operators and the phone authenticate in completely different ways, and it matters that you don't mix them.

  • Operators use a vmin client certificate (the same PKI as whisper.cajk.org). nginx terminates the TLS, checks the cert against the vmin CA, and passes the identity to the backend as headers (X-Operator-Cert plus a shared X-Proxy-Secret). The backend itself is plain HTTP on loopback behind nginx. Enforced by HELPDESK_OP_AUTH=1; checked in operator_authed? in server.rb.

    A valid signature is not by itself permission. ca.vmin.cz belongs to the hosting provider, so being signed by it proves who issued the cert, not that the holder works here. HELPDESK_OPERATOR_CNS lists the cert CNs allowed on the console, and it doubles as the revocation we own: no CRL is configured, so removing a CN is how a lost laptop's certificate is turned off. Left unset, any cert the CA verified is accepted, which is the older behaviour and what local dev relies on. The matching rule is pure and tested - Server.cert_cn and Server.operator_allowed?, grep anchor operator-allowlist.

    The identity is also kept, not just checked: resolutions an operator records carry a by field with their CN, so the ledger says who decided. nginx logs the same subject via the helpdesk_console log format, which is the only access attribution at the edge.

  • The phone has no certificate. It signs every request with an HMAC: a bearer token plus a signature over timestamp.nonce.body, with nonce replay protection and a clock-skew window. Enforced by HELPDESK_AUTH=1; checked in authed?. In production the phone posts to its own door, https://moje.al.army:8443/api/v1 - TLS without a client cert (the phone can't hold one), with nginx allowlisting only the device paths and blanking the operator-identity headers.

In dev mode (HELPDESK_DEV=1, the default) operator auth is off so you can open the console directly, and the device plane accepts the built-in dev credentials.

Both planes rest on certificates that expire, and the two ways they can fail are both silent: the vmin CA has no renewal at all, and the server certificate's automatic renewal writes its output to /dev/null. lib/helpdesk/cert_watch.rb makes that visible - it rides along on the /status poll the console already makes and puts a badge in the navbar when anything is inside the warning window, staying invisible otherwise. It reads the server certificate over TLS rather than from disk, both because the key directory is unreadable to the service user and because the served certificate is the honest one: a file check would not notice a certificate renewed on disk but never loaded by nginx. For a client certificate it reports the earliest date in the whole chain, since a leaf that outlives its CA stops working when the CA does.

Operators are a separate problem: there is one certificate per person, they live in browsers, and the server holds no copy, so none of them can be watched from a file. nginx knows the expiry of whichever certificate was just presented, and passes it as X-Operator-Cert-Days / X-Operator-Cert-Expires, so each operator is warned about their own and nobody needs an inventory of who holds what. Grep anchors: cert-watch-sources, cert-watch-severity, cert-watch-cache, cert-watch-presented.

Dual-channel transcription

When the phone records the two call legs as separate channels (persist.helpdesk.rec_dual_leg), the pipeline splits the file and transcribes each channel on its own, then merges the two transcripts on their timestamps. Speaker attribution stops being something a diariser infers from two similar voices on a narrowband line and becomes a fact about which channel the words came from.

Splitting is mandatory rather than an optimisation: Whisper resamples to 16 kHz mono before decoding, so handing it the stereo file puts both people back into one channel and discards exactly what the phone went to the trouble of capturing. Each channel is then transcribed with the diariser pinned to one speaker, since anything higher would invent a second voice out of crosstalk.

It costs two transcription jobs per call, so roughly double the turnaround on a CPU-only service.

Every failure path falls back to transcribing the original file: a leg that fails, a split that will not parse, or two channels whose transcripts come back identical (which means the legs were never really separated). A correctly attributed transcript is a bonus; a transcript is not. lib/helpdesk/wav_split.rb and lib/helpdesk/dual_transcript.rb, grep anchor dual-channel-transcribe.

The transcript glossary

Whisper has never seen "yamt" and the hosted service accepts no vocabulary hint - it takes exactly five form fields, none of them a prompt - so it writes the sound it heard, "jamt". Better audio cannot fix that, because nothing in the signal says how the word is spelt. lib/helpdesk/glossary.rb corrects a curated list of terms after transcription and before summarising, matching near-misses only.

Two limits are deliberate. Only distinctive terms belong in the list - system names, sites, operators. Ordinary words must not be added: "balíček" is one edit from "malíček" (little finger), and nothing in the text distinguishes a misheard word from a different one, so listing common words would eventually rewrite correct Czech. Common-word errors come from poor audio and are fixed at the recorder instead. And Czech inflection is left alone: "na Vodafonu" is the correct locative, so a difference confined to the ending is treated as grammar rather than a mistake.

Every substitution is returned and logged. A transcript records what a caller was told, so it is never edited silently. Grep anchors: glossary-load, glossary-correct, glossary-match.

Storage

Everything goes through the Store seam, so the backend doesn't care which storage it's on. Reads hit an in-memory cache; writes are explicit put_*/delete_* calls. There are two backends:

  • JSON file (default) - the whole store is dumped to ~/.config/helpdesk/state.json. Zero setup, good for local work.
  • PostgreSQL (set HELPDESK_DATABASE_URL) - a Sequel-backed write-through cache, PgStore. This is what production runs. The schema is db/migrate/001_init.sql, and that file is the source of truth for the data model. db/parity_check.rb proves the two backends behave identically.

One rule about the record shape: top-level fields are symbol keys (ev[:state]); nested JSON sub-documents are string keys (action_items[i]["text"]). The tests enforce it.

The AI subsystem

All model calls go through OpenRouter with one key (OPENROUTER_API_KEY). Three parts:

  • Per-call summary (summariser.rb). After transcription, the transcript goes to google/gemini-3.6-flash (override with SUMMARISER_MODEL) with a strict JSON schema. Out come a Czech summary, advisory action items, a context digest, durable caller facts, and yes/no decision deltas with verbatim quotes. The reply is validated client-side; if the strict schema isn't honoured the call is retried in a looser mode, and if that fails too the event is flagged for human review. Everything the AI produces is advisory - the operator sets the resolution.
  • Wiki grounding (wiki/). When an index file exists (WIKI_INDEX_PATH, default wiki/wiki_index.json), the pipeline retrieves the 15 wiki chunks closest to the transcript (cosine over OpenRouter embeddings) and passes them to the summariser. The prompt tells the model to state the concrete value or standard when the excerpts contain it, name the source wiki page in the action item's target, and never invent a number that isn't in the excerpts. Retrieval beat stuffing the whole wiki into the prompt in a side-by-side test: same grounded specifics, each one traceable to its page, at about 1.5% of the cost. The index is built offline by wiki/build_index.rb and is not in git.
  • Per-caller dossier (context_maintainer.rb). Each person carries an AI-maintained context: a short "who is this caller" summary plus other useful facts, stored as jsonb on the person with an optimistic-locking revision. After each summarised call (and on demand from the console) the backend rebuilds it - incrementally for small updates, from full history every few revisions or when the operator asks (both on google/gemini-3.6-flash; override with MAINTAINER_MODEL / MAINTAINER_REBUILD_MODEL). Operators can pin a section, and pinned text is never overwritten. Alongside the dossier sits a deterministic resolutions ledger (ledger.rb): yes/no decisions, where a reversal never deletes the old entry - it appends, marks the old one superseded, and flags the contradiction for a human. Entries come from two places: the summariser extracts them from call transcripts (ledger_deltas -> Helpdesk::Ledger.apply), and an operator can record one by hand from the console ("+ add" under Previous resolutions -> add_resolution in domain.rb, tagged source: operator) - useful when a policy is set outside a call and would otherwise only live in the dossier prose. Any active resolution is editable at any time (the pencil on its row -> edit_resolution), which supersedes it by id with the new decision or wording - so a decision that turned out wrong can be corrected without waiting for a conflict, and history is still kept. The operator resolves a flagged conflict from the console too: a review window shows both decisions with their quotes and source calls, and the operator sets the authoritative answer (which appends an operator-sourced entry, keeping the ledger append-only) or dismisses it (the current decision stands). resolve_contradiction, add_resolution, and edit_resolution in domain.rb are those seams.

Key decisions that still hold

  • Cellular, not landline. The audio is good enough for diarization and it's what technicians already use.
  • Record by patching the GrapheneOS Dialer inside a self-signed build. A normal app can't capture call audio.
  • The phone is the only audio endpoint. The server can't inject audio into the call, so there's no server-side greeting played to the caller.
  • Transcription is hosted Whisper over mTLS, with a local fallback. The summary is advisory: a human sets the resolution, the AI only suggests.
  • One dedicated inbound number per operator, so the dialed number says who should get the screen-pop.
  • Wiki grounding is retrieval of relevant chunks, not the whole wiki in the prompt. Cheaper, and every claim is traceable to a page.

Where things live

  • server.rb - HTTP routes, both auth checks (authed?, operator_authed?), the transcription trigger, the dossier refresh latch.
  • lib/helpdesk/domain.rb - the Store seam and the Service: call state machine, people, AI memory, reconcile.
  • lib/helpdesk/pg_store.rb - the PostgreSQL backend.
  • lib/helpdesk/cert_watch.rb - certificate expiry for the console's warning badge.
  • lib/helpdesk/glossary.rb + glossary_terms.txt - domain vocabulary applied to transcripts.
  • lib/helpdesk/wav_split.rb - splits a two-channel call recording into one file per speaker.
  • lib/helpdesk/dual_transcript.rb - merges the two per-speaker transcripts back into one.
  • db/migrate/001_init.sql - the schema.
  • summariser.rb, context_maintainer.rb, pipeline.rb, lib/helpdesk/ledger.rb, wiki/ - the AI path.
  • public/console.html - the whole operator console, one file.
  • sim/simulator.rb - drives the device API over HTTP; the quickest way to see the whole flow.

The device wire contract (endpoints, auth recipe, state machine, reconciliation timings) is in docs/device-api.md.