ArticleReadMain page

Nathan's Technology Wiki / Projects

Talkback

Live transcription, playback, and summarization that runs entirely on the local machine with no model provider.

Project identity

Talkback is a browser transcription, playback, and summarization application with a server-side recording library. Its defining decision is a negative one: there is no API key, no model provider, and no mock mode. Everything the server does with a transcript, it does with arithmetic on the machine it runs on.

Recorded releasev3.1.0, September 6, 2026
RuntimeOne container with a SQLite and audio volume
External model callsNone

Version 3.0.0 removed the language-model integration that earlier versions carried. Version 3.1.0 followed the same day and spent its whole scope on transcription quality rather than new surface area.

What it does

Capture and review

  • Records with a live level meter, timer, and sixteen recognition languages.
  • Shows committed words solid and the in-flight guess dimmed behind a caret, so the interface never pretends a provisional word is final.
  • Plays audio back with scrubbing, ten-second skips, and speed control.
  • Keeps a searchable library of every recording with favorites, rename, delete, and export.

Output

  • Summarizes as bullets, a paragraph, or action items, plus extracted keywords.
  • Exports plain text, Markdown, or JSON.
  • Light, dark, and system themes with documented keyboard shortcuts.

Local summarization

Summarization ranks the sentences the speaker actually said and selects the strongest, rather than generating new prose. The result is deterministic and returns in single-digit milliseconds.

VectorizeA term-frequency vector per sentence, compared by cosine similarity.
RankA centrality score over that similarity graph, in the PageRank family.
WeightPosition and length priors, which matter more for speech than for written text.
SelectMaximal marginal relevance, so the summary does not return three phrasings of one point.

Action items are found differently. Commitments are formulaic rather than central — modal verbs, ownership verbs, deadline patterns — so centrality does not surface them and cue-phrase scoring does.

Extractive, and the interface should say so
The summary selects real sentences from the recording. Nothing is invented, which is the point; but it also cannot paraphrase, cannot answer a question about the transcript, and cannot summarize what was implied rather than said. That tradeoff is recorded in an architecture decision record rather than left for a user to discover.

Transcript cleanup, added in v3.1.0

Raw browser speech output arrives lowercase, unpunctuated, and full of filler. Version 3.1.0 fixes that on the way in, as local string transformations with per-rule toggles that persist.

RuleBehavior
Filler removalRemoves filler sounds and parenthetical padding.
Stutter collapseCollapses accidental repetition while leaving deliberate repetition intact.
Spoken punctuationWrites punctuation the speaker dictated by name.
CapitalizationSentence starts, the first-person pronoun, weekdays, and months.
Spoken numbersConverts number runs to digits with thousands grouping, while leaving a lone number word as a word.
ParagraphingBreaks a wall of dictation into readable paragraphs.

Cleanup runs per utterance with the preceding text as context, so a phrase that continues a sentence is not wrongly capitalized, and a full-transcript cleanup action is available separately. The recognition engine now requests three alternatives and keeps the most confident, with a rolling confidence figure shown while recording and flagged amber below sixty percent.

What this release taught
  • Input quality is a feature. The summarizer was never the weak link; the text reaching it was.
  • A confidence figure shown during recording lets the user fix the room, the microphone, or their pace while it still matters.
  • Action-item extraction had to learn named owners, not only first-person forms, because a commitment is most often made about someone by name.

Architecture and data

The application is a single container. Persistence uses the SQLite driver built into the Node runtime rather than a native module, which keeps a build toolchain out of the image; the test bundler cannot resolve that built-in module statically, so the database layer reaches it through a runtime require instead of a static import.

DatabaseSQLite in WAL mode on a named volume
AudioClips on the same volume, roughly 180 KB per minute
PlaybackServed with HTTP range requests

Range support is what makes the scrub bar work at all; without partial-content responses the browser cannot seek within a clip it has not fully downloaded. Recording captures from the same audio stream the level meter already opened, because requesting a second stream makes the browser show a doubled recording indicator.

Backups
The database is WAL mode, so a backup must include the write-ahead log sidecars or be taken with the container stopped. Audio clips live beside it in the same volume and are part of the same backup unit.

Delivery and rollback

Code ships to the host as an incremental Git bundle rather than over a remote or a file sync, so the deployment is a fast-forward of a real repository and the deployed revision is verifiable after the fact. The deploy script asserts the expected commit hash, builds, waits for the container to report healthy, and hashes the environment file before and after to prove the deployment did not modify it.

BundlePackage only the commits the host does not already have.
AssertFail the deployment if the resulting revision is not the expected one.
Build and waitBuild the image and block until the health check passes.
ProveCompare the environment-file hash before and after the release.
Known cosmetic drift
The deploy script still prints two health fields that version 3 removed, so they render blank in the smoke test output. The blanks are stale script output rather than a failing check, and they are recorded here so the next reader does not treat them as a fault.

Security and privacy

The honest privacy boundary is worth stating precisely, because a claim of fully local is only true after the first step.

StageWhere it happens
Speech recognitionIn the browser, through its own speech interface. In some browsers that means the audio is sent to the browser vendor for recognition.
Transcript, storage, searchOn the local server only.
Summarization and keywordsOn the local server only, with no model provider involved.
Stored audioOn the local volume; it is never uploaded anywhere by the application.
  • The container runs with a read-only root filesystem, so every writable path is declared deliberately.
  • Microphone access requires a secure browsing context, which is why the service is reached through the HTTPS hostname rather than by address and port.
  • The application holds no provider credentials, because it has no provider.

Deployment hurdles

HurdleSymptomResolution and lesson
Upgrade-insecure-requestsA blank page with no useful console error, because the app rewrote its own same-origin asset URLs to HTTPS on a port with no TLS.Disable that directive for this deployment. It was already documented from another application in this lab, and it still caught this one — a lesson recorded is not a lesson applied.
Build-kit syntax directiveThe build failed over SSH because the builder tried to pull a frontend image and the credential helper failed non-interactively, even though base images were cached locally.Remove the syntax directive. A build that only works interactively is not a deployable build.
Volume ownershipA fresh named volume was created owned by root while the container runs as an unprivileged user, so the database could not be opened at all.Create and own the data directory in the image before dropping to the unprivileged user, so Docker copies that ownership into a new volume. An existing volume has to be recreated to pick it up.
Read-only filesystem and default pathsAudio uploads failed with a bare directory-creation error while everything else worked, because one writable path still defaulted to a relative location inside the application directory.On a read-only container, every writable path must be set explicitly. Configuration that works locally by falling back to a relative path fails silently in the hardened image.
Proxy TLS optionsThe proxy host issued a certificate but silently dropped forced HTTPS and HTTP/2, so the site served plain HTTP — which for this application means no secure context and a disabled microphone.Requesting a certificate and setting SSL options in one dialog does not persist both. Reopen the host and save the SSL options a second time, and check any host created that way.

Roadmap

  • Update the deploy script so its smoke test reports the health fields the current version actually publishes.
  • Add scheduled backup verification for the database and audio volume rather than relying on deployment-time copies.
  • Evaluate an offline recognition engine so the one remaining external step in the privacy chain can be closed.
  • Extend the cleanup rule set with per-language behavior, since the current rules are strongest for English dictation.