ArticleReadMain page

Nathan's Technology Wiki / Projects

Music Player

A self-hosted music library and player that identifies untagged recordings without calling an external model.

Project identity

Music Player is a self-hosted library and playback application for a personal audio collection held on local storage. It began as a browser-only player that kept everything in IndexedDB, and it became a server application once the library outgrew what a single browser session could scan and remember.

StatusRunning; published through the reverse proxy over HTTPS
RuntimeSingle Node container with a private SQLite volume
VersionUnversioned; see the release-hygiene gap below

The design constraint that shaped the whole project: most of the library carries no embedded tags at all. Making that collection browsable is the actual product problem, not playback.

Architecture and stack

LayerTechnology and responsibilityReason for the boundary
InterfaceReact 19, Vite, Tailwind, React Router, IndexedDB caching, Web Audio visualizerOwns browsing, playback, and presentation without holding library authority.
APIExpress with Helmet, Zod request validation, and rate limitingOwns authentication, validation, and every filesystem read.
ScannerWalks the media root and parses tags with music-metadataIsolated so a rebuild can be reasoned about and guarded independently.
InferenceDeterministic folder-name analysis with no network callsKeeps identification offline, repeatable, and free of provider cost.
StorageSQLite through the built-in Node driver, in WAL modeNo native build toolchain in the image and no separate database service.

The media library itself is never owned by the application. It is bind-mounted read-only, so the container can read audio files and can never modify, move, or delete the collection it is indexing.

Offline metadata inference

A large part of the library was ripped from video sources: the file names end in an eleven-character video identifier and there are no tags to read. The only signal available is the folder and file naming convention, so the server learns the convention of each folder from the files inside that same folder — a shared title prefix, a recurring composer credit, the way track numbers are written — and scores each guess by how many sibling files agree with it.

Why this is deliberately not an AI feature
Calling a model API for every unidentified track would add cost, latency, a network dependency, and non-repeatable answers to a job that folder structure already answers. The inference stays dependency-free and deterministic: the same library produces the same result every time, and it works with no internet connection at all.

The album-identity rule

The rule that matters most is how an album is keyed, because getting it wrong destroys the library view rather than merely degrading it.

CaseAlbum key
Track has a real album tagThe normalised album title, so the same album ripped twice still merges correctly.
Track has no album tagThe relative folder path, so each untagged release stays a separate album.

Before that rule existed, a missing album normalised to the literal text unknown album, and every untagged release in the library collapsed into a single enormous Unknown Album — unrelated film and series soundtracks merged into one record. The grouping key is album-scoped for the same reason, so a track title that repeats across two albums cannot silently re-merge them.

Honest limitation
  • Offline inference cannot distinguish a film score from a television score without a knowledge base; both classify as score, and the category does not carry that difference.
  • A per-folder manual override is the natural fix, and it is recorded as future work rather than described as if it already exists.

Operations and data

The application runs from a full source tree on the Docker host and is rebuilt in place, so the previous container keeps serving traffic while the new image builds. Library state lives in a named volume; the audio itself lives on an external drive that is mounted read-only.

BuildThe new image is built while the running container continues to serve.
StopThe container is stopped so the database is not written during the copy.
Back upThe SQLite file is copied with its write-ahead log into a timestamped backup folder.
StartThe new container starts and the library is checked through the public hostname.
Backups must be taken with the container stopped
The database runs in WAL mode with a multi-megabyte sidecar log. Copying the database file on its own, or copying it while the service is writing, does not produce a consistent snapshot. This applies to every WAL-mode service in the lab, not only this one.

The rebuild guard

A library rebuild begins by deleting the existing file and track rows. That is correct when the media root is present and wrong in every other case, because the audio lives on a removable drive. The scan therefore refuses to run when a configured directory is unreadable, or when no audio files are found while the library is not empty. Without that guard, unplugging the drive and triggering a scan would erase the entire catalogue.

Security posture

Application

  • Administrator passwords are stored as scrypt hashes using the built-in Node crypto module, so the image carries no native cryptography dependency.
  • Sessions are signed cookies with a twelve-hour lifetime.
  • Repeated failed logins lock the account for fifteen minutes after eight attempts.
  • Request bodies are schema-validated, and sensitive routes are rate limited.
  • Administrator-configured scan directories must resolve inside the mounted media root.

Container

  • The root filesystem is read-only, with a small writable temporary filesystem.
  • All Linux capabilities are dropped and privilege escalation is disabled.
  • The media mount is read-only, so an application fault cannot damage the collection.
  • Logs are size-capped and rotated so a noisy failure cannot fill the host disk.
  • Only the application port is published; the database is a private volume with no listener.

Hurdles and lessons

HurdleWhat actually happenedLesson
Blank page over plain HTTPHelmet sets an upgrade-insecure-requests policy by default. Served over plain HTTP on a non-standard port, the page rewrites its own asset URLs to HTTPS, both assets fail, and the browser shows an empty page with no useful error.Test through the real hostname and TLS path. This is a host-wide behavior in this lab, and a second application walked into it afterwards.
Artwork lookups timing outThe usual open metadata and cover-art services do not respond from this network, from the host as well as from inside the container, while other outbound calls succeed normally.Confirm an outage is network-specific before blaming Docker. The artwork provider order was changed to put a reachable service first and keep the unreachable one as a fallback.
Untagged albums mergingEvery release without an album tag normalised to the same placeholder string and collapsed into one album.An identity key built from a placeholder is not an identity. Key on something that is genuinely distinct — here, the folder path.
Removable media and destructive rebuildsThe rescan path starts by deleting rows, and the media sits on an external drive that the operating system reports as fixed.Any operation that deletes before it writes needs a precondition check that fails closed.
Release-hygiene gap, recorded rather than hidden
This project is still at the default package version and has no repository, tags, or changelog on the host. It therefore fails the versioning rule the rest of this wiki documents: a running service with no traceable version cannot be rolled back to a known-good release. Putting it under version control with a real version number is the first roadmap item.

Roadmap

  • Put the deployment tree under version control, adopt semantic versioning, and record a changelog so releases become traceable and reversible.
  • Add a per-folder metadata override so a human decision can beat inference where inference cannot win.
  • Add a health endpoint and a scheduled backup verification rather than backing up only at deployment time.
  • Record library statistics as release evidence so a scan regression is visible instead of silent.