Music Player
A self-hosted music library and player that identifies untagged recordings without calling an external model.
Project identity
Music Player is a self-hosted library and playback application for a personal audio collection held on local storage. It began as a browser-only player that kept everything in IndexedDB, and it became a server application once the library outgrew what a single browser session could scan and remember.
The design constraint that shaped the whole project: most of the library carries no embedded tags at all. Making that collection browsable is the actual product problem, not playback.
Architecture and stack
The media library itself is never owned by the application. It is bind-mounted read-only, so the container can read audio files and can never modify, move, or delete the collection it is indexing.
Offline metadata inference
A large part of the library was ripped from video sources: the file names end in an eleven-character video identifier and there are no tags to read. The only signal available is the folder and file naming convention, so the server learns the convention of each folder from the files inside that same folder — a shared title prefix, a recurring composer credit, the way track numbers are written — and scores each guess by how many sibling files agree with it.
Calling a model API for every unidentified track would add cost, latency, a network dependency, and non-repeatable answers to a job that folder structure already answers. The inference stays dependency-free and deterministic: the same library produces the same result every time, and it works with no internet connection at all.
The album-identity rule
The rule that matters most is how an album is keyed, because getting it wrong destroys the library view rather than merely degrading it.
Before that rule existed, a missing album normalised to the literal text unknown album, and every untagged release in the library collapsed into a single enormous Unknown Album — unrelated film and series soundtracks merged into one record. The grouping key is album-scoped for the same reason, so a track title that repeats across two albums cannot silently re-merge them.
- Offline inference cannot distinguish a film score from a television score without a knowledge base; both classify as score, and the category does not carry that difference.
- A per-folder manual override is the natural fix, and it is recorded as future work rather than described as if it already exists.
Operations and data
The application runs from a full source tree on the Docker host and is rebuilt in place, so the previous container keeps serving traffic while the new image builds. Library state lives in a named volume; the audio itself lives on an external drive that is mounted read-only.
The database runs in WAL mode with a multi-megabyte sidecar log. Copying the database file on its own, or copying it while the service is writing, does not produce a consistent snapshot. This applies to every WAL-mode service in the lab, not only this one.
The rebuild guard
A library rebuild begins by deleting the existing file and track rows. That is correct when the media root is present and wrong in every other case, because the audio lives on a removable drive. The scan therefore refuses to run when a configured directory is unreadable, or when no audio files are found while the library is not empty. Without that guard, unplugging the drive and triggering a scan would erase the entire catalogue.
Security posture
Application
- Administrator passwords are stored as scrypt hashes using the built-in Node crypto module, so the image carries no native cryptography dependency.
- Sessions are signed cookies with a twelve-hour lifetime.
- Repeated failed logins lock the account for fifteen minutes after eight attempts.
- Request bodies are schema-validated, and sensitive routes are rate limited.
- Administrator-configured scan directories must resolve inside the mounted media root.
Container
- The root filesystem is read-only, with a small writable temporary filesystem.
- All Linux capabilities are dropped and privilege escalation is disabled.
- The media mount is read-only, so an application fault cannot damage the collection.
- Logs are size-capped and rotated so a noisy failure cannot fill the host disk.
- Only the application port is published; the database is a private volume with no listener.
Hurdles and lessons
| Hurdle | What actually happened | Lesson |
|---|---|---|
| Blank page over plain HTTP | Helmet sets an upgrade-insecure-requests policy by default. Served over plain HTTP on a non-standard port, the page rewrites its own asset URLs to HTTPS, both assets fail, and the browser shows an empty page with no useful error. | Test through the real hostname and TLS path. This is a host-wide behavior in this lab, and a second application walked into it afterwards. |
| Artwork lookups timing out | The usual open metadata and cover-art services do not respond from this network, from the host as well as from inside the container, while other outbound calls succeed normally. | Confirm an outage is network-specific before blaming Docker. The artwork provider order was changed to put a reachable service first and keep the unreachable one as a fallback. |
| Untagged albums merging | Every release without an album tag normalised to the same placeholder string and collapsed into one album. | An identity key built from a placeholder is not an identity. Key on something that is genuinely distinct — here, the folder path. |
| Removable media and destructive rebuilds | The rescan path starts by deleting rows, and the media sits on an external drive that the operating system reports as fixed. | Any operation that deletes before it writes needs a precondition check that fails closed. |
This project is still at the default package version and has no repository, tags, or changelog on the host. It therefore fails the versioning rule the rest of this wiki documents: a running service with no traceable version cannot be rolled back to a known-good release. Putting it under version control with a real version number is the first roadmap item.
Roadmap
- Put the deployment tree under version control, adopt semantic versioning, and record a changelog so releases become traceable and reversible.
- Add a per-folder metadata override so a human decision can beat inference where inference cannot win.
- Add a health endpoint and a scheduled backup verification rather than backing up only at deployment time.
- Record library statistics as release evidence so a scan regression is visible instead of silent.