Danceseek parses messy DJ tracklists, normalizes track data using LLMs, matches tracks against Last.fm, and plays the set back with synced video while scrobbling your listens in real time.
Note: The deployed web UI is Read-Only. Tracklist ingestion and editing are performed locally via a maintainer application in
apps/ingest. Auto-matching and exporting playlists to Spotify and YouTube are currently work in progress.
maintainer's machine deployed
──────────────────── ────────
browser
LLM normalization Neon read API web app
platform search / matching Postgres (browse, (player,
ingest console + editor cues, scrobbling)
scrobble)
- Local Ingest: Scrapes 1001tracklists or YouTube descriptions, normalizes track data via LLMs, resolves platform IDs, and writes directly to Neon Postgres.
- Deployed App: Installs only light read dependencies (
apigroup).tests/test_api_surface.pyand CI enforce strict isolation so ingest/scraping packages (Playwright, LangChain, yt-dlp) are never imported in production.
uv sync # Install dependencies
uv run playwright install chromium
copy .env.example .env
npm ci # Run at root to hoist workspace packages
uv run alembic upgrade head # Apply DB schema- Required:
OPENROUTER_API_KEY(normalization),LASTFM_API_KEY+LASTFM_SECRET(matching/scrobbling),DATABASE_URL(Neon). - Optional:
SPOTIFY_CLIENT_ID/SECRETandGOOGLE_CLIENT_ID/SECRET(playlist exports; missing keys disable features gracefully).
Note: On first setup, run with SOUNDSEEK_HEADLESS=false to pass manual browser checks. Cookies persist in data/browser_profile/ for subsequent headless runs.
Launch the local maintainer console:
uv run soundseek console # Binds to http://127.0.0.1:8020 (no auth)- 1001tracklists: Paste a URL. The pipeline streams:
- Capture: Fetches HTML (cached in
data/raw_html/and stored in Postgres). - Normalize: Deterministically parses raw lines; an LLM structures artists, titles, remixes, IDs, and mashups.
- Resolve: Matches structured tracks against Last.fm (and optionally Spotify/YouTube).
- Capture: Fetches HTML (cached in
- YouTube Video: Paste a video link, title, and raw tracklist text. Timestamps, list numbering, and separators are parsed deterministically to guarantee accurate scrobble cues. Only track names are sent to the LLM.
raw_textis immutable to retain original provenance.- Identity edits invalidate matches: Editing a row's track info resets its match status to
pendingso outdated scrobble targets aren't kept (manually typed Last.fm targets are trusted). - Deleting rows or reordering automatically updates relative indexes (
played_with).
- Schema:
setlists(metadata + JSONB track array),tracks(global resolved registry),raw_pages(HTML history), andusers(Last.fm credentials). Seemigrations/versions/0001_initial_schema.py. - Precision over Recall: Matches are gated by
SOUNDSEEK_RESOLVE_MIN_CONFIDENCE. Unmatched tracks default tono_matchand fall back to normalized names. - Mashups: Resolved as single items on YouTube, but broken into component tracks for Last.fm scrobbling. Mashup credits (e.g.,
[Artist Mashup]) attach to the overall row.
Run the local dev servers:
uv run uvicorn apps.api.main:app --port 8010 # Read API
npm run dev # Web app (:3000)- Server-Validated Playhead:
scrobble/windows.pycontrols cue evaluation. The client sends playhead position; the server validates duration and listen thresholds (50% of track length or 4 mins, minimum 30s dwell time while playing). - Security: Last.fm session keys remain strictly server-side and are never exposed to the client (
db.get_userstrips them). - Configurable Logging: Choose whether to scrobble mashup components,
w/background tracks, unreleased tracks, or unmatched entries.
uv run soundseek publish <url> # Run full pipeline (capture → normalize → resolve → DB)
uv run soundseek publish <url> --lastfm-only --force
uv run soundseek publish <url> --reresolve
uv run soundseek show <id|url> # Inspect set in terminal
uv run soundseek list # View local catalog
uv run soundseek remote # View Neon DB catalog
uv run soundseek push <id|url> | --all # Backfill local sets to remote
uv run soundseek export <id|url> --target youtube|spotify [--dry-run]Debug pipeline issues by inspecting digest files generated for each set:
data/raw_html/<digest>.html— Fetched source HTML.data/llm_inputs/<digest>.json— Raw extracted text and LLM prompts.data/resolution_logs/<digest>.jsonl— Candidate matches, search queries, and confidence decisions.
src/soundseek/ Shared core logic
models.py Data schema definitions
extractor.py Deterministic HTML/text parsing
normalizer.py LLM text normalization logic
resolver/ Platform search and match ranking
edit.py State logic for maintainer edits
scrobble/windows.py Server-side scrobble evaluation
db.py Postgres database driver (raw SQL)
apps/api/ Deployed FastAPI read API
apps/ingest/ Local maintainer web engine
apps/web/ Next.js frontend
packages/api-client/ Generated TypeScript client from OpenAPI schema
Configuration uses SOUNDSEEK_* env vars (see .env.example). Key flags include SOUNDSEEK_STORE_BACKEND (json | postgres) and SOUNDSEEK_FETCH_BACKEND (local | stored).
uv run pytestTests run completely offline without database or browser dependencies (using static fixtures in tests/fixtures/). CI validates that the API package builds cleanly without ingestion dependencies.
