Skip to content
humansanPublic

About

AI powered normalization layer between 100tracklists, Youtube, and Last.fm

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Danceseek (formerly Soundseek)

Danceseek parses messy DJ tracklists, normalizes track data using LLMs, matches tracks against Last.fm, and plays the set back with synced video while scrobbling your listens in real time.

Note: The deployed web UI is Read-Only. Tracklist ingestion and editing are performed locally via a maintainer application in apps/ingest. Auto-matching and exporting playlists to Spotify and YouTube are currently work in progress.


Architecture

maintainer's machine                             deployed
────────────────────                             ────────
browser
LLM normalization                    Neon        read API       web app
platform search / matching           Postgres    (browse,       (player,
ingest console + editor                          cues,          scrobbling)
                                                 scrobble)
  • Local Ingest: Scrapes 1001tracklists or YouTube descriptions, normalizes track data via LLMs, resolves platform IDs, and writes directly to Neon Postgres.
  • Deployed App: Installs only light read dependencies (api group). tests/test_api_surface.py and CI enforce strict isolation so ingest/scraping packages (Playwright, LangChain, yt-dlp) are never imported in production.

Quick Setup

uv sync                          # Install dependencies
uv run playwright install chromium
copy .env.example .env
npm ci                           # Run at root to hoist workspace packages
uv run alembic upgrade head      # Apply DB schema

Environment Variables

  • Required: OPENROUTER_API_KEY (normalization), LASTFM_API_KEY + LASTFM_SECRET (matching/scrobbling), DATABASE_URL (Neon).
  • Optional: SPOTIFY_CLIENT_ID/SECRET and GOOGLE_CLIENT_ID/SECRET (playlist exports; missing keys disable features gracefully).

Note: On first setup, run with SOUNDSEEK_HEADLESS=false to pass manual browser checks. Cookies persist in data/browser_profile/ for subsequent headless runs.


Adding Sets — Ingest Console

Launch the local maintainer console:

uv run soundseek console      # Binds to http://127.0.0.1:8020 (no auth)

Ingestion Pipelines

  1. 1001tracklists: Paste a URL. The pipeline streams:
    • Capture: Fetches HTML (cached in data/raw_html/ and stored in Postgres).
    • Normalize: Deterministically parses raw lines; an LLM structures artists, titles, remixes, IDs, and mashups.
    • Resolve: Matches structured tracks against Last.fm (and optionally Spotify/YouTube).
  2. YouTube Video: Paste a video link, title, and raw tracklist text. Timestamps, list numbering, and separators are parsed deterministically to guarantee accurate scrobble cues. Only track names are sent to the LLM.

Editor Rules

  • raw_text is immutable to retain original provenance.
  • Identity edits invalidate matches: Editing a row's track info resets its match status to pending so outdated scrobble targets aren't kept (manually typed Last.fm targets are trusted).
  • Deleting rows or reordering automatically updates relative indexes (played_with).

Catalog & Matching Strategy

  • Schema: setlists (metadata + JSONB track array), tracks (global resolved registry), raw_pages (HTML history), and users (Last.fm credentials). See migrations/versions/0001_initial_schema.py.
  • Precision over Recall: Matches are gated by SOUNDSEEK_RESOLVE_MIN_CONFIDENCE. Unmatched tracks default to no_match and fall back to normalized names.
  • Mashups: Resolved as single items on YouTube, but broken into component tracks for Last.fm scrobbling. Mashup credits (e.g., [Artist Mashup]) attach to the overall row.

Web UI & Scrobbling

Run the local dev servers:

uv run uvicorn apps.api.main:app --port 8010     # Read API
npm run dev                                       # Web app (:3000)
  • Server-Validated Playhead: scrobble/windows.py controls cue evaluation. The client sends playhead position; the server validates duration and listen thresholds (50% of track length or 4 mins, minimum 30s dwell time while playing).
  • Security: Last.fm session keys remain strictly server-side and are never exposed to the client (db.get_user strips them).
  • Configurable Logging: Choose whether to scrobble mashup components, w/ background tracks, unreleased tracks, or unmatched entries.

CLI Reference

uv run soundseek publish <url>          # Run full pipeline (capture → normalize → resolve → DB)
uv run soundseek publish <url> --lastfm-only --force
uv run soundseek publish <url> --reresolve
uv run soundseek show <id|url>          # Inspect set in terminal
uv run soundseek list                   # View local catalog
uv run soundseek remote                 # View Neon DB catalog
uv run soundseek push <id|url> | --all  # Backfill local sets to remote
uv run soundseek export <id|url> --target youtube|spotify [--dry-run]

Debugging

Debug pipeline issues by inspecting digest files generated for each set:

  • data/raw_html/<digest>.html — Fetched source HTML.
  • data/llm_inputs/<digest>.json — Raw extracted text and LLM prompts.
  • data/resolution_logs/<digest>.jsonl — Candidate matches, search queries, and confidence decisions.

Repository Structure

src/soundseek/         Shared core logic
  models.py            Data schema definitions
  extractor.py         Deterministic HTML/text parsing
  normalizer.py        LLM text normalization logic
  resolver/            Platform search and match ranking
  edit.py              State logic for maintainer edits
  scrobble/windows.py  Server-side scrobble evaluation
  db.py                Postgres database driver (raw SQL)
apps/api/              Deployed FastAPI read API
apps/ingest/           Local maintainer web engine
apps/web/              Next.js frontend
packages/api-client/   Generated TypeScript client from OpenAPI schema

Configuration uses SOUNDSEEK_* env vars (see .env.example). Key flags include SOUNDSEEK_STORE_BACKEND (json | postgres) and SOUNDSEEK_FETCH_BACKEND (local | stored).


Running Tests

uv run pytest

Tests run completely offline without database or browser dependencies (using static fixtures in tests/fixtures/). CI validates that the API package builds cleanly without ingestion dependencies.

About

AI powered normalization layer between 100tracklists, Youtube, and Last.fm

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages