Compare commits

..

13 Commits

Author SHA1 Message Date
Philip Guzman 37786612b7 Add read-only database V2 parser 2026-07-01 08:13:20 -07:00
Philip Guzman 8ff4db4951 Merge feature/smart-crate-discovery into develop 2026-07-01 08:08:43 -07:00
Philip Guzman a384cdc88d Discover Serato smart crate definitions 2026-07-01 08:06:19 -07:00
Philip Guzman 71bca064ed Merge feature/web-interface into develop 2026-07-01 08:03:19 -07:00
Philip Guzman eda5d4e62b Document approved dashboard direction 2026-07-01 07:55:00 -07:00
Philip Guzman 9053b4ca3d Add local web analysis dashboard 2026-07-01 07:52:33 -07:00
Philip Guzman 251ab4d090 Merge feature/broken-symlinks into develop 2026-07-01 07:46:30 -07:00
Philip Guzman b8ea17450b Add broken symlink diagnostics 2026-07-01 07:32:07 -07:00
Philip Guzman c5449a278a Merge feature/duplicate-filenames into develop 2026-07-01 07:30:50 -07:00
Philip Guzman 6449c7de6a Add duplicate filename diagnostics 2026-06-30 18:52:16 -07:00
Philip Guzman d764f910e4 Merge feature/crate-classification into develop 2026-06-30 18:50:42 -07:00
Philip Guzman 35713bbbb3 Classify static and smart crates 2026-06-30 18:20:05 -07:00
Philip Guzman c92d929f37 Merge feature/analyze-command into develop 2026-06-30 18:17:20 -07:00
39 changed files with 1233 additions and 41 deletions
+1
View File
@@ -9,6 +9,7 @@ samples/small-library/generated/
*.db *.db
*.csv *.csv
*.html *.html
!serato_doctor/webui/*.html
# Never commit personal Serato data # Never commit personal Serato data
database V2 database V2
+11
View File
@@ -35,3 +35,14 @@ aggregate diagnostic log.
Serato Doctor never repairs files without an explicit future repair workflow, Serato Doctor never repairs files without an explicit future repair workflow,
preview, backup, and rollback path. preview, backup, and rollback path.
## Local Web Interface
Launch the responsive, local-only dashboard with:
```shell
serato-doctor-web
```
Then open `http://127.0.0.1:8765`. The interface exposes the same read-only health
analysis and never sends library paths or results to an external service.
+5 -5
View File
@@ -7,10 +7,10 @@
- [x] Serato crate parser - [x] Serato crate parser
- [x] Missing reference CSV report - [x] Missing reference CSV report
- [x] Grouped missing reference report - [x] Grouped missing reference report
- [ ] HTML health dashboard - [x] HTML health dashboard
- [x] Test suite - [x] Test suite
- [x] Sample library fixtures - [x] Sample library fixtures
- [ ] Database V2 read-only parser - [x] Database V2 read-only parser
- [x] Configuration - [x] Configuration
- [x] Logging - [x] Logging
- [x] Matching engine - [x] Matching engine
@@ -18,12 +18,12 @@
## v0.2 — Diagnostics ## v0.2 — Diagnostics
- [x] `serato-doctor analyze` command - [x] `serato-doctor analyze` command
- [ ] Duplicate filename detection - [x] Duplicate filename detection
- [ ] Duplicate audio hash detection - [ ] Duplicate audio hash detection
- [ ] Broken symlink detection - [x] Broken symlink detection
- [ ] Orphaned audio detection - [ ] Orphaned audio detection
- [ ] OneDrive rename detection - [ ] OneDrive rename detection
- [ ] Crate classification: static vs smart/dynamic - [x] Crate classification: static vs smart/dynamic
- [x] Library health score - [x] Library health score
## v0.3 — Safe Repair ## v0.3 — Safe Repair
+27
View File
@@ -0,0 +1,27 @@
# Broken Symlink Detection
## Problem
Library migrations may leave symbolic links pointing to files or folders that no
longer exist. The filesystem scanner previously skipped those links silently.
## Architecture
One read-only filesystem traversal now returns audio tracks and broken symbolic
links. Each finding preserves the link path and its raw target when the operating
system can read it. The library and health report retain only immutable findings.
Broken links are reported separately and do not affect the reference-integrity
score. The detector does not follow, recreate, remove, or rewrite any link.
## Edge Cases
- Relative and absolute link targets.
- Links to missing files and missing directories.
- Link targets that cannot be read due to an operating-system error.
- Valid symlinks, which remain eligible for ordinary audio scanning.
## Verification
Tests create synthetic valid and broken links in temporary directories, verify the
raw target, and confirm health integration. Existing scanner behavior is preserved.
+38
View File
@@ -0,0 +1,38 @@
# Crate Classification
## Problem
Static crates are manually maintained track lists. Smart crates are dynamic views
generated from rules, so stale-looking entries in them should not be presented as
broken manual references or given the same health-score weight.
## Architecture
Crates carry a `static`, `smart`, or `unknown` kind. Folder provenance is the
primary signal: Serato stores regular `.crate` files in `Subcrates` and smart
`.scrate` definitions in `SmartCrates`. `Compatible by key.crate` is also treated as smart
when encountered in `Subcrates`, based on the original migration case study.
Smart crate names use `≫≫` to encode hierarchy. The model preserves those segments,
so `Compatible by key≫≫10A.scrate` has a parent of `Compatible by key` and a display
name of `10A`. Smart definitions and dynamic `.crate` containers are counted
separately.
The library loader reads both folders. Health analysis reports all references but
scores only non-smart references. Unknown crates remain scoreable so incomplete
classification cannot silently hide potential problems.
Serato documents the folder distinction in [What is in the _Serato_ folder?](https://support.serato.com/hc/en-us/articles/204022904-What-is-in-the-Serato-folder)
and explains that smart crates are populated from rules in [Crates in Serato DJ](https://support.serato.com/hc/en-us/articles/227561407-Crates-in-Serato-DJ-Pro-Serato-DJ-Lite).
## Edge Cases
- The known `Compatible by key` dynamic crate in the `Subcrates` folder.
- Crate fixtures outside a recognized Serato folder.
- Libraries containing both static and smart references to the same track.
- Dynamic references whose current materialized paths appear missing.
## Verification
Tests cover all three kinds, both Serato folders, the known dynamic fallback, the
synthetic five-static/two-smart library, and exclusion from health scoring.
+28
View File
@@ -0,0 +1,28 @@
# Database V2 Read-only Parser
## Problem
Crates and files do not explain every orange track in Serato. The legacy
`database V2` contains Serato's library-level track paths and metadata, so it must
be inspected independently from crate references.
## Architecture
The parser reads the file as a big-endian tag-length-value stream. Top-level
`otrk` records contain nested fields including `pfil` (path), `tsng` (title),
`tart` (artist), `talb` (album), and `tgen` (genre). Text is UTF-16 big-endian.
Analysis reports total database entries, entries whose filenames occur in the
selected music scan, entries outside that scan, scanned tracks absent from the
database, and duplicate database paths. “Outside scan” is deliberately not called
missing because Serato databases can include samples and tracks from other roots.
## Safety
The parser calls only `read_bytes`; it never opens the database for writing. No
metadata values or personal paths are sent to logs or the dashboard.
## Verification
Synthetic TLV fixtures cover version, paths, metadata, incomplete records, and
health integration. The sample library includes a generated ten-entry database.
+32
View File
@@ -0,0 +1,32 @@
# Duplicate Filename Detection
## Problem
Two files with the same filename may be ordinary copies, while names such as
`Track.mp3` and `Track 2.mp3` may indicate a OneDrive conflict. These cases need
review, but neither is sufficient evidence for deletion.
## Architecture
The duplicate detector emits immutable groups of two kinds:
- `exact_name` groups filenames after case and Unicode normalization.
- `cloud_conflict` groups names after additionally removing a trailing numeric
suffix from the stem.
Groups contain every path, a stable comparison key, a display name, and the number
of files beyond the first. Health analysis reports exact and suspected-conflict
counts separately. Neither category changes the health score.
## Edge Cases
- Same filename in different folders.
- Case-only and Unicode representation differences.
- Multiple conflict suffixes such as `Track 2.mp3` and `Track 3.mp3`.
- Legitimate numbered titles, which remain explicitly labeled as suspected.
- A single file, which is never reported as a duplicate.
## Verification
Tests cover exact, normalized, conflict-family, unrelated, and deterministic-order
behavior, plus integration with health analysis. Detection is read-only.
+5 -4
View File
@@ -8,10 +8,11 @@ but an opaque blended score would imply confidence the current data cannot suppo
## Architecture ## Architecture
The health engine produces an immutable report from the core `Library`. Its score The health engine produces an immutable report from the core `Library`. Its score
is only the percentage of crate references resolved by exact filename. The report is only the percentage of non-dynamic crate references resolved by exact filename.
also exposes missing references, unique missing filenames, duplicate filename Smart-crate references are counted but excluded because their contents are derived
groups, extra duplicate files, unused tracks, and missing references with matching from rules. The report also exposes missing references, unique missing filenames,
candidates. duplicate filename groups, extra duplicate files, unused tracks, and missing
references with matching candidates.
Duplicate, unused, and candidate counts are informational. They do not affect the Duplicate, unused, and candidate counts are informational. They do not affect the
score until the project has a documented and validated weighting policy. An empty score until the project has a documented and validated weighting policy. An empty
+19
View File
@@ -0,0 +1,19 @@
# Smart Crate Discovery Correction
## Problem
The first classifier searched `SmartCrates` for `*.crate`. Real Serato smart-crate
definitions use `.scrate`, so a library with 26 definitions displayed only the
single name-based `Compatible by key.crate` fallback.
## Correction
Library discovery now reads `.scrate` definitions case-insensitively from the
`SmartCrates` folder. Dynamic `.crate` containers remain excluded from manual
reference scoring but are reported separately. The `≫≫` filename separator is
preserved as smart-crate hierarchy metadata.
## Verification
Synthetic tests cover `.scrate` discovery, case-correct folder names, hierarchy,
definition/container counts, and the existing five-static/two-smart sample.
Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

+37
View File
@@ -0,0 +1,37 @@
# Local Web Interface
![Approved Serato Doctor dashboard concept](web-dashboard-concept.png)
The generated concept above is the approved visual direction. The implemented UI
keeps its hierarchy, palette, safety emphasis, health ring, diagnostic cards, and
responsive behavior. Product truth takes precedence over mockup copy: a 77.8%
library is labeled for review rather than described as healthy.
## Problem
The command line is useful for automation but makes the growing diagnostic set
harder to explore. A visual dashboard lets users test analysis safely and understand
which findings affect health.
## Architecture
`serato-doctor-web` binds to `127.0.0.1:8765` by default and serves package-owned
HTML, CSS, and JavaScript with Python's standard library. A same-origin JSON endpoint
runs the existing read-only parser, scanner, matcher, duplicate detector, crate
classifier, and health engine. No web framework or external service is required.
The UI clearly labels read-only mode, separates scored health from informational
diagnostics, and adapts from a full sidebar layout to compact mobile navigation.
## Edge Cases
- Missing or invalid Serato and music directories.
- Empty libraries with no assessable health score.
- Large or malformed requests, capped at 64 KiB.
- HTML injection, avoided by rendering all results through `textContent`.
- Network exposure, avoided by a loopback-only default binding.
## Verification
Tests exercise the web analysis adapter and static package assets. Browser checks
cover real form submission, result rendering, error display, and responsive layout.
+4
View File
@@ -14,9 +14,13 @@ dev = ["pytest>=8,<9"]
[project.scripts] [project.scripts]
serato-doctor = "serato_doctor.cli:main" serato-doctor = "serato_doctor.cli:main"
serato-doctor-web = "serato_doctor.web:main"
[tool.pytest.ini_options] [tool.pytest.ini_options]
testpaths = ["tests"] testpaths = ["tests"]
[tool.setuptools.packages.find] [tool.setuptools.packages.find]
include = ["serato_doctor*"] include = ["serato_doctor*"]
[tool.setuptools.package-data]
"serato_doctor.webui" = ["*.html", "*.css", "*.js"]
+33 -3
View File
@@ -10,6 +10,10 @@ SAMPLE_ROOT = Path(__file__).parent
SERATO_PATH_PREFIX = "Users/sample-user/OneDrive/Jukebox/" SERATO_PATH_PREFIX = "Users/sample-user/OneDrive/Jukebox/"
def database_record(tag: bytes, payload: bytes) -> bytes:
return tag + len(payload).to_bytes(4, "big") + payload
def load_manifest() -> dict: def load_manifest() -> dict:
return json.loads((SAMPLE_ROOT / "manifest.json").read_text(encoding="utf-8")) return json.loads((SAMPLE_ROOT / "manifest.json").read_text(encoding="utf-8"))
@@ -18,8 +22,7 @@ def build_sample(output: Optional[Path] = None) -> Path:
root = output or SAMPLE_ROOT / "generated" root = output or SAMPLE_ROOT / "generated"
manifest = load_manifest() manifest = load_manifest()
music_root = root / "Music" music_root = root / "Music"
crate_root = root / "Serato" / "_Serato_" / "Subcrates" serato_root = root / "Serato" / "_Serato_"
crate_root.mkdir(parents=True, exist_ok=True)
for relative_path in manifest["tracks"]: for relative_path in manifest["tracks"]:
track_path = music_root / relative_path track_path = music_root / relative_path
@@ -29,11 +32,38 @@ def build_sample(output: Optional[Path] = None) -> Path:
) )
for crate in manifest["crates"]: for crate in manifest["crates"]:
is_smart = crate["type"] == "smart"
folder_name = "SmartCrates" if is_smart else "Subcrates"
crate_root = serato_root / folder_name
crate_root.mkdir(parents=True, exist_ok=True)
output_name = (
Path(crate["name"]).with_suffix(".scrate").name
if is_smart
else crate["name"]
)
for other_folder in ("Subcrates", "Smartcrates", "SmartCrates"):
for stale_name in (
crate["name"],
Path(crate["name"]).with_suffix(".scrate").name,
):
stale_path = serato_root / other_folder / stale_name
if stale_path.exists() and stale_path != crate_root / output_name:
stale_path.unlink()
records = "".join( records = "".join(
f"{SERATO_PATH_PREFIX}{relative_path}otrk" f"{SERATO_PATH_PREFIX}{relative_path}otrk"
for relative_path in crate["references"] for relative_path in crate["references"]
) )
(crate_root / crate["name"]).write_bytes(records.encode("utf-16-le")) (crate_root / output_name).write_bytes(records.encode("utf-16-le"))
database_records = [
database_record(b"vrsn", "2.0/Serato Doctor Fixture".encode("utf-16-be"))
]
for relative_path in manifest["tracks"]:
fields = database_record(
b"pfil", f"{SERATO_PATH_PREFIX}{relative_path}".encode("utf-16-be")
)
database_records.append(database_record(b"otrk", fields))
(serato_root / "database V2").write_bytes(b"".join(database_records))
return root return root
+33 -7
View File
@@ -2,11 +2,12 @@ from pathlib import Path
import argparse import argparse
from serato_doctor.config import ScanConfig from serato_doctor.config import ScanConfig
from serato_doctor.crate_parser import parse_crates from serato_doctor.crate_parser import load_library_crates
from serato_doctor.database_parser import parse_database
from serato_doctor.health import analyze_health from serato_doctor.health import analyze_health
from serato_doctor.logging import configure_logging from serato_doctor.logging import configure_logging
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
from serato_doctor.scanner import scan_audio from serato_doctor.scanner import scan_filesystem
from serato_doctor.report import write_csv, write_missing_report from serato_doctor.report import write_csv, write_missing_report
@@ -55,11 +56,15 @@ def main():
logger.debug("Serato directory: %s", config.serato) logger.debug("Serato directory: %s", config.serato)
logger.debug("Music directory: %s", config.music) logger.debug("Music directory: %s", config.music)
library = Library.build( crates = load_library_crates(config.serato, config.reference_roots)
references=parse_crates( filesystem = scan_filesystem(config.music)
config.serato / "Subcrates", config.reference_roots database_path = config.serato / "database V2"
), database = parse_database(database_path) if database_path.is_file() else None
tracks=scan_audio(config.music), library = Library.from_crates(
crates=crates,
tracks=filesystem.tracks,
broken_symlinks=filesystem.broken_symlinks,
database=database,
) )
results = library.reconcile_by_filename() results = library.reconcile_by_filename()
missing_count = sum(1 for result in results if not result.exists_by_filename) missing_count = sum(1 for result in results if not result.exists_by_filename)
@@ -75,13 +80,34 @@ def main():
print(f"Score Basis: {health.score_basis}") print(f"Score Basis: {health.score_basis}")
print(f"Tracks: {health.disk_tracks}") print(f"Tracks: {health.disk_tracks}")
print(f"Crate References: {health.total_references}") print(f"Crate References: {health.total_references}")
print(f"References Scored: {health.scored_references}")
print(f"Healthy References: {health.healthy_references}") print(f"Healthy References: {health.healthy_references}")
print(f"Broken References: {health.missing_references}") print(f"Broken References: {health.missing_references}")
print(f"Unique Missing Filenames: {health.unique_missing_filenames}") print(f"Unique Missing Filenames: {health.unique_missing_filenames}")
print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}") print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}")
print(f"Duplicate Files: {health.duplicate_files}") print(f"Duplicate Files: {health.duplicate_files}")
print(
"Suspected Cloud Conflict Groups: "
f"{health.suspected_cloud_conflict_groups}"
)
print(
"Suspected Cloud Conflict Files: "
f"{health.suspected_cloud_conflict_files}"
)
print(f"Unused Tracks: {health.unused_tracks}") print(f"Unused Tracks: {health.unused_tracks}")
print(f"Suggested Matches: {health.suggested_matches}") print(f"Suggested Matches: {health.suggested_matches}")
print(f"Broken Symlinks: {health.broken_symlinks}")
print(f"Static Crates: {health.static_crates}")
print(f"Smart Crates: {health.smart_crates}")
print(f"Smart Crate Containers: {health.smart_crate_containers}")
print(
f"Dynamic References Excluded: {health.dynamic_references_excluded}"
)
print(f"Database Entries: {health.database_entries}")
print(f"Database / Library Matches: {health.database_library_matches}")
print(f"Database Entries Outside Scan: {health.database_unmatched_entries}")
print(f"Tracks Missing From Database: {health.tracks_missing_from_database}")
print(f"Duplicate Database Paths: {health.duplicate_database_paths}")
else: else:
write_csv(results, config.out) write_csv(results, config.out)
write_missing_report(results, config.report) write_missing_report(results, config.report)
+39 -2
View File
@@ -1,7 +1,7 @@
from pathlib import Path from pathlib import Path
from typing import Iterable, Tuple from typing import Iterable, Tuple
from serato_doctor.models.crate import Crate from serato_doctor.models.crate import Crate, CrateKind
from serato_doctor.models.reference import TrackReference from serato_doctor.models.reference import TrackReference
@@ -32,6 +32,17 @@ def path_markers(reference_roots: Iterable[Path]) -> Tuple[str, ...]:
return configured or DEFAULT_PATH_MARKERS return configured or DEFAULT_PATH_MARKERS
def classify_crate(crate_path: Path) -> CrateKind:
parent_names = {parent.name.casefold() for parent in crate_path.parents}
if "smartcrates" in parent_names:
return CrateKind.SMART
if crate_path.name.casefold() == "compatible by key.crate":
return CrateKind.SMART
if "subcrates" in parent_names:
return CrateKind.STATIC
return CrateKind.UNKNOWN
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate: def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
text = read_crate_text(crate_path) text = read_crate_text(crate_path)
refs = [] refs = []
@@ -63,7 +74,11 @@ def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
) )
) )
return Crate(path=crate_path, references=tuple(refs)) return Crate(
path=crate_path,
references=tuple(refs),
kind=classify_crate(crate_path),
)
def parse_crate( def parse_crate(
@@ -82,3 +97,25 @@ def parse_crates(
for crate in root.rglob("*.crate"): for crate in root.rglob("*.crate"):
refs.extend(parse_crate(crate, reference_roots)) refs.extend(parse_crate(crate, reference_roots))
return refs return refs
def load_library_crates(
serato_root: Path, reference_roots: Iterable[Path] = ()
) -> Tuple[Crate, ...]:
reference_roots = tuple(reference_roots)
if not serato_root.is_dir():
return ()
crates = []
folder_patterns = {
"subcrates": ("*.crate",),
"smartcrates": ("*.scrate", "*.crate"),
}
for folder in serato_root.iterdir():
patterns = folder_patterns.get(folder.name.casefold())
if patterns is None or not folder.is_dir():
continue
for pattern in patterns:
for crate_path in folder.rglob(pattern):
crates.append(load_crate(crate_path, reference_roots))
return tuple(sorted(crates, key=lambda crate: str(crate.path)))
+67
View File
@@ -0,0 +1,67 @@
from pathlib import Path
from typing import Dict, Iterator, Optional, Tuple
from serato_doctor.models.database import DatabaseTrack, SeratoDatabase
TEXT_FIELDS = {
b"tsng": "title",
b"tart": "artist",
b"talb": "album",
b"tgen": "genre",
}
def iter_records(data: bytes) -> Iterator[Tuple[bytes, bytes]]:
"""Yield complete big-endian tag-length-value records."""
offset = 0
while offset + 8 <= len(data):
tag = data[offset : offset + 4]
length = int.from_bytes(data[offset + 4 : offset + 8], "big")
payload_start = offset + 8
payload_end = payload_start + length
if payload_end > len(data):
break
yield tag, data[payload_start:payload_end]
offset = payload_end
def decode_text(payload: bytes) -> Optional[str]:
value = payload.decode("utf-16-be", errors="ignore").strip("\x00").strip()
return value or None
def normalize_database_path(value: str) -> Path:
if value.startswith(("Users/", "Volumes/")):
value = "/" + value
return Path(value)
def parse_track(payload: bytes) -> Optional[DatabaseTrack]:
fields: Dict[str, Optional[str]] = {}
path = None
for tag, value in iter_records(payload):
if tag == b"pfil":
decoded_path = decode_text(value)
if decoded_path:
path = normalize_database_path(decoded_path)
elif tag in TEXT_FIELDS:
fields[TEXT_FIELDS[tag]] = decode_text(value)
if path is None:
return None
return DatabaseTrack(path=path, filename=path.name, **fields)
def parse_database(database_path: Path) -> SeratoDatabase:
data = database_path.read_bytes()
version = None
tracks = []
for tag, payload in iter_records(data):
if tag == b"vrsn":
version = decode_text(payload)
elif tag == b"otrk":
track = parse_track(payload)
if track is not None:
tracks.append(track)
return SeratoDatabase(database_path, version, tuple(tracks))
+41
View File
@@ -0,0 +1,41 @@
from collections import defaultdict
from typing import DefaultDict, Iterable, List, Tuple
from serato_doctor.matching import cloud_conflict_name, normalize
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
from serato_doctor.models.track import DiskTrack
def find_duplicate_groups(
tracks: Iterable[DiskTrack],
) -> Tuple[DuplicateGroup, ...]:
"""Find exact-name duplicates and suspected numeric conflict copies."""
track_tuple = tuple(tracks)
by_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
for track in track_tuple:
by_name[normalize(track.filename)].append(track)
by_conflict_name[cloud_conflict_name(track.filename)].append(track)
groups = []
for key, matches in by_name.items():
if len(matches) > 1:
groups.append(_group(DuplicateKind.EXACT_NAME, key, matches))
for key, matches in by_conflict_name.items():
distinct_names = {normalize(track.filename) for track in matches}
if len(distinct_names) > 1:
groups.append(_group(DuplicateKind.CLOUD_CONFLICT, key, matches))
return tuple(
sorted(groups, key=lambda group: (group.kind.value, group.comparison_key))
)
def _group(
kind: DuplicateKind, key: str, tracks: Iterable[DiskTrack]
) -> DuplicateGroup:
ordered = tuple(sorted(tracks, key=lambda track: str(track.path)))
return DuplicateGroup(kind, key, ordered)
+72 -7
View File
@@ -1,6 +1,9 @@
from collections import Counter from collections import Counter
from serato_doctor.matching import MatchingEngine from serato_doctor.duplicates import find_duplicate_groups
from serato_doctor.matching import MatchingEngine, normalize
from serato_doctor.models.crate import CrateKind
from serato_doctor.models.duplicate import DuplicateKind
from serato_doctor.models.health import HealthReport from serato_doctor.models.health import HealthReport
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
@@ -8,15 +11,31 @@ from serato_doctor.models.library import Library
def analyze_health(library: Library) -> HealthReport: def analyze_health(library: Library) -> HealthReport:
"""Calculate defensible health metrics without changing the library.""" """Calculate defensible health metrics without changing the library."""
results = library.reconcile_by_filename() dynamic_sources = {
crate.path for crate in library.crates if crate.kind is CrateKind.SMART
}
results = tuple(
result
for result in library.reconcile_by_filename()
if result.reference.source not in dynamic_sources
)
missing = [result for result in results if not result.exists_by_filename] missing = [result for result in results if not result.exists_by_filename]
healthy_count = len(results) - len(missing) healthy_count = len(results) - len(missing)
score = ( score = (
round(healthy_count / len(results) * 100, 1) if results else None round(healthy_count / len(results) * 100, 1) if results else None
) )
disk_name_counts = Counter(track.filename for track in library.tracks) duplicate_groups = find_duplicate_groups(library.tracks)
duplicate_counts = [count for count in disk_name_counts.values() if count > 1] exact_duplicates = [
group
for group in duplicate_groups
if group.kind is DuplicateKind.EXACT_NAME
]
cloud_conflicts = [
group
for group in duplicate_groups
if group.kind is DuplicateKind.CLOUD_CONFLICT
]
referenced_names = {reference.filename for reference in library.references} referenced_names = {reference.filename for reference in library.references}
unused_count = sum( unused_count = sum(
1 for track in library.tracks if track.filename not in referenced_names 1 for track in library.tracks if track.filename not in referenced_names
@@ -27,17 +46,63 @@ def analyze_health(library: Library) -> HealthReport:
bool(matcher.candidates_for(result.reference)) for result in missing bool(matcher.candidates_for(result.reference)) for result in missing
) )
database_tracks = library.database.tracks if library.database else ()
database_names = {normalize(track.filename) for track in database_tracks}
library_names = {normalize(track.filename) for track in library.tracks}
database_path_counts = Counter(
normalize(str(track.path)) for track in database_tracks
)
return HealthReport( return HealthReport(
score=score, score=score,
total_references=len(results), total_references=len(library.references),
scored_references=len(results),
healthy_references=healthy_count, healthy_references=healthy_count,
missing_references=len(missing), missing_references=len(missing),
unique_missing_filenames=len( unique_missing_filenames=len(
{result.reference.filename for result in missing} {result.reference.filename for result in missing}
), ),
disk_tracks=len(library.tracks), disk_tracks=len(library.tracks),
duplicate_filename_groups=len(duplicate_counts), duplicate_filename_groups=len(exact_duplicates),
duplicate_files=sum(count - 1 for count in duplicate_counts), duplicate_files=sum(group.extra_files for group in exact_duplicates),
suspected_cloud_conflict_groups=len(cloud_conflicts),
suspected_cloud_conflict_files=sum(
group.extra_files for group in cloud_conflicts
),
unused_tracks=unused_count, unused_tracks=unused_count,
suggested_matches=suggested_count, suggested_matches=suggested_count,
broken_symlinks=len(library.broken_symlinks),
static_crates=sum(
crate.kind is CrateKind.STATIC for crate in library.crates
),
smart_crates=sum(
crate.kind is CrateKind.SMART and crate.is_smart_definition
for crate in library.crates
),
smart_crate_containers=sum(
crate.kind is CrateKind.SMART and not crate.is_smart_definition
for crate in library.crates
),
unknown_crates=sum(
crate.kind is CrateKind.UNKNOWN for crate in library.crates
),
dynamic_references_excluded=sum(
len(crate.references)
for crate in library.crates
if crate.kind is CrateKind.SMART
),
database_present=library.database is not None,
database_entries=len(database_tracks),
database_library_matches=sum(
normalize(track.filename) in library_names for track in database_tracks
),
database_unmatched_entries=sum(
normalize(track.filename) not in library_names for track in database_tracks
),
tracks_missing_from_database=sum(
normalize(track.filename) not in database_names for track in library.tracks
),
duplicate_database_paths=sum(
count - 1 for count in database_path_counts.values() if count > 1
),
) )
+11 -1
View File
@@ -1,4 +1,7 @@
from serato_doctor.models.crate import Crate from serato_doctor.models.crate import Crate, CrateKind
from serato_doctor.models.database import DatabaseTrack, SeratoDatabase
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
from serato_doctor.models.filesystem import BrokenSymlink, FilesystemScan
from serato_doctor.models.health import HealthReport from serato_doctor.models.health import HealthReport
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
from serato_doctor.models.match import MatchEvidence, TrackMatch from serato_doctor.models.match import MatchEvidence, TrackMatch
@@ -7,11 +10,18 @@ from serato_doctor.models.track import DiskTrack
__all__ = [ __all__ = [
"Crate", "Crate",
"CrateKind",
"DatabaseTrack",
"DiskTrack", "DiskTrack",
"DuplicateGroup",
"DuplicateKind",
"BrokenSymlink",
"FilesystemScan",
"HealthReport", "HealthReport",
"Library", "Library",
"MatchEvidence", "MatchEvidence",
"ReferenceResult", "ReferenceResult",
"SeratoDatabase",
"TrackMatch", "TrackMatch",
"TrackReference", "TrackReference",
] ]
+20
View File
@@ -1,13 +1,33 @@
from dataclasses import dataclass from dataclasses import dataclass
from enum import Enum
from pathlib import Path from pathlib import Path
from typing import Tuple from typing import Tuple
from serato_doctor.models.reference import TrackReference from serato_doctor.models.reference import TrackReference
class CrateKind(str, Enum):
STATIC = "static"
SMART = "smart"
UNKNOWN = "unknown"
@dataclass(frozen=True) @dataclass(frozen=True)
class Crate: class Crate:
"""A Serato crate and the track references parsed from it.""" """A Serato crate and the track references parsed from it."""
path: Path path: Path
references: Tuple[TrackReference, ...] references: Tuple[TrackReference, ...]
kind: CrateKind = CrateKind.UNKNOWN
@property
def hierarchy(self) -> Tuple[str, ...]:
return tuple(self.path.stem.split("≫≫"))
@property
def display_name(self) -> str:
return self.hierarchy[-1]
@property
def is_smart_definition(self) -> bool:
return self.path.suffix.casefold() == ".scrate"
+22
View File
@@ -0,0 +1,22 @@
from dataclasses import dataclass
from pathlib import Path
from typing import Optional, Tuple
@dataclass(frozen=True)
class DatabaseTrack:
"""Read-only metadata extracted from one database V2 track record."""
path: Path
filename: str
title: Optional[str] = None
artist: Optional[str] = None
album: Optional[str] = None
genre: Optional[str] = None
@dataclass(frozen=True)
class SeratoDatabase:
path: Path
version: Optional[str]
tracks: Tuple[DatabaseTrack, ...]
+30
View File
@@ -0,0 +1,30 @@
from dataclasses import dataclass
from enum import Enum
from typing import Tuple
from serato_doctor.models.track import DiskTrack
class DuplicateKind(str, Enum):
EXACT_NAME = "exact_name"
CLOUD_CONFLICT = "cloud_conflict"
@dataclass(frozen=True)
class DuplicateGroup:
"""A deterministic group of files that warrants duplicate review."""
kind: DuplicateKind
comparison_key: str
tracks: Tuple[DiskTrack, ...]
@property
def extra_files(self) -> int:
return max(0, len(self.tracks) - 1)
@property
def display_name(self) -> str:
return min(
(track.filename for track in self.tracks),
key=lambda name: (len(name), name.casefold()),
)
+21
View File
@@ -0,0 +1,21 @@
from dataclasses import dataclass
from pathlib import Path
from typing import Optional, Tuple
from serato_doctor.models.track import DiskTrack
@dataclass(frozen=True)
class BrokenSymlink:
"""A symbolic link whose target cannot be resolved."""
path: Path
target: Optional[Path]
@dataclass(frozen=True)
class FilesystemScan:
"""Read-only findings from one traversal of a music folder."""
tracks: Tuple[DiskTrack, ...]
broken_symlinks: Tuple[BrokenSymlink, ...]
+16 -1
View File
@@ -8,15 +8,30 @@ class HealthReport:
score: Optional[float] score: Optional[float]
total_references: int total_references: int
scored_references: int
healthy_references: int healthy_references: int
missing_references: int missing_references: int
unique_missing_filenames: int unique_missing_filenames: int
disk_tracks: int disk_tracks: int
duplicate_filename_groups: int duplicate_filename_groups: int
duplicate_files: int duplicate_files: int
suspected_cloud_conflict_groups: int
suspected_cloud_conflict_files: int
unused_tracks: int unused_tracks: int
suggested_matches: int suggested_matches: int
broken_symlinks: int
static_crates: int
smart_crates: int
smart_crate_containers: int
unknown_crates: int
dynamic_references_excluded: int
database_present: bool
database_entries: int
database_library_matches: int
database_unmatched_entries: int
tracks_missing_from_database: int
duplicate_database_paths: int
@property @property
def score_basis(self) -> str: def score_basis(self) -> str:
return "Resolved crate references / total crate references" return "Resolved non-dynamic references / non-dynamic references scored"
+37 -2
View File
@@ -1,6 +1,9 @@
from dataclasses import dataclass from dataclasses import dataclass
from typing import Iterable, Tuple from typing import Iterable, Optional, Tuple
from serato_doctor.models.crate import Crate
from serato_doctor.models.database import SeratoDatabase
from serato_doctor.models.filesystem import BrokenSymlink
from serato_doctor.models.reference import ReferenceResult, TrackReference from serato_doctor.models.reference import ReferenceResult, TrackReference
from serato_doctor.models.track import DiskTrack from serato_doctor.models.track import DiskTrack
@@ -11,14 +14,46 @@ class Library:
references: Tuple[TrackReference, ...] references: Tuple[TrackReference, ...]
tracks: Tuple[DiskTrack, ...] tracks: Tuple[DiskTrack, ...]
crates: Tuple[Crate, ...] = ()
broken_symlinks: Tuple[BrokenSymlink, ...] = ()
database: Optional[SeratoDatabase] = None
@classmethod @classmethod
def build( def build(
cls, cls,
references: Iterable[TrackReference], references: Iterable[TrackReference],
tracks: Iterable[DiskTrack], tracks: Iterable[DiskTrack],
broken_symlinks: Iterable[BrokenSymlink] = (),
database: Optional[SeratoDatabase] = None,
) -> "Library": ) -> "Library":
return cls(tuple(references), tuple(tracks)) return cls(
tuple(references),
tuple(tracks),
broken_symlinks=tuple(broken_symlinks),
database=database,
)
@classmethod
def from_crates(
cls,
crates: Iterable[Crate],
tracks: Iterable[DiskTrack],
broken_symlinks: Iterable[BrokenSymlink] = (),
database: Optional[SeratoDatabase] = None,
) -> "Library":
crate_tuple = tuple(crates)
references = tuple(
reference
for crate in crate_tuple
for reference in crate.references
)
return cls(
references,
tuple(tracks),
crate_tuple,
tuple(broken_symlinks),
database,
)
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]: def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
disk_names = {track.filename for track in self.tracks} disk_names = {track.filename for track in self.tracks}
+20 -2
View File
@@ -1,14 +1,23 @@
from pathlib import Path from pathlib import Path
from serato_doctor.models.filesystem import BrokenSymlink, FilesystemScan
from serato_doctor.models.track import DiskTrack from serato_doctor.models.track import DiskTrack
AUDIO_SUFFIXES = {".mp3", ".m4a", ".wav", ".aif", ".aiff", ".flac"} AUDIO_SUFFIXES = {".mp3", ".m4a", ".wav", ".aif", ".aiff", ".flac"}
def scan_audio(folder: Path) -> list[DiskTrack]: def scan_filesystem(folder: Path) -> FilesystemScan:
tracks = [] tracks = []
broken_symlinks = []
for path in folder.rglob("*"): for path in folder.rglob("*"):
if path.is_symlink() and not path.exists():
try:
target = path.readlink()
except OSError:
target = None
broken_symlinks.append(BrokenSymlink(path=path, target=target))
continue
if not path.is_file(): if not path.is_file():
continue continue
if path.suffix.lower() not in AUDIO_SUFFIXES: if path.suffix.lower() not in AUDIO_SUFFIXES:
@@ -28,4 +37,13 @@ def scan_audio(folder: Path) -> list[DiskTrack]:
) )
) )
return tracks return FilesystemScan(
tracks=tuple(tracks),
broken_symlinks=tuple(broken_symlinks),
)
def scan_audio(folder: Path) -> list[DiskTrack]:
"""Scan audio files while preserving the prototype API."""
return list(scan_filesystem(folder).tracks)
+114
View File
@@ -0,0 +1,114 @@
import argparse
import json
from dataclasses import asdict
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
from importlib import resources
from pathlib import Path
from typing import Iterable
from serato_doctor.crate_parser import load_library_crates
from serato_doctor.database_parser import parse_database
from serato_doctor.health import analyze_health
from serato_doctor.models.library import Library
from serato_doctor.scanner import scan_filesystem
MAX_REQUEST_BYTES = 64 * 1024
STATIC_FILES = {
"/": ("index.html", "text/html; charset=utf-8"),
"/app.css": ("app.css", "text/css; charset=utf-8"),
"/app.js": ("app.js", "text/javascript; charset=utf-8"),
}
def analyze_paths(
serato: Path, music: Path, reference_roots: Iterable[Path] = ()
) -> dict:
serato = serato.expanduser()
music = music.expanduser()
reference_roots = tuple(root.expanduser() for root in reference_roots)
if not serato.is_dir():
raise ValueError(f"Serato folder does not exist: {serato}")
if not music.is_dir():
raise ValueError(f"Music folder does not exist: {music}")
crates = load_library_crates(serato, reference_roots)
filesystem = scan_filesystem(music)
database_path = serato / "database V2"
database = parse_database(database_path) if database_path.is_file() else None
library = Library.from_crates(
crates,
filesystem.tracks,
filesystem.broken_symlinks,
database,
)
report = analyze_health(library)
result = asdict(report)
result["score_basis"] = report.score_basis
return result
class SeratoDoctorHandler(BaseHTTPRequestHandler):
def do_GET(self) -> None:
asset = STATIC_FILES.get(self.path)
if asset is None:
self._json_response(404, {"error": "Not found"})
return
filename, content_type = asset
content = (
resources.files("serato_doctor.webui")
.joinpath(filename)
.read_bytes()
)
self.send_response(200)
self.send_header("Content-Type", content_type)
self.send_header("Content-Length", str(len(content)))
self.end_headers()
self.wfile.write(content)
def do_POST(self) -> None:
if self.path != "/api/analyze":
self._json_response(404, {"error": "Not found"})
return
try:
length = int(self.headers.get("Content-Length", "0"))
if length <= 0 or length > MAX_REQUEST_BYTES:
raise ValueError("Invalid request size")
payload = json.loads(self.rfile.read(length))
if not isinstance(payload, dict):
raise ValueError("Request body must be a JSON object")
roots = [Path(value) for value in payload.get("reference_roots", [])]
result = analyze_paths(
Path(payload["serato"]), Path(payload["music"]), roots
)
except (KeyError, TypeError, json.JSONDecodeError, ValueError) as error:
self._json_response(400, {"error": str(error)})
return
self._json_response(200, result)
def _json_response(self, status: int, payload: dict) -> None:
content = json.dumps(payload).encode("utf-8")
self.send_response(status)
self.send_header("Content-Type", "application/json; charset=utf-8")
self.send_header("Content-Length", str(len(content)))
self.end_headers()
self.wfile.write(content)
def log_message(self, format: str, *args: object) -> None:
return
def main() -> None:
parser = argparse.ArgumentParser(prog="serato-doctor-web")
parser.add_argument("--host", default="127.0.0.1")
parser.add_argument("--port", type=int, default=8765)
args = parser.parse_args()
server = ThreadingHTTPServer((args.host, args.port), SeratoDoctorHandler)
print(f"Serato Doctor web interface: http://{args.host}:{args.port}")
print("Press Ctrl+C to stop.")
try:
server.serve_forever()
except KeyboardInterrupt:
pass
finally:
server.server_close()
+1
View File
@@ -0,0 +1 @@
"""Static assets for the local Serato Doctor web interface."""
File diff suppressed because one or more lines are too long
+54
View File
@@ -0,0 +1,54 @@
const form = document.querySelector('#analysis-form');
const button = document.querySelector('#analyze-button');
const errorBox = document.querySelector('#error-message');
const results = document.querySelector('#dashboard');
function expandHome(path) {
return path.trim();
}
function render(data) {
document.querySelectorAll('[data-field]').forEach((element) => {
const value = data[element.dataset.field];
element.textContent = value ?? '—';
});
const score = data.score;
document.querySelector('#health-score').textContent = score == null ? '—' : `${score}%`;
document.querySelector('#score-ring').style.setProperty('--score', score ?? 0);
document.querySelector('#health-message').textContent = score == null
? 'Not enough data yet'
: score >= 95 ? 'Looking excellent' : score >= 80 ? 'A few things need attention' : 'Review recommended';
document.querySelector('#score-basis').textContent = data.score_basis;
document.querySelector('#analysis-time').textContent = `Completed ${new Date().toLocaleTimeString([], {hour: '2-digit', minute: '2-digit'})}`;
results.hidden = false;
results.scrollIntoView({behavior: 'smooth', block: 'start'});
}
form.addEventListener('submit', async (event) => {
event.preventDefault();
errorBox.hidden = true;
button.disabled = true;
button.querySelector('span').textContent = 'Analyzing safely…';
const roots = document.querySelector('#reference-roots').value
.split('\n').map((value) => value.trim()).filter(Boolean);
try {
const response = await fetch('/api/analyze', {
method: 'POST',
headers: {'Content-Type': 'application/json'},
body: JSON.stringify({
serato: expandHome(document.querySelector('#serato-path').value),
music: expandHome(document.querySelector('#music-path').value),
reference_roots: roots,
}),
});
const data = await response.json();
if (!response.ok) throw new Error(data.error || 'Analysis failed');
render(data);
} catch (error) {
errorBox.textContent = error.message;
errorBox.hidden = false;
} finally {
button.disabled = false;
button.querySelector('span').textContent = 'Analyze library';
}
});
+89
View File
@@ -0,0 +1,89 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="color-scheme" content="dark">
<title>Serato Doctor</title>
<link rel="stylesheet" href="/app.css">
</head>
<body>
<div class="ambient ambient-one"></div>
<div class="ambient ambient-two"></div>
<div class="shell">
<aside class="sidebar">
<a class="brand" href="/" aria-label="Serato Doctor home">
<span class="brand-mark">SD</span>
<span><strong>Serato</strong><small>Doctor</small></span>
</a>
<nav aria-label="Primary navigation">
<a class="nav-item active" href="#dashboard"><span></span> Dashboard</a>
<a class="nav-item" href="#scan"><span></span> New analysis</a>
<a class="nav-item" href="#diagnostics"><span></span> Diagnostics</a>
</nav>
<div class="safety-card">
<span class="safety-icon"></span>
<div><strong>Read-only mode</strong><p>Your library will not be modified.</p></div>
</div>
<div class="sidebar-foot">Local interface · v0.1</div>
</aside>
<main>
<header class="topbar">
<div><p class="eyebrow">Library intelligence</p><h1>Good evening.</h1></div>
<div class="status-pill"><span></span> Local &amp; private</div>
</header>
<section id="scan" class="scan-panel panel">
<div class="panel-copy">
<p class="eyebrow">Start here</p>
<h2>Analyze your library</h2>
<p>Point Serato Doctor at your Serato and music folders. We inspect references, duplicates, smart crates, and symlinks without changing a thing.</p>
</div>
<form id="analysis-form">
<label>Serato folder
<input id="serato-path" name="serato" value="~/Music/_Serato_" required>
</label>
<label>Music folder
<input id="music-path" name="music" value="~/Music/Jukebox" required>
</label>
<label class="wide">Historical reference roots <span>optional · one per line</span>
<textarea id="reference-roots" rows="2" placeholder="/Users/old-user/OneDrive/Jukebox"></textarea>
</label>
<button id="analyze-button" type="submit"><span>Analyze library</span><b></b></button>
</form>
<div id="error-message" class="error" role="alert" hidden></div>
</section>
<section id="dashboard" class="results" aria-live="polite" hidden>
<div class="section-heading"><div><p class="eyebrow">Latest analysis</p><h2>Library health</h2></div><span id="analysis-time"></span></div>
<div class="hero-grid">
<article class="score-card panel">
<div class="score-ring" id="score-ring"><div><strong id="health-score"></strong><span>health</span></div></div>
<div><p class="score-label">Reference integrity</p><h3 id="health-message">Ready to analyze</h3><p id="score-basis">We only score evidence we can defend.</p></div>
</article>
<div class="metrics-grid">
<article class="metric panel"><span>Tracks</span><strong data-field="disk_tracks"></strong><small>audio files found</small></article>
<article class="metric panel warning"><span>Broken references</span><strong data-field="missing_references"></strong><small>static crate entries</small></article>
<article class="metric panel"><span>Unused tracks</span><strong data-field="unused_tracks"></strong><small>not referenced by crates</small></article>
<article class="metric panel"><span>Suggested matches</span><strong data-field="suggested_matches"></strong><small>explainable candidates</small></article>
</div>
</div>
<div id="diagnostics" class="diagnostics panel">
<div class="section-heading"><div><p class="eyebrow">Full picture</p><h2>Diagnostics</h2></div><span class="read-only-tag">No changes made</span></div>
<div class="diagnostic-list">
<div><span class="diag-icon violet"></span><p><strong>Crates</strong><small><b data-field="static_crates"></b> static · <b data-field="smart_crates"></b> smart · <b data-field="smart_crate_containers"></b> dynamic containers · <b data-field="dynamic_references_excluded"></b> references excluded</small></p></div>
<div><span class="diag-icon amber"></span><p><strong>Duplicate filenames</strong><small><b data-field="duplicate_filename_groups"></b> exact groups · <b data-field="duplicate_files"></b> extra files</small></p></div>
<div><span class="diag-icon blue"></span><p><strong>Cloud conflicts</strong><small><b data-field="suspected_cloud_conflict_groups"></b> suspected groups · <b data-field="suspected_cloud_conflict_files"></b> extra files</small></p></div>
<div><span class="diag-icon red"></span><p><strong>Broken symlinks</strong><small><b data-field="broken_symlinks"></b> unresolved links</small></p></div>
<div><span class="diag-icon violet"></span><p><strong>Serato database V2</strong><small><b data-field="database_entries"></b> entries · <b data-field="database_library_matches"></b> match scanned filenames · <b data-field="tracks_missing_from_database"></b> scanned tracks absent</small></p></div>
<div><span class="diag-icon amber"></span><p><strong>Database review</strong><small><b data-field="database_unmatched_entries"></b> entries outside this music scan · <b data-field="duplicate_database_paths"></b> duplicate paths</small></p></div>
</div>
</div>
</section>
</main>
</div>
<script src="/app.js" defer></script>
</body>
</html>
+1
View File
@@ -123,6 +123,7 @@ def test_analyze_prints_health_without_writing_reports(tmp_path, monkeypatch, ca
assert "Overall Health: 50.0%" in output assert "Overall Health: 50.0%" in output
assert "Tracks: 1" in output assert "Tracks: 1" in output
assert "Crate References: 2" in output assert "Crate References: 2" in output
assert "References Scored: 2" in output
assert "Healthy References: 1" in output assert "Healthy References: 1" in output
assert "Broken References: 1" in output assert "Broken References: 1" in output
assert "Unused Tracks: 0" in output assert "Unused Tracks: 0" in output
+50
View File
@@ -3,12 +3,15 @@ from pathlib import Path
import pytest import pytest
from serato_doctor.crate_parser import ( from serato_doctor.crate_parser import (
classify_crate,
clean_path, clean_path,
load_crate, load_crate,
load_library_crates,
parse_crate, parse_crate,
parse_crates, parse_crates,
path_markers, path_markers,
) )
from serato_doctor.models.crate import CrateKind
@pytest.mark.parametrize( @pytest.mark.parametrize(
@@ -84,3 +87,50 @@ def test_parse_crates_reuses_configured_roots_for_every_crate(tmp_path):
"First.mp3", "First.mp3",
"Second.mp3", "Second.mp3",
} }
@pytest.mark.parametrize(
("relative_path", "expected"),
[
("_Serato_/Subcrates/House.crate", CrateKind.STATIC),
("_Serato_/SmartCrates/Warmup.scrate", CrateKind.SMART),
("_Serato_/Subcrates/Compatible by key.crate", CrateKind.SMART),
("fixtures/Unknown.crate", CrateKind.UNKNOWN),
],
)
def test_classify_crate_uses_provenance_and_known_dynamic_name(
tmp_path, relative_path, expected
):
assert classify_crate(tmp_path / relative_path) is expected
def test_smart_crate_preserves_encoded_hierarchy(tmp_path):
crate_path = (
tmp_path / "_Serato_" / "SmartCrates" / "Compatible by key≫≫10A.scrate"
)
crate_path.parent.mkdir(parents=True)
crate_path.write_bytes(b"")
crate = load_crate(crate_path)
assert crate.kind is CrateKind.SMART
assert crate.hierarchy == ("Compatible by key", "10A")
assert crate.display_name == "10A"
assert crate.is_smart_definition
def test_load_library_crates_discovers_scrate_definitions(tmp_path):
smart_folder = tmp_path / "SmartCrates"
static_folder = tmp_path / "Subcrates"
smart_folder.mkdir()
static_folder.mkdir()
(smart_folder / "New EDM.scrate").write_bytes(b"")
(smart_folder / "Re-Drums.scrate").write_bytes(b"")
(static_folder / "House.crate").write_bytes(b"")
crates = load_library_crates(tmp_path)
assert [(crate.path.name, crate.kind) for crate in crates] == [
("New EDM.scrate", CrateKind.SMART),
("Re-Drums.scrate", CrateKind.SMART),
("House.crate", CrateKind.STATIC),
]
+48
View File
@@ -0,0 +1,48 @@
from pathlib import Path
from serato_doctor.database_parser import iter_records, parse_database
def record(tag, payload):
return tag + len(payload).to_bytes(4, "big") + payload
def text_record(tag, value):
return record(tag, value.encode("utf-16-be"))
def test_parse_database_reads_track_paths_and_metadata(tmp_path):
track = b"".join(
[
text_record(b"pfil", "Users/sample/Music/Track.mp3"),
text_record(b"tsng", "Track title"),
text_record(b"tart", "Test artist"),
text_record(b"talb", "Test album"),
text_record(b"tgen", "House"),
]
)
database_path = tmp_path / "database V2"
database_path.write_bytes(
text_record(b"vrsn", "2.0/Test Database") + record(b"otrk", track)
)
database = parse_database(database_path)
assert database.version == "2.0/Test Database"
assert len(database.tracks) == 1
parsed = database.tracks[0]
assert parsed.path == Path("/Users/sample/Music/Track.mp3")
assert parsed.filename == "Track.mp3"
assert parsed.title == "Track title"
assert parsed.artist == "Test artist"
assert parsed.album == "Test album"
assert parsed.genre == "House"
def test_iter_records_ignores_incomplete_trailing_record():
complete = record(b"vrsn", "2.0".encode("utf-16-be"))
incomplete = b"otrk\x00\x00\x00\x10short"
records = list(iter_records(complete + incomplete))
assert records == [(b"vrsn", "2.0".encode("utf-16-be"))]
+57
View File
@@ -0,0 +1,57 @@
from pathlib import Path
from serato_doctor.duplicates import find_duplicate_groups
from serato_doctor.models.duplicate import DuplicateKind
from serato_doctor.models.track import DiskTrack
def track(path):
path = Path(path)
return DiskTrack(path, path.name, 100, path.suffix.lower())
def test_exact_names_in_different_folders_form_a_group():
groups = find_duplicate_groups(
[track("/A/Song.mp3"), track("/B/Song.mp3")]
)
assert len(groups) == 1
assert groups[0].kind is DuplicateKind.EXACT_NAME
assert groups[0].display_name == "Song.mp3"
assert groups[0].extra_files == 1
def test_case_only_difference_is_an_exact_name_duplicate():
groups = find_duplicate_groups(
[track("/A/SONG.MP3"), track("/B/song.mp3")]
)
assert groups[0].kind is DuplicateKind.EXACT_NAME
def test_numeric_suffixes_form_a_suspected_cloud_conflict_group():
groups = find_duplicate_groups(
[
track("/A/Track.mp3"),
track("/B/Track 2.mp3"),
track("/C/Track 3.mp3"),
]
)
assert len(groups) == 1
assert groups[0].kind is DuplicateKind.CLOUD_CONFLICT
assert groups[0].display_name == "Track.mp3"
assert groups[0].extra_files == 2
assert [item.path for item in groups[0].tracks] == [
Path("/A/Track.mp3"),
Path("/B/Track 2.mp3"),
Path("/C/Track 3.mp3"),
]
def test_unrelated_and_single_files_are_not_reported():
groups = find_duplicate_groups(
[track("/A/First.mp3"), track("/B/Second.mp3")]
)
assert groups == ()
+66 -1
View File
@@ -1,6 +1,8 @@
from pathlib import Path from pathlib import Path
from serato_doctor.health import analyze_health from serato_doctor.health import analyze_health
from serato_doctor.models.crate import Crate, CrateKind
from serato_doctor.models.filesystem import BrokenSymlink
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
from serato_doctor.models.reference import TrackReference from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack from serato_doctor.models.track import DiskTrack
@@ -36,9 +38,10 @@ def test_health_report_exposes_each_metric():
assert report.score == 33.3 assert report.score == 33.3
assert report.score_basis == ( assert report.score_basis == (
"Resolved crate references / total crate references" "Resolved non-dynamic references / non-dynamic references scored"
) )
assert report.total_references == 3 assert report.total_references == 3
assert report.scored_references == 3
assert report.healthy_references == 1 assert report.healthy_references == 1
assert report.missing_references == 2 assert report.missing_references == 2
assert report.unique_missing_filenames == 2 assert report.unique_missing_filenames == 2
@@ -54,3 +57,65 @@ def test_empty_library_has_no_health_score():
assert report.score is None assert report.score is None
assert report.total_references == 0 assert report.total_references == 0
assert report.scored_references == 0
def test_smart_crate_references_are_reported_but_not_scored():
static_reference = reference("Found.mp3")
smart_reference = TrackReference(
Path("SmartCrates/Dynamic.scrate"),
Path("/old/House/Dynamic.mp3"),
"Dynamic.mp3",
)
crates = [
Crate(Path("Subcrates/Static.crate"), (static_reference,), CrateKind.STATIC),
Crate(
Path("SmartCrates/Dynamic.scrate"),
(smart_reference,),
CrateKind.SMART,
),
]
library = Library.from_crates(crates, [track("Found.mp3")])
report = analyze_health(library)
assert report.score == 100.0
assert report.total_references == 2
assert report.scored_references == 1
assert report.missing_references == 0
assert report.static_crates == 1
assert report.smart_crates == 1
assert report.smart_crate_containers == 0
assert report.dynamic_references_excluded == 1
def test_health_reports_cloud_conflicts_separately_from_exact_duplicates():
library = Library.build(
[],
[
track("Track.mp3", "Original"),
track("Track 2.mp3", "Conflict"),
track("Copy.mp3", "First"),
track("Copy.mp3", "Second"),
],
)
report = analyze_health(library)
assert report.duplicate_filename_groups == 1
assert report.duplicate_files == 1
assert report.suspected_cloud_conflict_groups == 1
assert report.suspected_cloud_conflict_files == 1
def test_health_reports_broken_symlinks_without_changing_score():
library = Library.build(
[reference("Found.mp3")],
[track("Found.mp3")],
[BrokenSymlink(Path("/music/Broken.mp3"), Path("missing.mp3"))],
)
report = analyze_health(library)
assert report.score == 100.0
assert report.broken_symlinks == 1
+11 -5
View File
@@ -1,7 +1,8 @@
import runpy import runpy
from pathlib import Path from pathlib import Path
from serato_doctor.crate_parser import parse_crates from serato_doctor.crate_parser import load_library_crates
from serato_doctor.models.crate import CrateKind
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
from serato_doctor.scanner import scan_audio from serato_doctor.scanner import scan_audio
@@ -13,16 +14,21 @@ def test_generated_sample_library_has_expected_scenario(tmp_path):
build_sample = runpy.run_path(str(generator_path))["build_sample"] build_sample = runpy.run_path(str(generator_path))["build_sample"]
sample_root = build_sample(tmp_path / "sample") sample_root = build_sample(tmp_path / "sample")
references = parse_crates(sample_root / "Serato" / "_Serato_" / "Subcrates") crates = load_library_crates(sample_root / "Serato" / "_Serato_")
tracks = scan_audio(sample_root / "Music") tracks = scan_audio(sample_root / "Music")
results = Library.build(references, tracks).reconcile_by_filename() library = Library.from_crates(crates, tracks)
results = library.reconcile_by_filename()
missing = { missing = {
result.reference.filename result.reference.filename
for result in results for result in results
if not result.exists_by_filename if not result.exists_by_filename
} }
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7 crate_files = list((sample_root / "Serato").rglob("*.crate"))
assert len(references) == 13 smart_files = list((sample_root / "Serato").rglob("*.scrate"))
assert len(crate_files) + len(smart_files) == 7
assert len(library.references) == 13
assert len(tracks) == 10 assert len(tracks) == 10
assert missing == {"Missing.mp3", "Old Name.mp3"} assert missing == {"Missing.mp3", "Old Name.mp3"}
assert sum(crate.kind is CrateKind.STATIC for crate in crates) == 5
assert sum(crate.kind is CrateKind.SMART for crate in crates) == 2
+34 -1
View File
@@ -1,4 +1,6 @@
from serato_doctor.scanner import scan_audio from pathlib import Path
from serato_doctor.scanner import scan_audio, scan_filesystem
def test_scan_audio_finds_supported_files(tmp_path): def test_scan_audio_finds_supported_files(tmp_path):
@@ -15,3 +17,34 @@ def test_scan_audio_finds_supported_files(tmp_path):
first = next(track for track in tracks if track.filename == "First.MP3") first = next(track for track in tracks if track.filename == "First.MP3")
assert first.suffix == ".mp3" assert first.suffix == ".mp3"
assert first.size == len(b"synthetic audio") assert first.size == len(b"synthetic audio")
def test_scan_filesystem_reports_broken_symlink(tmp_path):
music = tmp_path / "music"
music.mkdir()
link = music / "Missing.mp3"
link.symlink_to("not-there.mp3")
result = scan_filesystem(music)
assert result.tracks == ()
assert len(result.broken_symlinks) == 1
assert result.broken_symlinks[0].path == link
assert result.broken_symlinks[0].target == Path("not-there.mp3")
def test_valid_audio_symlink_is_scanned_normally(tmp_path):
music = tmp_path / "music"
music.mkdir()
target = music / "Target.mp3"
target.write_bytes(b"synthetic audio")
link = music / "Linked.mp3"
link.symlink_to(target)
result = scan_filesystem(music)
assert {track.filename for track in result.tracks} == {
"Linked.mp3",
"Target.mp3",
}
assert result.broken_symlinks == ()
+38
View File
@@ -0,0 +1,38 @@
import runpy
from pathlib import Path
import pytest
from serato_doctor.web import STATIC_FILES, analyze_paths
def test_web_analysis_uses_production_health_pipeline(tmp_path):
generator = runpy.run_path(
str(Path(__file__).parents[1] / "samples/small-library/generate.py")
)
sample = generator["build_sample"](tmp_path / "sample")
result = analyze_paths(
sample / "Serato" / "_Serato_", sample / "Music"
)
assert result["score"] == 77.8
assert result["disk_tracks"] == 10
assert result["missing_references"] == 2
assert result["static_crates"] == 5
assert result["smart_crates"] == 2
assert result["database_entries"] == 10
assert result["database_library_matches"] == 10
assert result["tracks_missing_from_database"] == 0
def test_web_analysis_rejects_missing_folders(tmp_path):
with pytest.raises(ValueError, match="Serato folder does not exist"):
analyze_paths(tmp_path / "missing", tmp_path)
def test_web_static_assets_are_declared_and_packaged():
asset_root = Path(__file__).parents[1] / "serato_doctor" / "webui"
assert set(STATIC_FILES) == {"/", "/app.css", "/app.js"}
assert all((asset_root / filename).is_file() for filename, _ in STATIC_FILES.values())