Compare commits

..

7 Commits

Author SHA1 Message Date
Philip Guzman c92d929f37 Merge feature/analyze-command into develop 2026-06-30 18:17:20 -07:00
Philip Guzman e6f313d98a Add read-only analyze command 2026-06-30 17:55:45 -07:00
Philip Guzman 22424d0b5c Merge feature/health-engine into develop 2026-06-30 17:54:20 -07:00
Philip Guzman b8bf9ec6e1 Add transparent library health engine 2026-06-30 17:47:54 -07:00
Philip Guzman 167f029c28 Merge feature/matching-engine into develop 2026-06-30 17:46:56 -07:00
Philip Guzman 5a2edb8b94 Add explainable matching engine 2026-06-30 17:42:33 -07:00
Philip Guzman 29e53f7dc6 Merge feature/logging into develop 2026-06-30 17:40:57 -07:00
14 changed files with 533 additions and 15 deletions
+37
View File
@@ -0,0 +1,37 @@
# Serato Doctor
Serato Doctor is a read-only-first toolkit for inspecting and eventually repairing
DJ libraries. Serato is the first supported engine.
## Analyze a Library
```shell
serato-doctor analyze \
--serato ~/Music/_Serato_ \
--music ~/Music/Jukebox
```
Analysis prints a transparent reference-integrity score plus broken references,
duplicate filenames, unused tracks, and suggested filename matches. It does not
modify the library or create report files.
If crates contain an older library root, supply it explicitly:
```shell
serato-doctor analyze --reference-root /Users/old-user/OneDrive/Jukebox
```
## Generate Detailed Reports
Running without a command preserves the original scanner behavior and writes CSV
and text reports:
```shell
serato-doctor --serato ~/Music/_Serato_ --music ~/Music/Jukebox
```
Use `--verbose` for progress on standard error or `--log-file PATH` for an
aggregate diagnostic log.
Serato Doctor never repairs files without an explicit future repair workflow,
preview, backup, and rollback path.
+3 -1
View File
@@ -13,16 +13,18 @@
- [ ] Database V2 read-only parser - [ ] Database V2 read-only parser
- [x] Configuration - [x] Configuration
- [x] Logging - [x] Logging
- [x] Matching engine
## v0.2 — Diagnostics ## v0.2 — Diagnostics
- [x] `serato-doctor analyze` command
- [ ] Duplicate filename detection - [ ] Duplicate filename detection
- [ ] Duplicate audio hash detection - [ ] Duplicate audio hash detection
- [ ] Broken symlink detection - [ ] Broken symlink detection
- [ ] Orphaned audio detection - [ ] Orphaned audio detection
- [ ] OneDrive rename detection - [ ] OneDrive rename detection
- [ ] Crate classification: static vs smart/dynamic - [ ] Crate classification: static vs smart/dynamic
- [ ] Library health score - [x] Library health score
## v0.3 — Safe Repair ## v0.3 — Safe Repair
+27
View File
@@ -0,0 +1,27 @@
# Analyze Command
## Problem
The health and matching engines are only Python APIs. Users need one safe command
that summarizes library integrity without first interpreting CSV files.
## Architecture
`serato-doctor analyze` reuses the existing configured crate parser and filesystem
scanner, then prints the immutable health report. The original no-command mode is
retained as the report-producing scan workflow.
Analyze mode does not create CSV or text reports. Both modes remain read-only with
respect to Serato crates, databases, and music files.
## Edge Cases
- A library with no references displays `Not assessed` rather than a false score.
- Historical roots continue to work through repeatable `--reference-root` flags.
- Normal output remains separate from optional diagnostic logging.
- Existing scripts that invoke the CLI without a subcommand remain compatible.
## Verification
Tests prove the legacy five-line output and reports remain unchanged, while analyze
prints health metrics and does not create output files.
+31
View File
@@ -0,0 +1,31 @@
# Library Health Engine
## Problem
Raw missing-reference counts do not provide a compact view of library integrity,
but an opaque blended score would imply confidence the current data cannot support.
## Architecture
The health engine produces an immutable report from the core `Library`. Its score
is only the percentage of crate references resolved by exact filename. The report
also exposes missing references, unique missing filenames, duplicate filename
groups, extra duplicate files, unused tracks, and missing references with matching
candidates.
Duplicate, unused, and candidate counts are informational. They do not affect the
score until the project has a documented and validated weighting policy. An empty
library has no score rather than a misleading 0% or 100%.
## Edge Cases
- The same missing filename referenced by several crates.
- Several disk files sharing a filename.
- Conflict-suffixed files that are unused but may be match candidates.
- Libraries with no crate references.
- Tracks referenced by filename from more than one crate.
## Verification
Tests assert every metric, the disclosed score basis, candidate integration, and
empty-library behavior. Analysis remains entirely read-only.
+29
View File
@@ -0,0 +1,29 @@
# Explainable Matching Engine
## Problem
Missing references need ranked candidate files, but a filename-only yes/no check
cannot explain ambiguity or cloud-provider conflict names.
## Architecture
The read-only matching engine indexes normalized filenames and scores only related
candidates. Every score contains evidence for filename, extension, and parent
folder. Exact filenames earn 60 points, normalized names 55, numeric conflict-name
matches 50, extensions 10, and parent folders 20.
The displayed percentage is an evidence score, not a statistical probability.
Metadata, duration, hashes, and fingerprints can add stronger evidence later.
## Edge Cases
- Unicode and case differences.
- OneDrive-style names such as `Track 2.mp3`.
- Duplicate candidates in different folders.
- Legitimate numbered song titles, which remain candidates but are never repaired.
- Unrelated names, which are not emitted as candidates.
## Verification
Tests cover exact, normalized, conflict-suffix, ambiguous, and unrelated filenames.
Candidate ordering is deterministic. The engine never changes a track or reference.
+34 -5
View File
@@ -3,6 +3,7 @@ import argparse
from serato_doctor.config import ScanConfig from serato_doctor.config import ScanConfig
from serato_doctor.crate_parser import parse_crates from serato_doctor.crate_parser import parse_crates
from serato_doctor.health import analyze_health
from serato_doctor.logging import configure_logging from serato_doctor.logging import configure_logging
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
from serato_doctor.scanner import scan_audio from serato_doctor.scanner import scan_audio
@@ -11,10 +12,23 @@ from serato_doctor.report import write_csv, write_missing_report
def main(): def main():
parser = argparse.ArgumentParser(prog="serato-doctor") parser = argparse.ArgumentParser(prog="serato-doctor")
parser.add_argument(
"command", nargs="?", choices=("scan", "analyze"), default="scan"
)
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_")) parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox")) parser.add_argument(
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")) "--music",
parser.add_argument("--report", default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt")) default=str(
Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"
),
)
parser.add_argument(
"--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")
)
parser.add_argument(
"--report",
default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"),
)
parser.add_argument( parser.add_argument(
"--reference-root", "--reference-root",
action="append", action="append",
@@ -37,7 +51,7 @@ def main():
log_file=Path(args.log_file) if args.log_file else None, log_file=Path(args.log_file) if args.log_file else None,
) )
logger = configure_logging(config.verbose, config.log_file) logger = configure_logging(config.verbose, config.log_file)
logger.info("Starting read-only library scan") logger.info("Starting read-only library inspection")
logger.debug("Serato directory: %s", config.serato) logger.debug("Serato directory: %s", config.serato)
logger.debug("Music directory: %s", config.music) logger.debug("Music directory: %s", config.music)
@@ -53,10 +67,25 @@ def main():
logger.info("Scanned %d disk tracks", len(library.tracks)) logger.info("Scanned %d disk tracks", len(library.tracks))
logger.info("Found %d references missing by filename", missing_count) logger.info("Found %d references missing by filename", missing_count)
if args.command == "analyze":
health = analyze_health(library)
score = f"{health.score:.1f}%" if health.score is not None else "Not assessed"
logger.info("Calculated library health: %s", score)
print(f"Overall Health: {score}")
print(f"Score Basis: {health.score_basis}")
print(f"Tracks: {health.disk_tracks}")
print(f"Crate References: {health.total_references}")
print(f"Healthy References: {health.healthy_references}")
print(f"Broken References: {health.missing_references}")
print(f"Unique Missing Filenames: {health.unique_missing_filenames}")
print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}")
print(f"Duplicate Files: {health.duplicate_files}")
print(f"Unused Tracks: {health.unused_tracks}")
print(f"Suggested Matches: {health.suggested_matches}")
else:
write_csv(results, config.out) write_csv(results, config.out)
write_missing_report(results, config.report) write_missing_report(results, config.report)
logger.info("Wrote CSV and missing-reference reports") logger.info("Wrote CSV and missing-reference reports")
print(f"Crate references: {len(library.references)}") print(f"Crate references: {len(library.references)}")
print(f"Disk tracks: {len(library.tracks)}") print(f"Disk tracks: {len(library.tracks)}")
print(f"Missing by filename: {missing_count}") print(f"Missing by filename: {missing_count}")
+43
View File
@@ -0,0 +1,43 @@
from collections import Counter
from serato_doctor.matching import MatchingEngine
from serato_doctor.models.health import HealthReport
from serato_doctor.models.library import Library
def analyze_health(library: Library) -> HealthReport:
"""Calculate defensible health metrics without changing the library."""
results = library.reconcile_by_filename()
missing = [result for result in results if not result.exists_by_filename]
healthy_count = len(results) - len(missing)
score = (
round(healthy_count / len(results) * 100, 1) if results else None
)
disk_name_counts = Counter(track.filename for track in library.tracks)
duplicate_counts = [count for count in disk_name_counts.values() if count > 1]
referenced_names = {reference.filename for reference in library.references}
unused_count = sum(
1 for track in library.tracks if track.filename not in referenced_names
)
matcher = MatchingEngine(library.tracks)
suggested_count = sum(
bool(matcher.candidates_for(result.reference)) for result in missing
)
return HealthReport(
score=score,
total_references=len(results),
healthy_references=healthy_count,
missing_references=len(missing),
unique_missing_filenames=len(
{result.reference.filename for result in missing}
),
disk_tracks=len(library.tracks),
duplicate_filename_groups=len(duplicate_counts),
duplicate_files=sum(count - 1 for count in duplicate_counts),
unused_tracks=unused_count,
suggested_matches=suggested_count,
)
+86
View File
@@ -0,0 +1,86 @@
import re
import unicodedata
from collections import defaultdict
from pathlib import Path
from typing import DefaultDict, Iterable, List, Tuple
from serato_doctor.models.match import MatchEvidence, TrackMatch
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def normalize(value: str) -> str:
return unicodedata.normalize("NFKC", value).casefold()
def cloud_conflict_name(filename: str) -> str:
"""Remove a trailing numeric cloud-conflict suffix from a filename stem."""
path = Path(filename)
stem = re.sub(r" \d+$", "", path.stem)
return normalize(stem + path.suffix)
def score_candidate(reference: TrackReference, track: DiskTrack) -> TrackMatch:
reference_name = reference.filename
track_name = track.filename
if reference_name == track_name:
filename_points = 60
filename_reason = "Filename is identical"
elif normalize(reference_name) == normalize(track_name):
filename_points = 55
filename_reason = "Filename matches after case and Unicode normalization"
elif cloud_conflict_name(reference_name) == cloud_conflict_name(track_name):
filename_points = 50
filename_reason = "Filename matches after removing a numeric conflict suffix"
else:
filename_points = 0
filename_reason = "Filename does not match"
same_extension = normalize(reference.path.suffix) == normalize(track.suffix)
same_parent = normalize(reference.path.parent.name) == normalize(
track.path.parent.name
)
evidence = (
MatchEvidence(
"filename",
filename_points > 0,
filename_points,
60,
filename_reason,
),
MatchEvidence(
"extension",
same_extension,
10 if same_extension else 0,
10,
"File extension matches" if same_extension else "File extension differs",
),
MatchEvidence(
"parent_folder",
same_parent,
20 if same_parent else 0,
20,
"Parent folder matches" if same_parent else "Parent folder differs",
),
)
return TrackMatch(reference, track, evidence)
class MatchingEngine:
"""Find and rank filename-related disk candidates without modifying files."""
def __init__(self, tracks: Iterable[DiskTrack]):
self._by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
for track in tracks:
self._by_conflict_name[cloud_conflict_name(track.filename)].append(track)
def candidates_for(self, reference: TrackReference) -> Tuple[TrackMatch, ...]:
candidates = self._by_conflict_name.get(
cloud_conflict_name(reference.filename), []
)
matches = [score_candidate(reference, track) for track in candidates]
return tuple(
sorted(matches, key=lambda match: (-match.score, str(match.track.path)))
)
+12 -1
View File
@@ -1,6 +1,17 @@
from serato_doctor.models.crate import Crate from serato_doctor.models.crate import Crate
from serato_doctor.models.health import HealthReport
from serato_doctor.models.library import Library from serato_doctor.models.library import Library
from serato_doctor.models.match import MatchEvidence, TrackMatch
from serato_doctor.models.reference import ReferenceResult, TrackReference from serato_doctor.models.reference import ReferenceResult, TrackReference
from serato_doctor.models.track import DiskTrack from serato_doctor.models.track import DiskTrack
__all__ = ["Crate", "DiskTrack", "Library", "ReferenceResult", "TrackReference"] __all__ = [
"Crate",
"DiskTrack",
"HealthReport",
"Library",
"MatchEvidence",
"ReferenceResult",
"TrackMatch",
"TrackReference",
]
+22
View File
@@ -0,0 +1,22 @@
from dataclasses import dataclass
from typing import Optional
@dataclass(frozen=True)
class HealthReport:
"""Transparent aggregate findings from a read-only library analysis."""
score: Optional[float]
total_references: int
healthy_references: int
missing_references: int
unique_missing_filenames: int
disk_tracks: int
duplicate_filename_groups: int
duplicate_files: int
unused_tracks: int
suggested_matches: int
@property
def score_basis(self) -> str:
return "Resolved crate references / total crate references"
+39
View File
@@ -0,0 +1,39 @@
from dataclasses import dataclass
from typing import Tuple
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
@dataclass(frozen=True)
class MatchEvidence:
"""One explainable scoring decision for a candidate track."""
field: str
matched: bool
points: int
max_points: int
explanation: str
@dataclass(frozen=True)
class TrackMatch:
"""A ranked candidate backed by explicit, inspectable evidence."""
reference: TrackReference
track: DiskTrack
evidence: Tuple[MatchEvidence, ...]
@property
def score(self) -> int:
return sum(item.points for item in self.evidence)
@property
def max_score(self) -> int:
return sum(item.max_points for item in self.evidence)
@property
def score_percent(self) -> float:
if not self.max_score:
return 0.0
return round(self.score / self.max_score * 100, 1)
+44
View File
@@ -84,3 +84,47 @@ def test_cli_accepts_old_reference_root(tmp_path, monkeypatch, capsys):
output = capsys.readouterr().out output = capsys.readouterr().out
assert "Crate references: 1" in output assert "Crate references: 1" in output
assert "Missing by filename: 0" in output assert "Missing by filename: 0" in output
def test_analyze_prints_health_without_writing_reports(tmp_path, monkeypatch, capsys):
serato = tmp_path / "serato"
subcrates = serato / "Subcrates"
music = tmp_path / "music"
subcrates.mkdir(parents=True)
music.mkdir()
crate_text = (
"Users/sample-user/Jukebox/Found.mp3otrk"
"Users/sample-user/Jukebox/Missing.mp3otrk"
)
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
(music / "Found.mp3").write_bytes(b"synthetic audio")
csv_path = tmp_path / "scan.csv"
report_path = tmp_path / "report.txt"
monkeypatch.setattr(
sys,
"argv",
[
"serato-doctor",
"analyze",
"--serato",
str(serato),
"--music",
str(music),
"--out",
str(csv_path),
"--report",
str(report_path),
],
)
main()
output = capsys.readouterr().out
assert "Overall Health: 50.0%" in output
assert "Tracks: 1" in output
assert "Crate References: 2" in output
assert "Healthy References: 1" in output
assert "Broken References: 1" in output
assert "Unused Tracks: 0" in output
assert not csv_path.exists()
assert not report_path.exists()
+56
View File
@@ -0,0 +1,56 @@
from pathlib import Path
from serato_doctor.health import analyze_health
from serato_doctor.models.library import Library
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def reference(filename):
return TrackReference(
Path("Test.crate"), Path("/old/House") / filename, filename
)
def track(filename, folder="House"):
path = Path("/new") / folder / filename
return DiskTrack(path, filename, 100, path.suffix.lower())
def test_health_report_exposes_each_metric():
library = Library.build(
[
reference("Found.mp3"),
reference("Conflict.mp3"),
reference("Absent.mp3"),
],
[
track("Found.mp3"),
track("Conflict 2.mp3"),
track("Conflict 2.mp3", "Backup"),
track("Unused.mp3"),
],
)
report = analyze_health(library)
assert report.score == 33.3
assert report.score_basis == (
"Resolved crate references / total crate references"
)
assert report.total_references == 3
assert report.healthy_references == 1
assert report.missing_references == 2
assert report.unique_missing_filenames == 2
assert report.disk_tracks == 4
assert report.duplicate_filename_groups == 1
assert report.duplicate_files == 1
assert report.unused_tracks == 3
assert report.suggested_matches == 1
def test_empty_library_has_no_health_score():
report = analyze_health(Library.build([], []))
assert report.score is None
assert report.total_references == 0
+62
View File
@@ -0,0 +1,62 @@
from pathlib import Path
from serato_doctor.matching import MatchingEngine, score_candidate
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def reference(filename, folder="House"):
return TrackReference(
Path("Test.crate"), Path("/old") / folder / filename, filename
)
def track(filename, folder="House"):
path = Path("/new") / folder / filename
return DiskTrack(path, filename, 100, path.suffix.lower())
def test_exact_candidate_has_full_evidence_score():
match = score_candidate(reference("Track.mp3"), track("Track.mp3"))
assert match.score == 90
assert match.max_score == 90
assert match.score_percent == 100.0
assert all(item.matched for item in match.evidence)
def test_cloud_conflict_suffix_is_explained():
match = score_candidate(reference("Track.mp3"), track("Track 2.mp3"))
assert match.score == 80
assert match.score_percent == 88.9
assert match.evidence[0].explanation == (
"Filename matches after removing a numeric conflict suffix"
)
def test_case_normalized_match_scores_below_exact():
match = score_candidate(reference("TRACK.MP3"), track("track.mp3"))
assert match.score == 85
assert match.evidence[0].points == 55
def test_engine_omits_unrelated_filenames():
engine = MatchingEngine([track("Different.mp3")])
assert engine.candidates_for(reference("Missing.mp3")) == ()
def test_ambiguous_candidates_are_ranked_deterministically():
engine = MatchingEngine(
[track("Track 2.mp3", "Other"), track("Track 3.mp3", "House")]
)
matches = engine.candidates_for(reference("Track.mp3"))
assert [match.track.filename for match in matches] == [
"Track 3.mp3",
"Track 2.mp3",
]
assert [match.score for match in matches] == [80, 60]