Compare commits
5 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| c92d929f37 | |||
| e6f313d98a | |||
| 22424d0b5c | |||
| b8bf9ec6e1 | |||
| 167f029c28 |
@@ -0,0 +1,37 @@
|
||||
# Serato Doctor
|
||||
|
||||
Serato Doctor is a read-only-first toolkit for inspecting and eventually repairing
|
||||
DJ libraries. Serato is the first supported engine.
|
||||
|
||||
## Analyze a Library
|
||||
|
||||
```shell
|
||||
serato-doctor analyze \
|
||||
--serato ~/Music/_Serato_ \
|
||||
--music ~/Music/Jukebox
|
||||
```
|
||||
|
||||
Analysis prints a transparent reference-integrity score plus broken references,
|
||||
duplicate filenames, unused tracks, and suggested filename matches. It does not
|
||||
modify the library or create report files.
|
||||
|
||||
If crates contain an older library root, supply it explicitly:
|
||||
|
||||
```shell
|
||||
serato-doctor analyze --reference-root /Users/old-user/OneDrive/Jukebox
|
||||
```
|
||||
|
||||
## Generate Detailed Reports
|
||||
|
||||
Running without a command preserves the original scanner behavior and writes CSV
|
||||
and text reports:
|
||||
|
||||
```shell
|
||||
serato-doctor --serato ~/Music/_Serato_ --music ~/Music/Jukebox
|
||||
```
|
||||
|
||||
Use `--verbose` for progress on standard error or `--log-file PATH` for an
|
||||
aggregate diagnostic log.
|
||||
|
||||
Serato Doctor never repairs files without an explicit future repair workflow,
|
||||
preview, backup, and rollback path.
|
||||
|
||||
+2
-1
@@ -17,13 +17,14 @@
|
||||
|
||||
## v0.2 — Diagnostics
|
||||
|
||||
- [x] `serato-doctor analyze` command
|
||||
- [ ] Duplicate filename detection
|
||||
- [ ] Duplicate audio hash detection
|
||||
- [ ] Broken symlink detection
|
||||
- [ ] Orphaned audio detection
|
||||
- [ ] OneDrive rename detection
|
||||
- [ ] Crate classification: static vs smart/dynamic
|
||||
- [ ] Library health score
|
||||
- [x] Library health score
|
||||
|
||||
## v0.3 — Safe Repair
|
||||
|
||||
|
||||
@@ -0,0 +1,27 @@
|
||||
# Analyze Command
|
||||
|
||||
## Problem
|
||||
|
||||
The health and matching engines are only Python APIs. Users need one safe command
|
||||
that summarizes library integrity without first interpreting CSV files.
|
||||
|
||||
## Architecture
|
||||
|
||||
`serato-doctor analyze` reuses the existing configured crate parser and filesystem
|
||||
scanner, then prints the immutable health report. The original no-command mode is
|
||||
retained as the report-producing scan workflow.
|
||||
|
||||
Analyze mode does not create CSV or text reports. Both modes remain read-only with
|
||||
respect to Serato crates, databases, and music files.
|
||||
|
||||
## Edge Cases
|
||||
|
||||
- A library with no references displays `Not assessed` rather than a false score.
|
||||
- Historical roots continue to work through repeatable `--reference-root` flags.
|
||||
- Normal output remains separate from optional diagnostic logging.
|
||||
- Existing scripts that invoke the CLI without a subcommand remain compatible.
|
||||
|
||||
## Verification
|
||||
|
||||
Tests prove the legacy five-line output and reports remain unchanged, while analyze
|
||||
prints health metrics and does not create output files.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Library Health Engine
|
||||
|
||||
## Problem
|
||||
|
||||
Raw missing-reference counts do not provide a compact view of library integrity,
|
||||
but an opaque blended score would imply confidence the current data cannot support.
|
||||
|
||||
## Architecture
|
||||
|
||||
The health engine produces an immutable report from the core `Library`. Its score
|
||||
is only the percentage of crate references resolved by exact filename. The report
|
||||
also exposes missing references, unique missing filenames, duplicate filename
|
||||
groups, extra duplicate files, unused tracks, and missing references with matching
|
||||
candidates.
|
||||
|
||||
Duplicate, unused, and candidate counts are informational. They do not affect the
|
||||
score until the project has a documented and validated weighting policy. An empty
|
||||
library has no score rather than a misleading 0% or 100%.
|
||||
|
||||
## Edge Cases
|
||||
|
||||
- The same missing filename referenced by several crates.
|
||||
- Several disk files sharing a filename.
|
||||
- Conflict-suffixed files that are unused but may be match candidates.
|
||||
- Libraries with no crate references.
|
||||
- Tracks referenced by filename from more than one crate.
|
||||
|
||||
## Verification
|
||||
|
||||
Tests assert every metric, the disclosed score basis, candidate integration, and
|
||||
empty-library behavior. Analysis remains entirely read-only.
|
||||
+42
-13
@@ -3,6 +3,7 @@ import argparse
|
||||
|
||||
from serato_doctor.config import ScanConfig
|
||||
from serato_doctor.crate_parser import parse_crates
|
||||
from serato_doctor.health import analyze_health
|
||||
from serato_doctor.logging import configure_logging
|
||||
from serato_doctor.models.library import Library
|
||||
from serato_doctor.scanner import scan_audio
|
||||
@@ -11,10 +12,23 @@ from serato_doctor.report import write_csv, write_missing_report
|
||||
|
||||
def main():
|
||||
parser = argparse.ArgumentParser(prog="serato-doctor")
|
||||
parser.add_argument(
|
||||
"command", nargs="?", choices=("scan", "analyze"), default="scan"
|
||||
)
|
||||
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
||||
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"))
|
||||
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv"))
|
||||
parser.add_argument("--report", default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"))
|
||||
parser.add_argument(
|
||||
"--music",
|
||||
default=str(
|
||||
Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")
|
||||
)
|
||||
parser.add_argument(
|
||||
"--report",
|
||||
default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--reference-root",
|
||||
action="append",
|
||||
@@ -37,7 +51,7 @@ def main():
|
||||
log_file=Path(args.log_file) if args.log_file else None,
|
||||
)
|
||||
logger = configure_logging(config.verbose, config.log_file)
|
||||
logger.info("Starting read-only library scan")
|
||||
logger.info("Starting read-only library inspection")
|
||||
logger.debug("Serato directory: %s", config.serato)
|
||||
logger.debug("Music directory: %s", config.music)
|
||||
|
||||
@@ -53,15 +67,30 @@ def main():
|
||||
logger.info("Scanned %d disk tracks", len(library.tracks))
|
||||
logger.info("Found %d references missing by filename", missing_count)
|
||||
|
||||
write_csv(results, config.out)
|
||||
write_missing_report(results, config.report)
|
||||
logger.info("Wrote CSV and missing-reference reports")
|
||||
|
||||
print(f"Crate references: {len(library.references)}")
|
||||
print(f"Disk tracks: {len(library.tracks)}")
|
||||
print(f"Missing by filename: {missing_count}")
|
||||
print(f"CSV: {config.out}")
|
||||
print(f"Report: {config.report}")
|
||||
if args.command == "analyze":
|
||||
health = analyze_health(library)
|
||||
score = f"{health.score:.1f}%" if health.score is not None else "Not assessed"
|
||||
logger.info("Calculated library health: %s", score)
|
||||
print(f"Overall Health: {score}")
|
||||
print(f"Score Basis: {health.score_basis}")
|
||||
print(f"Tracks: {health.disk_tracks}")
|
||||
print(f"Crate References: {health.total_references}")
|
||||
print(f"Healthy References: {health.healthy_references}")
|
||||
print(f"Broken References: {health.missing_references}")
|
||||
print(f"Unique Missing Filenames: {health.unique_missing_filenames}")
|
||||
print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}")
|
||||
print(f"Duplicate Files: {health.duplicate_files}")
|
||||
print(f"Unused Tracks: {health.unused_tracks}")
|
||||
print(f"Suggested Matches: {health.suggested_matches}")
|
||||
else:
|
||||
write_csv(results, config.out)
|
||||
write_missing_report(results, config.report)
|
||||
logger.info("Wrote CSV and missing-reference reports")
|
||||
print(f"Crate references: {len(library.references)}")
|
||||
print(f"Disk tracks: {len(library.tracks)}")
|
||||
print(f"Missing by filename: {missing_count}")
|
||||
print(f"CSV: {config.out}")
|
||||
print(f"Report: {config.report}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
from collections import Counter
|
||||
|
||||
from serato_doctor.matching import MatchingEngine
|
||||
from serato_doctor.models.health import HealthReport
|
||||
from serato_doctor.models.library import Library
|
||||
|
||||
|
||||
def analyze_health(library: Library) -> HealthReport:
|
||||
"""Calculate defensible health metrics without changing the library."""
|
||||
|
||||
results = library.reconcile_by_filename()
|
||||
missing = [result for result in results if not result.exists_by_filename]
|
||||
healthy_count = len(results) - len(missing)
|
||||
score = (
|
||||
round(healthy_count / len(results) * 100, 1) if results else None
|
||||
)
|
||||
|
||||
disk_name_counts = Counter(track.filename for track in library.tracks)
|
||||
duplicate_counts = [count for count in disk_name_counts.values() if count > 1]
|
||||
referenced_names = {reference.filename for reference in library.references}
|
||||
unused_count = sum(
|
||||
1 for track in library.tracks if track.filename not in referenced_names
|
||||
)
|
||||
|
||||
matcher = MatchingEngine(library.tracks)
|
||||
suggested_count = sum(
|
||||
bool(matcher.candidates_for(result.reference)) for result in missing
|
||||
)
|
||||
|
||||
return HealthReport(
|
||||
score=score,
|
||||
total_references=len(results),
|
||||
healthy_references=healthy_count,
|
||||
missing_references=len(missing),
|
||||
unique_missing_filenames=len(
|
||||
{result.reference.filename for result in missing}
|
||||
),
|
||||
disk_tracks=len(library.tracks),
|
||||
duplicate_filename_groups=len(duplicate_counts),
|
||||
duplicate_files=sum(count - 1 for count in duplicate_counts),
|
||||
unused_tracks=unused_count,
|
||||
suggested_matches=suggested_count,
|
||||
)
|
||||
@@ -1,4 +1,5 @@
|
||||
from serato_doctor.models.crate import Crate
|
||||
from serato_doctor.models.health import HealthReport
|
||||
from serato_doctor.models.library import Library
|
||||
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
||||
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
||||
@@ -7,6 +8,7 @@ from serato_doctor.models.track import DiskTrack
|
||||
__all__ = [
|
||||
"Crate",
|
||||
"DiskTrack",
|
||||
"HealthReport",
|
||||
"Library",
|
||||
"MatchEvidence",
|
||||
"ReferenceResult",
|
||||
|
||||
@@ -0,0 +1,22 @@
|
||||
from dataclasses import dataclass
|
||||
from typing import Optional
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class HealthReport:
|
||||
"""Transparent aggregate findings from a read-only library analysis."""
|
||||
|
||||
score: Optional[float]
|
||||
total_references: int
|
||||
healthy_references: int
|
||||
missing_references: int
|
||||
unique_missing_filenames: int
|
||||
disk_tracks: int
|
||||
duplicate_filename_groups: int
|
||||
duplicate_files: int
|
||||
unused_tracks: int
|
||||
suggested_matches: int
|
||||
|
||||
@property
|
||||
def score_basis(self) -> str:
|
||||
return "Resolved crate references / total crate references"
|
||||
@@ -84,3 +84,47 @@ def test_cli_accepts_old_reference_root(tmp_path, monkeypatch, capsys):
|
||||
output = capsys.readouterr().out
|
||||
assert "Crate references: 1" in output
|
||||
assert "Missing by filename: 0" in output
|
||||
|
||||
|
||||
def test_analyze_prints_health_without_writing_reports(tmp_path, monkeypatch, capsys):
|
||||
serato = tmp_path / "serato"
|
||||
subcrates = serato / "Subcrates"
|
||||
music = tmp_path / "music"
|
||||
subcrates.mkdir(parents=True)
|
||||
music.mkdir()
|
||||
crate_text = (
|
||||
"Users/sample-user/Jukebox/Found.mp3otrk"
|
||||
"Users/sample-user/Jukebox/Missing.mp3otrk"
|
||||
)
|
||||
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
|
||||
(music / "Found.mp3").write_bytes(b"synthetic audio")
|
||||
csv_path = tmp_path / "scan.csv"
|
||||
report_path = tmp_path / "report.txt"
|
||||
monkeypatch.setattr(
|
||||
sys,
|
||||
"argv",
|
||||
[
|
||||
"serato-doctor",
|
||||
"analyze",
|
||||
"--serato",
|
||||
str(serato),
|
||||
"--music",
|
||||
str(music),
|
||||
"--out",
|
||||
str(csv_path),
|
||||
"--report",
|
||||
str(report_path),
|
||||
],
|
||||
)
|
||||
|
||||
main()
|
||||
|
||||
output = capsys.readouterr().out
|
||||
assert "Overall Health: 50.0%" in output
|
||||
assert "Tracks: 1" in output
|
||||
assert "Crate References: 2" in output
|
||||
assert "Healthy References: 1" in output
|
||||
assert "Broken References: 1" in output
|
||||
assert "Unused Tracks: 0" in output
|
||||
assert not csv_path.exists()
|
||||
assert not report_path.exists()
|
||||
|
||||
@@ -0,0 +1,56 @@
|
||||
from pathlib import Path
|
||||
|
||||
from serato_doctor.health import analyze_health
|
||||
from serato_doctor.models.library import Library
|
||||
from serato_doctor.models.reference import TrackReference
|
||||
from serato_doctor.models.track import DiskTrack
|
||||
|
||||
|
||||
def reference(filename):
|
||||
return TrackReference(
|
||||
Path("Test.crate"), Path("/old/House") / filename, filename
|
||||
)
|
||||
|
||||
|
||||
def track(filename, folder="House"):
|
||||
path = Path("/new") / folder / filename
|
||||
return DiskTrack(path, filename, 100, path.suffix.lower())
|
||||
|
||||
|
||||
def test_health_report_exposes_each_metric():
|
||||
library = Library.build(
|
||||
[
|
||||
reference("Found.mp3"),
|
||||
reference("Conflict.mp3"),
|
||||
reference("Absent.mp3"),
|
||||
],
|
||||
[
|
||||
track("Found.mp3"),
|
||||
track("Conflict 2.mp3"),
|
||||
track("Conflict 2.mp3", "Backup"),
|
||||
track("Unused.mp3"),
|
||||
],
|
||||
)
|
||||
|
||||
report = analyze_health(library)
|
||||
|
||||
assert report.score == 33.3
|
||||
assert report.score_basis == (
|
||||
"Resolved crate references / total crate references"
|
||||
)
|
||||
assert report.total_references == 3
|
||||
assert report.healthy_references == 1
|
||||
assert report.missing_references == 2
|
||||
assert report.unique_missing_filenames == 2
|
||||
assert report.disk_tracks == 4
|
||||
assert report.duplicate_filename_groups == 1
|
||||
assert report.duplicate_files == 1
|
||||
assert report.unused_tracks == 3
|
||||
assert report.suggested_matches == 1
|
||||
|
||||
|
||||
def test_empty_library_has_no_health_score():
|
||||
report = analyze_health(Library.build([], []))
|
||||
|
||||
assert report.score is None
|
||||
assert report.total_references == 0
|
||||
Reference in New Issue
Block a user