Compare commits
7 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| c5449a278a | |||
| 6449c7de6a | |||
| d764f910e4 | |||
| 35713bbbb3 | |||
| c92d929f37 | |||
| e6f313d98a | |||
| 22424d0b5c |
@@ -0,0 +1,37 @@
|
|||||||
|
# Serato Doctor
|
||||||
|
|
||||||
|
Serato Doctor is a read-only-first toolkit for inspecting and eventually repairing
|
||||||
|
DJ libraries. Serato is the first supported engine.
|
||||||
|
|
||||||
|
## Analyze a Library
|
||||||
|
|
||||||
|
```shell
|
||||||
|
serato-doctor analyze \
|
||||||
|
--serato ~/Music/_Serato_ \
|
||||||
|
--music ~/Music/Jukebox
|
||||||
|
```
|
||||||
|
|
||||||
|
Analysis prints a transparent reference-integrity score plus broken references,
|
||||||
|
duplicate filenames, unused tracks, and suggested filename matches. It does not
|
||||||
|
modify the library or create report files.
|
||||||
|
|
||||||
|
If crates contain an older library root, supply it explicitly:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
serato-doctor analyze --reference-root /Users/old-user/OneDrive/Jukebox
|
||||||
|
```
|
||||||
|
|
||||||
|
## Generate Detailed Reports
|
||||||
|
|
||||||
|
Running without a command preserves the original scanner behavior and writes CSV
|
||||||
|
and text reports:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
serato-doctor --serato ~/Music/_Serato_ --music ~/Music/Jukebox
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `--verbose` for progress on standard error or `--log-file PATH` for an
|
||||||
|
aggregate diagnostic log.
|
||||||
|
|
||||||
|
Serato Doctor never repairs files without an explicit future repair workflow,
|
||||||
|
preview, backup, and rollback path.
|
||||||
|
|||||||
+3
-2
@@ -17,12 +17,13 @@
|
|||||||
|
|
||||||
## v0.2 — Diagnostics
|
## v0.2 — Diagnostics
|
||||||
|
|
||||||
- [ ] Duplicate filename detection
|
- [x] `serato-doctor analyze` command
|
||||||
|
- [x] Duplicate filename detection
|
||||||
- [ ] Duplicate audio hash detection
|
- [ ] Duplicate audio hash detection
|
||||||
- [ ] Broken symlink detection
|
- [ ] Broken symlink detection
|
||||||
- [ ] Orphaned audio detection
|
- [ ] Orphaned audio detection
|
||||||
- [ ] OneDrive rename detection
|
- [ ] OneDrive rename detection
|
||||||
- [ ] Crate classification: static vs smart/dynamic
|
- [x] Crate classification: static vs smart/dynamic
|
||||||
- [x] Library health score
|
- [x] Library health score
|
||||||
|
|
||||||
## v0.3 — Safe Repair
|
## v0.3 — Safe Repair
|
||||||
|
|||||||
@@ -0,0 +1,27 @@
|
|||||||
|
# Analyze Command
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The health and matching engines are only Python APIs. Users need one safe command
|
||||||
|
that summarizes library integrity without first interpreting CSV files.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
`serato-doctor analyze` reuses the existing configured crate parser and filesystem
|
||||||
|
scanner, then prints the immutable health report. The original no-command mode is
|
||||||
|
retained as the report-producing scan workflow.
|
||||||
|
|
||||||
|
Analyze mode does not create CSV or text reports. Both modes remain read-only with
|
||||||
|
respect to Serato crates, databases, and music files.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- A library with no references displays `Not assessed` rather than a false score.
|
||||||
|
- Historical roots continue to work through repeatable `--reference-root` flags.
|
||||||
|
- Normal output remains separate from optional diagnostic logging.
|
||||||
|
- Existing scripts that invoke the CLI without a subcommand remain compatible.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests prove the legacy five-line output and reports remain unchanged, while analyze
|
||||||
|
prints health metrics and does not create output files.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# Crate Classification
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Static crates are manually maintained track lists. Smart crates are dynamic views
|
||||||
|
generated from rules, so stale-looking entries in them should not be presented as
|
||||||
|
broken manual references or given the same health-score weight.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Crates carry a `static`, `smart`, or `unknown` kind. Folder provenance is the
|
||||||
|
primary signal: Serato stores regular definitions in `Subcrates` and smart
|
||||||
|
definitions in `Smartcrates`. `Compatible by key.crate` is also treated as smart
|
||||||
|
when encountered in `Subcrates`, based on the original migration case study.
|
||||||
|
|
||||||
|
The library loader reads both folders. Health analysis reports all references but
|
||||||
|
scores only non-smart references. Unknown crates remain scoreable so incomplete
|
||||||
|
classification cannot silently hide potential problems.
|
||||||
|
|
||||||
|
Serato documents the folder distinction in [What is in the _Serato_ folder?](https://support.serato.com/hc/en-us/articles/204022904-What-is-in-the-Serato-folder)
|
||||||
|
and explains that smart crates are populated from rules in [Crates in Serato DJ](https://support.serato.com/hc/en-us/articles/227561407-Crates-in-Serato-DJ-Pro-Serato-DJ-Lite).
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- The known `Compatible by key` dynamic crate in the `Subcrates` folder.
|
||||||
|
- Crate fixtures outside a recognized Serato folder.
|
||||||
|
- Libraries containing both static and smart references to the same track.
|
||||||
|
- Dynamic references whose current materialized paths appear missing.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover all three kinds, both Serato folders, the known dynamic fallback, the
|
||||||
|
synthetic five-static/two-smart library, and exclusion from health scoring.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Duplicate Filename Detection
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Two files with the same filename may be ordinary copies, while names such as
|
||||||
|
`Track.mp3` and `Track 2.mp3` may indicate a OneDrive conflict. These cases need
|
||||||
|
review, but neither is sufficient evidence for deletion.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The duplicate detector emits immutable groups of two kinds:
|
||||||
|
|
||||||
|
- `exact_name` groups filenames after case and Unicode normalization.
|
||||||
|
- `cloud_conflict` groups names after additionally removing a trailing numeric
|
||||||
|
suffix from the stem.
|
||||||
|
|
||||||
|
Groups contain every path, a stable comparison key, a display name, and the number
|
||||||
|
of files beyond the first. Health analysis reports exact and suspected-conflict
|
||||||
|
counts separately. Neither category changes the health score.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Same filename in different folders.
|
||||||
|
- Case-only and Unicode representation differences.
|
||||||
|
- Multiple conflict suffixes such as `Track 2.mp3` and `Track 3.mp3`.
|
||||||
|
- Legitimate numbered titles, which remain explicitly labeled as suspected.
|
||||||
|
- A single file, which is never reported as a duplicate.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover exact, normalized, conflict-family, unrelated, and deterministic-order
|
||||||
|
behavior, plus integration with health analysis. Detection is read-only.
|
||||||
@@ -8,10 +8,11 @@ but an opaque blended score would imply confidence the current data cannot suppo
|
|||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
The health engine produces an immutable report from the core `Library`. Its score
|
The health engine produces an immutable report from the core `Library`. Its score
|
||||||
is only the percentage of crate references resolved by exact filename. The report
|
is only the percentage of non-dynamic crate references resolved by exact filename.
|
||||||
also exposes missing references, unique missing filenames, duplicate filename
|
Smart-crate references are counted but excluded because their contents are derived
|
||||||
groups, extra duplicate files, unused tracks, and missing references with matching
|
from rules. The report also exposes missing references, unique missing filenames,
|
||||||
candidates.
|
duplicate filename groups, extra duplicate files, unused tracks, and missing
|
||||||
|
references with matching candidates.
|
||||||
|
|
||||||
Duplicate, unused, and candidate counts are informational. They do not affect the
|
Duplicate, unused, and candidate counts are informational. They do not affect the
|
||||||
score until the project has a documented and validated weighting policy. An empty
|
score until the project has a documented and validated weighting policy. An empty
|
||||||
|
|||||||
@@ -18,8 +18,7 @@ def build_sample(output: Optional[Path] = None) -> Path:
|
|||||||
root = output or SAMPLE_ROOT / "generated"
|
root = output or SAMPLE_ROOT / "generated"
|
||||||
manifest = load_manifest()
|
manifest = load_manifest()
|
||||||
music_root = root / "Music"
|
music_root = root / "Music"
|
||||||
crate_root = root / "Serato" / "_Serato_" / "Subcrates"
|
serato_root = root / "Serato" / "_Serato_"
|
||||||
crate_root.mkdir(parents=True, exist_ok=True)
|
|
||||||
|
|
||||||
for relative_path in manifest["tracks"]:
|
for relative_path in manifest["tracks"]:
|
||||||
track_path = music_root / relative_path
|
track_path = music_root / relative_path
|
||||||
@@ -29,6 +28,13 @@ def build_sample(output: Optional[Path] = None) -> Path:
|
|||||||
)
|
)
|
||||||
|
|
||||||
for crate in manifest["crates"]:
|
for crate in manifest["crates"]:
|
||||||
|
folder_name = "Smartcrates" if crate["type"] == "smart" else "Subcrates"
|
||||||
|
crate_root = serato_root / folder_name
|
||||||
|
crate_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
other_folder = "Subcrates" if folder_name == "Smartcrates" else "Smartcrates"
|
||||||
|
stale_path = serato_root / other_folder / crate["name"]
|
||||||
|
if stale_path.exists():
|
||||||
|
stale_path.unlink()
|
||||||
records = "".join(
|
records = "".join(
|
||||||
f"{SERATO_PATH_PREFIX}{relative_path}otrk"
|
f"{SERATO_PATH_PREFIX}{relative_path}otrk"
|
||||||
for relative_path in crate["references"]
|
for relative_path in crate["references"]
|
||||||
|
|||||||
+60
-18
@@ -2,7 +2,8 @@ from pathlib import Path
|
|||||||
import argparse
|
import argparse
|
||||||
|
|
||||||
from serato_doctor.config import ScanConfig
|
from serato_doctor.config import ScanConfig
|
||||||
from serato_doctor.crate_parser import parse_crates
|
from serato_doctor.crate_parser import load_library_crates
|
||||||
|
from serato_doctor.health import analyze_health
|
||||||
from serato_doctor.logging import configure_logging
|
from serato_doctor.logging import configure_logging
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.scanner import scan_audio
|
from serato_doctor.scanner import scan_audio
|
||||||
@@ -11,10 +12,23 @@ from serato_doctor.report import write_csv, write_missing_report
|
|||||||
|
|
||||||
def main():
|
def main():
|
||||||
parser = argparse.ArgumentParser(prog="serato-doctor")
|
parser = argparse.ArgumentParser(prog="serato-doctor")
|
||||||
|
parser.add_argument(
|
||||||
|
"command", nargs="?", choices=("scan", "analyze"), default="scan"
|
||||||
|
)
|
||||||
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
||||||
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"))
|
parser.add_argument(
|
||||||
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv"))
|
"--music",
|
||||||
parser.add_argument("--report", default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"))
|
default=str(
|
||||||
|
Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"
|
||||||
|
),
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--report",
|
||||||
|
default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"),
|
||||||
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--reference-root",
|
"--reference-root",
|
||||||
action="append",
|
action="append",
|
||||||
@@ -37,14 +51,13 @@ def main():
|
|||||||
log_file=Path(args.log_file) if args.log_file else None,
|
log_file=Path(args.log_file) if args.log_file else None,
|
||||||
)
|
)
|
||||||
logger = configure_logging(config.verbose, config.log_file)
|
logger = configure_logging(config.verbose, config.log_file)
|
||||||
logger.info("Starting read-only library scan")
|
logger.info("Starting read-only library inspection")
|
||||||
logger.debug("Serato directory: %s", config.serato)
|
logger.debug("Serato directory: %s", config.serato)
|
||||||
logger.debug("Music directory: %s", config.music)
|
logger.debug("Music directory: %s", config.music)
|
||||||
|
|
||||||
library = Library.build(
|
crates = load_library_crates(config.serato, config.reference_roots)
|
||||||
references=parse_crates(
|
library = Library.from_crates(
|
||||||
config.serato / "Subcrates", config.reference_roots
|
crates=crates,
|
||||||
),
|
|
||||||
tracks=scan_audio(config.music),
|
tracks=scan_audio(config.music),
|
||||||
)
|
)
|
||||||
results = library.reconcile_by_filename()
|
results = library.reconcile_by_filename()
|
||||||
@@ -53,15 +66,44 @@ def main():
|
|||||||
logger.info("Scanned %d disk tracks", len(library.tracks))
|
logger.info("Scanned %d disk tracks", len(library.tracks))
|
||||||
logger.info("Found %d references missing by filename", missing_count)
|
logger.info("Found %d references missing by filename", missing_count)
|
||||||
|
|
||||||
write_csv(results, config.out)
|
if args.command == "analyze":
|
||||||
write_missing_report(results, config.report)
|
health = analyze_health(library)
|
||||||
logger.info("Wrote CSV and missing-reference reports")
|
score = f"{health.score:.1f}%" if health.score is not None else "Not assessed"
|
||||||
|
logger.info("Calculated library health: %s", score)
|
||||||
print(f"Crate references: {len(library.references)}")
|
print(f"Overall Health: {score}")
|
||||||
print(f"Disk tracks: {len(library.tracks)}")
|
print(f"Score Basis: {health.score_basis}")
|
||||||
print(f"Missing by filename: {missing_count}")
|
print(f"Tracks: {health.disk_tracks}")
|
||||||
print(f"CSV: {config.out}")
|
print(f"Crate References: {health.total_references}")
|
||||||
print(f"Report: {config.report}")
|
print(f"References Scored: {health.scored_references}")
|
||||||
|
print(f"Healthy References: {health.healthy_references}")
|
||||||
|
print(f"Broken References: {health.missing_references}")
|
||||||
|
print(f"Unique Missing Filenames: {health.unique_missing_filenames}")
|
||||||
|
print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}")
|
||||||
|
print(f"Duplicate Files: {health.duplicate_files}")
|
||||||
|
print(
|
||||||
|
"Suspected Cloud Conflict Groups: "
|
||||||
|
f"{health.suspected_cloud_conflict_groups}"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
"Suspected Cloud Conflict Files: "
|
||||||
|
f"{health.suspected_cloud_conflict_files}"
|
||||||
|
)
|
||||||
|
print(f"Unused Tracks: {health.unused_tracks}")
|
||||||
|
print(f"Suggested Matches: {health.suggested_matches}")
|
||||||
|
print(f"Static Crates: {health.static_crates}")
|
||||||
|
print(f"Smart Crates: {health.smart_crates}")
|
||||||
|
print(
|
||||||
|
f"Dynamic References Excluded: {health.dynamic_references_excluded}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
write_csv(results, config.out)
|
||||||
|
write_missing_report(results, config.report)
|
||||||
|
logger.info("Wrote CSV and missing-reference reports")
|
||||||
|
print(f"Crate references: {len(library.references)}")
|
||||||
|
print(f"Disk tracks: {len(library.tracks)}")
|
||||||
|
print(f"Missing by filename: {missing_count}")
|
||||||
|
print(f"CSV: {config.out}")
|
||||||
|
print(f"Report: {config.report}")
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Iterable, Tuple
|
from typing import Iterable, Tuple
|
||||||
|
|
||||||
from serato_doctor.models.crate import Crate
|
from serato_doctor.models.crate import Crate, CrateKind
|
||||||
from serato_doctor.models.reference import TrackReference
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
|
||||||
|
|
||||||
@@ -32,6 +32,17 @@ def path_markers(reference_roots: Iterable[Path]) -> Tuple[str, ...]:
|
|||||||
return configured or DEFAULT_PATH_MARKERS
|
return configured or DEFAULT_PATH_MARKERS
|
||||||
|
|
||||||
|
|
||||||
|
def classify_crate(crate_path: Path) -> CrateKind:
|
||||||
|
parent_names = {parent.name.casefold() for parent in crate_path.parents}
|
||||||
|
if "smartcrates" in parent_names:
|
||||||
|
return CrateKind.SMART
|
||||||
|
if crate_path.name.casefold() == "compatible by key.crate":
|
||||||
|
return CrateKind.SMART
|
||||||
|
if "subcrates" in parent_names:
|
||||||
|
return CrateKind.STATIC
|
||||||
|
return CrateKind.UNKNOWN
|
||||||
|
|
||||||
|
|
||||||
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
||||||
text = read_crate_text(crate_path)
|
text = read_crate_text(crate_path)
|
||||||
refs = []
|
refs = []
|
||||||
@@ -63,7 +74,11 @@ def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
|||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
||||||
return Crate(path=crate_path, references=tuple(refs))
|
return Crate(
|
||||||
|
path=crate_path,
|
||||||
|
references=tuple(refs),
|
||||||
|
kind=classify_crate(crate_path),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def parse_crate(
|
def parse_crate(
|
||||||
@@ -82,3 +97,15 @@ def parse_crates(
|
|||||||
for crate in root.rglob("*.crate"):
|
for crate in root.rglob("*.crate"):
|
||||||
refs.extend(parse_crate(crate, reference_roots))
|
refs.extend(parse_crate(crate, reference_roots))
|
||||||
return refs
|
return refs
|
||||||
|
|
||||||
|
|
||||||
|
def load_library_crates(
|
||||||
|
serato_root: Path, reference_roots: Iterable[Path] = ()
|
||||||
|
) -> Tuple[Crate, ...]:
|
||||||
|
reference_roots = tuple(reference_roots)
|
||||||
|
crates = []
|
||||||
|
for folder_name in ("Subcrates", "Smartcrates"):
|
||||||
|
folder = serato_root / folder_name
|
||||||
|
for crate_path in folder.rglob("*.crate"):
|
||||||
|
crates.append(load_crate(crate_path, reference_roots))
|
||||||
|
return tuple(sorted(crates, key=lambda crate: str(crate.path)))
|
||||||
|
|||||||
@@ -0,0 +1,41 @@
|
|||||||
|
from collections import defaultdict
|
||||||
|
from typing import DefaultDict, Iterable, List, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.matching import cloud_conflict_name, normalize
|
||||||
|
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def find_duplicate_groups(
|
||||||
|
tracks: Iterable[DiskTrack],
|
||||||
|
) -> Tuple[DuplicateGroup, ...]:
|
||||||
|
"""Find exact-name duplicates and suspected numeric conflict copies."""
|
||||||
|
|
||||||
|
track_tuple = tuple(tracks)
|
||||||
|
by_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
|
||||||
|
by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
|
||||||
|
|
||||||
|
for track in track_tuple:
|
||||||
|
by_name[normalize(track.filename)].append(track)
|
||||||
|
by_conflict_name[cloud_conflict_name(track.filename)].append(track)
|
||||||
|
|
||||||
|
groups = []
|
||||||
|
for key, matches in by_name.items():
|
||||||
|
if len(matches) > 1:
|
||||||
|
groups.append(_group(DuplicateKind.EXACT_NAME, key, matches))
|
||||||
|
|
||||||
|
for key, matches in by_conflict_name.items():
|
||||||
|
distinct_names = {normalize(track.filename) for track in matches}
|
||||||
|
if len(distinct_names) > 1:
|
||||||
|
groups.append(_group(DuplicateKind.CLOUD_CONFLICT, key, matches))
|
||||||
|
|
||||||
|
return tuple(
|
||||||
|
sorted(groups, key=lambda group: (group.kind.value, group.comparison_key))
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _group(
|
||||||
|
kind: DuplicateKind, key: str, tracks: Iterable[DiskTrack]
|
||||||
|
) -> DuplicateGroup:
|
||||||
|
ordered = tuple(sorted(tracks, key=lambda track: str(track.path)))
|
||||||
|
return DuplicateGroup(kind, key, ordered)
|
||||||
+44
-8
@@ -1,6 +1,7 @@
|
|||||||
from collections import Counter
|
from serato_doctor.duplicates import find_duplicate_groups
|
||||||
|
|
||||||
from serato_doctor.matching import MatchingEngine
|
from serato_doctor.matching import MatchingEngine
|
||||||
|
from serato_doctor.models.crate import CrateKind
|
||||||
|
from serato_doctor.models.duplicate import DuplicateKind
|
||||||
from serato_doctor.models.health import HealthReport
|
from serato_doctor.models.health import HealthReport
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
|
|
||||||
@@ -8,15 +9,31 @@ from serato_doctor.models.library import Library
|
|||||||
def analyze_health(library: Library) -> HealthReport:
|
def analyze_health(library: Library) -> HealthReport:
|
||||||
"""Calculate defensible health metrics without changing the library."""
|
"""Calculate defensible health metrics without changing the library."""
|
||||||
|
|
||||||
results = library.reconcile_by_filename()
|
dynamic_sources = {
|
||||||
|
crate.path for crate in library.crates if crate.kind is CrateKind.SMART
|
||||||
|
}
|
||||||
|
results = tuple(
|
||||||
|
result
|
||||||
|
for result in library.reconcile_by_filename()
|
||||||
|
if result.reference.source not in dynamic_sources
|
||||||
|
)
|
||||||
missing = [result for result in results if not result.exists_by_filename]
|
missing = [result for result in results if not result.exists_by_filename]
|
||||||
healthy_count = len(results) - len(missing)
|
healthy_count = len(results) - len(missing)
|
||||||
score = (
|
score = (
|
||||||
round(healthy_count / len(results) * 100, 1) if results else None
|
round(healthy_count / len(results) * 100, 1) if results else None
|
||||||
)
|
)
|
||||||
|
|
||||||
disk_name_counts = Counter(track.filename for track in library.tracks)
|
duplicate_groups = find_duplicate_groups(library.tracks)
|
||||||
duplicate_counts = [count for count in disk_name_counts.values() if count > 1]
|
exact_duplicates = [
|
||||||
|
group
|
||||||
|
for group in duplicate_groups
|
||||||
|
if group.kind is DuplicateKind.EXACT_NAME
|
||||||
|
]
|
||||||
|
cloud_conflicts = [
|
||||||
|
group
|
||||||
|
for group in duplicate_groups
|
||||||
|
if group.kind is DuplicateKind.CLOUD_CONFLICT
|
||||||
|
]
|
||||||
referenced_names = {reference.filename for reference in library.references}
|
referenced_names = {reference.filename for reference in library.references}
|
||||||
unused_count = sum(
|
unused_count = sum(
|
||||||
1 for track in library.tracks if track.filename not in referenced_names
|
1 for track in library.tracks if track.filename not in referenced_names
|
||||||
@@ -29,15 +46,34 @@ def analyze_health(library: Library) -> HealthReport:
|
|||||||
|
|
||||||
return HealthReport(
|
return HealthReport(
|
||||||
score=score,
|
score=score,
|
||||||
total_references=len(results),
|
total_references=len(library.references),
|
||||||
|
scored_references=len(results),
|
||||||
healthy_references=healthy_count,
|
healthy_references=healthy_count,
|
||||||
missing_references=len(missing),
|
missing_references=len(missing),
|
||||||
unique_missing_filenames=len(
|
unique_missing_filenames=len(
|
||||||
{result.reference.filename for result in missing}
|
{result.reference.filename for result in missing}
|
||||||
),
|
),
|
||||||
disk_tracks=len(library.tracks),
|
disk_tracks=len(library.tracks),
|
||||||
duplicate_filename_groups=len(duplicate_counts),
|
duplicate_filename_groups=len(exact_duplicates),
|
||||||
duplicate_files=sum(count - 1 for count in duplicate_counts),
|
duplicate_files=sum(group.extra_files for group in exact_duplicates),
|
||||||
|
suspected_cloud_conflict_groups=len(cloud_conflicts),
|
||||||
|
suspected_cloud_conflict_files=sum(
|
||||||
|
group.extra_files for group in cloud_conflicts
|
||||||
|
),
|
||||||
unused_tracks=unused_count,
|
unused_tracks=unused_count,
|
||||||
suggested_matches=suggested_count,
|
suggested_matches=suggested_count,
|
||||||
|
static_crates=sum(
|
||||||
|
crate.kind is CrateKind.STATIC for crate in library.crates
|
||||||
|
),
|
||||||
|
smart_crates=sum(
|
||||||
|
crate.kind is CrateKind.SMART for crate in library.crates
|
||||||
|
),
|
||||||
|
unknown_crates=sum(
|
||||||
|
crate.kind is CrateKind.UNKNOWN for crate in library.crates
|
||||||
|
),
|
||||||
|
dynamic_references_excluded=sum(
|
||||||
|
len(crate.references)
|
||||||
|
for crate in library.crates
|
||||||
|
if crate.kind is CrateKind.SMART
|
||||||
|
),
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -1,4 +1,5 @@
|
|||||||
from serato_doctor.models.crate import Crate
|
from serato_doctor.models.crate import Crate, CrateKind
|
||||||
|
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
|
||||||
from serato_doctor.models.health import HealthReport
|
from serato_doctor.models.health import HealthReport
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
||||||
@@ -7,7 +8,10 @@ from serato_doctor.models.track import DiskTrack
|
|||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
"Crate",
|
"Crate",
|
||||||
|
"CrateKind",
|
||||||
"DiskTrack",
|
"DiskTrack",
|
||||||
|
"DuplicateGroup",
|
||||||
|
"DuplicateKind",
|
||||||
"HealthReport",
|
"HealthReport",
|
||||||
"Library",
|
"Library",
|
||||||
"MatchEvidence",
|
"MatchEvidence",
|
||||||
|
|||||||
@@ -1,13 +1,21 @@
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
|
from enum import Enum
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Tuple
|
from typing import Tuple
|
||||||
|
|
||||||
from serato_doctor.models.reference import TrackReference
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
|
||||||
|
|
||||||
|
class CrateKind(str, Enum):
|
||||||
|
STATIC = "static"
|
||||||
|
SMART = "smart"
|
||||||
|
UNKNOWN = "unknown"
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
class Crate:
|
class Crate:
|
||||||
"""A Serato crate and the track references parsed from it."""
|
"""A Serato crate and the track references parsed from it."""
|
||||||
|
|
||||||
path: Path
|
path: Path
|
||||||
references: Tuple[TrackReference, ...]
|
references: Tuple[TrackReference, ...]
|
||||||
|
kind: CrateKind = CrateKind.UNKNOWN
|
||||||
|
|||||||
@@ -0,0 +1,30 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from enum import Enum
|
||||||
|
from typing import Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
class DuplicateKind(str, Enum):
|
||||||
|
EXACT_NAME = "exact_name"
|
||||||
|
CLOUD_CONFLICT = "cloud_conflict"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class DuplicateGroup:
|
||||||
|
"""A deterministic group of files that warrants duplicate review."""
|
||||||
|
|
||||||
|
kind: DuplicateKind
|
||||||
|
comparison_key: str
|
||||||
|
tracks: Tuple[DiskTrack, ...]
|
||||||
|
|
||||||
|
@property
|
||||||
|
def extra_files(self) -> int:
|
||||||
|
return max(0, len(self.tracks) - 1)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def display_name(self) -> str:
|
||||||
|
return min(
|
||||||
|
(track.filename for track in self.tracks),
|
||||||
|
key=lambda name: (len(name), name.casefold()),
|
||||||
|
)
|
||||||
@@ -8,15 +8,22 @@ class HealthReport:
|
|||||||
|
|
||||||
score: Optional[float]
|
score: Optional[float]
|
||||||
total_references: int
|
total_references: int
|
||||||
|
scored_references: int
|
||||||
healthy_references: int
|
healthy_references: int
|
||||||
missing_references: int
|
missing_references: int
|
||||||
unique_missing_filenames: int
|
unique_missing_filenames: int
|
||||||
disk_tracks: int
|
disk_tracks: int
|
||||||
duplicate_filename_groups: int
|
duplicate_filename_groups: int
|
||||||
duplicate_files: int
|
duplicate_files: int
|
||||||
|
suspected_cloud_conflict_groups: int
|
||||||
|
suspected_cloud_conflict_files: int
|
||||||
unused_tracks: int
|
unused_tracks: int
|
||||||
suggested_matches: int
|
suggested_matches: int
|
||||||
|
static_crates: int
|
||||||
|
smart_crates: int
|
||||||
|
unknown_crates: int
|
||||||
|
dynamic_references_excluded: int
|
||||||
|
|
||||||
@property
|
@property
|
||||||
def score_basis(self) -> str:
|
def score_basis(self) -> str:
|
||||||
return "Resolved crate references / total crate references"
|
return "Resolved non-dynamic references / non-dynamic references scored"
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from typing import Iterable, Tuple
|
from typing import Iterable, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.crate import Crate
|
||||||
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
||||||
from serato_doctor.models.track import DiskTrack
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
@@ -11,6 +12,7 @@ class Library:
|
|||||||
|
|
||||||
references: Tuple[TrackReference, ...]
|
references: Tuple[TrackReference, ...]
|
||||||
tracks: Tuple[DiskTrack, ...]
|
tracks: Tuple[DiskTrack, ...]
|
||||||
|
crates: Tuple[Crate, ...] = ()
|
||||||
|
|
||||||
@classmethod
|
@classmethod
|
||||||
def build(
|
def build(
|
||||||
@@ -20,6 +22,18 @@ class Library:
|
|||||||
) -> "Library":
|
) -> "Library":
|
||||||
return cls(tuple(references), tuple(tracks))
|
return cls(tuple(references), tuple(tracks))
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def from_crates(
|
||||||
|
cls, crates: Iterable[Crate], tracks: Iterable[DiskTrack]
|
||||||
|
) -> "Library":
|
||||||
|
crate_tuple = tuple(crates)
|
||||||
|
references = tuple(
|
||||||
|
reference
|
||||||
|
for crate in crate_tuple
|
||||||
|
for reference in crate.references
|
||||||
|
)
|
||||||
|
return cls(references, tuple(tracks), crate_tuple)
|
||||||
|
|
||||||
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
|
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
|
||||||
disk_names = {track.filename for track in self.tracks}
|
disk_names = {track.filename for track in self.tracks}
|
||||||
return tuple(
|
return tuple(
|
||||||
|
|||||||
@@ -84,3 +84,48 @@ def test_cli_accepts_old_reference_root(tmp_path, monkeypatch, capsys):
|
|||||||
output = capsys.readouterr().out
|
output = capsys.readouterr().out
|
||||||
assert "Crate references: 1" in output
|
assert "Crate references: 1" in output
|
||||||
assert "Missing by filename: 0" in output
|
assert "Missing by filename: 0" in output
|
||||||
|
|
||||||
|
|
||||||
|
def test_analyze_prints_health_without_writing_reports(tmp_path, monkeypatch, capsys):
|
||||||
|
serato = tmp_path / "serato"
|
||||||
|
subcrates = serato / "Subcrates"
|
||||||
|
music = tmp_path / "music"
|
||||||
|
subcrates.mkdir(parents=True)
|
||||||
|
music.mkdir()
|
||||||
|
crate_text = (
|
||||||
|
"Users/sample-user/Jukebox/Found.mp3otrk"
|
||||||
|
"Users/sample-user/Jukebox/Missing.mp3otrk"
|
||||||
|
)
|
||||||
|
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
|
||||||
|
(music / "Found.mp3").write_bytes(b"synthetic audio")
|
||||||
|
csv_path = tmp_path / "scan.csv"
|
||||||
|
report_path = tmp_path / "report.txt"
|
||||||
|
monkeypatch.setattr(
|
||||||
|
sys,
|
||||||
|
"argv",
|
||||||
|
[
|
||||||
|
"serato-doctor",
|
||||||
|
"analyze",
|
||||||
|
"--serato",
|
||||||
|
str(serato),
|
||||||
|
"--music",
|
||||||
|
str(music),
|
||||||
|
"--out",
|
||||||
|
str(csv_path),
|
||||||
|
"--report",
|
||||||
|
str(report_path),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
main()
|
||||||
|
|
||||||
|
output = capsys.readouterr().out
|
||||||
|
assert "Overall Health: 50.0%" in output
|
||||||
|
assert "Tracks: 1" in output
|
||||||
|
assert "Crate References: 2" in output
|
||||||
|
assert "References Scored: 2" in output
|
||||||
|
assert "Healthy References: 1" in output
|
||||||
|
assert "Broken References: 1" in output
|
||||||
|
assert "Unused Tracks: 0" in output
|
||||||
|
assert not csv_path.exists()
|
||||||
|
assert not report_path.exists()
|
||||||
|
|||||||
@@ -3,12 +3,14 @@ from pathlib import Path
|
|||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
from serato_doctor.crate_parser import (
|
from serato_doctor.crate_parser import (
|
||||||
|
classify_crate,
|
||||||
clean_path,
|
clean_path,
|
||||||
load_crate,
|
load_crate,
|
||||||
parse_crate,
|
parse_crate,
|
||||||
parse_crates,
|
parse_crates,
|
||||||
path_markers,
|
path_markers,
|
||||||
)
|
)
|
||||||
|
from serato_doctor.models.crate import CrateKind
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize(
|
@pytest.mark.parametrize(
|
||||||
@@ -84,3 +86,18 @@ def test_parse_crates_reuses_configured_roots_for_every_crate(tmp_path):
|
|||||||
"First.mp3",
|
"First.mp3",
|
||||||
"Second.mp3",
|
"Second.mp3",
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("relative_path", "expected"),
|
||||||
|
[
|
||||||
|
("_Serato_/Subcrates/House.crate", CrateKind.STATIC),
|
||||||
|
("_Serato_/Smartcrates/Warmup.crate", CrateKind.SMART),
|
||||||
|
("_Serato_/Subcrates/Compatible by key.crate", CrateKind.SMART),
|
||||||
|
("fixtures/Unknown.crate", CrateKind.UNKNOWN),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_classify_crate_uses_provenance_and_known_dynamic_name(
|
||||||
|
tmp_path, relative_path, expected
|
||||||
|
):
|
||||||
|
assert classify_crate(tmp_path / relative_path) is expected
|
||||||
|
|||||||
@@ -0,0 +1,57 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.duplicates import find_duplicate_groups
|
||||||
|
from serato_doctor.models.duplicate import DuplicateKind
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def track(path):
|
||||||
|
path = Path(path)
|
||||||
|
return DiskTrack(path, path.name, 100, path.suffix.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def test_exact_names_in_different_folders_form_a_group():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[track("/A/Song.mp3"), track("/B/Song.mp3")]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert len(groups) == 1
|
||||||
|
assert groups[0].kind is DuplicateKind.EXACT_NAME
|
||||||
|
assert groups[0].display_name == "Song.mp3"
|
||||||
|
assert groups[0].extra_files == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_case_only_difference_is_an_exact_name_duplicate():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[track("/A/SONG.MP3"), track("/B/song.mp3")]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert groups[0].kind is DuplicateKind.EXACT_NAME
|
||||||
|
|
||||||
|
|
||||||
|
def test_numeric_suffixes_form_a_suspected_cloud_conflict_group():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[
|
||||||
|
track("/A/Track.mp3"),
|
||||||
|
track("/B/Track 2.mp3"),
|
||||||
|
track("/C/Track 3.mp3"),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert len(groups) == 1
|
||||||
|
assert groups[0].kind is DuplicateKind.CLOUD_CONFLICT
|
||||||
|
assert groups[0].display_name == "Track.mp3"
|
||||||
|
assert groups[0].extra_files == 2
|
||||||
|
assert [item.path for item in groups[0].tracks] == [
|
||||||
|
Path("/A/Track.mp3"),
|
||||||
|
Path("/B/Track 2.mp3"),
|
||||||
|
Path("/C/Track 3.mp3"),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_unrelated_and_single_files_are_not_reported():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[track("/A/First.mp3"), track("/B/Second.mp3")]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert groups == ()
|
||||||
+51
-1
@@ -1,6 +1,7 @@
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from serato_doctor.health import analyze_health
|
from serato_doctor.health import analyze_health
|
||||||
|
from serato_doctor.models.crate import Crate, CrateKind
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.models.reference import TrackReference
|
from serato_doctor.models.reference import TrackReference
|
||||||
from serato_doctor.models.track import DiskTrack
|
from serato_doctor.models.track import DiskTrack
|
||||||
@@ -36,9 +37,10 @@ def test_health_report_exposes_each_metric():
|
|||||||
|
|
||||||
assert report.score == 33.3
|
assert report.score == 33.3
|
||||||
assert report.score_basis == (
|
assert report.score_basis == (
|
||||||
"Resolved crate references / total crate references"
|
"Resolved non-dynamic references / non-dynamic references scored"
|
||||||
)
|
)
|
||||||
assert report.total_references == 3
|
assert report.total_references == 3
|
||||||
|
assert report.scored_references == 3
|
||||||
assert report.healthy_references == 1
|
assert report.healthy_references == 1
|
||||||
assert report.missing_references == 2
|
assert report.missing_references == 2
|
||||||
assert report.unique_missing_filenames == 2
|
assert report.unique_missing_filenames == 2
|
||||||
@@ -54,3 +56,51 @@ def test_empty_library_has_no_health_score():
|
|||||||
|
|
||||||
assert report.score is None
|
assert report.score is None
|
||||||
assert report.total_references == 0
|
assert report.total_references == 0
|
||||||
|
assert report.scored_references == 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_smart_crate_references_are_reported_but_not_scored():
|
||||||
|
static_reference = reference("Found.mp3")
|
||||||
|
smart_reference = TrackReference(
|
||||||
|
Path("Smartcrates/Dynamic.crate"),
|
||||||
|
Path("/old/House/Dynamic.mp3"),
|
||||||
|
"Dynamic.mp3",
|
||||||
|
)
|
||||||
|
crates = [
|
||||||
|
Crate(Path("Subcrates/Static.crate"), (static_reference,), CrateKind.STATIC),
|
||||||
|
Crate(
|
||||||
|
Path("Smartcrates/Dynamic.crate"),
|
||||||
|
(smart_reference,),
|
||||||
|
CrateKind.SMART,
|
||||||
|
),
|
||||||
|
]
|
||||||
|
library = Library.from_crates(crates, [track("Found.mp3")])
|
||||||
|
|
||||||
|
report = analyze_health(library)
|
||||||
|
|
||||||
|
assert report.score == 100.0
|
||||||
|
assert report.total_references == 2
|
||||||
|
assert report.scored_references == 1
|
||||||
|
assert report.missing_references == 0
|
||||||
|
assert report.static_crates == 1
|
||||||
|
assert report.smart_crates == 1
|
||||||
|
assert report.dynamic_references_excluded == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_health_reports_cloud_conflicts_separately_from_exact_duplicates():
|
||||||
|
library = Library.build(
|
||||||
|
[],
|
||||||
|
[
|
||||||
|
track("Track.mp3", "Original"),
|
||||||
|
track("Track 2.mp3", "Conflict"),
|
||||||
|
track("Copy.mp3", "First"),
|
||||||
|
track("Copy.mp3", "Second"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
report = analyze_health(library)
|
||||||
|
|
||||||
|
assert report.duplicate_filename_groups == 1
|
||||||
|
assert report.duplicate_files == 1
|
||||||
|
assert report.suspected_cloud_conflict_groups == 1
|
||||||
|
assert report.suspected_cloud_conflict_files == 1
|
||||||
|
|||||||
@@ -1,7 +1,8 @@
|
|||||||
import runpy
|
import runpy
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from serato_doctor.crate_parser import parse_crates
|
from serato_doctor.crate_parser import load_library_crates
|
||||||
|
from serato_doctor.models.crate import CrateKind
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.scanner import scan_audio
|
from serato_doctor.scanner import scan_audio
|
||||||
|
|
||||||
@@ -13,9 +14,10 @@ def test_generated_sample_library_has_expected_scenario(tmp_path):
|
|||||||
build_sample = runpy.run_path(str(generator_path))["build_sample"]
|
build_sample = runpy.run_path(str(generator_path))["build_sample"]
|
||||||
sample_root = build_sample(tmp_path / "sample")
|
sample_root = build_sample(tmp_path / "sample")
|
||||||
|
|
||||||
references = parse_crates(sample_root / "Serato" / "_Serato_" / "Subcrates")
|
crates = load_library_crates(sample_root / "Serato" / "_Serato_")
|
||||||
tracks = scan_audio(sample_root / "Music")
|
tracks = scan_audio(sample_root / "Music")
|
||||||
results = Library.build(references, tracks).reconcile_by_filename()
|
library = Library.from_crates(crates, tracks)
|
||||||
|
results = library.reconcile_by_filename()
|
||||||
missing = {
|
missing = {
|
||||||
result.reference.filename
|
result.reference.filename
|
||||||
for result in results
|
for result in results
|
||||||
@@ -23,6 +25,8 @@ def test_generated_sample_library_has_expected_scenario(tmp_path):
|
|||||||
}
|
}
|
||||||
|
|
||||||
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7
|
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7
|
||||||
assert len(references) == 13
|
assert len(library.references) == 13
|
||||||
assert len(tracks) == 10
|
assert len(tracks) == 10
|
||||||
assert missing == {"Missing.mp3", "Old Name.mp3"}
|
assert missing == {"Missing.mp3", "Old Name.mp3"}
|
||||||
|
assert sum(crate.kind is CrateKind.STATIC for crate in crates) == 5
|
||||||
|
assert sum(crate.kind is CrateKind.SMART for crate in crates) == 2
|
||||||
|
|||||||
Reference in New Issue
Block a user