Compare commits
9 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 6449c7de6a | |||
| d764f910e4 | |||
| 35713bbbb3 | |||
| c92d929f37 | |||
| e6f313d98a | |||
| 22424d0b5c | |||
| b8bf9ec6e1 | |||
| 167f029c28 | |||
| 5a2edb8b94 |
@@ -0,0 +1,37 @@
|
|||||||
|
# Serato Doctor
|
||||||
|
|
||||||
|
Serato Doctor is a read-only-first toolkit for inspecting and eventually repairing
|
||||||
|
DJ libraries. Serato is the first supported engine.
|
||||||
|
|
||||||
|
## Analyze a Library
|
||||||
|
|
||||||
|
```shell
|
||||||
|
serato-doctor analyze \
|
||||||
|
--serato ~/Music/_Serato_ \
|
||||||
|
--music ~/Music/Jukebox
|
||||||
|
```
|
||||||
|
|
||||||
|
Analysis prints a transparent reference-integrity score plus broken references,
|
||||||
|
duplicate filenames, unused tracks, and suggested filename matches. It does not
|
||||||
|
modify the library or create report files.
|
||||||
|
|
||||||
|
If crates contain an older library root, supply it explicitly:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
serato-doctor analyze --reference-root /Users/old-user/OneDrive/Jukebox
|
||||||
|
```
|
||||||
|
|
||||||
|
## Generate Detailed Reports
|
||||||
|
|
||||||
|
Running without a command preserves the original scanner behavior and writes CSV
|
||||||
|
and text reports:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
serato-doctor --serato ~/Music/_Serato_ --music ~/Music/Jukebox
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `--verbose` for progress on standard error or `--log-file PATH` for an
|
||||||
|
aggregate diagnostic log.
|
||||||
|
|
||||||
|
Serato Doctor never repairs files without an explicit future repair workflow,
|
||||||
|
preview, backup, and rollback path.
|
||||||
|
|||||||
+5
-3
@@ -13,16 +13,18 @@
|
|||||||
- [ ] Database V2 read-only parser
|
- [ ] Database V2 read-only parser
|
||||||
- [x] Configuration
|
- [x] Configuration
|
||||||
- [x] Logging
|
- [x] Logging
|
||||||
|
- [x] Matching engine
|
||||||
|
|
||||||
## v0.2 — Diagnostics
|
## v0.2 — Diagnostics
|
||||||
|
|
||||||
- [ ] Duplicate filename detection
|
- [x] `serato-doctor analyze` command
|
||||||
|
- [x] Duplicate filename detection
|
||||||
- [ ] Duplicate audio hash detection
|
- [ ] Duplicate audio hash detection
|
||||||
- [ ] Broken symlink detection
|
- [ ] Broken symlink detection
|
||||||
- [ ] Orphaned audio detection
|
- [ ] Orphaned audio detection
|
||||||
- [ ] OneDrive rename detection
|
- [ ] OneDrive rename detection
|
||||||
- [ ] Crate classification: static vs smart/dynamic
|
- [x] Crate classification: static vs smart/dynamic
|
||||||
- [ ] Library health score
|
- [x] Library health score
|
||||||
|
|
||||||
## v0.3 — Safe Repair
|
## v0.3 — Safe Repair
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,27 @@
|
|||||||
|
# Analyze Command
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The health and matching engines are only Python APIs. Users need one safe command
|
||||||
|
that summarizes library integrity without first interpreting CSV files.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
`serato-doctor analyze` reuses the existing configured crate parser and filesystem
|
||||||
|
scanner, then prints the immutable health report. The original no-command mode is
|
||||||
|
retained as the report-producing scan workflow.
|
||||||
|
|
||||||
|
Analyze mode does not create CSV or text reports. Both modes remain read-only with
|
||||||
|
respect to Serato crates, databases, and music files.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- A library with no references displays `Not assessed` rather than a false score.
|
||||||
|
- Historical roots continue to work through repeatable `--reference-root` flags.
|
||||||
|
- Normal output remains separate from optional diagnostic logging.
|
||||||
|
- Existing scripts that invoke the CLI without a subcommand remain compatible.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests prove the legacy five-line output and reports remain unchanged, while analyze
|
||||||
|
prints health metrics and does not create output files.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
# Crate Classification
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Static crates are manually maintained track lists. Smart crates are dynamic views
|
||||||
|
generated from rules, so stale-looking entries in them should not be presented as
|
||||||
|
broken manual references or given the same health-score weight.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Crates carry a `static`, `smart`, or `unknown` kind. Folder provenance is the
|
||||||
|
primary signal: Serato stores regular definitions in `Subcrates` and smart
|
||||||
|
definitions in `Smartcrates`. `Compatible by key.crate` is also treated as smart
|
||||||
|
when encountered in `Subcrates`, based on the original migration case study.
|
||||||
|
|
||||||
|
The library loader reads both folders. Health analysis reports all references but
|
||||||
|
scores only non-smart references. Unknown crates remain scoreable so incomplete
|
||||||
|
classification cannot silently hide potential problems.
|
||||||
|
|
||||||
|
Serato documents the folder distinction in [What is in the _Serato_ folder?](https://support.serato.com/hc/en-us/articles/204022904-What-is-in-the-Serato-folder)
|
||||||
|
and explains that smart crates are populated from rules in [Crates in Serato DJ](https://support.serato.com/hc/en-us/articles/227561407-Crates-in-Serato-DJ-Pro-Serato-DJ-Lite).
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- The known `Compatible by key` dynamic crate in the `Subcrates` folder.
|
||||||
|
- Crate fixtures outside a recognized Serato folder.
|
||||||
|
- Libraries containing both static and smart references to the same track.
|
||||||
|
- Dynamic references whose current materialized paths appear missing.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover all three kinds, both Serato folders, the known dynamic fallback, the
|
||||||
|
synthetic five-static/two-smart library, and exclusion from health scoring.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Duplicate Filename Detection
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Two files with the same filename may be ordinary copies, while names such as
|
||||||
|
`Track.mp3` and `Track 2.mp3` may indicate a OneDrive conflict. These cases need
|
||||||
|
review, but neither is sufficient evidence for deletion.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The duplicate detector emits immutable groups of two kinds:
|
||||||
|
|
||||||
|
- `exact_name` groups filenames after case and Unicode normalization.
|
||||||
|
- `cloud_conflict` groups names after additionally removing a trailing numeric
|
||||||
|
suffix from the stem.
|
||||||
|
|
||||||
|
Groups contain every path, a stable comparison key, a display name, and the number
|
||||||
|
of files beyond the first. Health analysis reports exact and suspected-conflict
|
||||||
|
counts separately. Neither category changes the health score.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Same filename in different folders.
|
||||||
|
- Case-only and Unicode representation differences.
|
||||||
|
- Multiple conflict suffixes such as `Track 2.mp3` and `Track 3.mp3`.
|
||||||
|
- Legitimate numbered titles, which remain explicitly labeled as suspected.
|
||||||
|
- A single file, which is never reported as a duplicate.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover exact, normalized, conflict-family, unrelated, and deterministic-order
|
||||||
|
behavior, plus integration with health analysis. Detection is read-only.
|
||||||
@@ -0,0 +1,32 @@
|
|||||||
|
# Library Health Engine
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Raw missing-reference counts do not provide a compact view of library integrity,
|
||||||
|
but an opaque blended score would imply confidence the current data cannot support.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The health engine produces an immutable report from the core `Library`. Its score
|
||||||
|
is only the percentage of non-dynamic crate references resolved by exact filename.
|
||||||
|
Smart-crate references are counted but excluded because their contents are derived
|
||||||
|
from rules. The report also exposes missing references, unique missing filenames,
|
||||||
|
duplicate filename groups, extra duplicate files, unused tracks, and missing
|
||||||
|
references with matching candidates.
|
||||||
|
|
||||||
|
Duplicate, unused, and candidate counts are informational. They do not affect the
|
||||||
|
score until the project has a documented and validated weighting policy. An empty
|
||||||
|
library has no score rather than a misleading 0% or 100%.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- The same missing filename referenced by several crates.
|
||||||
|
- Several disk files sharing a filename.
|
||||||
|
- Conflict-suffixed files that are unused but may be match candidates.
|
||||||
|
- Libraries with no crate references.
|
||||||
|
- Tracks referenced by filename from more than one crate.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests assert every metric, the disclosed score basis, candidate integration, and
|
||||||
|
empty-library behavior. Analysis remains entirely read-only.
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Explainable Matching Engine
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Missing references need ranked candidate files, but a filename-only yes/no check
|
||||||
|
cannot explain ambiguity or cloud-provider conflict names.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The read-only matching engine indexes normalized filenames and scores only related
|
||||||
|
candidates. Every score contains evidence for filename, extension, and parent
|
||||||
|
folder. Exact filenames earn 60 points, normalized names 55, numeric conflict-name
|
||||||
|
matches 50, extensions 10, and parent folders 20.
|
||||||
|
|
||||||
|
The displayed percentage is an evidence score, not a statistical probability.
|
||||||
|
Metadata, duration, hashes, and fingerprints can add stronger evidence later.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Unicode and case differences.
|
||||||
|
- OneDrive-style names such as `Track 2.mp3`.
|
||||||
|
- Duplicate candidates in different folders.
|
||||||
|
- Legitimate numbered song titles, which remain candidates but are never repaired.
|
||||||
|
- Unrelated names, which are not emitted as candidates.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover exact, normalized, conflict-suffix, ambiguous, and unrelated filenames.
|
||||||
|
Candidate ordering is deterministic. The engine never changes a track or reference.
|
||||||
@@ -18,8 +18,7 @@ def build_sample(output: Optional[Path] = None) -> Path:
|
|||||||
root = output or SAMPLE_ROOT / "generated"
|
root = output or SAMPLE_ROOT / "generated"
|
||||||
manifest = load_manifest()
|
manifest = load_manifest()
|
||||||
music_root = root / "Music"
|
music_root = root / "Music"
|
||||||
crate_root = root / "Serato" / "_Serato_" / "Subcrates"
|
serato_root = root / "Serato" / "_Serato_"
|
||||||
crate_root.mkdir(parents=True, exist_ok=True)
|
|
||||||
|
|
||||||
for relative_path in manifest["tracks"]:
|
for relative_path in manifest["tracks"]:
|
||||||
track_path = music_root / relative_path
|
track_path = music_root / relative_path
|
||||||
@@ -29,6 +28,13 @@ def build_sample(output: Optional[Path] = None) -> Path:
|
|||||||
)
|
)
|
||||||
|
|
||||||
for crate in manifest["crates"]:
|
for crate in manifest["crates"]:
|
||||||
|
folder_name = "Smartcrates" if crate["type"] == "smart" else "Subcrates"
|
||||||
|
crate_root = serato_root / folder_name
|
||||||
|
crate_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
other_folder = "Subcrates" if folder_name == "Smartcrates" else "Smartcrates"
|
||||||
|
stale_path = serato_root / other_folder / crate["name"]
|
||||||
|
if stale_path.exists():
|
||||||
|
stale_path.unlink()
|
||||||
records = "".join(
|
records = "".join(
|
||||||
f"{SERATO_PATH_PREFIX}{relative_path}otrk"
|
f"{SERATO_PATH_PREFIX}{relative_path}otrk"
|
||||||
for relative_path in crate["references"]
|
for relative_path in crate["references"]
|
||||||
|
|||||||
+52
-10
@@ -2,7 +2,8 @@ from pathlib import Path
|
|||||||
import argparse
|
import argparse
|
||||||
|
|
||||||
from serato_doctor.config import ScanConfig
|
from serato_doctor.config import ScanConfig
|
||||||
from serato_doctor.crate_parser import parse_crates
|
from serato_doctor.crate_parser import load_library_crates
|
||||||
|
from serato_doctor.health import analyze_health
|
||||||
from serato_doctor.logging import configure_logging
|
from serato_doctor.logging import configure_logging
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.scanner import scan_audio
|
from serato_doctor.scanner import scan_audio
|
||||||
@@ -11,10 +12,23 @@ from serato_doctor.report import write_csv, write_missing_report
|
|||||||
|
|
||||||
def main():
|
def main():
|
||||||
parser = argparse.ArgumentParser(prog="serato-doctor")
|
parser = argparse.ArgumentParser(prog="serato-doctor")
|
||||||
|
parser.add_argument(
|
||||||
|
"command", nargs="?", choices=("scan", "analyze"), default="scan"
|
||||||
|
)
|
||||||
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
||||||
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"))
|
parser.add_argument(
|
||||||
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv"))
|
"--music",
|
||||||
parser.add_argument("--report", default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"))
|
default=str(
|
||||||
|
Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"
|
||||||
|
),
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--report",
|
||||||
|
default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"),
|
||||||
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--reference-root",
|
"--reference-root",
|
||||||
action="append",
|
action="append",
|
||||||
@@ -37,14 +51,13 @@ def main():
|
|||||||
log_file=Path(args.log_file) if args.log_file else None,
|
log_file=Path(args.log_file) if args.log_file else None,
|
||||||
)
|
)
|
||||||
logger = configure_logging(config.verbose, config.log_file)
|
logger = configure_logging(config.verbose, config.log_file)
|
||||||
logger.info("Starting read-only library scan")
|
logger.info("Starting read-only library inspection")
|
||||||
logger.debug("Serato directory: %s", config.serato)
|
logger.debug("Serato directory: %s", config.serato)
|
||||||
logger.debug("Music directory: %s", config.music)
|
logger.debug("Music directory: %s", config.music)
|
||||||
|
|
||||||
library = Library.build(
|
crates = load_library_crates(config.serato, config.reference_roots)
|
||||||
references=parse_crates(
|
library = Library.from_crates(
|
||||||
config.serato / "Subcrates", config.reference_roots
|
crates=crates,
|
||||||
),
|
|
||||||
tracks=scan_audio(config.music),
|
tracks=scan_audio(config.music),
|
||||||
)
|
)
|
||||||
results = library.reconcile_by_filename()
|
results = library.reconcile_by_filename()
|
||||||
@@ -53,10 +66,39 @@ def main():
|
|||||||
logger.info("Scanned %d disk tracks", len(library.tracks))
|
logger.info("Scanned %d disk tracks", len(library.tracks))
|
||||||
logger.info("Found %d references missing by filename", missing_count)
|
logger.info("Found %d references missing by filename", missing_count)
|
||||||
|
|
||||||
|
if args.command == "analyze":
|
||||||
|
health = analyze_health(library)
|
||||||
|
score = f"{health.score:.1f}%" if health.score is not None else "Not assessed"
|
||||||
|
logger.info("Calculated library health: %s", score)
|
||||||
|
print(f"Overall Health: {score}")
|
||||||
|
print(f"Score Basis: {health.score_basis}")
|
||||||
|
print(f"Tracks: {health.disk_tracks}")
|
||||||
|
print(f"Crate References: {health.total_references}")
|
||||||
|
print(f"References Scored: {health.scored_references}")
|
||||||
|
print(f"Healthy References: {health.healthy_references}")
|
||||||
|
print(f"Broken References: {health.missing_references}")
|
||||||
|
print(f"Unique Missing Filenames: {health.unique_missing_filenames}")
|
||||||
|
print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}")
|
||||||
|
print(f"Duplicate Files: {health.duplicate_files}")
|
||||||
|
print(
|
||||||
|
"Suspected Cloud Conflict Groups: "
|
||||||
|
f"{health.suspected_cloud_conflict_groups}"
|
||||||
|
)
|
||||||
|
print(
|
||||||
|
"Suspected Cloud Conflict Files: "
|
||||||
|
f"{health.suspected_cloud_conflict_files}"
|
||||||
|
)
|
||||||
|
print(f"Unused Tracks: {health.unused_tracks}")
|
||||||
|
print(f"Suggested Matches: {health.suggested_matches}")
|
||||||
|
print(f"Static Crates: {health.static_crates}")
|
||||||
|
print(f"Smart Crates: {health.smart_crates}")
|
||||||
|
print(
|
||||||
|
f"Dynamic References Excluded: {health.dynamic_references_excluded}"
|
||||||
|
)
|
||||||
|
else:
|
||||||
write_csv(results, config.out)
|
write_csv(results, config.out)
|
||||||
write_missing_report(results, config.report)
|
write_missing_report(results, config.report)
|
||||||
logger.info("Wrote CSV and missing-reference reports")
|
logger.info("Wrote CSV and missing-reference reports")
|
||||||
|
|
||||||
print(f"Crate references: {len(library.references)}")
|
print(f"Crate references: {len(library.references)}")
|
||||||
print(f"Disk tracks: {len(library.tracks)}")
|
print(f"Disk tracks: {len(library.tracks)}")
|
||||||
print(f"Missing by filename: {missing_count}")
|
print(f"Missing by filename: {missing_count}")
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Iterable, Tuple
|
from typing import Iterable, Tuple
|
||||||
|
|
||||||
from serato_doctor.models.crate import Crate
|
from serato_doctor.models.crate import Crate, CrateKind
|
||||||
from serato_doctor.models.reference import TrackReference
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
|
||||||
|
|
||||||
@@ -32,6 +32,17 @@ def path_markers(reference_roots: Iterable[Path]) -> Tuple[str, ...]:
|
|||||||
return configured or DEFAULT_PATH_MARKERS
|
return configured or DEFAULT_PATH_MARKERS
|
||||||
|
|
||||||
|
|
||||||
|
def classify_crate(crate_path: Path) -> CrateKind:
|
||||||
|
parent_names = {parent.name.casefold() for parent in crate_path.parents}
|
||||||
|
if "smartcrates" in parent_names:
|
||||||
|
return CrateKind.SMART
|
||||||
|
if crate_path.name.casefold() == "compatible by key.crate":
|
||||||
|
return CrateKind.SMART
|
||||||
|
if "subcrates" in parent_names:
|
||||||
|
return CrateKind.STATIC
|
||||||
|
return CrateKind.UNKNOWN
|
||||||
|
|
||||||
|
|
||||||
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
||||||
text = read_crate_text(crate_path)
|
text = read_crate_text(crate_path)
|
||||||
refs = []
|
refs = []
|
||||||
@@ -63,7 +74,11 @@ def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
|||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
||||||
return Crate(path=crate_path, references=tuple(refs))
|
return Crate(
|
||||||
|
path=crate_path,
|
||||||
|
references=tuple(refs),
|
||||||
|
kind=classify_crate(crate_path),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def parse_crate(
|
def parse_crate(
|
||||||
@@ -82,3 +97,15 @@ def parse_crates(
|
|||||||
for crate in root.rglob("*.crate"):
|
for crate in root.rglob("*.crate"):
|
||||||
refs.extend(parse_crate(crate, reference_roots))
|
refs.extend(parse_crate(crate, reference_roots))
|
||||||
return refs
|
return refs
|
||||||
|
|
||||||
|
|
||||||
|
def load_library_crates(
|
||||||
|
serato_root: Path, reference_roots: Iterable[Path] = ()
|
||||||
|
) -> Tuple[Crate, ...]:
|
||||||
|
reference_roots = tuple(reference_roots)
|
||||||
|
crates = []
|
||||||
|
for folder_name in ("Subcrates", "Smartcrates"):
|
||||||
|
folder = serato_root / folder_name
|
||||||
|
for crate_path in folder.rglob("*.crate"):
|
||||||
|
crates.append(load_crate(crate_path, reference_roots))
|
||||||
|
return tuple(sorted(crates, key=lambda crate: str(crate.path)))
|
||||||
|
|||||||
@@ -0,0 +1,41 @@
|
|||||||
|
from collections import defaultdict
|
||||||
|
from typing import DefaultDict, Iterable, List, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.matching import cloud_conflict_name, normalize
|
||||||
|
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def find_duplicate_groups(
|
||||||
|
tracks: Iterable[DiskTrack],
|
||||||
|
) -> Tuple[DuplicateGroup, ...]:
|
||||||
|
"""Find exact-name duplicates and suspected numeric conflict copies."""
|
||||||
|
|
||||||
|
track_tuple = tuple(tracks)
|
||||||
|
by_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
|
||||||
|
by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
|
||||||
|
|
||||||
|
for track in track_tuple:
|
||||||
|
by_name[normalize(track.filename)].append(track)
|
||||||
|
by_conflict_name[cloud_conflict_name(track.filename)].append(track)
|
||||||
|
|
||||||
|
groups = []
|
||||||
|
for key, matches in by_name.items():
|
||||||
|
if len(matches) > 1:
|
||||||
|
groups.append(_group(DuplicateKind.EXACT_NAME, key, matches))
|
||||||
|
|
||||||
|
for key, matches in by_conflict_name.items():
|
||||||
|
distinct_names = {normalize(track.filename) for track in matches}
|
||||||
|
if len(distinct_names) > 1:
|
||||||
|
groups.append(_group(DuplicateKind.CLOUD_CONFLICT, key, matches))
|
||||||
|
|
||||||
|
return tuple(
|
||||||
|
sorted(groups, key=lambda group: (group.kind.value, group.comparison_key))
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _group(
|
||||||
|
kind: DuplicateKind, key: str, tracks: Iterable[DiskTrack]
|
||||||
|
) -> DuplicateGroup:
|
||||||
|
ordered = tuple(sorted(tracks, key=lambda track: str(track.path)))
|
||||||
|
return DuplicateGroup(kind, key, ordered)
|
||||||
@@ -0,0 +1,79 @@
|
|||||||
|
from serato_doctor.duplicates import find_duplicate_groups
|
||||||
|
from serato_doctor.matching import MatchingEngine
|
||||||
|
from serato_doctor.models.crate import CrateKind
|
||||||
|
from serato_doctor.models.duplicate import DuplicateKind
|
||||||
|
from serato_doctor.models.health import HealthReport
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
|
||||||
|
|
||||||
|
def analyze_health(library: Library) -> HealthReport:
|
||||||
|
"""Calculate defensible health metrics without changing the library."""
|
||||||
|
|
||||||
|
dynamic_sources = {
|
||||||
|
crate.path for crate in library.crates if crate.kind is CrateKind.SMART
|
||||||
|
}
|
||||||
|
results = tuple(
|
||||||
|
result
|
||||||
|
for result in library.reconcile_by_filename()
|
||||||
|
if result.reference.source not in dynamic_sources
|
||||||
|
)
|
||||||
|
missing = [result for result in results if not result.exists_by_filename]
|
||||||
|
healthy_count = len(results) - len(missing)
|
||||||
|
score = (
|
||||||
|
round(healthy_count / len(results) * 100, 1) if results else None
|
||||||
|
)
|
||||||
|
|
||||||
|
duplicate_groups = find_duplicate_groups(library.tracks)
|
||||||
|
exact_duplicates = [
|
||||||
|
group
|
||||||
|
for group in duplicate_groups
|
||||||
|
if group.kind is DuplicateKind.EXACT_NAME
|
||||||
|
]
|
||||||
|
cloud_conflicts = [
|
||||||
|
group
|
||||||
|
for group in duplicate_groups
|
||||||
|
if group.kind is DuplicateKind.CLOUD_CONFLICT
|
||||||
|
]
|
||||||
|
referenced_names = {reference.filename for reference in library.references}
|
||||||
|
unused_count = sum(
|
||||||
|
1 for track in library.tracks if track.filename not in referenced_names
|
||||||
|
)
|
||||||
|
|
||||||
|
matcher = MatchingEngine(library.tracks)
|
||||||
|
suggested_count = sum(
|
||||||
|
bool(matcher.candidates_for(result.reference)) for result in missing
|
||||||
|
)
|
||||||
|
|
||||||
|
return HealthReport(
|
||||||
|
score=score,
|
||||||
|
total_references=len(library.references),
|
||||||
|
scored_references=len(results),
|
||||||
|
healthy_references=healthy_count,
|
||||||
|
missing_references=len(missing),
|
||||||
|
unique_missing_filenames=len(
|
||||||
|
{result.reference.filename for result in missing}
|
||||||
|
),
|
||||||
|
disk_tracks=len(library.tracks),
|
||||||
|
duplicate_filename_groups=len(exact_duplicates),
|
||||||
|
duplicate_files=sum(group.extra_files for group in exact_duplicates),
|
||||||
|
suspected_cloud_conflict_groups=len(cloud_conflicts),
|
||||||
|
suspected_cloud_conflict_files=sum(
|
||||||
|
group.extra_files for group in cloud_conflicts
|
||||||
|
),
|
||||||
|
unused_tracks=unused_count,
|
||||||
|
suggested_matches=suggested_count,
|
||||||
|
static_crates=sum(
|
||||||
|
crate.kind is CrateKind.STATIC for crate in library.crates
|
||||||
|
),
|
||||||
|
smart_crates=sum(
|
||||||
|
crate.kind is CrateKind.SMART for crate in library.crates
|
||||||
|
),
|
||||||
|
unknown_crates=sum(
|
||||||
|
crate.kind is CrateKind.UNKNOWN for crate in library.crates
|
||||||
|
),
|
||||||
|
dynamic_references_excluded=sum(
|
||||||
|
len(crate.references)
|
||||||
|
for crate in library.crates
|
||||||
|
if crate.kind is CrateKind.SMART
|
||||||
|
),
|
||||||
|
)
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
import re
|
||||||
|
import unicodedata
|
||||||
|
from collections import defaultdict
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import DefaultDict, Iterable, List, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def normalize(value: str) -> str:
|
||||||
|
return unicodedata.normalize("NFKC", value).casefold()
|
||||||
|
|
||||||
|
|
||||||
|
def cloud_conflict_name(filename: str) -> str:
|
||||||
|
"""Remove a trailing numeric cloud-conflict suffix from a filename stem."""
|
||||||
|
|
||||||
|
path = Path(filename)
|
||||||
|
stem = re.sub(r" \d+$", "", path.stem)
|
||||||
|
return normalize(stem + path.suffix)
|
||||||
|
|
||||||
|
|
||||||
|
def score_candidate(reference: TrackReference, track: DiskTrack) -> TrackMatch:
|
||||||
|
reference_name = reference.filename
|
||||||
|
track_name = track.filename
|
||||||
|
|
||||||
|
if reference_name == track_name:
|
||||||
|
filename_points = 60
|
||||||
|
filename_reason = "Filename is identical"
|
||||||
|
elif normalize(reference_name) == normalize(track_name):
|
||||||
|
filename_points = 55
|
||||||
|
filename_reason = "Filename matches after case and Unicode normalization"
|
||||||
|
elif cloud_conflict_name(reference_name) == cloud_conflict_name(track_name):
|
||||||
|
filename_points = 50
|
||||||
|
filename_reason = "Filename matches after removing a numeric conflict suffix"
|
||||||
|
else:
|
||||||
|
filename_points = 0
|
||||||
|
filename_reason = "Filename does not match"
|
||||||
|
|
||||||
|
same_extension = normalize(reference.path.suffix) == normalize(track.suffix)
|
||||||
|
same_parent = normalize(reference.path.parent.name) == normalize(
|
||||||
|
track.path.parent.name
|
||||||
|
)
|
||||||
|
evidence = (
|
||||||
|
MatchEvidence(
|
||||||
|
"filename",
|
||||||
|
filename_points > 0,
|
||||||
|
filename_points,
|
||||||
|
60,
|
||||||
|
filename_reason,
|
||||||
|
),
|
||||||
|
MatchEvidence(
|
||||||
|
"extension",
|
||||||
|
same_extension,
|
||||||
|
10 if same_extension else 0,
|
||||||
|
10,
|
||||||
|
"File extension matches" if same_extension else "File extension differs",
|
||||||
|
),
|
||||||
|
MatchEvidence(
|
||||||
|
"parent_folder",
|
||||||
|
same_parent,
|
||||||
|
20 if same_parent else 0,
|
||||||
|
20,
|
||||||
|
"Parent folder matches" if same_parent else "Parent folder differs",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
return TrackMatch(reference, track, evidence)
|
||||||
|
|
||||||
|
|
||||||
|
class MatchingEngine:
|
||||||
|
"""Find and rank filename-related disk candidates without modifying files."""
|
||||||
|
|
||||||
|
def __init__(self, tracks: Iterable[DiskTrack]):
|
||||||
|
self._by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
|
||||||
|
for track in tracks:
|
||||||
|
self._by_conflict_name[cloud_conflict_name(track.filename)].append(track)
|
||||||
|
|
||||||
|
def candidates_for(self, reference: TrackReference) -> Tuple[TrackMatch, ...]:
|
||||||
|
candidates = self._by_conflict_name.get(
|
||||||
|
cloud_conflict_name(reference.filename), []
|
||||||
|
)
|
||||||
|
matches = [score_candidate(reference, track) for track in candidates]
|
||||||
|
return tuple(
|
||||||
|
sorted(matches, key=lambda match: (-match.score, str(match.track.path)))
|
||||||
|
)
|
||||||
@@ -1,6 +1,21 @@
|
|||||||
from serato_doctor.models.crate import Crate
|
from serato_doctor.models.crate import Crate, CrateKind
|
||||||
|
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
|
||||||
|
from serato_doctor.models.health import HealthReport
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
|
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
||||||
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
||||||
from serato_doctor.models.track import DiskTrack
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
__all__ = ["Crate", "DiskTrack", "Library", "ReferenceResult", "TrackReference"]
|
__all__ = [
|
||||||
|
"Crate",
|
||||||
|
"CrateKind",
|
||||||
|
"DiskTrack",
|
||||||
|
"DuplicateGroup",
|
||||||
|
"DuplicateKind",
|
||||||
|
"HealthReport",
|
||||||
|
"Library",
|
||||||
|
"MatchEvidence",
|
||||||
|
"ReferenceResult",
|
||||||
|
"TrackMatch",
|
||||||
|
"TrackReference",
|
||||||
|
]
|
||||||
|
|||||||
@@ -1,13 +1,21 @@
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
|
from enum import Enum
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Tuple
|
from typing import Tuple
|
||||||
|
|
||||||
from serato_doctor.models.reference import TrackReference
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
|
||||||
|
|
||||||
|
class CrateKind(str, Enum):
|
||||||
|
STATIC = "static"
|
||||||
|
SMART = "smart"
|
||||||
|
UNKNOWN = "unknown"
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
class Crate:
|
class Crate:
|
||||||
"""A Serato crate and the track references parsed from it."""
|
"""A Serato crate and the track references parsed from it."""
|
||||||
|
|
||||||
path: Path
|
path: Path
|
||||||
references: Tuple[TrackReference, ...]
|
references: Tuple[TrackReference, ...]
|
||||||
|
kind: CrateKind = CrateKind.UNKNOWN
|
||||||
|
|||||||
@@ -0,0 +1,30 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from enum import Enum
|
||||||
|
from typing import Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
class DuplicateKind(str, Enum):
|
||||||
|
EXACT_NAME = "exact_name"
|
||||||
|
CLOUD_CONFLICT = "cloud_conflict"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class DuplicateGroup:
|
||||||
|
"""A deterministic group of files that warrants duplicate review."""
|
||||||
|
|
||||||
|
kind: DuplicateKind
|
||||||
|
comparison_key: str
|
||||||
|
tracks: Tuple[DiskTrack, ...]
|
||||||
|
|
||||||
|
@property
|
||||||
|
def extra_files(self) -> int:
|
||||||
|
return max(0, len(self.tracks) - 1)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def display_name(self) -> str:
|
||||||
|
return min(
|
||||||
|
(track.filename for track in self.tracks),
|
||||||
|
key=lambda name: (len(name), name.casefold()),
|
||||||
|
)
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Optional
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class HealthReport:
|
||||||
|
"""Transparent aggregate findings from a read-only library analysis."""
|
||||||
|
|
||||||
|
score: Optional[float]
|
||||||
|
total_references: int
|
||||||
|
scored_references: int
|
||||||
|
healthy_references: int
|
||||||
|
missing_references: int
|
||||||
|
unique_missing_filenames: int
|
||||||
|
disk_tracks: int
|
||||||
|
duplicate_filename_groups: int
|
||||||
|
duplicate_files: int
|
||||||
|
suspected_cloud_conflict_groups: int
|
||||||
|
suspected_cloud_conflict_files: int
|
||||||
|
unused_tracks: int
|
||||||
|
suggested_matches: int
|
||||||
|
static_crates: int
|
||||||
|
smart_crates: int
|
||||||
|
unknown_crates: int
|
||||||
|
dynamic_references_excluded: int
|
||||||
|
|
||||||
|
@property
|
||||||
|
def score_basis(self) -> str:
|
||||||
|
return "Resolved non-dynamic references / non-dynamic references scored"
|
||||||
@@ -1,6 +1,7 @@
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from typing import Iterable, Tuple
|
from typing import Iterable, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.crate import Crate
|
||||||
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
||||||
from serato_doctor.models.track import DiskTrack
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
@@ -11,6 +12,7 @@ class Library:
|
|||||||
|
|
||||||
references: Tuple[TrackReference, ...]
|
references: Tuple[TrackReference, ...]
|
||||||
tracks: Tuple[DiskTrack, ...]
|
tracks: Tuple[DiskTrack, ...]
|
||||||
|
crates: Tuple[Crate, ...] = ()
|
||||||
|
|
||||||
@classmethod
|
@classmethod
|
||||||
def build(
|
def build(
|
||||||
@@ -20,6 +22,18 @@ class Library:
|
|||||||
) -> "Library":
|
) -> "Library":
|
||||||
return cls(tuple(references), tuple(tracks))
|
return cls(tuple(references), tuple(tracks))
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def from_crates(
|
||||||
|
cls, crates: Iterable[Crate], tracks: Iterable[DiskTrack]
|
||||||
|
) -> "Library":
|
||||||
|
crate_tuple = tuple(crates)
|
||||||
|
references = tuple(
|
||||||
|
reference
|
||||||
|
for crate in crate_tuple
|
||||||
|
for reference in crate.references
|
||||||
|
)
|
||||||
|
return cls(references, tuple(tracks), crate_tuple)
|
||||||
|
|
||||||
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
|
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
|
||||||
disk_names = {track.filename for track in self.tracks}
|
disk_names = {track.filename for track in self.tracks}
|
||||||
return tuple(
|
return tuple(
|
||||||
|
|||||||
@@ -0,0 +1,39 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class MatchEvidence:
|
||||||
|
"""One explainable scoring decision for a candidate track."""
|
||||||
|
|
||||||
|
field: str
|
||||||
|
matched: bool
|
||||||
|
points: int
|
||||||
|
max_points: int
|
||||||
|
explanation: str
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class TrackMatch:
|
||||||
|
"""A ranked candidate backed by explicit, inspectable evidence."""
|
||||||
|
|
||||||
|
reference: TrackReference
|
||||||
|
track: DiskTrack
|
||||||
|
evidence: Tuple[MatchEvidence, ...]
|
||||||
|
|
||||||
|
@property
|
||||||
|
def score(self) -> int:
|
||||||
|
return sum(item.points for item in self.evidence)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def max_score(self) -> int:
|
||||||
|
return sum(item.max_points for item in self.evidence)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def score_percent(self) -> float:
|
||||||
|
if not self.max_score:
|
||||||
|
return 0.0
|
||||||
|
return round(self.score / self.max_score * 100, 1)
|
||||||
@@ -84,3 +84,48 @@ def test_cli_accepts_old_reference_root(tmp_path, monkeypatch, capsys):
|
|||||||
output = capsys.readouterr().out
|
output = capsys.readouterr().out
|
||||||
assert "Crate references: 1" in output
|
assert "Crate references: 1" in output
|
||||||
assert "Missing by filename: 0" in output
|
assert "Missing by filename: 0" in output
|
||||||
|
|
||||||
|
|
||||||
|
def test_analyze_prints_health_without_writing_reports(tmp_path, monkeypatch, capsys):
|
||||||
|
serato = tmp_path / "serato"
|
||||||
|
subcrates = serato / "Subcrates"
|
||||||
|
music = tmp_path / "music"
|
||||||
|
subcrates.mkdir(parents=True)
|
||||||
|
music.mkdir()
|
||||||
|
crate_text = (
|
||||||
|
"Users/sample-user/Jukebox/Found.mp3otrk"
|
||||||
|
"Users/sample-user/Jukebox/Missing.mp3otrk"
|
||||||
|
)
|
||||||
|
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
|
||||||
|
(music / "Found.mp3").write_bytes(b"synthetic audio")
|
||||||
|
csv_path = tmp_path / "scan.csv"
|
||||||
|
report_path = tmp_path / "report.txt"
|
||||||
|
monkeypatch.setattr(
|
||||||
|
sys,
|
||||||
|
"argv",
|
||||||
|
[
|
||||||
|
"serato-doctor",
|
||||||
|
"analyze",
|
||||||
|
"--serato",
|
||||||
|
str(serato),
|
||||||
|
"--music",
|
||||||
|
str(music),
|
||||||
|
"--out",
|
||||||
|
str(csv_path),
|
||||||
|
"--report",
|
||||||
|
str(report_path),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
main()
|
||||||
|
|
||||||
|
output = capsys.readouterr().out
|
||||||
|
assert "Overall Health: 50.0%" in output
|
||||||
|
assert "Tracks: 1" in output
|
||||||
|
assert "Crate References: 2" in output
|
||||||
|
assert "References Scored: 2" in output
|
||||||
|
assert "Healthy References: 1" in output
|
||||||
|
assert "Broken References: 1" in output
|
||||||
|
assert "Unused Tracks: 0" in output
|
||||||
|
assert not csv_path.exists()
|
||||||
|
assert not report_path.exists()
|
||||||
|
|||||||
@@ -3,12 +3,14 @@ from pathlib import Path
|
|||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
from serato_doctor.crate_parser import (
|
from serato_doctor.crate_parser import (
|
||||||
|
classify_crate,
|
||||||
clean_path,
|
clean_path,
|
||||||
load_crate,
|
load_crate,
|
||||||
parse_crate,
|
parse_crate,
|
||||||
parse_crates,
|
parse_crates,
|
||||||
path_markers,
|
path_markers,
|
||||||
)
|
)
|
||||||
|
from serato_doctor.models.crate import CrateKind
|
||||||
|
|
||||||
|
|
||||||
@pytest.mark.parametrize(
|
@pytest.mark.parametrize(
|
||||||
@@ -84,3 +86,18 @@ def test_parse_crates_reuses_configured_roots_for_every_crate(tmp_path):
|
|||||||
"First.mp3",
|
"First.mp3",
|
||||||
"Second.mp3",
|
"Second.mp3",
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("relative_path", "expected"),
|
||||||
|
[
|
||||||
|
("_Serato_/Subcrates/House.crate", CrateKind.STATIC),
|
||||||
|
("_Serato_/Smartcrates/Warmup.crate", CrateKind.SMART),
|
||||||
|
("_Serato_/Subcrates/Compatible by key.crate", CrateKind.SMART),
|
||||||
|
("fixtures/Unknown.crate", CrateKind.UNKNOWN),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_classify_crate_uses_provenance_and_known_dynamic_name(
|
||||||
|
tmp_path, relative_path, expected
|
||||||
|
):
|
||||||
|
assert classify_crate(tmp_path / relative_path) is expected
|
||||||
|
|||||||
@@ -0,0 +1,57 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.duplicates import find_duplicate_groups
|
||||||
|
from serato_doctor.models.duplicate import DuplicateKind
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def track(path):
|
||||||
|
path = Path(path)
|
||||||
|
return DiskTrack(path, path.name, 100, path.suffix.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def test_exact_names_in_different_folders_form_a_group():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[track("/A/Song.mp3"), track("/B/Song.mp3")]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert len(groups) == 1
|
||||||
|
assert groups[0].kind is DuplicateKind.EXACT_NAME
|
||||||
|
assert groups[0].display_name == "Song.mp3"
|
||||||
|
assert groups[0].extra_files == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_case_only_difference_is_an_exact_name_duplicate():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[track("/A/SONG.MP3"), track("/B/song.mp3")]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert groups[0].kind is DuplicateKind.EXACT_NAME
|
||||||
|
|
||||||
|
|
||||||
|
def test_numeric_suffixes_form_a_suspected_cloud_conflict_group():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[
|
||||||
|
track("/A/Track.mp3"),
|
||||||
|
track("/B/Track 2.mp3"),
|
||||||
|
track("/C/Track 3.mp3"),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert len(groups) == 1
|
||||||
|
assert groups[0].kind is DuplicateKind.CLOUD_CONFLICT
|
||||||
|
assert groups[0].display_name == "Track.mp3"
|
||||||
|
assert groups[0].extra_files == 2
|
||||||
|
assert [item.path for item in groups[0].tracks] == [
|
||||||
|
Path("/A/Track.mp3"),
|
||||||
|
Path("/B/Track 2.mp3"),
|
||||||
|
Path("/C/Track 3.mp3"),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_unrelated_and_single_files_are_not_reported():
|
||||||
|
groups = find_duplicate_groups(
|
||||||
|
[track("/A/First.mp3"), track("/B/Second.mp3")]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert groups == ()
|
||||||
@@ -0,0 +1,106 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.health import analyze_health
|
||||||
|
from serato_doctor.models.crate import Crate, CrateKind
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def reference(filename):
|
||||||
|
return TrackReference(
|
||||||
|
Path("Test.crate"), Path("/old/House") / filename, filename
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def track(filename, folder="House"):
|
||||||
|
path = Path("/new") / folder / filename
|
||||||
|
return DiskTrack(path, filename, 100, path.suffix.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def test_health_report_exposes_each_metric():
|
||||||
|
library = Library.build(
|
||||||
|
[
|
||||||
|
reference("Found.mp3"),
|
||||||
|
reference("Conflict.mp3"),
|
||||||
|
reference("Absent.mp3"),
|
||||||
|
],
|
||||||
|
[
|
||||||
|
track("Found.mp3"),
|
||||||
|
track("Conflict 2.mp3"),
|
||||||
|
track("Conflict 2.mp3", "Backup"),
|
||||||
|
track("Unused.mp3"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
report = analyze_health(library)
|
||||||
|
|
||||||
|
assert report.score == 33.3
|
||||||
|
assert report.score_basis == (
|
||||||
|
"Resolved non-dynamic references / non-dynamic references scored"
|
||||||
|
)
|
||||||
|
assert report.total_references == 3
|
||||||
|
assert report.scored_references == 3
|
||||||
|
assert report.healthy_references == 1
|
||||||
|
assert report.missing_references == 2
|
||||||
|
assert report.unique_missing_filenames == 2
|
||||||
|
assert report.disk_tracks == 4
|
||||||
|
assert report.duplicate_filename_groups == 1
|
||||||
|
assert report.duplicate_files == 1
|
||||||
|
assert report.unused_tracks == 3
|
||||||
|
assert report.suggested_matches == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_empty_library_has_no_health_score():
|
||||||
|
report = analyze_health(Library.build([], []))
|
||||||
|
|
||||||
|
assert report.score is None
|
||||||
|
assert report.total_references == 0
|
||||||
|
assert report.scored_references == 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_smart_crate_references_are_reported_but_not_scored():
|
||||||
|
static_reference = reference("Found.mp3")
|
||||||
|
smart_reference = TrackReference(
|
||||||
|
Path("Smartcrates/Dynamic.crate"),
|
||||||
|
Path("/old/House/Dynamic.mp3"),
|
||||||
|
"Dynamic.mp3",
|
||||||
|
)
|
||||||
|
crates = [
|
||||||
|
Crate(Path("Subcrates/Static.crate"), (static_reference,), CrateKind.STATIC),
|
||||||
|
Crate(
|
||||||
|
Path("Smartcrates/Dynamic.crate"),
|
||||||
|
(smart_reference,),
|
||||||
|
CrateKind.SMART,
|
||||||
|
),
|
||||||
|
]
|
||||||
|
library = Library.from_crates(crates, [track("Found.mp3")])
|
||||||
|
|
||||||
|
report = analyze_health(library)
|
||||||
|
|
||||||
|
assert report.score == 100.0
|
||||||
|
assert report.total_references == 2
|
||||||
|
assert report.scored_references == 1
|
||||||
|
assert report.missing_references == 0
|
||||||
|
assert report.static_crates == 1
|
||||||
|
assert report.smart_crates == 1
|
||||||
|
assert report.dynamic_references_excluded == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_health_reports_cloud_conflicts_separately_from_exact_duplicates():
|
||||||
|
library = Library.build(
|
||||||
|
[],
|
||||||
|
[
|
||||||
|
track("Track.mp3", "Original"),
|
||||||
|
track("Track 2.mp3", "Conflict"),
|
||||||
|
track("Copy.mp3", "First"),
|
||||||
|
track("Copy.mp3", "Second"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
report = analyze_health(library)
|
||||||
|
|
||||||
|
assert report.duplicate_filename_groups == 1
|
||||||
|
assert report.duplicate_files == 1
|
||||||
|
assert report.suspected_cloud_conflict_groups == 1
|
||||||
|
assert report.suspected_cloud_conflict_files == 1
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.matching import MatchingEngine, score_candidate
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def reference(filename, folder="House"):
|
||||||
|
return TrackReference(
|
||||||
|
Path("Test.crate"), Path("/old") / folder / filename, filename
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def track(filename, folder="House"):
|
||||||
|
path = Path("/new") / folder / filename
|
||||||
|
return DiskTrack(path, filename, 100, path.suffix.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def test_exact_candidate_has_full_evidence_score():
|
||||||
|
match = score_candidate(reference("Track.mp3"), track("Track.mp3"))
|
||||||
|
|
||||||
|
assert match.score == 90
|
||||||
|
assert match.max_score == 90
|
||||||
|
assert match.score_percent == 100.0
|
||||||
|
assert all(item.matched for item in match.evidence)
|
||||||
|
|
||||||
|
|
||||||
|
def test_cloud_conflict_suffix_is_explained():
|
||||||
|
match = score_candidate(reference("Track.mp3"), track("Track 2.mp3"))
|
||||||
|
|
||||||
|
assert match.score == 80
|
||||||
|
assert match.score_percent == 88.9
|
||||||
|
assert match.evidence[0].explanation == (
|
||||||
|
"Filename matches after removing a numeric conflict suffix"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_case_normalized_match_scores_below_exact():
|
||||||
|
match = score_candidate(reference("TRACK.MP3"), track("track.mp3"))
|
||||||
|
|
||||||
|
assert match.score == 85
|
||||||
|
assert match.evidence[0].points == 55
|
||||||
|
|
||||||
|
|
||||||
|
def test_engine_omits_unrelated_filenames():
|
||||||
|
engine = MatchingEngine([track("Different.mp3")])
|
||||||
|
|
||||||
|
assert engine.candidates_for(reference("Missing.mp3")) == ()
|
||||||
|
|
||||||
|
|
||||||
|
def test_ambiguous_candidates_are_ranked_deterministically():
|
||||||
|
engine = MatchingEngine(
|
||||||
|
[track("Track 2.mp3", "Other"), track("Track 3.mp3", "House")]
|
||||||
|
)
|
||||||
|
|
||||||
|
matches = engine.candidates_for(reference("Track.mp3"))
|
||||||
|
|
||||||
|
assert [match.track.filename for match in matches] == [
|
||||||
|
"Track 3.mp3",
|
||||||
|
"Track 2.mp3",
|
||||||
|
]
|
||||||
|
assert [match.score for match in matches] == [80, 60]
|
||||||
@@ -1,7 +1,8 @@
|
|||||||
import runpy
|
import runpy
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from serato_doctor.crate_parser import parse_crates
|
from serato_doctor.crate_parser import load_library_crates
|
||||||
|
from serato_doctor.models.crate import CrateKind
|
||||||
from serato_doctor.models.library import Library
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.scanner import scan_audio
|
from serato_doctor.scanner import scan_audio
|
||||||
|
|
||||||
@@ -13,9 +14,10 @@ def test_generated_sample_library_has_expected_scenario(tmp_path):
|
|||||||
build_sample = runpy.run_path(str(generator_path))["build_sample"]
|
build_sample = runpy.run_path(str(generator_path))["build_sample"]
|
||||||
sample_root = build_sample(tmp_path / "sample")
|
sample_root = build_sample(tmp_path / "sample")
|
||||||
|
|
||||||
references = parse_crates(sample_root / "Serato" / "_Serato_" / "Subcrates")
|
crates = load_library_crates(sample_root / "Serato" / "_Serato_")
|
||||||
tracks = scan_audio(sample_root / "Music")
|
tracks = scan_audio(sample_root / "Music")
|
||||||
results = Library.build(references, tracks).reconcile_by_filename()
|
library = Library.from_crates(crates, tracks)
|
||||||
|
results = library.reconcile_by_filename()
|
||||||
missing = {
|
missing = {
|
||||||
result.reference.filename
|
result.reference.filename
|
||||||
for result in results
|
for result in results
|
||||||
@@ -23,6 +25,8 @@ def test_generated_sample_library_has_expected_scenario(tmp_path):
|
|||||||
}
|
}
|
||||||
|
|
||||||
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7
|
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7
|
||||||
assert len(references) == 13
|
assert len(library.references) == 13
|
||||||
assert len(tracks) == 10
|
assert len(tracks) == 10
|
||||||
assert missing == {"Missing.mp3", "Old Name.mp3"}
|
assert missing == {"Missing.mp3", "Old Name.mp3"}
|
||||||
|
assert sum(crate.kind is CrateKind.STATIC for crate in crates) == 5
|
||||||
|
assert sum(crate.kind is CrateKind.SMART for crate in crates) == 2
|
||||||
|
|||||||
Reference in New Issue
Block a user