Compare commits

...

24 Commits

Author SHA1 Message Date
Philip Guzman c5449a278a Merge feature/duplicate-filenames into develop 2026-07-01 07:30:50 -07:00
Philip Guzman 6449c7de6a Add duplicate filename diagnostics 2026-06-30 18:52:16 -07:00
Philip Guzman d764f910e4 Merge feature/crate-classification into develop 2026-06-30 18:50:42 -07:00
Philip Guzman 35713bbbb3 Classify static and smart crates 2026-06-30 18:20:05 -07:00
Philip Guzman c92d929f37 Merge feature/analyze-command into develop 2026-06-30 18:17:20 -07:00
Philip Guzman e6f313d98a Add read-only analyze command 2026-06-30 17:55:45 -07:00
Philip Guzman 22424d0b5c Merge feature/health-engine into develop 2026-06-30 17:54:20 -07:00
Philip Guzman b8bf9ec6e1 Add transparent library health engine 2026-06-30 17:47:54 -07:00
Philip Guzman 167f029c28 Merge feature/matching-engine into develop 2026-06-30 17:46:56 -07:00
Philip Guzman 5a2edb8b94 Add explainable matching engine 2026-06-30 17:42:33 -07:00
Philip Guzman 29e53f7dc6 Merge feature/logging into develop 2026-06-30 17:40:57 -07:00
Philip Guzman 4fde93613f Add opt-in diagnostic logging 2026-06-30 17:36:58 -07:00
Philip Guzman e4d5a32de2 Merge feature/configuration into develop 2026-06-30 17:35:59 -07:00
Philip Guzman ec671079ba Add configurable library reference roots 2026-06-30 17:20:08 -07:00
Philip Guzman 4e85b5912b Merge feature/sample-library into develop 2026-06-30 17:18:06 -07:00
Philip Guzman d4fe61cf8c Merge feature/test-foundation into develop 2026-06-30 17:18:06 -07:00
Philip Guzman c4d08bfe27 Merge feature/library-model into develop 2026-06-30 17:18:06 -07:00
Philip Guzman bb74c14742 Add synthetic sample library 2026-06-30 17:15:56 -07:00
Philip Guzman 71fa64a960 Establish automated test foundation 2026-06-30 10:31:34 -07:00
Philip Guzman 09167e44c7 Introduce core library model 2026-06-30 10:10:14 -07:00
Philip Guzman 508eecd0bd Add project roadmap and architecture docs 2026-06-30 09:53:09 -07:00
Philip Guzman 577fe1f7a7 Merge improved parser into missing report branch 2026-06-30 09:18:22 -07:00
Philip Guzman c3bf111107 Add grouped missing reference report 2026-06-30 09:16:27 -07:00
Philip Guzman 826428f380 Improve Serato crate parser for UTF-16 LE path records 2026-06-30 08:43:11 -07:00
46 changed files with 1877 additions and 67 deletions
+5
View File
@@ -2,7 +2,9 @@ __pycache__/
*.pyc *.pyc
.env .env
.venv/ .venv/
*.egg-info/
reports/ reports/
samples/small-library/generated/
*.sqlite *.sqlite
*.db *.db
*.csv *.csv
@@ -11,3 +13,6 @@ reports/
# Never commit personal Serato data # Never commit personal Serato data
database V2 database V2
*.crate *.crate
# macOS
.DS_Store
+19
View File
@@ -0,0 +1,19 @@
# Contributing
## Branching
- `main` is stable.
- `develop` is the integration branch.
- Feature branches use: `feature/<name>`.
## Safety Rules
Never commit personal music library files, Serato databases, or crates.
Never write repair code without dry-run mode, backup plan, and rollback log.
## Testing
Install development dependencies with `python3 -m pip install -e '.[dev]'`.
Run `pytest` before committing. Tests and fixtures must use synthetic library data.
+37
View File
@@ -0,0 +1,37 @@
# Serato Doctor
Serato Doctor is a read-only-first toolkit for inspecting and eventually repairing
DJ libraries. Serato is the first supported engine.
## Analyze a Library
```shell
serato-doctor analyze \
--serato ~/Music/_Serato_ \
--music ~/Music/Jukebox
```
Analysis prints a transparent reference-integrity score plus broken references,
duplicate filenames, unused tracks, and suggested filename matches. It does not
modify the library or create report files.
If crates contain an older library root, supply it explicitly:
```shell
serato-doctor analyze --reference-root /Users/old-user/OneDrive/Jukebox
```
## Generate Detailed Reports
Running without a command preserves the original scanner behavior and writes CSV
and text reports:
```shell
serato-doctor --serato ~/Music/_Serato_ --music ~/Music/Jukebox
```
Use `--verbose` for progress on standard error or `--log-file PATH` for an
aggregate diagnostic log.
Serato Doctor never repairs files without an explicit future repair workflow,
preview, backup, and rollback path.
+52
View File
@@ -0,0 +1,52 @@
# Serato Doctor Roadmap
## v0.1 — Library Inspector
- [x] Project repository
- [x] Filesystem scanner
- [x] Serato crate parser
- [x] Missing reference CSV report
- [x] Grouped missing reference report
- [ ] HTML health dashboard
- [x] Test suite
- [x] Sample library fixtures
- [ ] Database V2 read-only parser
- [x] Configuration
- [x] Logging
- [x] Matching engine
## v0.2 — Diagnostics
- [x] `serato-doctor analyze` command
- [x] Duplicate filename detection
- [ ] Duplicate audio hash detection
- [ ] Broken symlink detection
- [ ] Orphaned audio detection
- [ ] OneDrive rename detection
- [x] Crate classification: static vs smart/dynamic
- [x] Library health score
## v0.3 — Safe Repair
- [ ] Dry-run repair plan
- [ ] Backup before repair
- [ ] Compatibility symlink creation
- [ ] Compatibility copy creation
- [ ] Rename repair
- [ ] Rollback log
## v0.4 — Migration Wizard
- [ ] Move library root
- [ ] Cloud provider migration
- [ ] External drive migration
- [ ] Verify moved library
- [ ] Update application references
## v1.0 — DJ Library Doctor
- [ ] Desktop UI
- [ ] Serato support
- [ ] Rekordbox support
- [ ] VirtualDJ support
- [ ] Engine DJ support
+11
View File
@@ -0,0 +1,11 @@
# Architecture
Serato Doctor is designed as a DJ library inspection, repair, and migration platform.
## Design Principles
1. Read-only by default.
2. Every repair must support preview/dry-run.
3. Every repair must create a backup or rollback path.
4. Application-specific logic lives in engines.
5. Core matching and scanning logic should be application-agnostic.
@@ -0,0 +1,26 @@
# Case Study: OneDrive Mac Migration
## Scenario
A large Serato DJ library was migrated from an older Mac to a newer Mac using OneDrive.
## Symptoms
- OneDrive client stuck syncing
- Duplicate OneDrive folders
- Thousands of files renamed with trailing ` 2`
- Serato reported many tracks as missing
- Some files existed on disk but still appeared orange in Serato
## Findings
- OneDrive sync state was rebuilt successfully
- Thousands of orphaned filename conflicts were repaired
- Some Serato references were stale database objects, not missing files
- Smart/dynamic crates should be classified separately from static user crates
## Lessons
- Filesystem health and Serato database health are separate problems
- Smart crates should not be treated the same as static crates
- Repair tools must be read-only by default and generate a plan before changing anything
+27
View File
@@ -0,0 +1,27 @@
# Analyze Command
## Problem
The health and matching engines are only Python APIs. Users need one safe command
that summarizes library integrity without first interpreting CSV files.
## Architecture
`serato-doctor analyze` reuses the existing configured crate parser and filesystem
scanner, then prints the immutable health report. The original no-command mode is
retained as the report-producing scan workflow.
Analyze mode does not create CSV or text reports. Both modes remain read-only with
respect to Serato crates, databases, and music files.
## Edge Cases
- A library with no references displays `Not assessed` rather than a false score.
- Historical roots continue to work through repeatable `--reference-root` flags.
- Normal output remains separate from optional diagnostic logging.
- Existing scripts that invoke the CLI without a subcommand remain compatible.
## Verification
Tests prove the legacy five-line output and reports remain unchanged, while analyze
prints health metrics and does not create output files.
+25
View File
@@ -0,0 +1,25 @@
# Scan Configuration
## Problem
The prototype recognized crate paths by matching one developer's historical
OneDrive path. That made otherwise valid libraries invisible to the parser.
## Architecture
`ScanConfig` owns the paths for a read-only scan. The CLI accepts repeatable
`--reference-root` options for roots embedded in crates before a migration. The
parser uses those roots as record boundaries. Without explicit roots, it recognizes
generic macOS home and mounted-volume paths without embedding a username.
## Edge Cases
- A library may have references from more than one historical root.
- Root paths may contain spaces or have leading/trailing separators.
- Existing `/Users/...` and `/Volumes/...` crates must work without new flags.
- A record without a recognized terminator remains ignored.
## Verification
Tests cover default root discovery, a custom migrated root, the CLI option, and
the existing synthetic sample library. The scanner remains read-only.
+33
View File
@@ -0,0 +1,33 @@
# Crate Classification
## Problem
Static crates are manually maintained track lists. Smart crates are dynamic views
generated from rules, so stale-looking entries in them should not be presented as
broken manual references or given the same health-score weight.
## Architecture
Crates carry a `static`, `smart`, or `unknown` kind. Folder provenance is the
primary signal: Serato stores regular definitions in `Subcrates` and smart
definitions in `Smartcrates`. `Compatible by key.crate` is also treated as smart
when encountered in `Subcrates`, based on the original migration case study.
The library loader reads both folders. Health analysis reports all references but
scores only non-smart references. Unknown crates remain scoreable so incomplete
classification cannot silently hide potential problems.
Serato documents the folder distinction in [What is in the _Serato_ folder?](https://support.serato.com/hc/en-us/articles/204022904-What-is-in-the-Serato-folder)
and explains that smart crates are populated from rules in [Crates in Serato DJ](https://support.serato.com/hc/en-us/articles/227561407-Crates-in-Serato-DJ-Pro-Serato-DJ-Lite).
## Edge Cases
- The known `Compatible by key` dynamic crate in the `Subcrates` folder.
- Crate fixtures outside a recognized Serato folder.
- Libraries containing both static and smart references to the same track.
- Dynamic references whose current materialized paths appear missing.
## Verification
Tests cover all three kinds, both Serato folders, the known dynamic fallback, the
synthetic five-static/two-smart library, and exclusion from health scoring.
+32
View File
@@ -0,0 +1,32 @@
# Duplicate Filename Detection
## Problem
Two files with the same filename may be ordinary copies, while names such as
`Track.mp3` and `Track 2.mp3` may indicate a OneDrive conflict. These cases need
review, but neither is sufficient evidence for deletion.
## Architecture
The duplicate detector emits immutable groups of two kinds:
- `exact_name` groups filenames after case and Unicode normalization.
- `cloud_conflict` groups names after additionally removing a trailing numeric
suffix from the stem.
Groups contain every path, a stable comparison key, a display name, and the number
of files beyond the first. Health analysis reports exact and suspected-conflict
counts separately. Neither category changes the health score.
## Edge Cases
- Same filename in different folders.
- Case-only and Unicode representation differences.
- Multiple conflict suffixes such as `Track 2.mp3` and `Track 3.mp3`.
- Legitimate numbered titles, which remain explicitly labeled as suspected.
- A single file, which is never reported as a duplicate.
## Verification
Tests cover exact, normalized, conflict-family, unrelated, and deterministic-order
behavior, plus integration with health analysis. Detection is read-only.
+32
View File
@@ -0,0 +1,32 @@
# Library Health Engine
## Problem
Raw missing-reference counts do not provide a compact view of library integrity,
but an opaque blended score would imply confidence the current data cannot support.
## Architecture
The health engine produces an immutable report from the core `Library`. Its score
is only the percentage of non-dynamic crate references resolved by exact filename.
Smart-crate references are counted but excluded because their contents are derived
from rules. The report also exposes missing references, unique missing filenames,
duplicate filename groups, extra duplicate files, unused tracks, and missing
references with matching candidates.
Duplicate, unused, and candidate counts are informational. They do not affect the
score until the project has a documented and validated weighting policy. An empty
library has no score rather than a misleading 0% or 100%.
## Edge Cases
- The same missing filename referenced by several crates.
- Several disk files sharing a filename.
- Conflict-suffixed files that are unused but may be match candidates.
- Libraries with no crate references.
- Tracks referenced by filename from more than one crate.
## Verification
Tests assert every metric, the disclosed score basis, candidate integration, and
empty-library behavior. Analysis remains entirely read-only.
+25
View File
@@ -0,0 +1,25 @@
# Application Logging
## Problem
The CLI reports final counts but provides no diagnostic trail when a scan behaves
unexpectedly. Troubleshooting should not require adding print statements or expose
library contents by default.
## Architecture
The project uses an isolated standard-library logger. It has no visible output by
default. `--verbose` writes progress to standard error, while `--log-file PATH`
writes an informational audit trail. Normal result lines remain on standard output.
## Edge Cases
- Reconfiguring logging in the same process must not duplicate handlers.
- Console and file logging may be enabled together.
- Log messages contain aggregate counts, not track names or crate contents.
- A default scan must remain quiet except for its established result output.
## Verification
Tests verify quiet defaults, file output, handler replacement, and unchanged CLI
result lines. The full sample-library scan remains read-only.
+29
View File
@@ -0,0 +1,29 @@
# Explainable Matching Engine
## Problem
Missing references need ranked candidate files, but a filename-only yes/no check
cannot explain ambiguity or cloud-provider conflict names.
## Architecture
The read-only matching engine indexes normalized filenames and scores only related
candidates. Every score contains evidence for filename, extension, and parent
folder. Exact filenames earn 60 points, normalized names 55, numeric conflict-name
matches 50, extensions 10, and parent folders 20.
The displayed percentage is an evidence score, not a statistical probability.
Metadata, duration, hashes, and fingerprints can add stronger evidence later.
## Edge Cases
- Unicode and case differences.
- OneDrive-style names such as `Track 2.mp3`.
- Duplicate candidates in different folders.
- Legitimate numbered song titles, which remain candidates but are never repaired.
- Unrelated names, which are not emitted as candidates.
## Verification
Tests cover exact, normalized, conflict-suffix, ambiguous, and unrelated filenames.
Candidate ordering is deterministic. The engine never changes a track or reference.
+26
View File
@@ -0,0 +1,26 @@
# Synthetic Sample Library
## Problem
Real Serato libraries contain private paths, listening history, and copyrighted
music. Contributors still need a repeatable library for tests and demonstrations.
## Architecture
`samples/small-library/manifest.json` describes a compact migration scenario.
`generate.py` turns that manifest into fake audio files and UTF-16-LE crate files
inside an ignored `generated` directory. The generated files are disposable and
contain no real audio or user library data.
## Edge Cases
- Five manually maintained crates and two smart/dynamic crates.
- Mixed supported audio extensions.
- One reference whose file is absent.
- One stale reference whose file has been renamed.
- References shared by static and smart crates.
## Verification
`pytest` generates the sample in a temporary directory and verifies its crate,
track, and missing-reference counts through the production parser and scanner.
+24
View File
@@ -0,0 +1,24 @@
# Test Foundation
## Problem
The crate parser handles unusual binary text and known extension artifacts. Without
automated tests, a small refactor could silently change library counts or reports.
## Architecture
Tests cover the read-only pipeline from crate and filesystem inputs through the
core library model to CSV, text, and CLI output. All test data is synthetic and is
created in temporary directories.
## Edge Cases
- Known MP3, M4A, WAV, and AIF decode artifacts.
- Crate records without a recognized stop marker.
- Supported extensions with mixed case and unsupported files.
- References that exist by filename and references that remain missing.
## Verification
Run `pytest`. No test reads a real Serato library or writes outside pytest's
temporary directory.
+22
View File
@@ -0,0 +1,22 @@
[build-system]
requires = ["setuptools>=61"]
build-backend = "setuptools.build_meta"
[project]
name = "serato-doctor"
version = "0.1.0"
description = "Inspect, diagnose, repair, and migrate DJ libraries."
requires-python = ">=3.9"
dependencies = []
[project.optional-dependencies]
dev = ["pytest>=8,<9"]
[project.scripts]
serato-doctor = "serato_doctor.cli:main"
[tool.pytest.ini_options]
testpaths = ["tests"]
[tool.setuptools.packages.find]
include = ["serato_doctor*"]
+25
View File
@@ -0,0 +1,25 @@
# Small Synthetic Library
This fixture models a small Serato migration without containing music or personal
library data. It has 10 fake audio files, five static crates, two smart crates,
one absent song, and one renamed song.
Generate it with:
```shell
python3 samples/small-library/generate.py
```
Analyze it with:
```shell
python3 -m serato_doctor.cli \
--serato samples/small-library/generated/Serato/_Serato_ \
--music samples/small-library/generated/Music \
--out samples/small-library/generated/scan.csv \
--report samples/small-library/generated/report.txt
```
Expected CLI counts are 13 crate references, 10 disk tracks, and 2 references
missing by filename. Everything under `generated/` is disposable and ignored by
Git.
+56
View File
@@ -0,0 +1,56 @@
"""Generate a disposable, synthetic Serato library from the sample manifest."""
import argparse
import json
from pathlib import Path
from typing import Optional
SAMPLE_ROOT = Path(__file__).parent
SERATO_PATH_PREFIX = "Users/sample-user/OneDrive/Jukebox/"
def load_manifest() -> dict:
return json.loads((SAMPLE_ROOT / "manifest.json").read_text(encoding="utf-8"))
def build_sample(output: Optional[Path] = None) -> Path:
root = output or SAMPLE_ROOT / "generated"
manifest = load_manifest()
music_root = root / "Music"
serato_root = root / "Serato" / "_Serato_"
for relative_path in manifest["tracks"]:
track_path = music_root / relative_path
track_path.parent.mkdir(parents=True, exist_ok=True)
track_path.write_bytes(
f"Synthetic Serato Doctor fixture: {relative_path}\n".encode("utf-8")
)
for crate in manifest["crates"]:
folder_name = "Smartcrates" if crate["type"] == "smart" else "Subcrates"
crate_root = serato_root / folder_name
crate_root.mkdir(parents=True, exist_ok=True)
other_folder = "Subcrates" if folder_name == "Smartcrates" else "Smartcrates"
stale_path = serato_root / other_folder / crate["name"]
if stale_path.exists():
stale_path.unlink()
records = "".join(
f"{SERATO_PATH_PREFIX}{relative_path}otrk"
for relative_path in crate["references"]
)
(crate_root / crate["name"]).write_bytes(records.encode("utf-16-le"))
return root
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--output", type=Path)
args = parser.parse_args()
root = build_sample(args.output)
print(f"Generated synthetic library: {root}")
if __name__ == "__main__":
main()
+56
View File
@@ -0,0 +1,56 @@
{
"tracks": [
"House/First.mp3",
"House/Second.m4a",
"Open Format/Third.wav",
"Open Format/Fourth.aif",
"Classics/Fifth.mp3",
"Classics/Sixth.flac",
"Warmup/Seventh.mp3",
"Warmup/Eighth.mp3",
"Renamed/New Name.mp3",
"Bonus/Ninth.MP3"
],
"crates": [
{
"name": "House.crate",
"type": "static",
"references": ["House/First.mp3", "House/Second.m4a", "House/Missing.mp3"]
},
{
"name": "Open Format.crate",
"type": "static",
"references": ["Open Format/Third.wav", "Open Format/Fourth.aif"]
},
{
"name": "Classics.crate",
"type": "static",
"references": ["Classics/Fifth.mp3", "Classics/Sixth.flac"]
},
{
"name": "Renamed Tracks.crate",
"type": "static",
"references": ["Renamed/Old Name.mp3"]
},
{
"name": "Bonus.crate",
"type": "static",
"references": ["Bonus/Ninth.MP3"]
},
{
"name": "Compatible by key.crate",
"type": "smart",
"references": ["House/First.mp3", "House/Second.m4a"]
},
{
"name": "Smart Warmup.crate",
"type": "smart",
"references": ["Warmup/Seventh.mp3", "Warmup/Eighth.mp3"]
}
],
"expected": {
"disk_tracks": 10,
"crate_references": 13,
"missing_by_filename": ["Missing.mp3", "Old Name.mp3"]
}
}
+93 -29
View File
@@ -1,45 +1,109 @@
from pathlib import Path from pathlib import Path
import argparse import argparse
import csv
from serato_doctor.crate_parser import parse_crates from serato_doctor.config import ScanConfig
from serato_doctor.crate_parser import load_library_crates
from serato_doctor.health import analyze_health
from serato_doctor.logging import configure_logging
from serato_doctor.models.library import Library
from serato_doctor.scanner import scan_audio from serato_doctor.scanner import scan_audio
from serato_doctor.report import write_csv, write_missing_report
def main(): def main():
parser = argparse.ArgumentParser(prog="serato-doctor") parser = argparse.ArgumentParser(prog="serato-doctor")
parser.add_argument(
"command", nargs="?", choices=("scan", "analyze"), default="scan"
)
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_")) parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox")) parser.add_argument(
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")) "--music",
default=str(
Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"
),
)
parser.add_argument(
"--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv")
)
parser.add_argument(
"--report",
default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"),
)
parser.add_argument(
"--reference-root",
action="append",
default=[],
help="Old library root stored in crates; may be supplied more than once",
)
parser.add_argument(
"--verbose", action="store_true", help="Write diagnostic progress to stderr"
)
parser.add_argument("--log-file", help="Write scan progress to a log file")
args = parser.parse_args() args = parser.parse_args()
serato = Path(args.serato) config = ScanConfig.build(
music = Path(args.music) serato=Path(args.serato),
out = Path(args.out) music=Path(args.music),
out=Path(args.out),
report=Path(args.report),
reference_roots=(Path(root) for root in args.reference_root),
verbose=args.verbose,
log_file=Path(args.log_file) if args.log_file else None,
)
logger = configure_logging(config.verbose, config.log_file)
logger.info("Starting read-only library inspection")
logger.debug("Serato directory: %s", config.serato)
logger.debug("Music directory: %s", config.music)
refs = parse_crates(serato / "Subcrates") crates = load_library_crates(config.serato, config.reference_roots)
disk = scan_audio(music) library = Library.from_crates(
crates=crates,
tracks=scan_audio(config.music),
)
results = library.reconcile_by_filename()
missing_count = sum(1 for result in results if not result.exists_by_filename)
logger.info("Parsed %d crate references", len(library.references))
logger.info("Scanned %d disk tracks", len(library.tracks))
logger.info("Found %d references missing by filename", missing_count)
disk_names = {t.filename for t in disk} if args.command == "analyze":
health = analyze_health(library)
rows = [] score = f"{health.score:.1f}%" if health.score is not None else "Not assessed"
for ref in refs: logger.info("Calculated library health: %s", score)
rows.append({ print(f"Overall Health: {score}")
"crate": str(ref.source), print(f"Score Basis: {health.score_basis}")
"serato_path": str(ref.path), print(f"Tracks: {health.disk_tracks}")
"filename": ref.filename, print(f"Crate References: {health.total_references}")
"exists_by_filename": ref.filename in disk_names, print(f"References Scored: {health.scored_references}")
}) print(f"Healthy References: {health.healthy_references}")
print(f"Broken References: {health.missing_references}")
with out.open("w", newline="", encoding="utf-8") as f: print(f"Unique Missing Filenames: {health.unique_missing_filenames}")
writer = csv.DictWriter(f, fieldnames=["crate", "serato_path", "filename", "exists_by_filename"]) print(f"Duplicate Filename Groups: {health.duplicate_filename_groups}")
writer.writeheader() print(f"Duplicate Files: {health.duplicate_files}")
writer.writerows(rows) print(
"Suspected Cloud Conflict Groups: "
print(f"Crate references: {len(refs)}") f"{health.suspected_cloud_conflict_groups}"
print(f"Disk tracks: {len(disk)}") )
print(f"Missing by filename: {sum(1 for r in rows if not r['exists_by_filename'])}") print(
print(f"Wrote: {out}") "Suspected Cloud Conflict Files: "
f"{health.suspected_cloud_conflict_files}"
)
print(f"Unused Tracks: {health.unused_tracks}")
print(f"Suggested Matches: {health.suggested_matches}")
print(f"Static Crates: {health.static_crates}")
print(f"Smart Crates: {health.smart_crates}")
print(
f"Dynamic References Excluded: {health.dynamic_references_excluded}"
)
else:
write_csv(results, config.out)
write_missing_report(results, config.report)
logger.info("Wrote CSV and missing-reference reports")
print(f"Crate references: {len(library.references)}")
print(f"Disk tracks: {len(library.tracks)}")
print(f"Missing by filename: {missing_count}")
print(f"CSV: {config.out}")
print(f"Report: {config.report}")
if __name__ == "__main__": if __name__ == "__main__":
+37
View File
@@ -0,0 +1,37 @@
from dataclasses import dataclass
from pathlib import Path
from typing import Iterable, Optional, Tuple
@dataclass(frozen=True)
class ScanConfig:
"""Read-only paths used for one library scan."""
serato: Path
music: Path
out: Path
report: Path
reference_roots: Tuple[Path, ...] = ()
verbose: bool = False
log_file: Optional[Path] = None
@classmethod
def build(
cls,
serato: Path,
music: Path,
out: Path,
report: Path,
reference_roots: Iterable[Path] = (),
verbose: bool = False,
log_file: Optional[Path] = None,
) -> "ScanConfig":
return cls(
serato,
music,
out,
report,
tuple(reference_roots),
verbose,
log_file,
)
+94 -30
View File
@@ -1,47 +1,111 @@
from pathlib import Path from pathlib import Path
import re from typing import Iterable, Tuple
from serato_doctor.models import TrackReference from serato_doctor.models.crate import Crate, CrateKind
from serato_doctor.models.reference import TrackReference
AUDIO_EXTS = "mp3|m4a|wav|aif|aiff|flac|MP3|M4A|WAV|AIF|AIFF|FLAC"
def read_crate_text(crate_path: Path) -> str: def read_crate_text(crate_path: Path) -> str:
raw = crate_path.read_bytes() return crate_path.read_bytes().decode("utf-16-le", errors="ignore").replace("\x00", "")
for enc in ("utf-16-be", "utf-16-le", "utf-8", "latin1"):
text = raw.decode(enc, errors="ignore")
if "Users" in text or "Jukebox" in text:
return text
return raw.decode("latin1", errors="ignore")
def parse_crate(crate_path: Path) -> list[TrackReference]: def clean_path(raw: str) -> str:
# Common Serato decode artifacts where final extension char gets merged.
raw = raw.replace(".mp漳", ".mp3")
raw = raw.replace(".MP漳", ".MP3")
raw = raw.replace(".m4愠", ".m4a")
raw = raw.replace(".M4愠", ".M4A")
raw = raw.replace(".wa瘠", ".wav")
raw = raw.replace(".WA瘠", ".WAV")
raw = raw.replace(".ai映", ".aif")
raw = raw.replace(".AI映", ".AIF")
return raw.strip()
DEFAULT_PATH_MARKERS = ("Users/", "Volumes/")
def path_markers(reference_roots: Iterable[Path]) -> Tuple[str, ...]:
configured = tuple(
root.expanduser().as_posix().strip("/") + "/" for root in reference_roots
)
return configured or DEFAULT_PATH_MARKERS
def classify_crate(crate_path: Path) -> CrateKind:
parent_names = {parent.name.casefold() for parent in crate_path.parents}
if "smartcrates" in parent_names:
return CrateKind.SMART
if crate_path.name.casefold() == "compatible by key.crate":
return CrateKind.SMART
if "subcrates" in parent_names:
return CrateKind.STATIC
return CrateKind.UNKNOWN
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
text = read_crate_text(crate_path) text = read_crate_text(crate_path)
# Serato crate files often decode with weird spacing/null-ish characters.
# This finds paths from /Users/... through the audio extension without
# greedily scanning the entire file.
pattern = rf"/?Users/[^\r\n]+?\.(?:{AUDIO_EXTS})"
refs = [] refs = []
for match in re.finditer(pattern, text):
raw_path = "/" + match.group(0).lstrip("/")
raw_path = raw_path.replace("\x00", "")
path = Path(raw_path)
refs.append( # Serato record markers seen after paths in UTF-16-LE decoded crate data.
TrackReference( stop_markers = ["牴k", "otrk", "ptrk", "tvcn", "ovct"]
source=crate_path,
path=path, for marker in path_markers(reference_roots):
filename=path.name, for part in text.split(marker)[1:]:
candidate = marker + part
stops = [
candidate.find(stop)
for stop in stop_markers
if candidate.find(stop) != -1
]
if not stops:
continue
raw_path = "/" + candidate[: min(stops)]
raw_path = clean_path(raw_path)
path = Path(raw_path)
refs.append(
TrackReference(
source=crate_path,
path=path,
filename=path.name,
)
) )
)
return refs return Crate(
path=crate_path,
references=tuple(refs),
kind=classify_crate(crate_path),
)
def parse_crates(root: Path) -> list[TrackReference]: def parse_crate(
crate_path: Path, reference_roots: Iterable[Path] = ()
) -> list[TrackReference]:
"""Parse references from one crate, preserving the prototype API."""
return list(load_crate(crate_path, reference_roots).references)
def parse_crates(
root: Path, reference_roots: Iterable[Path] = ()
) -> list[TrackReference]:
refs = [] refs = []
reference_roots = tuple(reference_roots)
for crate in root.rglob("*.crate"): for crate in root.rglob("*.crate"):
refs.extend(parse_crate(crate)) refs.extend(parse_crate(crate, reference_roots))
return refs return refs
def load_library_crates(
serato_root: Path, reference_roots: Iterable[Path] = ()
) -> Tuple[Crate, ...]:
reference_roots = tuple(reference_roots)
crates = []
for folder_name in ("Subcrates", "Smartcrates"):
folder = serato_root / folder_name
for crate_path in folder.rglob("*.crate"):
crates.append(load_crate(crate_path, reference_roots))
return tuple(sorted(crates, key=lambda crate: str(crate.path)))
+41
View File
@@ -0,0 +1,41 @@
from collections import defaultdict
from typing import DefaultDict, Iterable, List, Tuple
from serato_doctor.matching import cloud_conflict_name, normalize
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
from serato_doctor.models.track import DiskTrack
def find_duplicate_groups(
tracks: Iterable[DiskTrack],
) -> Tuple[DuplicateGroup, ...]:
"""Find exact-name duplicates and suspected numeric conflict copies."""
track_tuple = tuple(tracks)
by_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
for track in track_tuple:
by_name[normalize(track.filename)].append(track)
by_conflict_name[cloud_conflict_name(track.filename)].append(track)
groups = []
for key, matches in by_name.items():
if len(matches) > 1:
groups.append(_group(DuplicateKind.EXACT_NAME, key, matches))
for key, matches in by_conflict_name.items():
distinct_names = {normalize(track.filename) for track in matches}
if len(distinct_names) > 1:
groups.append(_group(DuplicateKind.CLOUD_CONFLICT, key, matches))
return tuple(
sorted(groups, key=lambda group: (group.kind.value, group.comparison_key))
)
def _group(
kind: DuplicateKind, key: str, tracks: Iterable[DiskTrack]
) -> DuplicateGroup:
ordered = tuple(sorted(tracks, key=lambda track: str(track.path)))
return DuplicateGroup(kind, key, ordered)
+79
View File
@@ -0,0 +1,79 @@
from serato_doctor.duplicates import find_duplicate_groups
from serato_doctor.matching import MatchingEngine
from serato_doctor.models.crate import CrateKind
from serato_doctor.models.duplicate import DuplicateKind
from serato_doctor.models.health import HealthReport
from serato_doctor.models.library import Library
def analyze_health(library: Library) -> HealthReport:
"""Calculate defensible health metrics without changing the library."""
dynamic_sources = {
crate.path for crate in library.crates if crate.kind is CrateKind.SMART
}
results = tuple(
result
for result in library.reconcile_by_filename()
if result.reference.source not in dynamic_sources
)
missing = [result for result in results if not result.exists_by_filename]
healthy_count = len(results) - len(missing)
score = (
round(healthy_count / len(results) * 100, 1) if results else None
)
duplicate_groups = find_duplicate_groups(library.tracks)
exact_duplicates = [
group
for group in duplicate_groups
if group.kind is DuplicateKind.EXACT_NAME
]
cloud_conflicts = [
group
for group in duplicate_groups
if group.kind is DuplicateKind.CLOUD_CONFLICT
]
referenced_names = {reference.filename for reference in library.references}
unused_count = sum(
1 for track in library.tracks if track.filename not in referenced_names
)
matcher = MatchingEngine(library.tracks)
suggested_count = sum(
bool(matcher.candidates_for(result.reference)) for result in missing
)
return HealthReport(
score=score,
total_references=len(library.references),
scored_references=len(results),
healthy_references=healthy_count,
missing_references=len(missing),
unique_missing_filenames=len(
{result.reference.filename for result in missing}
),
disk_tracks=len(library.tracks),
duplicate_filename_groups=len(exact_duplicates),
duplicate_files=sum(group.extra_files for group in exact_duplicates),
suspected_cloud_conflict_groups=len(cloud_conflicts),
suspected_cloud_conflict_files=sum(
group.extra_files for group in cloud_conflicts
),
unused_tracks=unused_count,
suggested_matches=suggested_count,
static_crates=sum(
crate.kind is CrateKind.STATIC for crate in library.crates
),
smart_crates=sum(
crate.kind is CrateKind.SMART for crate in library.crates
),
unknown_crates=sum(
crate.kind is CrateKind.UNKNOWN for crate in library.crates
),
dynamic_references_excluded=sum(
len(crate.references)
for crate in library.crates
if crate.kind is CrateKind.SMART
),
)
+40
View File
@@ -0,0 +1,40 @@
import logging
from pathlib import Path
from typing import Optional
LOGGER_NAME = "serato_doctor"
LOG_FORMAT = "%(asctime)s %(levelname)s %(message)s"
def configure_logging(
verbose: bool = False, log_file: Optional[Path] = None
) -> logging.Logger:
"""Configure isolated application logging and return the project logger."""
logger = logging.getLogger(LOGGER_NAME)
logger.setLevel(logging.DEBUG)
logger.propagate = False
for handler in logger.handlers[:]:
handler.close()
logger.removeHandler(handler)
formatter = logging.Formatter(LOG_FORMAT)
if verbose:
console = logging.StreamHandler()
console.setLevel(logging.DEBUG)
console.setFormatter(formatter)
logger.addHandler(console)
if log_file is not None:
file_handler = logging.FileHandler(log_file, encoding="utf-8")
file_handler.setLevel(logging.INFO)
file_handler.setFormatter(formatter)
logger.addHandler(file_handler)
if not logger.handlers:
logger.addHandler(logging.NullHandler())
return logger
+86
View File
@@ -0,0 +1,86 @@
import re
import unicodedata
from collections import defaultdict
from pathlib import Path
from typing import DefaultDict, Iterable, List, Tuple
from serato_doctor.models.match import MatchEvidence, TrackMatch
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def normalize(value: str) -> str:
return unicodedata.normalize("NFKC", value).casefold()
def cloud_conflict_name(filename: str) -> str:
"""Remove a trailing numeric cloud-conflict suffix from a filename stem."""
path = Path(filename)
stem = re.sub(r" \d+$", "", path.stem)
return normalize(stem + path.suffix)
def score_candidate(reference: TrackReference, track: DiskTrack) -> TrackMatch:
reference_name = reference.filename
track_name = track.filename
if reference_name == track_name:
filename_points = 60
filename_reason = "Filename is identical"
elif normalize(reference_name) == normalize(track_name):
filename_points = 55
filename_reason = "Filename matches after case and Unicode normalization"
elif cloud_conflict_name(reference_name) == cloud_conflict_name(track_name):
filename_points = 50
filename_reason = "Filename matches after removing a numeric conflict suffix"
else:
filename_points = 0
filename_reason = "Filename does not match"
same_extension = normalize(reference.path.suffix) == normalize(track.suffix)
same_parent = normalize(reference.path.parent.name) == normalize(
track.path.parent.name
)
evidence = (
MatchEvidence(
"filename",
filename_points > 0,
filename_points,
60,
filename_reason,
),
MatchEvidence(
"extension",
same_extension,
10 if same_extension else 0,
10,
"File extension matches" if same_extension else "File extension differs",
),
MatchEvidence(
"parent_folder",
same_parent,
20 if same_parent else 0,
20,
"Parent folder matches" if same_parent else "Parent folder differs",
),
)
return TrackMatch(reference, track, evidence)
class MatchingEngine:
"""Find and rank filename-related disk candidates without modifying files."""
def __init__(self, tracks: Iterable[DiskTrack]):
self._by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
for track in tracks:
self._by_conflict_name[cloud_conflict_name(track.filename)].append(track)
def candidates_for(self, reference: TrackReference) -> Tuple[TrackMatch, ...]:
candidates = self._by_conflict_name.get(
cloud_conflict_name(reference.filename), []
)
matches = [score_candidate(reference, track) for track in candidates]
return tuple(
sorted(matches, key=lambda match: (-match.score, str(match.track.path)))
)
+21
View File
@@ -0,0 +1,21 @@
from serato_doctor.models.crate import Crate, CrateKind
from serato_doctor.models.duplicate import DuplicateGroup, DuplicateKind
from serato_doctor.models.health import HealthReport
from serato_doctor.models.library import Library
from serato_doctor.models.match import MatchEvidence, TrackMatch
from serato_doctor.models.reference import ReferenceResult, TrackReference
from serato_doctor.models.track import DiskTrack
__all__ = [
"Crate",
"CrateKind",
"DiskTrack",
"DuplicateGroup",
"DuplicateKind",
"HealthReport",
"Library",
"MatchEvidence",
"ReferenceResult",
"TrackMatch",
"TrackReference",
]
+21
View File
@@ -0,0 +1,21 @@
from dataclasses import dataclass
from enum import Enum
from pathlib import Path
from typing import Tuple
from serato_doctor.models.reference import TrackReference
class CrateKind(str, Enum):
STATIC = "static"
SMART = "smart"
UNKNOWN = "unknown"
@dataclass(frozen=True)
class Crate:
"""A Serato crate and the track references parsed from it."""
path: Path
references: Tuple[TrackReference, ...]
kind: CrateKind = CrateKind.UNKNOWN
+30
View File
@@ -0,0 +1,30 @@
from dataclasses import dataclass
from enum import Enum
from typing import Tuple
from serato_doctor.models.track import DiskTrack
class DuplicateKind(str, Enum):
EXACT_NAME = "exact_name"
CLOUD_CONFLICT = "cloud_conflict"
@dataclass(frozen=True)
class DuplicateGroup:
"""A deterministic group of files that warrants duplicate review."""
kind: DuplicateKind
comparison_key: str
tracks: Tuple[DiskTrack, ...]
@property
def extra_files(self) -> int:
return max(0, len(self.tracks) - 1)
@property
def display_name(self) -> str:
return min(
(track.filename for track in self.tracks),
key=lambda name: (len(name), name.casefold()),
)
+29
View File
@@ -0,0 +1,29 @@
from dataclasses import dataclass
from typing import Optional
@dataclass(frozen=True)
class HealthReport:
"""Transparent aggregate findings from a read-only library analysis."""
score: Optional[float]
total_references: int
scored_references: int
healthy_references: int
missing_references: int
unique_missing_filenames: int
disk_tracks: int
duplicate_filename_groups: int
duplicate_files: int
suspected_cloud_conflict_groups: int
suspected_cloud_conflict_files: int
unused_tracks: int
suggested_matches: int
static_crates: int
smart_crates: int
unknown_crates: int
dynamic_references_excluded: int
@property
def score_basis(self) -> str:
return "Resolved non-dynamic references / non-dynamic references scored"
+45
View File
@@ -0,0 +1,45 @@
from dataclasses import dataclass
from typing import Iterable, Tuple
from serato_doctor.models.crate import Crate
from serato_doctor.models.reference import ReferenceResult, TrackReference
from serato_doctor.models.track import DiskTrack
@dataclass(frozen=True)
class Library:
"""The read-only view of crate references and audio found on disk."""
references: Tuple[TrackReference, ...]
tracks: Tuple[DiskTrack, ...]
crates: Tuple[Crate, ...] = ()
@classmethod
def build(
cls,
references: Iterable[TrackReference],
tracks: Iterable[DiskTrack],
) -> "Library":
return cls(tuple(references), tuple(tracks))
@classmethod
def from_crates(
cls, crates: Iterable[Crate], tracks: Iterable[DiskTrack]
) -> "Library":
crate_tuple = tuple(crates)
references = tuple(
reference
for crate in crate_tuple
for reference in crate.references
)
return cls(references, tuple(tracks), crate_tuple)
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
disk_names = {track.filename for track in self.tracks}
return tuple(
ReferenceResult(
reference=reference,
exists_by_filename=reference.filename in disk_names,
)
for reference in self.references
)
+39
View File
@@ -0,0 +1,39 @@
from dataclasses import dataclass
from typing import Tuple
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
@dataclass(frozen=True)
class MatchEvidence:
"""One explainable scoring decision for a candidate track."""
field: str
matched: bool
points: int
max_points: int
explanation: str
@dataclass(frozen=True)
class TrackMatch:
"""A ranked candidate backed by explicit, inspectable evidence."""
reference: TrackReference
track: DiskTrack
evidence: Tuple[MatchEvidence, ...]
@property
def score(self) -> int:
return sum(item.points for item in self.evidence)
@property
def max_score(self) -> int:
return sum(item.max_points for item in self.evidence)
@property
def score_percent(self) -> float:
if not self.max_score:
return 0.0
return round(self.score / self.max_score * 100, 1)
+27
View File
@@ -0,0 +1,27 @@
from dataclasses import dataclass
from pathlib import Path
@dataclass(frozen=True)
class TrackReference:
"""A track path referenced by a Serato crate."""
source: Path
path: Path
filename: str
@dataclass(frozen=True)
class ReferenceResult:
"""The filename-level reconciliation result for a crate reference."""
reference: TrackReference
exists_by_filename: bool
def as_row(self) -> dict:
return {
"crate": str(self.reference.source),
"serato_path": str(self.reference.path),
"filename": self.reference.filename,
"exists_by_filename": self.exists_by_filename,
}
@@ -2,15 +2,10 @@ from dataclasses import dataclass
from pathlib import Path from pathlib import Path
@dataclass(frozen=True)
class TrackReference:
source: Path
path: Path
filename: str
@dataclass(frozen=True) @dataclass(frozen=True)
class DiskTrack: class DiskTrack:
"""An audio file discovered on disk."""
path: Path path: Path
filename: str filename: str
size: int size: int
+47
View File
@@ -0,0 +1,47 @@
from collections import Counter, defaultdict
from pathlib import Path
from typing import Iterable
import csv
from serato_doctor.models.reference import ReferenceResult
def write_missing_report(results: Iterable[ReferenceResult], out: Path) -> None:
rows = [result.as_row() for result in results]
missing = [r for r in rows if not r["exists_by_filename"]]
crate_counts = Counter(r["crate"] for r in missing)
filename_counts = Counter(r["filename"] for r in missing)
with out.open("w", encoding="utf-8") as f:
f.write("# Serato Doctor Missing Report\n\n")
f.write(f"Total missing references: {len(missing)}\n\n")
f.write("## Missing by crate\n\n")
for crate, count in crate_counts.most_common():
f.write(f"{count:5} {crate}\n")
f.write("\n## Most common missing filenames\n\n")
for filename, count in filename_counts.most_common(100):
f.write(f"{count:5} {filename}\n")
f.write("\n## Detail\n\n")
by_crate = defaultdict(list)
for r in missing:
by_crate[r["crate"]].append(r["filename"])
for crate, names in sorted(by_crate.items()):
f.write(f"\n### {crate}\n")
for name in sorted(set(names)):
f.write(f"- {name}\n")
def write_csv(results: Iterable[ReferenceResult], out: Path) -> None:
rows = [result.as_row() for result in results]
with out.open("w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(
f,
fieldnames=["crate", "serato_path", "filename", "exists_by_filename"],
)
writer.writeheader()
writer.writerows(rows)
+1 -1
View File
@@ -1,6 +1,6 @@
from pathlib import Path from pathlib import Path
from serato_doctor.models import DiskTrack from serato_doctor.models.track import DiskTrack
AUDIO_SUFFIXES = {".mp3", ".m4a", ".wav", ".aif", ".aiff", ".flac"} AUDIO_SUFFIXES = {".mp3", ".m4a", ".wav", ".aif", ".aiff", ".flac"}
+131
View File
@@ -0,0 +1,131 @@
import csv
import sys
from serato_doctor.cli import main
def test_cli_writes_reports_and_prints_counts(tmp_path, monkeypatch, capsys):
serato = tmp_path / "serato"
subcrates = serato / "Subcrates"
music = tmp_path / "music"
subcrates.mkdir(parents=True)
music.mkdir()
crate_text = (
"Users/sample-user/OneDrive/Jukebox/Found.mp漳牴k"
"Users/sample-user/OneDrive/Jukebox/Missing.mp漳牴k"
)
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
(music / "Found.mp3").write_bytes(b"synthetic audio")
csv_path = tmp_path / "scan.csv"
report_path = tmp_path / "report.txt"
monkeypatch.setattr(
sys,
"argv",
[
"serato-doctor",
"--serato",
str(serato),
"--music",
str(music),
"--out",
str(csv_path),
"--report",
str(report_path),
],
)
main()
captured = capsys.readouterr()
output = captured.out
assert captured.err == ""
assert "Crate references: 2" in output
assert "Disk tracks: 1" in output
assert "Missing by filename: 1" in output
assert f"CSV: {csv_path}" in output
assert f"Report: {report_path}" in output
with csv_path.open(newline="", encoding="utf-8") as csv_file:
rows = list(csv.DictReader(csv_file))
assert [row["exists_by_filename"] for row in rows] == ["True", "False"]
assert "Total missing references: 1" in report_path.read_text(encoding="utf-8")
def test_cli_accepts_old_reference_root(tmp_path, monkeypatch, capsys):
serato = tmp_path / "serato"
subcrates = serato / "Subcrates"
music = tmp_path / "music"
subcrates.mkdir(parents=True)
music.mkdir()
(subcrates / "Test.crate").write_bytes(
"Archive/Jukebox/Found.mp3otrk".encode("utf-16-le")
)
(music / "Found.mp3").write_bytes(b"synthetic audio")
monkeypatch.setattr(
sys,
"argv",
[
"serato-doctor",
"--serato",
str(serato),
"--music",
str(music),
"--out",
str(tmp_path / "scan.csv"),
"--report",
str(tmp_path / "report.txt"),
"--reference-root",
"/Archive/Jukebox",
],
)
main()
output = capsys.readouterr().out
assert "Crate references: 1" in output
assert "Missing by filename: 0" in output
def test_analyze_prints_health_without_writing_reports(tmp_path, monkeypatch, capsys):
serato = tmp_path / "serato"
subcrates = serato / "Subcrates"
music = tmp_path / "music"
subcrates.mkdir(parents=True)
music.mkdir()
crate_text = (
"Users/sample-user/Jukebox/Found.mp3otrk"
"Users/sample-user/Jukebox/Missing.mp3otrk"
)
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
(music / "Found.mp3").write_bytes(b"synthetic audio")
csv_path = tmp_path / "scan.csv"
report_path = tmp_path / "report.txt"
monkeypatch.setattr(
sys,
"argv",
[
"serato-doctor",
"analyze",
"--serato",
str(serato),
"--music",
str(music),
"--out",
str(csv_path),
"--report",
str(report_path),
],
)
main()
output = capsys.readouterr().out
assert "Overall Health: 50.0%" in output
assert "Tracks: 1" in output
assert "Crate References: 2" in output
assert "References Scored: 2" in output
assert "Healthy References: 1" in output
assert "Broken References: 1" in output
assert "Unused Tracks: 0" in output
assert not csv_path.exists()
assert not report_path.exists()
+17
View File
@@ -0,0 +1,17 @@
from pathlib import Path
from serato_doctor.config import ScanConfig
def test_scan_config_freezes_reference_roots():
roots = (Path(root) for root in ["/old/one", "/old/two"])
config = ScanConfig.build(
serato=Path("serato"),
music=Path("music"),
out=Path("scan.csv"),
report=Path("report.txt"),
reference_roots=roots,
)
assert config.reference_roots == (Path("/old/one"), Path("/old/two"))
+103
View File
@@ -0,0 +1,103 @@
from pathlib import Path
import pytest
from serato_doctor.crate_parser import (
classify_crate,
clean_path,
load_crate,
parse_crate,
parse_crates,
path_markers,
)
from serato_doctor.models.crate import CrateKind
@pytest.mark.parametrize(
("artifact", "expected"),
[
("song.mp漳", "song.mp3"),
("song.m4愠", "song.m4a"),
("song.wa瘠", "song.wav"),
("song.ai映", "song.aif"),
],
)
def test_clean_path_repairs_known_extension_artifacts(artifact, expected):
assert clean_path(artifact) == expected
def test_parse_crate_extracts_references(tmp_path):
crate_path = tmp_path / "House.crate"
crate_text = (
"header"
"Users/sample-user/OneDrive/Jukebox/House/First.mp漳牴k"
"metadata"
"Users/sample-user/OneDrive/Jukebox/House/Second.m4愠otrk"
)
crate_path.write_bytes(crate_text.encode("utf-16-le"))
crate = load_crate(crate_path)
assert crate.path == crate_path
assert [reference.filename for reference in crate.references] == [
"First.mp3",
"Second.m4a",
]
assert parse_crate(crate_path) == list(crate.references)
def test_parse_crate_ignores_record_without_stop_marker(tmp_path):
crate_path = Path(tmp_path) / "Incomplete.crate"
crate_path.write_bytes(
"Users/sample-user/OneDrive/Jukebox/House/Incomplete.mp3".encode(
"utf-16-le"
)
)
assert parse_crate(crate_path) == []
def test_parse_crate_uses_configured_reference_root(tmp_path):
crate_path = tmp_path / "Migrated.crate"
crate_path.write_bytes(
"Archive/Old Library/House/Track.mp3otrk".encode("utf-16-le")
)
references = parse_crate(crate_path, [Path("/Archive/Old Library")])
assert references[0].path == Path("/Archive/Old Library/House/Track.mp3")
def test_default_path_markers_do_not_contain_a_username():
assert path_markers([]) == ("Users/", "Volumes/")
def test_parse_crates_reuses_configured_roots_for_every_crate(tmp_path):
for name in ("First", "Second"):
(tmp_path / f"{name}.crate").write_bytes(
f"Archive/Jukebox/{name}.mp3otrk".encode("utf-16-le")
)
references = parse_crates(
tmp_path, (root for root in [Path("/Archive/Jukebox")])
)
assert {reference.filename for reference in references} == {
"First.mp3",
"Second.mp3",
}
@pytest.mark.parametrize(
("relative_path", "expected"),
[
("_Serato_/Subcrates/House.crate", CrateKind.STATIC),
("_Serato_/Smartcrates/Warmup.crate", CrateKind.SMART),
("_Serato_/Subcrates/Compatible by key.crate", CrateKind.SMART),
("fixtures/Unknown.crate", CrateKind.UNKNOWN),
],
)
def test_classify_crate_uses_provenance_and_known_dynamic_name(
tmp_path, relative_path, expected
):
assert classify_crate(tmp_path / relative_path) is expected
+57
View File
@@ -0,0 +1,57 @@
from pathlib import Path
from serato_doctor.duplicates import find_duplicate_groups
from serato_doctor.models.duplicate import DuplicateKind
from serato_doctor.models.track import DiskTrack
def track(path):
path = Path(path)
return DiskTrack(path, path.name, 100, path.suffix.lower())
def test_exact_names_in_different_folders_form_a_group():
groups = find_duplicate_groups(
[track("/A/Song.mp3"), track("/B/Song.mp3")]
)
assert len(groups) == 1
assert groups[0].kind is DuplicateKind.EXACT_NAME
assert groups[0].display_name == "Song.mp3"
assert groups[0].extra_files == 1
def test_case_only_difference_is_an_exact_name_duplicate():
groups = find_duplicate_groups(
[track("/A/SONG.MP3"), track("/B/song.mp3")]
)
assert groups[0].kind is DuplicateKind.EXACT_NAME
def test_numeric_suffixes_form_a_suspected_cloud_conflict_group():
groups = find_duplicate_groups(
[
track("/A/Track.mp3"),
track("/B/Track 2.mp3"),
track("/C/Track 3.mp3"),
]
)
assert len(groups) == 1
assert groups[0].kind is DuplicateKind.CLOUD_CONFLICT
assert groups[0].display_name == "Track.mp3"
assert groups[0].extra_files == 2
assert [item.path for item in groups[0].tracks] == [
Path("/A/Track.mp3"),
Path("/B/Track 2.mp3"),
Path("/C/Track 3.mp3"),
]
def test_unrelated_and_single_files_are_not_reported():
groups = find_duplicate_groups(
[track("/A/First.mp3"), track("/B/Second.mp3")]
)
assert groups == ()
+106
View File
@@ -0,0 +1,106 @@
from pathlib import Path
from serato_doctor.health import analyze_health
from serato_doctor.models.crate import Crate, CrateKind
from serato_doctor.models.library import Library
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def reference(filename):
return TrackReference(
Path("Test.crate"), Path("/old/House") / filename, filename
)
def track(filename, folder="House"):
path = Path("/new") / folder / filename
return DiskTrack(path, filename, 100, path.suffix.lower())
def test_health_report_exposes_each_metric():
library = Library.build(
[
reference("Found.mp3"),
reference("Conflict.mp3"),
reference("Absent.mp3"),
],
[
track("Found.mp3"),
track("Conflict 2.mp3"),
track("Conflict 2.mp3", "Backup"),
track("Unused.mp3"),
],
)
report = analyze_health(library)
assert report.score == 33.3
assert report.score_basis == (
"Resolved non-dynamic references / non-dynamic references scored"
)
assert report.total_references == 3
assert report.scored_references == 3
assert report.healthy_references == 1
assert report.missing_references == 2
assert report.unique_missing_filenames == 2
assert report.disk_tracks == 4
assert report.duplicate_filename_groups == 1
assert report.duplicate_files == 1
assert report.unused_tracks == 3
assert report.suggested_matches == 1
def test_empty_library_has_no_health_score():
report = analyze_health(Library.build([], []))
assert report.score is None
assert report.total_references == 0
assert report.scored_references == 0
def test_smart_crate_references_are_reported_but_not_scored():
static_reference = reference("Found.mp3")
smart_reference = TrackReference(
Path("Smartcrates/Dynamic.crate"),
Path("/old/House/Dynamic.mp3"),
"Dynamic.mp3",
)
crates = [
Crate(Path("Subcrates/Static.crate"), (static_reference,), CrateKind.STATIC),
Crate(
Path("Smartcrates/Dynamic.crate"),
(smart_reference,),
CrateKind.SMART,
),
]
library = Library.from_crates(crates, [track("Found.mp3")])
report = analyze_health(library)
assert report.score == 100.0
assert report.total_references == 2
assert report.scored_references == 1
assert report.missing_references == 0
assert report.static_crates == 1
assert report.smart_crates == 1
assert report.dynamic_references_excluded == 1
def test_health_reports_cloud_conflicts_separately_from_exact_duplicates():
library = Library.build(
[],
[
track("Track.mp3", "Original"),
track("Track 2.mp3", "Conflict"),
track("Copy.mp3", "First"),
track("Copy.mp3", "Second"),
],
)
report = analyze_health(library)
assert report.duplicate_filename_groups == 1
assert report.duplicate_files == 1
assert report.suspected_cloud_conflict_groups == 1
assert report.suspected_cloud_conflict_files == 1
+24
View File
@@ -0,0 +1,24 @@
from pathlib import Path
from serato_doctor.models.library import Library
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def test_library_reconciles_references_by_filename():
crate = Path("House.crate")
references = [
TrackReference(crate, Path("/old/Found.mp3"), "Found.mp3"),
TrackReference(crate, Path("/old/Missing.mp3"), "Missing.mp3"),
]
tracks = [DiskTrack(Path("/new/Found.mp3"), "Found.mp3", 10, ".mp3")]
results = Library.build(references, tracks).reconcile_by_filename()
assert [result.exists_by_filename for result in results] == [True, False]
assert results[1].as_row() == {
"crate": "House.crate",
"serato_path": "/old/Missing.mp3",
"filename": "Missing.mp3",
"exists_by_filename": False,
}
+34
View File
@@ -0,0 +1,34 @@
import logging
from serato_doctor.logging import LOGGER_NAME, configure_logging
def test_logging_is_quiet_by_default(capsys):
logger = configure_logging()
logger.info("not visible")
assert capsys.readouterr().err == ""
def test_logging_writes_aggregate_progress_to_file(tmp_path):
log_path = tmp_path / "scan.log"
logger = configure_logging(log_file=log_path)
logger.info("Scanned %d disk tracks", 10)
contents = log_path.read_text(encoding="utf-8")
assert "INFO Scanned 10 disk tracks" in contents
def test_reconfiguring_logging_replaces_handlers():
configure_logging(verbose=True)
logger = configure_logging(verbose=True)
active_handlers = [
handler
for handler in logger.handlers
if not isinstance(handler, logging.NullHandler)
]
assert logger.name == LOGGER_NAME
assert len(active_handlers) == 1
+62
View File
@@ -0,0 +1,62 @@
from pathlib import Path
from serato_doctor.matching import MatchingEngine, score_candidate
from serato_doctor.models.reference import TrackReference
from serato_doctor.models.track import DiskTrack
def reference(filename, folder="House"):
return TrackReference(
Path("Test.crate"), Path("/old") / folder / filename, filename
)
def track(filename, folder="House"):
path = Path("/new") / folder / filename
return DiskTrack(path, filename, 100, path.suffix.lower())
def test_exact_candidate_has_full_evidence_score():
match = score_candidate(reference("Track.mp3"), track("Track.mp3"))
assert match.score == 90
assert match.max_score == 90
assert match.score_percent == 100.0
assert all(item.matched for item in match.evidence)
def test_cloud_conflict_suffix_is_explained():
match = score_candidate(reference("Track.mp3"), track("Track 2.mp3"))
assert match.score == 80
assert match.score_percent == 88.9
assert match.evidence[0].explanation == (
"Filename matches after removing a numeric conflict suffix"
)
def test_case_normalized_match_scores_below_exact():
match = score_candidate(reference("TRACK.MP3"), track("track.mp3"))
assert match.score == 85
assert match.evidence[0].points == 55
def test_engine_omits_unrelated_filenames():
engine = MatchingEngine([track("Different.mp3")])
assert engine.candidates_for(reference("Missing.mp3")) == ()
def test_ambiguous_candidates_are_ranked_deterministically():
engine = MatchingEngine(
[track("Track 2.mp3", "Other"), track("Track 3.mp3", "House")]
)
matches = engine.candidates_for(reference("Track.mp3"))
assert [match.track.filename for match in matches] == [
"Track 3.mp3",
"Track 2.mp3",
]
assert [match.score for match in matches] == [80, 60]
+32
View File
@@ -0,0 +1,32 @@
import runpy
from pathlib import Path
from serato_doctor.crate_parser import load_library_crates
from serato_doctor.models.crate import CrateKind
from serato_doctor.models.library import Library
from serato_doctor.scanner import scan_audio
def test_generated_sample_library_has_expected_scenario(tmp_path):
generator_path = (
Path(__file__).parents[1] / "samples" / "small-library" / "generate.py"
)
build_sample = runpy.run_path(str(generator_path))["build_sample"]
sample_root = build_sample(tmp_path / "sample")
crates = load_library_crates(sample_root / "Serato" / "_Serato_")
tracks = scan_audio(sample_root / "Music")
library = Library.from_crates(crates, tracks)
results = library.reconcile_by_filename()
missing = {
result.reference.filename
for result in results
if not result.exists_by_filename
}
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7
assert len(library.references) == 13
assert len(tracks) == 10
assert missing == {"Missing.mp3", "Old Name.mp3"}
assert sum(crate.kind is CrateKind.STATIC for crate in crates) == 5
assert sum(crate.kind is CrateKind.SMART for crate in crates) == 2
+17
View File
@@ -0,0 +1,17 @@
from serato_doctor.scanner import scan_audio
def test_scan_audio_finds_supported_files(tmp_path):
music = tmp_path / "music"
nested = music / "House"
nested.mkdir(parents=True)
(nested / "First.MP3").write_bytes(b"synthetic audio")
(nested / "Second.flac").write_bytes(b"fixture")
(nested / "notes.txt").write_text("not audio", encoding="utf-8")
tracks = scan_audio(music)
assert {track.filename for track in tracks} == {"First.MP3", "Second.flac"}
first = next(track for track in tracks if track.filename == "First.MP3")
assert first.suffix == ".mp3"
assert first.size == len(b"synthetic audio")