Compare commits
16 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| b8bf9ec6e1 | |||
| 167f029c28 | |||
| 5a2edb8b94 | |||
| 29e53f7dc6 | |||
| 4fde93613f | |||
| e4d5a32de2 | |||
| ec671079ba | |||
| 4e85b5912b | |||
| d4fe61cf8c | |||
| c4d08bfe27 | |||
| bb74c14742 | |||
| 71fa64a960 | |||
| 09167e44c7 | |||
| 508eecd0bd | |||
| 577fe1f7a7 | |||
| c3bf111107 |
@@ -2,7 +2,9 @@ __pycache__/
|
|||||||
*.pyc
|
*.pyc
|
||||||
.env
|
.env
|
||||||
.venv/
|
.venv/
|
||||||
|
*.egg-info/
|
||||||
reports/
|
reports/
|
||||||
|
samples/small-library/generated/
|
||||||
*.sqlite
|
*.sqlite
|
||||||
*.db
|
*.db
|
||||||
*.csv
|
*.csv
|
||||||
|
|||||||
@@ -0,0 +1,19 @@
|
|||||||
|
# Contributing
|
||||||
|
|
||||||
|
## Branching
|
||||||
|
|
||||||
|
- `main` is stable.
|
||||||
|
- `develop` is the integration branch.
|
||||||
|
- Feature branches use: `feature/<name>`.
|
||||||
|
|
||||||
|
## Safety Rules
|
||||||
|
|
||||||
|
Never commit personal music library files, Serato databases, or crates.
|
||||||
|
|
||||||
|
Never write repair code without dry-run mode, backup plan, and rollback log.
|
||||||
|
|
||||||
|
## Testing
|
||||||
|
|
||||||
|
Install development dependencies with `python3 -m pip install -e '.[dev]'`.
|
||||||
|
|
||||||
|
Run `pytest` before committing. Tests and fixtures must use synthetic library data.
|
||||||
+51
@@ -0,0 +1,51 @@
|
|||||||
|
# Serato Doctor Roadmap
|
||||||
|
|
||||||
|
## v0.1 — Library Inspector
|
||||||
|
|
||||||
|
- [x] Project repository
|
||||||
|
- [x] Filesystem scanner
|
||||||
|
- [x] Serato crate parser
|
||||||
|
- [x] Missing reference CSV report
|
||||||
|
- [x] Grouped missing reference report
|
||||||
|
- [ ] HTML health dashboard
|
||||||
|
- [x] Test suite
|
||||||
|
- [x] Sample library fixtures
|
||||||
|
- [ ] Database V2 read-only parser
|
||||||
|
- [x] Configuration
|
||||||
|
- [x] Logging
|
||||||
|
- [x] Matching engine
|
||||||
|
|
||||||
|
## v0.2 — Diagnostics
|
||||||
|
|
||||||
|
- [ ] Duplicate filename detection
|
||||||
|
- [ ] Duplicate audio hash detection
|
||||||
|
- [ ] Broken symlink detection
|
||||||
|
- [ ] Orphaned audio detection
|
||||||
|
- [ ] OneDrive rename detection
|
||||||
|
- [ ] Crate classification: static vs smart/dynamic
|
||||||
|
- [x] Library health score
|
||||||
|
|
||||||
|
## v0.3 — Safe Repair
|
||||||
|
|
||||||
|
- [ ] Dry-run repair plan
|
||||||
|
- [ ] Backup before repair
|
||||||
|
- [ ] Compatibility symlink creation
|
||||||
|
- [ ] Compatibility copy creation
|
||||||
|
- [ ] Rename repair
|
||||||
|
- [ ] Rollback log
|
||||||
|
|
||||||
|
## v0.4 — Migration Wizard
|
||||||
|
|
||||||
|
- [ ] Move library root
|
||||||
|
- [ ] Cloud provider migration
|
||||||
|
- [ ] External drive migration
|
||||||
|
- [ ] Verify moved library
|
||||||
|
- [ ] Update application references
|
||||||
|
|
||||||
|
## v1.0 — DJ Library Doctor
|
||||||
|
|
||||||
|
- [ ] Desktop UI
|
||||||
|
- [ ] Serato support
|
||||||
|
- [ ] Rekordbox support
|
||||||
|
- [ ] VirtualDJ support
|
||||||
|
- [ ] Engine DJ support
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
# Architecture
|
||||||
|
|
||||||
|
Serato Doctor is designed as a DJ library inspection, repair, and migration platform.
|
||||||
|
|
||||||
|
## Design Principles
|
||||||
|
|
||||||
|
1. Read-only by default.
|
||||||
|
2. Every repair must support preview/dry-run.
|
||||||
|
3. Every repair must create a backup or rollback path.
|
||||||
|
4. Application-specific logic lives in engines.
|
||||||
|
5. Core matching and scanning logic should be application-agnostic.
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
# Case Study: OneDrive Mac Migration
|
||||||
|
|
||||||
|
## Scenario
|
||||||
|
|
||||||
|
A large Serato DJ library was migrated from an older Mac to a newer Mac using OneDrive.
|
||||||
|
|
||||||
|
## Symptoms
|
||||||
|
|
||||||
|
- OneDrive client stuck syncing
|
||||||
|
- Duplicate OneDrive folders
|
||||||
|
- Thousands of files renamed with trailing ` 2`
|
||||||
|
- Serato reported many tracks as missing
|
||||||
|
- Some files existed on disk but still appeared orange in Serato
|
||||||
|
|
||||||
|
## Findings
|
||||||
|
|
||||||
|
- OneDrive sync state was rebuilt successfully
|
||||||
|
- Thousands of orphaned filename conflicts were repaired
|
||||||
|
- Some Serato references were stale database objects, not missing files
|
||||||
|
- Smart/dynamic crates should be classified separately from static user crates
|
||||||
|
|
||||||
|
## Lessons
|
||||||
|
|
||||||
|
- Filesystem health and Serato database health are separate problems
|
||||||
|
- Smart crates should not be treated the same as static crates
|
||||||
|
- Repair tools must be read-only by default and generate a plan before changing anything
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Scan Configuration
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The prototype recognized crate paths by matching one developer's historical
|
||||||
|
OneDrive path. That made otherwise valid libraries invisible to the parser.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
`ScanConfig` owns the paths for a read-only scan. The CLI accepts repeatable
|
||||||
|
`--reference-root` options for roots embedded in crates before a migration. The
|
||||||
|
parser uses those roots as record boundaries. Without explicit roots, it recognizes
|
||||||
|
generic macOS home and mounted-volume paths without embedding a username.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- A library may have references from more than one historical root.
|
||||||
|
- Root paths may contain spaces or have leading/trailing separators.
|
||||||
|
- Existing `/Users/...` and `/Volumes/...` crates must work without new flags.
|
||||||
|
- A record without a recognized terminator remains ignored.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover default root discovery, a custom migrated root, the CLI option, and
|
||||||
|
the existing synthetic sample library. The scanner remains read-only.
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
# Library Health Engine
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Raw missing-reference counts do not provide a compact view of library integrity,
|
||||||
|
but an opaque blended score would imply confidence the current data cannot support.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The health engine produces an immutable report from the core `Library`. Its score
|
||||||
|
is only the percentage of crate references resolved by exact filename. The report
|
||||||
|
also exposes missing references, unique missing filenames, duplicate filename
|
||||||
|
groups, extra duplicate files, unused tracks, and missing references with matching
|
||||||
|
candidates.
|
||||||
|
|
||||||
|
Duplicate, unused, and candidate counts are informational. They do not affect the
|
||||||
|
score until the project has a documented and validated weighting policy. An empty
|
||||||
|
library has no score rather than a misleading 0% or 100%.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- The same missing filename referenced by several crates.
|
||||||
|
- Several disk files sharing a filename.
|
||||||
|
- Conflict-suffixed files that are unused but may be match candidates.
|
||||||
|
- Libraries with no crate references.
|
||||||
|
- Tracks referenced by filename from more than one crate.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests assert every metric, the disclosed score basis, candidate integration, and
|
||||||
|
empty-library behavior. Analysis remains entirely read-only.
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Application Logging
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The CLI reports final counts but provides no diagnostic trail when a scan behaves
|
||||||
|
unexpectedly. Troubleshooting should not require adding print statements or expose
|
||||||
|
library contents by default.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The project uses an isolated standard-library logger. It has no visible output by
|
||||||
|
default. `--verbose` writes progress to standard error, while `--log-file PATH`
|
||||||
|
writes an informational audit trail. Normal result lines remain on standard output.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Reconfiguring logging in the same process must not duplicate handlers.
|
||||||
|
- Console and file logging may be enabled together.
|
||||||
|
- Log messages contain aggregate counts, not track names or crate contents.
|
||||||
|
- A default scan must remain quiet except for its established result output.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests verify quiet defaults, file output, handler replacement, and unchanged CLI
|
||||||
|
result lines. The full sample-library scan remains read-only.
|
||||||
@@ -0,0 +1,29 @@
|
|||||||
|
# Explainable Matching Engine
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Missing references need ranked candidate files, but a filename-only yes/no check
|
||||||
|
cannot explain ambiguity or cloud-provider conflict names.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
The read-only matching engine indexes normalized filenames and scores only related
|
||||||
|
candidates. Every score contains evidence for filename, extension, and parent
|
||||||
|
folder. Exact filenames earn 60 points, normalized names 55, numeric conflict-name
|
||||||
|
matches 50, extensions 10, and parent folders 20.
|
||||||
|
|
||||||
|
The displayed percentage is an evidence score, not a statistical probability.
|
||||||
|
Metadata, duration, hashes, and fingerprints can add stronger evidence later.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Unicode and case differences.
|
||||||
|
- OneDrive-style names such as `Track 2.mp3`.
|
||||||
|
- Duplicate candidates in different folders.
|
||||||
|
- Legitimate numbered song titles, which remain candidates but are never repaired.
|
||||||
|
- Unrelated names, which are not emitted as candidates.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Tests cover exact, normalized, conflict-suffix, ambiguous, and unrelated filenames.
|
||||||
|
Candidate ordering is deterministic. The engine never changes a track or reference.
|
||||||
@@ -0,0 +1,26 @@
|
|||||||
|
# Synthetic Sample Library
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
Real Serato libraries contain private paths, listening history, and copyrighted
|
||||||
|
music. Contributors still need a repeatable library for tests and demonstrations.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
`samples/small-library/manifest.json` describes a compact migration scenario.
|
||||||
|
`generate.py` turns that manifest into fake audio files and UTF-16-LE crate files
|
||||||
|
inside an ignored `generated` directory. The generated files are disposable and
|
||||||
|
contain no real audio or user library data.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Five manually maintained crates and two smart/dynamic crates.
|
||||||
|
- Mixed supported audio extensions.
|
||||||
|
- One reference whose file is absent.
|
||||||
|
- One stale reference whose file has been renamed.
|
||||||
|
- References shared by static and smart crates.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
`pytest` generates the sample in a temporary directory and verifies its crate,
|
||||||
|
track, and missing-reference counts through the production parser and scanner.
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
# Test Foundation
|
||||||
|
|
||||||
|
## Problem
|
||||||
|
|
||||||
|
The crate parser handles unusual binary text and known extension artifacts. Without
|
||||||
|
automated tests, a small refactor could silently change library counts or reports.
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Tests cover the read-only pipeline from crate and filesystem inputs through the
|
||||||
|
core library model to CSV, text, and CLI output. All test data is synthetic and is
|
||||||
|
created in temporary directories.
|
||||||
|
|
||||||
|
## Edge Cases
|
||||||
|
|
||||||
|
- Known MP3, M4A, WAV, and AIF decode artifacts.
|
||||||
|
- Crate records without a recognized stop marker.
|
||||||
|
- Supported extensions with mixed case and unsupported files.
|
||||||
|
- References that exist by filename and references that remain missing.
|
||||||
|
|
||||||
|
## Verification
|
||||||
|
|
||||||
|
Run `pytest`. No test reads a real Serato library or writes outside pytest's
|
||||||
|
temporary directory.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
[build-system]
|
||||||
|
requires = ["setuptools>=61"]
|
||||||
|
build-backend = "setuptools.build_meta"
|
||||||
|
|
||||||
|
[project]
|
||||||
|
name = "serato-doctor"
|
||||||
|
version = "0.1.0"
|
||||||
|
description = "Inspect, diagnose, repair, and migrate DJ libraries."
|
||||||
|
requires-python = ">=3.9"
|
||||||
|
dependencies = []
|
||||||
|
|
||||||
|
[project.optional-dependencies]
|
||||||
|
dev = ["pytest>=8,<9"]
|
||||||
|
|
||||||
|
[project.scripts]
|
||||||
|
serato-doctor = "serato_doctor.cli:main"
|
||||||
|
|
||||||
|
[tool.pytest.ini_options]
|
||||||
|
testpaths = ["tests"]
|
||||||
|
|
||||||
|
[tool.setuptools.packages.find]
|
||||||
|
include = ["serato_doctor*"]
|
||||||
@@ -0,0 +1,25 @@
|
|||||||
|
# Small Synthetic Library
|
||||||
|
|
||||||
|
This fixture models a small Serato migration without containing music or personal
|
||||||
|
library data. It has 10 fake audio files, five static crates, two smart crates,
|
||||||
|
one absent song, and one renamed song.
|
||||||
|
|
||||||
|
Generate it with:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
python3 samples/small-library/generate.py
|
||||||
|
```
|
||||||
|
|
||||||
|
Analyze it with:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
python3 -m serato_doctor.cli \
|
||||||
|
--serato samples/small-library/generated/Serato/_Serato_ \
|
||||||
|
--music samples/small-library/generated/Music \
|
||||||
|
--out samples/small-library/generated/scan.csv \
|
||||||
|
--report samples/small-library/generated/report.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
Expected CLI counts are 13 crate references, 10 disk tracks, and 2 references
|
||||||
|
missing by filename. Everything under `generated/` is disposable and ignored by
|
||||||
|
Git.
|
||||||
@@ -0,0 +1,50 @@
|
|||||||
|
"""Generate a disposable, synthetic Serato library from the sample manifest."""
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Optional
|
||||||
|
|
||||||
|
|
||||||
|
SAMPLE_ROOT = Path(__file__).parent
|
||||||
|
SERATO_PATH_PREFIX = "Users/sample-user/OneDrive/Jukebox/"
|
||||||
|
|
||||||
|
|
||||||
|
def load_manifest() -> dict:
|
||||||
|
return json.loads((SAMPLE_ROOT / "manifest.json").read_text(encoding="utf-8"))
|
||||||
|
|
||||||
|
|
||||||
|
def build_sample(output: Optional[Path] = None) -> Path:
|
||||||
|
root = output or SAMPLE_ROOT / "generated"
|
||||||
|
manifest = load_manifest()
|
||||||
|
music_root = root / "Music"
|
||||||
|
crate_root = root / "Serato" / "_Serato_" / "Subcrates"
|
||||||
|
crate_root.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
for relative_path in manifest["tracks"]:
|
||||||
|
track_path = music_root / relative_path
|
||||||
|
track_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
track_path.write_bytes(
|
||||||
|
f"Synthetic Serato Doctor fixture: {relative_path}\n".encode("utf-8")
|
||||||
|
)
|
||||||
|
|
||||||
|
for crate in manifest["crates"]:
|
||||||
|
records = "".join(
|
||||||
|
f"{SERATO_PATH_PREFIX}{relative_path}otrk"
|
||||||
|
for relative_path in crate["references"]
|
||||||
|
)
|
||||||
|
(crate_root / crate["name"]).write_bytes(records.encode("utf-16-le"))
|
||||||
|
|
||||||
|
return root
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("--output", type=Path)
|
||||||
|
args = parser.parse_args()
|
||||||
|
root = build_sample(args.output)
|
||||||
|
print(f"Generated synthetic library: {root}")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
{
|
||||||
|
"tracks": [
|
||||||
|
"House/First.mp3",
|
||||||
|
"House/Second.m4a",
|
||||||
|
"Open Format/Third.wav",
|
||||||
|
"Open Format/Fourth.aif",
|
||||||
|
"Classics/Fifth.mp3",
|
||||||
|
"Classics/Sixth.flac",
|
||||||
|
"Warmup/Seventh.mp3",
|
||||||
|
"Warmup/Eighth.mp3",
|
||||||
|
"Renamed/New Name.mp3",
|
||||||
|
"Bonus/Ninth.MP3"
|
||||||
|
],
|
||||||
|
"crates": [
|
||||||
|
{
|
||||||
|
"name": "House.crate",
|
||||||
|
"type": "static",
|
||||||
|
"references": ["House/First.mp3", "House/Second.m4a", "House/Missing.mp3"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Open Format.crate",
|
||||||
|
"type": "static",
|
||||||
|
"references": ["Open Format/Third.wav", "Open Format/Fourth.aif"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Classics.crate",
|
||||||
|
"type": "static",
|
||||||
|
"references": ["Classics/Fifth.mp3", "Classics/Sixth.flac"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Renamed Tracks.crate",
|
||||||
|
"type": "static",
|
||||||
|
"references": ["Renamed/Old Name.mp3"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Bonus.crate",
|
||||||
|
"type": "static",
|
||||||
|
"references": ["Bonus/Ninth.MP3"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Compatible by key.crate",
|
||||||
|
"type": "smart",
|
||||||
|
"references": ["House/First.mp3", "House/Second.m4a"]
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Smart Warmup.crate",
|
||||||
|
"type": "smart",
|
||||||
|
"references": ["Warmup/Seventh.mp3", "Warmup/Eighth.mp3"]
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"expected": {
|
||||||
|
"disk_tracks": 10,
|
||||||
|
"crate_references": 13,
|
||||||
|
"missing_by_filename": ["Missing.mp3", "Old Name.mp3"]
|
||||||
|
}
|
||||||
|
}
|
||||||
+47
-25
@@ -1,9 +1,12 @@
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
import argparse
|
import argparse
|
||||||
import csv
|
|
||||||
|
|
||||||
|
from serato_doctor.config import ScanConfig
|
||||||
from serato_doctor.crate_parser import parse_crates
|
from serato_doctor.crate_parser import parse_crates
|
||||||
|
from serato_doctor.logging import configure_logging
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
from serato_doctor.scanner import scan_audio
|
from serato_doctor.scanner import scan_audio
|
||||||
|
from serato_doctor.report import write_csv, write_missing_report
|
||||||
|
|
||||||
|
|
||||||
def main():
|
def main():
|
||||||
@@ -11,35 +14,54 @@ def main():
|
|||||||
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
parser.add_argument("--serato", default=str(Path.home() / "Music/_Serato_"))
|
||||||
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"))
|
parser.add_argument("--music", default=str(Path.home() / "Library/CloudStorage/OneDrive-Personal/Jukebox"))
|
||||||
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv"))
|
parser.add_argument("--out", default=str(Path.home() / "Desktop/serato_doctor_scan.csv"))
|
||||||
|
parser.add_argument("--report", default=str(Path.home() / "Desktop/serato_doctor_missing_report.txt"))
|
||||||
|
parser.add_argument(
|
||||||
|
"--reference-root",
|
||||||
|
action="append",
|
||||||
|
default=[],
|
||||||
|
help="Old library root stored in crates; may be supplied more than once",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--verbose", action="store_true", help="Write diagnostic progress to stderr"
|
||||||
|
)
|
||||||
|
parser.add_argument("--log-file", help="Write scan progress to a log file")
|
||||||
args = parser.parse_args()
|
args = parser.parse_args()
|
||||||
|
|
||||||
serato = Path(args.serato)
|
config = ScanConfig.build(
|
||||||
music = Path(args.music)
|
serato=Path(args.serato),
|
||||||
out = Path(args.out)
|
music=Path(args.music),
|
||||||
|
out=Path(args.out),
|
||||||
|
report=Path(args.report),
|
||||||
|
reference_roots=(Path(root) for root in args.reference_root),
|
||||||
|
verbose=args.verbose,
|
||||||
|
log_file=Path(args.log_file) if args.log_file else None,
|
||||||
|
)
|
||||||
|
logger = configure_logging(config.verbose, config.log_file)
|
||||||
|
logger.info("Starting read-only library scan")
|
||||||
|
logger.debug("Serato directory: %s", config.serato)
|
||||||
|
logger.debug("Music directory: %s", config.music)
|
||||||
|
|
||||||
refs = parse_crates(serato / "Subcrates")
|
library = Library.build(
|
||||||
disk = scan_audio(music)
|
references=parse_crates(
|
||||||
|
config.serato / "Subcrates", config.reference_roots
|
||||||
|
),
|
||||||
|
tracks=scan_audio(config.music),
|
||||||
|
)
|
||||||
|
results = library.reconcile_by_filename()
|
||||||
|
missing_count = sum(1 for result in results if not result.exists_by_filename)
|
||||||
|
logger.info("Parsed %d crate references", len(library.references))
|
||||||
|
logger.info("Scanned %d disk tracks", len(library.tracks))
|
||||||
|
logger.info("Found %d references missing by filename", missing_count)
|
||||||
|
|
||||||
disk_names = {t.filename for t in disk}
|
write_csv(results, config.out)
|
||||||
|
write_missing_report(results, config.report)
|
||||||
|
logger.info("Wrote CSV and missing-reference reports")
|
||||||
|
|
||||||
rows = []
|
print(f"Crate references: {len(library.references)}")
|
||||||
for ref in refs:
|
print(f"Disk tracks: {len(library.tracks)}")
|
||||||
rows.append({
|
print(f"Missing by filename: {missing_count}")
|
||||||
"crate": str(ref.source),
|
print(f"CSV: {config.out}")
|
||||||
"serato_path": str(ref.path),
|
print(f"Report: {config.report}")
|
||||||
"filename": ref.filename,
|
|
||||||
"exists_by_filename": ref.filename in disk_names,
|
|
||||||
})
|
|
||||||
|
|
||||||
with out.open("w", newline="", encoding="utf-8") as f:
|
|
||||||
writer = csv.DictWriter(f, fieldnames=["crate", "serato_path", "filename", "exists_by_filename"])
|
|
||||||
writer.writeheader()
|
|
||||||
writer.writerows(rows)
|
|
||||||
|
|
||||||
print(f"Crate references: {len(refs)}")
|
|
||||||
print(f"Disk tracks: {len(disk)}")
|
|
||||||
print(f"Missing by filename: {sum(1 for r in rows if not r['exists_by_filename'])}")
|
|
||||||
print(f"Wrote: {out}")
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
|
|||||||
@@ -0,0 +1,37 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Iterable, Optional, Tuple
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ScanConfig:
|
||||||
|
"""Read-only paths used for one library scan."""
|
||||||
|
|
||||||
|
serato: Path
|
||||||
|
music: Path
|
||||||
|
out: Path
|
||||||
|
report: Path
|
||||||
|
reference_roots: Tuple[Path, ...] = ()
|
||||||
|
verbose: bool = False
|
||||||
|
log_file: Optional[Path] = None
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def build(
|
||||||
|
cls,
|
||||||
|
serato: Path,
|
||||||
|
music: Path,
|
||||||
|
out: Path,
|
||||||
|
report: Path,
|
||||||
|
reference_roots: Iterable[Path] = (),
|
||||||
|
verbose: bool = False,
|
||||||
|
log_file: Optional[Path] = None,
|
||||||
|
) -> "ScanConfig":
|
||||||
|
return cls(
|
||||||
|
serato,
|
||||||
|
music,
|
||||||
|
out,
|
||||||
|
report,
|
||||||
|
tuple(reference_roots),
|
||||||
|
verbose,
|
||||||
|
log_file,
|
||||||
|
)
|
||||||
@@ -1,6 +1,8 @@
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
from typing import Iterable, Tuple
|
||||||
|
|
||||||
from serato_doctor.models import TrackReference
|
from serato_doctor.models.crate import Crate
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
|
||||||
|
|
||||||
def read_crate_text(crate_path: Path) -> str:
|
def read_crate_text(crate_path: Path) -> str:
|
||||||
@@ -20,19 +22,32 @@ def clean_path(raw: str) -> str:
|
|||||||
return raw.strip()
|
return raw.strip()
|
||||||
|
|
||||||
|
|
||||||
def parse_crate(crate_path: Path) -> list[TrackReference]:
|
DEFAULT_PATH_MARKERS = ("Users/", "Volumes/")
|
||||||
|
|
||||||
|
|
||||||
|
def path_markers(reference_roots: Iterable[Path]) -> Tuple[str, ...]:
|
||||||
|
configured = tuple(
|
||||||
|
root.expanduser().as_posix().strip("/") + "/" for root in reference_roots
|
||||||
|
)
|
||||||
|
return configured or DEFAULT_PATH_MARKERS
|
||||||
|
|
||||||
|
|
||||||
|
def load_crate(crate_path: Path, reference_roots: Iterable[Path] = ()) -> Crate:
|
||||||
text = read_crate_text(crate_path)
|
text = read_crate_text(crate_path)
|
||||||
refs = []
|
refs = []
|
||||||
|
|
||||||
marker = "Users/djsplice/OneDrive/Jukebox/"
|
|
||||||
|
|
||||||
# Serato record markers seen after paths in UTF-16-LE decoded crate data.
|
# Serato record markers seen after paths in UTF-16-LE decoded crate data.
|
||||||
stop_markers = ["牴k", "otrk", "ptrk", "tvcn", "ovct"]
|
stop_markers = ["牴k", "otrk", "ptrk", "tvcn", "ovct"]
|
||||||
|
|
||||||
|
for marker in path_markers(reference_roots):
|
||||||
for part in text.split(marker)[1:]:
|
for part in text.split(marker)[1:]:
|
||||||
candidate = marker + part
|
candidate = marker + part
|
||||||
|
|
||||||
stops = [candidate.find(m) for m in stop_markers if candidate.find(m) != -1]
|
stops = [
|
||||||
|
candidate.find(stop)
|
||||||
|
for stop in stop_markers
|
||||||
|
if candidate.find(stop) != -1
|
||||||
|
]
|
||||||
if not stops:
|
if not stops:
|
||||||
continue
|
continue
|
||||||
|
|
||||||
@@ -48,11 +63,22 @@ def parse_crate(crate_path: Path) -> list[TrackReference]:
|
|||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|
||||||
return refs
|
return Crate(path=crate_path, references=tuple(refs))
|
||||||
|
|
||||||
|
|
||||||
def parse_crates(root: Path) -> list[TrackReference]:
|
def parse_crate(
|
||||||
|
crate_path: Path, reference_roots: Iterable[Path] = ()
|
||||||
|
) -> list[TrackReference]:
|
||||||
|
"""Parse references from one crate, preserving the prototype API."""
|
||||||
|
|
||||||
|
return list(load_crate(crate_path, reference_roots).references)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_crates(
|
||||||
|
root: Path, reference_roots: Iterable[Path] = ()
|
||||||
|
) -> list[TrackReference]:
|
||||||
refs = []
|
refs = []
|
||||||
|
reference_roots = tuple(reference_roots)
|
||||||
for crate in root.rglob("*.crate"):
|
for crate in root.rglob("*.crate"):
|
||||||
refs.extend(parse_crate(crate))
|
refs.extend(parse_crate(crate, reference_roots))
|
||||||
return refs
|
return refs
|
||||||
|
|||||||
@@ -0,0 +1,43 @@
|
|||||||
|
from collections import Counter
|
||||||
|
|
||||||
|
from serato_doctor.matching import MatchingEngine
|
||||||
|
from serato_doctor.models.health import HealthReport
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
|
||||||
|
|
||||||
|
def analyze_health(library: Library) -> HealthReport:
|
||||||
|
"""Calculate defensible health metrics without changing the library."""
|
||||||
|
|
||||||
|
results = library.reconcile_by_filename()
|
||||||
|
missing = [result for result in results if not result.exists_by_filename]
|
||||||
|
healthy_count = len(results) - len(missing)
|
||||||
|
score = (
|
||||||
|
round(healthy_count / len(results) * 100, 1) if results else None
|
||||||
|
)
|
||||||
|
|
||||||
|
disk_name_counts = Counter(track.filename for track in library.tracks)
|
||||||
|
duplicate_counts = [count for count in disk_name_counts.values() if count > 1]
|
||||||
|
referenced_names = {reference.filename for reference in library.references}
|
||||||
|
unused_count = sum(
|
||||||
|
1 for track in library.tracks if track.filename not in referenced_names
|
||||||
|
)
|
||||||
|
|
||||||
|
matcher = MatchingEngine(library.tracks)
|
||||||
|
suggested_count = sum(
|
||||||
|
bool(matcher.candidates_for(result.reference)) for result in missing
|
||||||
|
)
|
||||||
|
|
||||||
|
return HealthReport(
|
||||||
|
score=score,
|
||||||
|
total_references=len(results),
|
||||||
|
healthy_references=healthy_count,
|
||||||
|
missing_references=len(missing),
|
||||||
|
unique_missing_filenames=len(
|
||||||
|
{result.reference.filename for result in missing}
|
||||||
|
),
|
||||||
|
disk_tracks=len(library.tracks),
|
||||||
|
duplicate_filename_groups=len(duplicate_counts),
|
||||||
|
duplicate_files=sum(count - 1 for count in duplicate_counts),
|
||||||
|
unused_tracks=unused_count,
|
||||||
|
suggested_matches=suggested_count,
|
||||||
|
)
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
import logging
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Optional
|
||||||
|
|
||||||
|
|
||||||
|
LOGGER_NAME = "serato_doctor"
|
||||||
|
LOG_FORMAT = "%(asctime)s %(levelname)s %(message)s"
|
||||||
|
|
||||||
|
|
||||||
|
def configure_logging(
|
||||||
|
verbose: bool = False, log_file: Optional[Path] = None
|
||||||
|
) -> logging.Logger:
|
||||||
|
"""Configure isolated application logging and return the project logger."""
|
||||||
|
|
||||||
|
logger = logging.getLogger(LOGGER_NAME)
|
||||||
|
logger.setLevel(logging.DEBUG)
|
||||||
|
logger.propagate = False
|
||||||
|
|
||||||
|
for handler in logger.handlers[:]:
|
||||||
|
handler.close()
|
||||||
|
logger.removeHandler(handler)
|
||||||
|
|
||||||
|
formatter = logging.Formatter(LOG_FORMAT)
|
||||||
|
|
||||||
|
if verbose:
|
||||||
|
console = logging.StreamHandler()
|
||||||
|
console.setLevel(logging.DEBUG)
|
||||||
|
console.setFormatter(formatter)
|
||||||
|
logger.addHandler(console)
|
||||||
|
|
||||||
|
if log_file is not None:
|
||||||
|
file_handler = logging.FileHandler(log_file, encoding="utf-8")
|
||||||
|
file_handler.setLevel(logging.INFO)
|
||||||
|
file_handler.setFormatter(formatter)
|
||||||
|
logger.addHandler(file_handler)
|
||||||
|
|
||||||
|
if not logger.handlers:
|
||||||
|
logger.addHandler(logging.NullHandler())
|
||||||
|
|
||||||
|
return logger
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
import re
|
||||||
|
import unicodedata
|
||||||
|
from collections import defaultdict
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import DefaultDict, Iterable, List, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def normalize(value: str) -> str:
|
||||||
|
return unicodedata.normalize("NFKC", value).casefold()
|
||||||
|
|
||||||
|
|
||||||
|
def cloud_conflict_name(filename: str) -> str:
|
||||||
|
"""Remove a trailing numeric cloud-conflict suffix from a filename stem."""
|
||||||
|
|
||||||
|
path = Path(filename)
|
||||||
|
stem = re.sub(r" \d+$", "", path.stem)
|
||||||
|
return normalize(stem + path.suffix)
|
||||||
|
|
||||||
|
|
||||||
|
def score_candidate(reference: TrackReference, track: DiskTrack) -> TrackMatch:
|
||||||
|
reference_name = reference.filename
|
||||||
|
track_name = track.filename
|
||||||
|
|
||||||
|
if reference_name == track_name:
|
||||||
|
filename_points = 60
|
||||||
|
filename_reason = "Filename is identical"
|
||||||
|
elif normalize(reference_name) == normalize(track_name):
|
||||||
|
filename_points = 55
|
||||||
|
filename_reason = "Filename matches after case and Unicode normalization"
|
||||||
|
elif cloud_conflict_name(reference_name) == cloud_conflict_name(track_name):
|
||||||
|
filename_points = 50
|
||||||
|
filename_reason = "Filename matches after removing a numeric conflict suffix"
|
||||||
|
else:
|
||||||
|
filename_points = 0
|
||||||
|
filename_reason = "Filename does not match"
|
||||||
|
|
||||||
|
same_extension = normalize(reference.path.suffix) == normalize(track.suffix)
|
||||||
|
same_parent = normalize(reference.path.parent.name) == normalize(
|
||||||
|
track.path.parent.name
|
||||||
|
)
|
||||||
|
evidence = (
|
||||||
|
MatchEvidence(
|
||||||
|
"filename",
|
||||||
|
filename_points > 0,
|
||||||
|
filename_points,
|
||||||
|
60,
|
||||||
|
filename_reason,
|
||||||
|
),
|
||||||
|
MatchEvidence(
|
||||||
|
"extension",
|
||||||
|
same_extension,
|
||||||
|
10 if same_extension else 0,
|
||||||
|
10,
|
||||||
|
"File extension matches" if same_extension else "File extension differs",
|
||||||
|
),
|
||||||
|
MatchEvidence(
|
||||||
|
"parent_folder",
|
||||||
|
same_parent,
|
||||||
|
20 if same_parent else 0,
|
||||||
|
20,
|
||||||
|
"Parent folder matches" if same_parent else "Parent folder differs",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
return TrackMatch(reference, track, evidence)
|
||||||
|
|
||||||
|
|
||||||
|
class MatchingEngine:
|
||||||
|
"""Find and rank filename-related disk candidates without modifying files."""
|
||||||
|
|
||||||
|
def __init__(self, tracks: Iterable[DiskTrack]):
|
||||||
|
self._by_conflict_name: DefaultDict[str, List[DiskTrack]] = defaultdict(list)
|
||||||
|
for track in tracks:
|
||||||
|
self._by_conflict_name[cloud_conflict_name(track.filename)].append(track)
|
||||||
|
|
||||||
|
def candidates_for(self, reference: TrackReference) -> Tuple[TrackMatch, ...]:
|
||||||
|
candidates = self._by_conflict_name.get(
|
||||||
|
cloud_conflict_name(reference.filename), []
|
||||||
|
)
|
||||||
|
matches = [score_candidate(reference, track) for track in candidates]
|
||||||
|
return tuple(
|
||||||
|
sorted(matches, key=lambda match: (-match.score, str(match.track.path)))
|
||||||
|
)
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
from serato_doctor.models.crate import Crate
|
||||||
|
from serato_doctor.models.health import HealthReport
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
from serato_doctor.models.match import MatchEvidence, TrackMatch
|
||||||
|
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
__all__ = [
|
||||||
|
"Crate",
|
||||||
|
"DiskTrack",
|
||||||
|
"HealthReport",
|
||||||
|
"Library",
|
||||||
|
"MatchEvidence",
|
||||||
|
"ReferenceResult",
|
||||||
|
"TrackMatch",
|
||||||
|
"TrackReference",
|
||||||
|
]
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class Crate:
|
||||||
|
"""A Serato crate and the track references parsed from it."""
|
||||||
|
|
||||||
|
path: Path
|
||||||
|
references: Tuple[TrackReference, ...]
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Optional
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class HealthReport:
|
||||||
|
"""Transparent aggregate findings from a read-only library analysis."""
|
||||||
|
|
||||||
|
score: Optional[float]
|
||||||
|
total_references: int
|
||||||
|
healthy_references: int
|
||||||
|
missing_references: int
|
||||||
|
unique_missing_filenames: int
|
||||||
|
disk_tracks: int
|
||||||
|
duplicate_filename_groups: int
|
||||||
|
duplicate_files: int
|
||||||
|
unused_tracks: int
|
||||||
|
suggested_matches: int
|
||||||
|
|
||||||
|
@property
|
||||||
|
def score_basis(self) -> str:
|
||||||
|
return "Resolved crate references / total crate references"
|
||||||
@@ -0,0 +1,31 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Iterable, Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.reference import ReferenceResult, TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class Library:
|
||||||
|
"""The read-only view of crate references and audio found on disk."""
|
||||||
|
|
||||||
|
references: Tuple[TrackReference, ...]
|
||||||
|
tracks: Tuple[DiskTrack, ...]
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def build(
|
||||||
|
cls,
|
||||||
|
references: Iterable[TrackReference],
|
||||||
|
tracks: Iterable[DiskTrack],
|
||||||
|
) -> "Library":
|
||||||
|
return cls(tuple(references), tuple(tracks))
|
||||||
|
|
||||||
|
def reconcile_by_filename(self) -> Tuple[ReferenceResult, ...]:
|
||||||
|
disk_names = {track.filename for track in self.tracks}
|
||||||
|
return tuple(
|
||||||
|
ReferenceResult(
|
||||||
|
reference=reference,
|
||||||
|
exists_by_filename=reference.filename in disk_names,
|
||||||
|
)
|
||||||
|
for reference in self.references
|
||||||
|
)
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Tuple
|
||||||
|
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class MatchEvidence:
|
||||||
|
"""One explainable scoring decision for a candidate track."""
|
||||||
|
|
||||||
|
field: str
|
||||||
|
matched: bool
|
||||||
|
points: int
|
||||||
|
max_points: int
|
||||||
|
explanation: str
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class TrackMatch:
|
||||||
|
"""A ranked candidate backed by explicit, inspectable evidence."""
|
||||||
|
|
||||||
|
reference: TrackReference
|
||||||
|
track: DiskTrack
|
||||||
|
evidence: Tuple[MatchEvidence, ...]
|
||||||
|
|
||||||
|
@property
|
||||||
|
def score(self) -> int:
|
||||||
|
return sum(item.points for item in self.evidence)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def max_score(self) -> int:
|
||||||
|
return sum(item.max_points for item in self.evidence)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def score_percent(self) -> float:
|
||||||
|
if not self.max_score:
|
||||||
|
return 0.0
|
||||||
|
return round(self.score / self.max_score * 100, 1)
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class TrackReference:
|
||||||
|
"""A track path referenced by a Serato crate."""
|
||||||
|
|
||||||
|
source: Path
|
||||||
|
path: Path
|
||||||
|
filename: str
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ReferenceResult:
|
||||||
|
"""The filename-level reconciliation result for a crate reference."""
|
||||||
|
|
||||||
|
reference: TrackReference
|
||||||
|
exists_by_filename: bool
|
||||||
|
|
||||||
|
def as_row(self) -> dict:
|
||||||
|
return {
|
||||||
|
"crate": str(self.reference.source),
|
||||||
|
"serato_path": str(self.reference.path),
|
||||||
|
"filename": self.reference.filename,
|
||||||
|
"exists_by_filename": self.exists_by_filename,
|
||||||
|
}
|
||||||
@@ -2,15 +2,10 @@ from dataclasses import dataclass
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
|
||||||
class TrackReference:
|
|
||||||
source: Path
|
|
||||||
path: Path
|
|
||||||
filename: str
|
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
class DiskTrack:
|
class DiskTrack:
|
||||||
|
"""An audio file discovered on disk."""
|
||||||
|
|
||||||
path: Path
|
path: Path
|
||||||
filename: str
|
filename: str
|
||||||
size: int
|
size: int
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
from collections import Counter, defaultdict
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Iterable
|
||||||
|
import csv
|
||||||
|
|
||||||
|
from serato_doctor.models.reference import ReferenceResult
|
||||||
|
|
||||||
|
|
||||||
|
def write_missing_report(results: Iterable[ReferenceResult], out: Path) -> None:
|
||||||
|
rows = [result.as_row() for result in results]
|
||||||
|
missing = [r for r in rows if not r["exists_by_filename"]]
|
||||||
|
|
||||||
|
crate_counts = Counter(r["crate"] for r in missing)
|
||||||
|
filename_counts = Counter(r["filename"] for r in missing)
|
||||||
|
|
||||||
|
with out.open("w", encoding="utf-8") as f:
|
||||||
|
f.write("# Serato Doctor Missing Report\n\n")
|
||||||
|
f.write(f"Total missing references: {len(missing)}\n\n")
|
||||||
|
|
||||||
|
f.write("## Missing by crate\n\n")
|
||||||
|
for crate, count in crate_counts.most_common():
|
||||||
|
f.write(f"{count:5} {crate}\n")
|
||||||
|
|
||||||
|
f.write("\n## Most common missing filenames\n\n")
|
||||||
|
for filename, count in filename_counts.most_common(100):
|
||||||
|
f.write(f"{count:5} {filename}\n")
|
||||||
|
|
||||||
|
f.write("\n## Detail\n\n")
|
||||||
|
by_crate = defaultdict(list)
|
||||||
|
for r in missing:
|
||||||
|
by_crate[r["crate"]].append(r["filename"])
|
||||||
|
|
||||||
|
for crate, names in sorted(by_crate.items()):
|
||||||
|
f.write(f"\n### {crate}\n")
|
||||||
|
for name in sorted(set(names)):
|
||||||
|
f.write(f"- {name}\n")
|
||||||
|
|
||||||
|
|
||||||
|
def write_csv(results: Iterable[ReferenceResult], out: Path) -> None:
|
||||||
|
rows = [result.as_row() for result in results]
|
||||||
|
with out.open("w", newline="", encoding="utf-8") as f:
|
||||||
|
writer = csv.DictWriter(
|
||||||
|
f,
|
||||||
|
fieldnames=["crate", "serato_path", "filename", "exists_by_filename"],
|
||||||
|
)
|
||||||
|
writer.writeheader()
|
||||||
|
writer.writerows(rows)
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from serato_doctor.models import DiskTrack
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
AUDIO_SUFFIXES = {".mp3", ".m4a", ".wav", ".aif", ".aiff", ".flac"}
|
AUDIO_SUFFIXES = {".mp3", ".m4a", ".wav", ".aif", ".aiff", ".flac"}
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,86 @@
|
|||||||
|
import csv
|
||||||
|
import sys
|
||||||
|
|
||||||
|
from serato_doctor.cli import main
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_writes_reports_and_prints_counts(tmp_path, monkeypatch, capsys):
|
||||||
|
serato = tmp_path / "serato"
|
||||||
|
subcrates = serato / "Subcrates"
|
||||||
|
music = tmp_path / "music"
|
||||||
|
subcrates.mkdir(parents=True)
|
||||||
|
music.mkdir()
|
||||||
|
|
||||||
|
crate_text = (
|
||||||
|
"Users/sample-user/OneDrive/Jukebox/Found.mp漳牴k"
|
||||||
|
"Users/sample-user/OneDrive/Jukebox/Missing.mp漳牴k"
|
||||||
|
)
|
||||||
|
(subcrates / "Test.crate").write_bytes(crate_text.encode("utf-16-le"))
|
||||||
|
(music / "Found.mp3").write_bytes(b"synthetic audio")
|
||||||
|
csv_path = tmp_path / "scan.csv"
|
||||||
|
report_path = tmp_path / "report.txt"
|
||||||
|
monkeypatch.setattr(
|
||||||
|
sys,
|
||||||
|
"argv",
|
||||||
|
[
|
||||||
|
"serato-doctor",
|
||||||
|
"--serato",
|
||||||
|
str(serato),
|
||||||
|
"--music",
|
||||||
|
str(music),
|
||||||
|
"--out",
|
||||||
|
str(csv_path),
|
||||||
|
"--report",
|
||||||
|
str(report_path),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
main()
|
||||||
|
|
||||||
|
captured = capsys.readouterr()
|
||||||
|
output = captured.out
|
||||||
|
assert captured.err == ""
|
||||||
|
assert "Crate references: 2" in output
|
||||||
|
assert "Disk tracks: 1" in output
|
||||||
|
assert "Missing by filename: 1" in output
|
||||||
|
assert f"CSV: {csv_path}" in output
|
||||||
|
assert f"Report: {report_path}" in output
|
||||||
|
with csv_path.open(newline="", encoding="utf-8") as csv_file:
|
||||||
|
rows = list(csv.DictReader(csv_file))
|
||||||
|
assert [row["exists_by_filename"] for row in rows] == ["True", "False"]
|
||||||
|
assert "Total missing references: 1" in report_path.read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def test_cli_accepts_old_reference_root(tmp_path, monkeypatch, capsys):
|
||||||
|
serato = tmp_path / "serato"
|
||||||
|
subcrates = serato / "Subcrates"
|
||||||
|
music = tmp_path / "music"
|
||||||
|
subcrates.mkdir(parents=True)
|
||||||
|
music.mkdir()
|
||||||
|
(subcrates / "Test.crate").write_bytes(
|
||||||
|
"Archive/Jukebox/Found.mp3otrk".encode("utf-16-le")
|
||||||
|
)
|
||||||
|
(music / "Found.mp3").write_bytes(b"synthetic audio")
|
||||||
|
monkeypatch.setattr(
|
||||||
|
sys,
|
||||||
|
"argv",
|
||||||
|
[
|
||||||
|
"serato-doctor",
|
||||||
|
"--serato",
|
||||||
|
str(serato),
|
||||||
|
"--music",
|
||||||
|
str(music),
|
||||||
|
"--out",
|
||||||
|
str(tmp_path / "scan.csv"),
|
||||||
|
"--report",
|
||||||
|
str(tmp_path / "report.txt"),
|
||||||
|
"--reference-root",
|
||||||
|
"/Archive/Jukebox",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
main()
|
||||||
|
|
||||||
|
output = capsys.readouterr().out
|
||||||
|
assert "Crate references: 1" in output
|
||||||
|
assert "Missing by filename: 0" in output
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.config import ScanConfig
|
||||||
|
|
||||||
|
|
||||||
|
def test_scan_config_freezes_reference_roots():
|
||||||
|
roots = (Path(root) for root in ["/old/one", "/old/two"])
|
||||||
|
|
||||||
|
config = ScanConfig.build(
|
||||||
|
serato=Path("serato"),
|
||||||
|
music=Path("music"),
|
||||||
|
out=Path("scan.csv"),
|
||||||
|
report=Path("report.txt"),
|
||||||
|
reference_roots=roots,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert config.reference_roots == (Path("/old/one"), Path("/old/two"))
|
||||||
@@ -0,0 +1,86 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from serato_doctor.crate_parser import (
|
||||||
|
clean_path,
|
||||||
|
load_crate,
|
||||||
|
parse_crate,
|
||||||
|
parse_crates,
|
||||||
|
path_markers,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("artifact", "expected"),
|
||||||
|
[
|
||||||
|
("song.mp漳", "song.mp3"),
|
||||||
|
("song.m4愠", "song.m4a"),
|
||||||
|
("song.wa瘠", "song.wav"),
|
||||||
|
("song.ai映", "song.aif"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_clean_path_repairs_known_extension_artifacts(artifact, expected):
|
||||||
|
assert clean_path(artifact) == expected
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_crate_extracts_references(tmp_path):
|
||||||
|
crate_path = tmp_path / "House.crate"
|
||||||
|
crate_text = (
|
||||||
|
"header"
|
||||||
|
"Users/sample-user/OneDrive/Jukebox/House/First.mp漳牴k"
|
||||||
|
"metadata"
|
||||||
|
"Users/sample-user/OneDrive/Jukebox/House/Second.m4愠otrk"
|
||||||
|
)
|
||||||
|
crate_path.write_bytes(crate_text.encode("utf-16-le"))
|
||||||
|
|
||||||
|
crate = load_crate(crate_path)
|
||||||
|
|
||||||
|
assert crate.path == crate_path
|
||||||
|
assert [reference.filename for reference in crate.references] == [
|
||||||
|
"First.mp3",
|
||||||
|
"Second.m4a",
|
||||||
|
]
|
||||||
|
assert parse_crate(crate_path) == list(crate.references)
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_crate_ignores_record_without_stop_marker(tmp_path):
|
||||||
|
crate_path = Path(tmp_path) / "Incomplete.crate"
|
||||||
|
crate_path.write_bytes(
|
||||||
|
"Users/sample-user/OneDrive/Jukebox/House/Incomplete.mp3".encode(
|
||||||
|
"utf-16-le"
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert parse_crate(crate_path) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_crate_uses_configured_reference_root(tmp_path):
|
||||||
|
crate_path = tmp_path / "Migrated.crate"
|
||||||
|
crate_path.write_bytes(
|
||||||
|
"Archive/Old Library/House/Track.mp3otrk".encode("utf-16-le")
|
||||||
|
)
|
||||||
|
|
||||||
|
references = parse_crate(crate_path, [Path("/Archive/Old Library")])
|
||||||
|
|
||||||
|
assert references[0].path == Path("/Archive/Old Library/House/Track.mp3")
|
||||||
|
|
||||||
|
|
||||||
|
def test_default_path_markers_do_not_contain_a_username():
|
||||||
|
assert path_markers([]) == ("Users/", "Volumes/")
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_crates_reuses_configured_roots_for_every_crate(tmp_path):
|
||||||
|
for name in ("First", "Second"):
|
||||||
|
(tmp_path / f"{name}.crate").write_bytes(
|
||||||
|
f"Archive/Jukebox/{name}.mp3otrk".encode("utf-16-le")
|
||||||
|
)
|
||||||
|
|
||||||
|
references = parse_crates(
|
||||||
|
tmp_path, (root for root in [Path("/Archive/Jukebox")])
|
||||||
|
)
|
||||||
|
|
||||||
|
assert {reference.filename for reference in references} == {
|
||||||
|
"First.mp3",
|
||||||
|
"Second.mp3",
|
||||||
|
}
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.health import analyze_health
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def reference(filename):
|
||||||
|
return TrackReference(
|
||||||
|
Path("Test.crate"), Path("/old/House") / filename, filename
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def track(filename, folder="House"):
|
||||||
|
path = Path("/new") / folder / filename
|
||||||
|
return DiskTrack(path, filename, 100, path.suffix.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def test_health_report_exposes_each_metric():
|
||||||
|
library = Library.build(
|
||||||
|
[
|
||||||
|
reference("Found.mp3"),
|
||||||
|
reference("Conflict.mp3"),
|
||||||
|
reference("Absent.mp3"),
|
||||||
|
],
|
||||||
|
[
|
||||||
|
track("Found.mp3"),
|
||||||
|
track("Conflict 2.mp3"),
|
||||||
|
track("Conflict 2.mp3", "Backup"),
|
||||||
|
track("Unused.mp3"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
report = analyze_health(library)
|
||||||
|
|
||||||
|
assert report.score == 33.3
|
||||||
|
assert report.score_basis == (
|
||||||
|
"Resolved crate references / total crate references"
|
||||||
|
)
|
||||||
|
assert report.total_references == 3
|
||||||
|
assert report.healthy_references == 1
|
||||||
|
assert report.missing_references == 2
|
||||||
|
assert report.unique_missing_filenames == 2
|
||||||
|
assert report.disk_tracks == 4
|
||||||
|
assert report.duplicate_filename_groups == 1
|
||||||
|
assert report.duplicate_files == 1
|
||||||
|
assert report.unused_tracks == 3
|
||||||
|
assert report.suggested_matches == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_empty_library_has_no_health_score():
|
||||||
|
report = analyze_health(Library.build([], []))
|
||||||
|
|
||||||
|
assert report.score is None
|
||||||
|
assert report.total_references == 0
|
||||||
@@ -0,0 +1,24 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def test_library_reconciles_references_by_filename():
|
||||||
|
crate = Path("House.crate")
|
||||||
|
references = [
|
||||||
|
TrackReference(crate, Path("/old/Found.mp3"), "Found.mp3"),
|
||||||
|
TrackReference(crate, Path("/old/Missing.mp3"), "Missing.mp3"),
|
||||||
|
]
|
||||||
|
tracks = [DiskTrack(Path("/new/Found.mp3"), "Found.mp3", 10, ".mp3")]
|
||||||
|
|
||||||
|
results = Library.build(references, tracks).reconcile_by_filename()
|
||||||
|
|
||||||
|
assert [result.exists_by_filename for result in results] == [True, False]
|
||||||
|
assert results[1].as_row() == {
|
||||||
|
"crate": "House.crate",
|
||||||
|
"serato_path": "/old/Missing.mp3",
|
||||||
|
"filename": "Missing.mp3",
|
||||||
|
"exists_by_filename": False,
|
||||||
|
}
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
import logging
|
||||||
|
|
||||||
|
from serato_doctor.logging import LOGGER_NAME, configure_logging
|
||||||
|
|
||||||
|
|
||||||
|
def test_logging_is_quiet_by_default(capsys):
|
||||||
|
logger = configure_logging()
|
||||||
|
|
||||||
|
logger.info("not visible")
|
||||||
|
|
||||||
|
assert capsys.readouterr().err == ""
|
||||||
|
|
||||||
|
|
||||||
|
def test_logging_writes_aggregate_progress_to_file(tmp_path):
|
||||||
|
log_path = tmp_path / "scan.log"
|
||||||
|
logger = configure_logging(log_file=log_path)
|
||||||
|
|
||||||
|
logger.info("Scanned %d disk tracks", 10)
|
||||||
|
|
||||||
|
contents = log_path.read_text(encoding="utf-8")
|
||||||
|
assert "INFO Scanned 10 disk tracks" in contents
|
||||||
|
|
||||||
|
|
||||||
|
def test_reconfiguring_logging_replaces_handlers():
|
||||||
|
configure_logging(verbose=True)
|
||||||
|
logger = configure_logging(verbose=True)
|
||||||
|
|
||||||
|
active_handlers = [
|
||||||
|
handler
|
||||||
|
for handler in logger.handlers
|
||||||
|
if not isinstance(handler, logging.NullHandler)
|
||||||
|
]
|
||||||
|
assert logger.name == LOGGER_NAME
|
||||||
|
assert len(active_handlers) == 1
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.matching import MatchingEngine, score_candidate
|
||||||
|
from serato_doctor.models.reference import TrackReference
|
||||||
|
from serato_doctor.models.track import DiskTrack
|
||||||
|
|
||||||
|
|
||||||
|
def reference(filename, folder="House"):
|
||||||
|
return TrackReference(
|
||||||
|
Path("Test.crate"), Path("/old") / folder / filename, filename
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def track(filename, folder="House"):
|
||||||
|
path = Path("/new") / folder / filename
|
||||||
|
return DiskTrack(path, filename, 100, path.suffix.lower())
|
||||||
|
|
||||||
|
|
||||||
|
def test_exact_candidate_has_full_evidence_score():
|
||||||
|
match = score_candidate(reference("Track.mp3"), track("Track.mp3"))
|
||||||
|
|
||||||
|
assert match.score == 90
|
||||||
|
assert match.max_score == 90
|
||||||
|
assert match.score_percent == 100.0
|
||||||
|
assert all(item.matched for item in match.evidence)
|
||||||
|
|
||||||
|
|
||||||
|
def test_cloud_conflict_suffix_is_explained():
|
||||||
|
match = score_candidate(reference("Track.mp3"), track("Track 2.mp3"))
|
||||||
|
|
||||||
|
assert match.score == 80
|
||||||
|
assert match.score_percent == 88.9
|
||||||
|
assert match.evidence[0].explanation == (
|
||||||
|
"Filename matches after removing a numeric conflict suffix"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_case_normalized_match_scores_below_exact():
|
||||||
|
match = score_candidate(reference("TRACK.MP3"), track("track.mp3"))
|
||||||
|
|
||||||
|
assert match.score == 85
|
||||||
|
assert match.evidence[0].points == 55
|
||||||
|
|
||||||
|
|
||||||
|
def test_engine_omits_unrelated_filenames():
|
||||||
|
engine = MatchingEngine([track("Different.mp3")])
|
||||||
|
|
||||||
|
assert engine.candidates_for(reference("Missing.mp3")) == ()
|
||||||
|
|
||||||
|
|
||||||
|
def test_ambiguous_candidates_are_ranked_deterministically():
|
||||||
|
engine = MatchingEngine(
|
||||||
|
[track("Track 2.mp3", "Other"), track("Track 3.mp3", "House")]
|
||||||
|
)
|
||||||
|
|
||||||
|
matches = engine.candidates_for(reference("Track.mp3"))
|
||||||
|
|
||||||
|
assert [match.track.filename for match in matches] == [
|
||||||
|
"Track 3.mp3",
|
||||||
|
"Track 2.mp3",
|
||||||
|
]
|
||||||
|
assert [match.score for match in matches] == [80, 60]
|
||||||
@@ -0,0 +1,28 @@
|
|||||||
|
import runpy
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from serato_doctor.crate_parser import parse_crates
|
||||||
|
from serato_doctor.models.library import Library
|
||||||
|
from serato_doctor.scanner import scan_audio
|
||||||
|
|
||||||
|
|
||||||
|
def test_generated_sample_library_has_expected_scenario(tmp_path):
|
||||||
|
generator_path = (
|
||||||
|
Path(__file__).parents[1] / "samples" / "small-library" / "generate.py"
|
||||||
|
)
|
||||||
|
build_sample = runpy.run_path(str(generator_path))["build_sample"]
|
||||||
|
sample_root = build_sample(tmp_path / "sample")
|
||||||
|
|
||||||
|
references = parse_crates(sample_root / "Serato" / "_Serato_" / "Subcrates")
|
||||||
|
tracks = scan_audio(sample_root / "Music")
|
||||||
|
results = Library.build(references, tracks).reconcile_by_filename()
|
||||||
|
missing = {
|
||||||
|
result.reference.filename
|
||||||
|
for result in results
|
||||||
|
if not result.exists_by_filename
|
||||||
|
}
|
||||||
|
|
||||||
|
assert len(list((sample_root / "Serato").rglob("*.crate"))) == 7
|
||||||
|
assert len(references) == 13
|
||||||
|
assert len(tracks) == 10
|
||||||
|
assert missing == {"Missing.mp3", "Old Name.mp3"}
|
||||||
@@ -0,0 +1,17 @@
|
|||||||
|
from serato_doctor.scanner import scan_audio
|
||||||
|
|
||||||
|
|
||||||
|
def test_scan_audio_finds_supported_files(tmp_path):
|
||||||
|
music = tmp_path / "music"
|
||||||
|
nested = music / "House"
|
||||||
|
nested.mkdir(parents=True)
|
||||||
|
(nested / "First.MP3").write_bytes(b"synthetic audio")
|
||||||
|
(nested / "Second.flac").write_bytes(b"fixture")
|
||||||
|
(nested / "notes.txt").write_text("not audio", encoding="utf-8")
|
||||||
|
|
||||||
|
tracks = scan_audio(music)
|
||||||
|
|
||||||
|
assert {track.filename for track in tracks} == {"First.MP3", "Second.flac"}
|
||||||
|
first = next(track for track in tracks if track.filename == "First.MP3")
|
||||||
|
assert first.suffix == ".mp3"
|
||||||
|
assert first.size == len(b"synthetic audio")
|
||||||
Reference in New Issue
Block a user