Add duplicate audio hash detection
This commit is contained in:
@@ -0,0 +1,34 @@
|
||||
# Duplicate audio hashing
|
||||
|
||||
## Problem
|
||||
|
||||
Matching filenames do not prove matching content. Two files can be different DJ
|
||||
edits, masters, encodes, or entirely different tracks. Removing either without
|
||||
stronger evidence is unsafe.
|
||||
|
||||
## Design
|
||||
|
||||
Serato Doctor computes SHA-256 only for the duplicate group currently being
|
||||
reviewed. This avoids hashing an entire 22,000-track library during every scan.
|
||||
The review queue labels a group as byte-for-byte identical or warns that its
|
||||
files differ. File size and a short fingerprint remain available as supporting
|
||||
detail.
|
||||
|
||||
The final batch dry-run recomputes every approved comparison server-side and
|
||||
includes the status beside each keeper decision. The hash is evidence, not an
|
||||
automatic repair decision; the DJ remains in control.
|
||||
|
||||
## Edge cases
|
||||
|
||||
- Identical audio stored under different filenames is recognized.
|
||||
- Metadata changes inside an audio container produce a different byte hash and
|
||||
therefore a conservative warning.
|
||||
- Missing or unreadable files fail verification rather than being treated as
|
||||
identical.
|
||||
- Hashing is streamed in chunks rather than loading whole tracks into memory.
|
||||
|
||||
## Tests
|
||||
|
||||
- Equal byte content produces identical SHA-256 fingerprints.
|
||||
- Different content is labeled different.
|
||||
- Batch previews carry the server-computed confidence status.
|
||||
Reference in New Issue
Block a user