The Leakage Ledger, Vol. 1
- Henry Marsden

- Aug 12
- 5 min read

I write a lot about data quality in music rights. I tend to deliberately keep this high level- the issues so often are sustained by political inertia more than lack of technology options to solve. A side effect of this is it creates intangibles- arguments, frameworks and principles rather than practical advice from living it out in the trenches. This is the start of something different to address that.
The Leakage Ledger is a recurring edition where I’ll publish actual numbers from actual catalog work. Anonymised (obviously) and stripped of anything identifying- but real. One catalog, the findings, and what they were worth. No thesis to defend, just the pure maths (= math for my US friends).
The reason is straightforward- the industry has become comfortable talking about "messy data" as a general condition, a bit like the weather- universally acknowledged, nobody's fault but nothing to be done today. Tangible numbers- those that you can see, touch and feel, change that conversation. A publisher can't benchmark themselves against a sentiment, but they can absolutely benchmark themselves against a percentage.
So here is Volume 1.
The Catalog
A publishing catalog of 100,000+ works. Established business, with an experienced team and functioning systems- and, importantly for what follows, not a business anyone would describe as badly run. Every finding below came from a catalog with experienced owners who were already taking proactive steps to cleanup data.
Nobody commissions cleanup work expecting a true disaster. They commission it because something doesn't quite reconcile, or because a transaction is coming, or increasingly because external forces like the 2027 MLC market share distributions bring focus to efforts. The findings then tend to be a surprise- not because the team was careless, but because too often these problems are structurally invisible from the inside, or too vast in volume to tackle economically without tooling. Your own system will happily tell you what you own- but it doesn’t tend to know what the outside world thinks you own.
Finding 1: Duplicate Registrations
Just under 11% of the catalog existed as duplicate registrations at the MLC.
Not near-duplicates in the philosophical sense that keeps metadata people up at night- but the same underlying composition, present more than once, with royalties and matches distributed across the copies.
Duplicates are the most underrated problem in publishing data because their damage is indirect. A duplicate doesn't announce itself as lost revenue, but rather splits the work's identity, so recordings match to one instance while claims sit on another. It fragments the claim and ownership history. It creates situations where a work looks correctly registered from one angle and incompletely registered from another- with both views seemingly ‘accurate’.
They are also, awkwardly, mostly self-inflicted by industry level infrastructure and working practices. Registrations sent by different parties at different times with partial information, migrations that ran twice, versions and alternative titles registered as new works. As we discussed in The Metadata Iceberg, the work layer is where this all compounds- and often it isn’t looked at until something downstream breaks.
Finding 2: Works underclaimed at the MLC
3% of the catalog didn’t have the same % claim registered as measured against the client's own internal record of what they controlled.
Though some were overclaims, most of these were underclaims. Not missing share %s, but shares where the amount registered was lower than it should have been- sometimes by a lot, and sometimes on significant and high value works.
This figure specifically excludes missing shares (‘unclaimed’ in MLC parlance)- works where claimed shares do not add up to 100%. That's a separate, larger, messier exercise- and not necessarily a revenue issue for the catalog of interest.
The only reason these sorts of issues persist is that the reconciliation between internal truth and the external registrations is tricky- hampered either by access to 3rd party data (at scale), or by the ability to generate the required export from internal tooling. This is typically either because of a system bottle neck, or lack of comprehensive and trusted works data in the system in the first place.
Together, findings 1 and 2 meant 11% of the entire catalog was carrying material work-level issues- and this only at one society.
Finding 3: Over half of works with no ISRCs matched at all
10,000s of works had no recordings linked to them at the MLC. Not partially matched. Not matched to the wrong recording. No ISRCs attached whatsoever. In practical terms this means those works were earning nothing from US digital mechanicals, quietly, while sitting in a copyright database as apparently productive assets.
Some of that is legitimate. Every catalog contains works that were never commercially released, or whose recordings genuinely don't stream. The long tail is real and some of it is genuinely inert.
But 50%+ is more than inertness. It's the recording-work link being a critical and yet fragile joint in the entire royalty chain. We all know the history here- publishers manage works, labels manage recordings and the link between them… well, that’s for another newsletter.
This also echoes precisely what the public MLC data showed when we took a look back in 2025: only 22% of ISRCs in the entire dataset is attributed to a work. The industry-level number and the catalog-level number are telling the same story, but from opposite ends.
What's It All Worth?
Across the cleanup and claiming work we do, a minimum 10-15% uplift in collections is the general pattern. Remember this isn’t from new repertoire, better deals or harder negotiation- it’s just from money that was always owed finally finding its way home.
For another view- a large PRO recently told us the client’s revenue had doubled over the period we were doing cleanup work on their catalog. To their credit, they did note that not all of that could be attributed directly to our work. Fair- other things were happening in the market at the time, but nobody involved thought the two were unrelated either.
The honest framing is this: the uplift is real, it is material, and it is almost entirely recovery rather than growth. A strange thing for an industry to leave on the table for as long as it has (again, a story for another time).
The Point of Publishing a Leakage Ledger
Three reasons.
First, benchmarking. If you run a catalog, you now have numbers to hold yours against. Is your duplicate rate above or below 10%? Do you know? Could you find out this quarter?
Second, the invisibility problem. Every one of these findings was invisible from inside a well-run business using competent systems. The issue isn't capability, it's that internal systems are designed to tell you what you own- and every one of these problems lives in the gap between what you own and what the outside world has registered.
Third, timing. From early 2027, MLC’s unclaimed/unmatched money starts being permanently redistributed on a rolling monthly basis. The work described here has always been worth doing- it just has acquired a deadline.
Data quality has been an argument in this industry for over a decade- it is time it became a number.
I'll publish the next Ledger with a different catalog and a different set of findings. In the meantime: what's your duplicate rate? What’s your number of registrations with no ISRCs matched? If you don't know, that itself is the finding.




Comments