Zones of Distinguishability

Artists and philosophers have always been drawn to zones of indistinguishability, i.e. the places where one thing shades into another and the boundary quietly dissolves. Music is full of them. When does a remix stop being the recording it started from and become a new work? When is a cover “the same song,” and when is it something genuinely different? Is a track that merely feels like another one a copy, or just a cousin? These are wonderful, slippery questions, and the honest answer is that music lives in the blur. There is no line painted on the ground where sameness ends and difference begins.

But if you run a digital media platform, a rights organization, or a label, you don’t get to enjoy the blur. You have to make the call. Is this upload the original master? A sped-up edit? A live version? A cover by an unknown artist? An AI clone of a voice? Something entirely unrelated that only sounds similar? Attribution, licensing, monetization, and takedowns all depend on a clear, defensible answer at a scale of billions of files.

For more than two decades, the music and technology community has asked Audible Magic to do one deceptively simple thing: turn zones of indistinguishability into zones of distinguishability. To draw a clean, reliable boundary through a continuum that, left to itself, has none. That is our whole job, and it has become more interesting every single year.

First, what do you mean by “the same”?

Before you can identify anything, you have to decide which kind of “same” you care about, because a song is not one object. There can be many choices of axes, but here are four, each of which can vary on their own:

The recording:  the specific master: this exact performance captured in this exact studio.
The composition:  the underlying work: the melody, harmony, and structure, no matter who plays it.
The lyrics: the words, which can be translated, rewritten, or dropped entirely.
The voice: the identity of the performer, which used to be welded to the recording and is now, in the age of AI, floating free.

Almost every “hard case” becomes easy to describe once you separate these. A karaoke track keeps the composition but removes the lead voice. A translation keeps the music and swaps the words. An instrumental keeps everything but the vocal. An AI voice clone keeps a voice and attaches it to a brand-new song. Each is a different problem, and you can explore the whole space below. Drag to rotate it; hover any point; and use the menus to remap the axes and see, for instance, how the karaoke, instrumental, and translation cases separate along the lyric dimension they quietly share.

The space of “sameness.” Three axes are spatial – how much of the recording, the composition, and the original voice survives – while color marks the voice’s provenance (genuine, cloned, or other) and shape marks whether the lyrics are kept. Drag to rotate, scroll to zoom, hover for a name, and remap any axis.

Each of these problems has needed a different generation of technology to solve. Here is how that story unfolded.

It starts with the master: Audio ID

The most familiar kind of identification is a one-to-one match to a specific recording. You’ve done it a hundred times: hold up your phone in a bar, and a few seconds later you have the title and artist. Our Audio ID has powered this kind of recognition for the B2B world for over twenty years, using compact acoustic fingerprints – in our case built on MFCCs, much as Shazam’s consumer app uses spectral peaks – to match an incoming file against a reference catalog.

Crucially, Audio ID was built from the start to survive the ordinary indignities of real-world audio: ambient noise, EQ, compression artifacts, channel effects, and small shifts in pitch or tempo of a few percent. This was the technology that grew up alongside Napster, the DMCA, the explosion of user-generated content, and is the tool that let platforms reduce takedown notices and administer licenses instead of fighting fires by hand. On this near shore, the boundary is genuinely clean, which is exactly why almost no one argues about it.

When the master gets bent: Broad Spectrum

Then listeners and creators started reshaping recordings on purpose. Speeding a track up turns a ballad into a dance number and hits the hook faster: the entire sped-up trend that took over TikTok, so pervasive that labels now release official sped-up editions. Slowing down gave us DJ Screw’s chopped-and-screwed hip-hop; speeding up even more gave us Nightcore. DJs beat-match and pitch-match neighboring tracks as a matter of course, and some uploaders nudge the key or tempo specifically to slip past identification systems. In our own data, as much as half of the music on UGC platforms is transformed somewhat in pitch and/or tempo, and a meaningful share is altered by 20% or more.

Small changes were always within Audio ID’s tolerance. For the rest – the heavy transforms, the deliberate evasions – we built Broad Spectrum, which identifies recordings that have been stretched, shifted, and reshaped far beyond what any conventional matcher can follow. The boundary of “the same recording” had moved, so we followed it out.

When there is no master at all: Version ID

The hardest leap comes when the original recording simply isn’t there anymore. A live performance, a cover in a completely different style, a parody, an AI-cloned voice.  None of these shares a single second of audio with the master. What connects them is the composition and, if any, the lyrics. Matching them requires musical and linguistic understanding, not signal comparison: recognizing a tune in a new key, at a new tempo, on new instruments, with new singers or none at all.

This is what Version ID does. It is a capability we introduced last year, built on two foundational patents: one that works from melody, harmony, sound, and structure, and one that works from lyrics and phonemes, combined so each can catch what the other misses. The range is remarkable, and worth hearing in the wild:

An AI voice clone of Johnny Cash singing Aqua’s Barbie Girl recast as a Folsom-Prison-style country tune. The melody and genre are transformed and the voice is synthetic, yet the lyrics and structure tie it straight back to the original.
A parody like Weird Al’s Amish Paradise: the words are entirely new, but the music is deliberately intact, so Coolio’s Gangsta’s Paradise comes through on melody, chords, and structure alone.
A purely instrumental arrangement such Classern Quartet’s version of Doja Cat’s Say So with no voice and no words, is matched by its harmony, melody, and structure.
An interpolation, where a new song quietly borrows an older one: Ariana Grande’s 7 rings is built on the melody of My Favorite Things. Version ID reports the timing, so you can see exactly which section is borrowed.
A radically reworked live version, like the kind Bob Dylan is famous for, where an artist reinvents their own song on stage until it’s barely recognizable, is still anchored by enough melody, structure, or lyric to connect it home.

Covers, live sets, and AI versions are no longer a niche. They are an onslaught, and every one of them is a rights and attribution event.

The new Version ID: powered by VIBE

Today we’re announcing the next generation of Version ID, and it is a substantial step forward.

At its core is VIBE (Version Identification By Embedding) our neural-network embedding model. Rather than hand-designed features, VIBE learns to place every track as a point in a high-dimensional space arranged so that versions of the same work land near one another, regardless of key, tempo, arrangement, instrumentation, or singer, while unrelated music lands far away. Identification becomes a search for near neighbors in that space. The practical gains are exactly the ones that matter:

A much broader reference catalog, so far more works are covered and far more versions are caught.
Much faster matching, which is what makes it viable at platform transaction volumes.
Much higher accuracy meaning more true versions found, and fewer false matches to clean up.

More recall means fewer covers and clones slipping through unattributed. More precision means fewer legitimate uploads wrongly flagged. Both translate directly into proper attribution for the rights holders and creators being treated according to the rights picture that actually applies.

A note on philosophy

It’s worth saying plainly why this problem never simply “gets solved.” Similarity does not behave like equality. Equality is transitive: if A equals B and B equals C, then A equals C. Similarity isn’t. This is a fact that Poincaré noted long ago: you can be unable to tell A from B, or B from C, and yet tell A from C at a glance. A remaster is nearly identical to its master; a remix is nearly identical to the remaster; a cover of the remix is nearly identical to that, and by the end of the chain you arrive at something that shares almost nothing with where you started. Each step is small; the sum is a stranger. It’s the old puzzle of Theseus’ ship, whose planks are replaced one by one until none of the original remains: at what point is it no longer the same ship? Music asks that question millions of times a day.

A brief aside for the curious: the very idea of a musical “work”, i.e. a fixed, ownable object that exists apart from any performance, is surprisingly recent. The philosopher Lydia Goehr traces it to around 1800, the era of Beethoven, and copyright law grew up around that idea and inherited its assumptions. But most of the world’s music has never behaved that way: a raga, a maqam, a gamelan piece, a jazz solo lives in the performance, not on a page. Wilhelm von Humboldt’s old distinction fits well here: some music is ergon, a finished work; much of it is energeia, an activity. Even modern pop, typically composed in the studio as a recording rather than written as a score, has to be gently shoehorned into the work-shaped box the law provides. Identification is, in part, the ongoing negotiation between how music actually behaves and the categories we’ve inherited.

Which means there is no natural boundary waiting to be discovered.  There is only a boundary to be drawn. Every identification system chooses an operating point: reach further and catch more versions at the cost of the occasional false match, or stay strict and miss a few to stay clean. That choice looks like a technical threshold, but it is really a business and editorial decision about whose content gets flagged and whose royalties flow. The value we provide is not a fantasy of perfect certainty; it’s a trustworthy dial, set where you need it, backed by the best accuracy in the field.

And the frontier keeps moving. Generative AI has begun to invert the question itself, from “is this upload a copy of a recording?” to “was this model trained on our catalog?” and “is that the artist, or a convincing clone of a voice that copyright was never designed to protect?” Increasingly, identification is becoming a question of provenance: not only what is this? but where did it come from? That is the next zone of indistinguishability, and it’s the one we’re building for now.


For more than 20 years, Audible Magic has delivered B2B solutions in content identification, for both music and TV/film, along with music catalog fulfillment and rights administration. Our systems power billions of transactions every month for customers around the world, helping platforms reduce DMCA takedown notices and administer licenses for user-generated content. The new VIBE-powered Version ID is the latest step in that work. We’d love to hear about your use case and figure out how to solve it together. Contact us to learn more.