Zones of Distinguishability

Zones of Distinguishability

Artists and philosophers have always been drawn to zones of indistinguishability, i.e. the places where one thing shades into another and the boundary quietly dissolves. Music is full of them. When does a remix stop being the recording it started from and become a new work? When is a cover “the same song,” and when is it something genuinely different? Is a track that merely feels like another one a copy, or just a cousin? These are wonderful, slippery questions, and the honest answer is that music lives in the blur. There is no line painted on the ground where sameness ends and difference begins.

But if you run a digital media platform, a rights organization, or a label, you don’t get to enjoy the blur. You have to make the call. Is this upload the original master? A sped-up edit? A live version? A cover by an unknown artist? An AI clone of a voice? Something entirely unrelated that only sounds similar? Attribution, licensing, monetization, and takedowns all depend on a clear, defensible answer at a scale of billions of files.

For more than two decades, the music and technology community has asked Audible Magic to do one deceptively simple thing: turn zones of indistinguishability into zones of distinguishability. To draw a clean, reliable boundary through a continuum that, left to itself, has none. That is our whole job, and it has become more interesting every single year.

First, what do you mean by “the same”?

Before you can identify anything, you have to decide which kind of “same” you care about, because a song is not one object. There can be many choices of axes, but here are four, each of which can vary on their own:

The recording:  the specific master: this exact performance captured in this exact studio.
The composition:  the underlying work: the melody, harmony, and structure, no matter who plays it.
The lyrics: the words, which can be translated, rewritten, or dropped entirely.
The voice: the identity of the performer, which used to be welded to the recording and is now, in the age of AI, floating free.

Almost every “hard case” becomes easy to describe once you separate these. A karaoke track keeps the composition but removes the lead voice. A translation keeps the music and swaps the words. An instrumental keeps everything but the vocal. An AI voice clone keeps a voice and attaches it to a brand-new song. Each is a different problem, and you can explore the whole space below. Drag to rotate it; hover any point; and use the menus to remap the axes and see, for instance, how the karaoke, instrumental, and translation cases separate along the lyric dimension they quietly share.

The space of “sameness.” Three axes are spatial - how much of the recording, the composition, and the original voice survives - while color marks the voice’s provenance (genuine, cloned, or other) and shape marks whether the lyrics are kept. Drag to rotate, scroll to zoom, hover for a name, and remap any axis.

Each of these problems has needed a different generation of technology to solve. Here is how that story unfolded.

It starts with the master: Audio ID

The most familiar kind of identification is a one-to-one match to a specific recording. You’ve done it a hundred times: hold up your phone in a bar, and a few seconds later you have the title and artist. Our Audio ID has powered this kind of recognition for the B2B world for over twenty years, using compact acoustic fingerprints - in our case built on MFCCs, much as Shazam’s consumer app uses spectral peaks - to match an incoming file against a reference catalog.

Crucially, Audio ID was built from the start to survive the ordinary indignities of real-world audio: ambient noise, EQ, compression artifacts, channel effects, and small shifts in pitch or tempo of a few percent. This was the technology that grew up alongside Napster, the DMCA, the explosion of user-generated content, and is the tool that let platforms reduce takedown notices and administer licenses instead of fighting fires by hand. On this near shore, the boundary is genuinely clean, which is exactly why almost no one argues about it.

When the master gets bent: Broad Spectrum

Then listeners and creators started reshaping recordings on purpose. Speeding a track up turns a ballad into a dance number and hits the hook faster: the entire sped-up trend that took over TikTok, so pervasive that labels now release official sped-up editions. Slowing down gave us DJ Screw’s chopped-and-screwed hip-hop; speeding up even more gave us Nightcore. DJs beat-match and pitch-match neighboring tracks as a matter of course, and some uploaders nudge the key or tempo specifically to slip past identification systems. In our own data, as much as half of the music on UGC platforms is transformed somewhat in pitch and/or tempo, and a meaningful share is altered by 20% or more.

Small changes were always within Audio ID’s tolerance. For the rest - the heavy transforms, the deliberate evasions - we built Broad Spectrum, which identifies recordings that have been stretched, shifted, and reshaped far beyond what any conventional matcher can follow. The boundary of “the same recording” had moved, so we followed it out.

When there is no master at all: Version ID

The hardest leap comes when the original recording simply isn’t there anymore. A live performance, a cover in a completely different style, a parody, an AI-cloned voice.  None of these shares a single second of audio with the master. What connects them is the composition and, if any, the lyrics. Matching them requires musical and linguistic understanding, not signal comparison: recognizing a tune in a new key, at a new tempo, on new instruments, with new singers or none at all.

This is what Version ID does. It is a capability we introduced last year, built on two foundational patents: one that works from melody, harmony, sound, and structure, and one that works from lyrics and phonemes, combined so each can catch what the other misses. The range is remarkable, and worth hearing in the wild:

An AI voice clone of Johnny Cash singing Aqua’s Barbie Girl recast as a Folsom-Prison-style country tune. The melody and genre are transformed and the voice is synthetic, yet the lyrics and structure tie it straight back to the original.
A parody like Weird Al’s Amish Paradise: the words are entirely new, but the music is deliberately intact, so Coolio’s Gangsta’s Paradise comes through on melody, chords, and structure alone.
A purely instrumental arrangement such Classern Quartet's version of Doja Cat's Say So with no voice and no words, is matched by its harmony, melody, and structure.
An interpolation, where a new song quietly borrows an older one: Ariana Grande’s 7 rings is built on the melody of My Favorite Things. Version ID reports the timing, so you can see exactly which section is borrowed.
A radically reworked live version, like the kind Bob Dylan is famous for, where an artist reinvents their own song on stage until it’s barely recognizable, is still anchored by enough melody, structure, or lyric to connect it home.

Covers, live sets, and AI versions are no longer a niche. They are an onslaught, and every one of them is a rights and attribution event.

The new Version ID: powered by VIBE

Today we’re announcing the next generation of Version ID, and it is a substantial step forward.

At its core is VIBE (Version Identification By Embedding) our neural-network embedding model. Rather than hand-designed features, VIBE learns to place every track as a point in a high-dimensional space arranged so that versions of the same work land near one another, regardless of key, tempo, arrangement, instrumentation, or singer, while unrelated music lands far away. Identification becomes a search for near neighbors in that space. The practical gains are exactly the ones that matter:

A much broader reference catalog, so far more works are covered and far more versions are caught.
Much faster matching, which is what makes it viable at platform transaction volumes.
Much higher accuracy meaning more true versions found, and fewer false matches to clean up.

More recall means fewer covers and clones slipping through unattributed. More precision means fewer legitimate uploads wrongly flagged. Both translate directly into proper attribution for the rights holders and creators being treated according to the rights picture that actually applies.

A note on philosophy

It’s worth saying plainly why this problem never simply “gets solved.” Similarity does not behave like equality. Equality is transitive: if A equals B and B equals C, then A equals C. Similarity isn’t. This is a fact that Poincaré noted long ago: you can be unable to tell A from B, or B from C, and yet tell A from C at a glance. A remaster is nearly identical to its master; a remix is nearly identical to the remaster; a cover of the remix is nearly identical to that, and by the end of the chain you arrive at something that shares almost nothing with where you started. Each step is small; the sum is a stranger. It’s the old puzzle of Theseus’ ship, whose planks are replaced one by one until none of the original remains: at what point is it no longer the same ship? Music asks that question millions of times a day.

A brief aside for the curious: the very idea of a musical “work”, i.e. a fixed, ownable object that exists apart from any performance, is surprisingly recent. The philosopher Lydia Goehr traces it to around 1800, the era of Beethoven, and copyright law grew up around that idea and inherited its assumptions. But most of the world’s music has never behaved that way: a raga, a maqam, a gamelan piece, a jazz solo lives in the performance, not on a page. Wilhelm von Humboldt’s old distinction fits well here: some music is ergon, a finished work; much of it is energeia, an activity. Even modern pop, typically composed in the studio as a recording rather than written as a score, has to be gently shoehorned into the work-shaped box the law provides. Identification is, in part, the ongoing negotiation between how music actually behaves and the categories we’ve inherited.

Which means there is no natural boundary waiting to be discovered.  There is only a boundary to be drawn. Every identification system chooses an operating point: reach further and catch more versions at the cost of the occasional false match, or stay strict and miss a few to stay clean. That choice looks like a technical threshold, but it is really a business and editorial decision about whose content gets flagged and whose royalties flow. The value we provide is not a fantasy of perfect certainty; it’s a trustworthy dial, set where you need it, backed by the best accuracy in the field.

And the frontier keeps moving. Generative AI has begun to invert the question itself, from “is this upload a copy of a recording?” to “was this model trained on our catalog?” and “is that the artist, or a convincing clone of a voice that copyright was never designed to protect?” Increasingly, identification is becoming a question of provenance: not only what is this? but where did it come from? That is the next zone of indistinguishability, and it’s the one we’re building for now.


For more than 20 years, Audible Magic has delivered B2B solutions in content identification, for both music and TV/film, along with music catalog fulfillment and rights administration. Our systems power billions of transactions every month for customers around the world, helping platforms reduce DMCA takedown notices and administer licenses for user-generated content. The new VIBE-powered Version ID is the latest step in that work. We’d love to hear about your use case and figure out how to solve it together. Contact us to learn more.


Ensuring Content Integrity: Suno Partners with Audible Magic for User Uploads

Suno released their Audio Inputs and Covers features, which allow users to upload their own creative content to create songs from any sound, be that their own original full-scale productions or small ideas straight from their mobile device or even jam sessions amongst collaborators, fostering an assistive, creative and collaborative environment for budding or experienced artists alike.

We are excited to announce a partnership between Suno and Audible Magic. Suno’s user upload feature is central to its mission of fostering a vibrant and collaborative creative community. By integrating Audible Magic’s content identification technology to block copyrighted recordings directly into the user upload process, Suno is committed to maintaining a trustworthy environment for sharing and collaboration and keeping its platform focused on creating original music.

A Word from Audible Magic

“We are thrilled to partner with Suno to enhance their Audio Inputs and Covers features. This is a significant step in fostering a trustworthy creator community while preventing unauthorized user uploads.” – Kuni Takahashi, Audible Magic – CEO

A Word from Suno

“Audible Magic is the trusted partner for copyright compliance. By leveraging the best in content identification and rights management technology, Suno ensures that it remains a trusted and vibrant community for our creators.” Mikey Shulman, Suno – Co-Founder and CEO


Audible Magic's Broad Spectrum: detecting transformed audio

Speeding up...

Social media creators and DJs love to put their own spin on released recordings, and one of the simplest ways is to speed up a track.  Speeding up turns a slow tune into an uptempo dance number, and faster music hits the hook or chorus more quickly.  The current trend on TikTok, where all pop songs seem to exist only in their sped-up versions, has caused record labels and artists across the spectrum to release official speed ups to capitalize on this phenomenon†, from Benson Boon's Beautiful Things, Iñigo Quintero's Si No Estás, and Tate McRae's Greedy, to Lizzy McAlpine's ceilings, which achieved hit status almost entirely due to the unofficial and official sped-up version on TikTok:

and slowing down...

But altering recordings is not new. In the early days of hip hop, DJ Screw and others reversed the above, slowing down tempi to 60-70 beats per minute to create a mellow, bass-and-lyric-heavy grooving laid-back version of what was normally uptempo music. In the 2000s, Nightcore did the opposite with Eurodance, speeding the tempo up to 160-180 beats per minute, chipmunking the vocals into a happy hardcore style of club music. And DJs as a matter of course beat and pitch-match tracks to keep the energy level or to avoid a jarring change of key.  The ability to make these alterations has been democratized by the availability of easy-to-use tools, meaning that anyone with a phone or a computer can produce altered recordings to suit any creative impulse.

and changing the key

That being said, sometimes alterations are done for completely pedestrian reasons. Since at least the 1950s, radio stations have played music faster to fit in more discs per hour and also more commercials and, more recently, uploaders who wanted to bypass content ID systems have been tempted to change the sound just enough that it isn't noticeable to the average listener, but to get it past the automatic filters.  Or, in this case of this Weezer track, pitch shifting up a semitone so that guitarists don't have to retune to play along:

The solution: Broad Spectrum

But in all these cases everyone wants correct attribution and, especially where there are royalties due, for the money to go to the proper artists. Audible Magic's standard Audio ID system was designed from the beginning to easily handle pitch shifts and time scalings in the range of a few percent, in addition to resistance to ambient noise, EQ and other channel effects.  To identify audio that has been altered beyond that - including modifications far beyond what we see in the wild - customers use our Broad Spectrum service.

Designed to tackle the complex issue of manipulated music, Audible Magic's Broad Spectrum is a significant advance in content identification technology. Unlike most Audio Content Recognition (ACR) systems, Broad Spectrum employs an innovative approach to detect alterations.  It gives platforms the confidence that they will be able to identify and attribute content correctly, and also ensures rights holders' content is treated according to their license agreements.

While monitoring what is identified by our system, we see that as much as 50% of the music on UGC platforms is transformed in pitch and/or tempo.  While small alterations are more common, a significant number are altered by 20% or more.  The current trend at TikTok and other platforms is clear in the data: about 4x more of those with larger alterations are sped up vs slowed down.  And changes to pitch alone, while they occur, do not occur as often as playing the music faster or slower, either with or without changes to pitch.

For more than 20 years, Audible Magic has provided B2B solutions in content identification (both music and TV/Film), music catalog fulfillment, and rights administration. Audible Magic powers billions of monthly transactions to serve customers all around the world. Its music identification services are used to help reduce the number of DMCA takedown notices and administer licenses for user generated content (UGC). Broad Spectrum is one of many technological innovations we have pioneered. Contact us to learn more.

– Erling Wold, Chief Scientist and Technologist

† more examples:

https://open.spotify.com/playlist/4sQRT2zDzaMk92Fw75TZHn

https://open.spotify.com/playlist/37i9dQZF1DX0h2LvJ7ZJ15