For 20 years, we’ve been turning our love of music into technology that helps make sense of what’s happening behind every song. And the industry keeps giving us many new things to make sense of, a challenge we’re more than happy to take on.
Our in-house R&D team keeps pushing that work forward, testing what’s next and sharing what we learn along the way. Latest case in point: a new study from the UPF-BMAT Chair on AI and Music takes AI detection beyond the usual yes-or-no answer, exploring whether we can quantify the amount of AI-generated audio inside a piece of music.
“How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures” will be presented at ISMIR 2026 this November.
And with AI-generated music everywhere (invited, assisted or otherwise just there) the timing feels pretty spot on.
It’s not over until the generated lady sings
Fully AI-generated tracks reportedly made up 50% of new daily uploads on one of the major streaming platforms at its peak in June 2026, with around 90,000 AI tracks arriving per day on average that month.
Elsewhere, another major platform said in April 2026 that more than a 3rd of the music it received each month was fully AI-generated.
For some scale, public estimates put the number of new recordings entering digital distribution at around 93,400 new tracks per day in 2022 and 103,500 in 2023.
Now compare that with the supposedly 90,000 fully AI-generated tracks arriving on a single platform in a day.
Not quite apples to apples: one counts unique recordings entering distribution, the other tracks landing on one service, but that’s still one spectacularly large pile of synthetic fruit.
Absolutely full of AI
Fully generated tracks are only one part of the music AI picture. AI can also be present inside otherwise human-made productions: a generated drum part here, an accompaniment there, perhaps a vocal or instrumental layer somewhere in between.
Whether AI use is obvious or barely noticeable, the line between human-made and AI-made music is getting blurry, at least to some extent.
It seems average listeners can’t tell for sure anymore. In a blind test involving 9,000 people across 8 countries in late 2025, 97% failed to reliably distinguish fully AI-generated music from human-made recordings.
The remaining 3% all work here, but they forgot to mention this in the study.
Puns aside, even a well-trained human auditory system can be deceived. In a 2026 blind-listening study with 71 musically trained participants, only 31.2% of AI-generated excerpts were correctly identified as AI, while 54.7% of genuinely human-composed excerpts were mistaken for involving AI.
Seems like we might need machines to catch the other machines in the act.
AI-generated, AI-assisted and everything in between
The industry is also getting more nuanced about the labelling and terminology. IFPI, RIAA and other music organisations have proposed separate “AI-Generated” and “AI-Assisted” labels, recognising that AI can play very different roles in a recording.
That distinction is already appearing across popular DSPs. Some use metadata to disclose where AI contributed, others automatically identify fully generated music, with those classifications sometimes affecting recommendations or monetisation.
At the infrastructure level, DDEX is actively defining what metadata is needed to communicate AI-generated music through the digital music supply chain.
The question “AI or not?” is starting to look underqualified for the job. That is where things get interesting for anyone whose job involves turning messy music reality into usable data (that’s us).
There’s also a regulatory reason to care. Since August 2026, Article 50 of the EU AI Act has required certain AI-generated audio to be machine-readable and identifiable as synthetic or manipulated.
The harder question is what happens in between: when only part of a track is AI-generated. That means being able to detect it, quantify it and describe it in a way that can actually travel through music systems.
Which brings us back to the paper at hand.
Dissecting an AI-assisted track
Today’s AI music detectors can exceed 99% accuracy when separating fully human from fully AI audio.
Part of what they learn to spot is an accidental spectral fingerprint.
Neural audio decoders can leave regular patterns in the frequency spectrum as they reconstruct sound. Think of it as a tiny technological barcode: mostly invisible to the ear, but something a machine can learn to recognise.
The new study starts with a fairly logical assumption. Most studies do, but hear us out:
If those fingerprints survive when AI-generated and human-made elements are mixed together, perhaps a detector can estimate how much AI-related audio is present, rather than forcing the entire track into one “Made by AI: yes or no” box.
This is where stems get to shine. A stem is an individual component of a production, like vocals, drums, bass, guitar and so on. If you combine AI-generated stems with human-performed ones, you get a hybrid music mixture.
Instead of giving that mixture a binary label, the paper introduces an AI energy ratio, α, between 0 and 1. It measures how much of the mix’s signal energy comes from AI-related stems.

An AI keyboard whispering somewhere in the background therefore does not count the same as AI drums taking over the chorus. α is not a percentage of creativity, songwriting or artistic angst, as we are measuring audio, not existential ownership of the bridge or the expression of human creativity.
Stuck in the middle with 21,212 hybrid mixtures
There is one important experimental detail.
The researchers did not use commercial GenAI outputs. Instead, they take human-performed stems and run selected ones through EnCodec, a neural audio codec.
The performance remains essentially the same, but the reconstruction introduces the decoder fingerprints the experiment needs. It needs the same music, with a different technological footprint.
That gave the researchers the means to create mixtures where they know exactly which stems have been reconstructed and exactly how much signal energy those stems contribute.
The recipe has 3 steps: reconstruct the selected stems, generate every possible original/reconstructed combination, then calculate the AI energy ratio for each mix.
A four-stem song gives 16 combinations. Across 240 MoisesDB tracks, that became 21,212 hybrid mixtures.
That’s enough grey area to make a binary classifier feel slightly overwhelmed.
Binary works great, until the answer is 0.46
The first CNN (for anyone not fluent in machine learning, we unpacked what a convolutional neural network is in an earlier article) was very good at the job it already knew, reaching 99.97% accuracy on its held-out binary evaluation.
But asking the same model to estimate a percentage exposed the problem. For mixtures containing 40–50% reconstructed energy, its median score was only 0.10.
So the researchers here trained another CNN specifically to predict α.
That worked considerably better. Using five-second windows, the regression model achieved an MAE of 0.076 and an R² of 0.85: roughly 7.6 percentage points of error on average across the controlled test mixtures.
Drums snitch while the bass keeps a low profile
Apparently, not all instruments are equally good at hiding the Gen-AI evidence.
Drums and guitar carried stronger detectable fingerprints, while vocals were subtler and bass proved particularly elusive. The same proportion of AI-related signal can therefore be easier or harder to detect depending on what is actually carrying it.

It turns out drums are terrible at keeping secrets. Although anyone familiar with a certain very recognisable 6-second break may have suspected that already. No shade to drum and bass, we love all music equally.
There is, naturally, an asterisk the size of your favourite GenAI music platform’s T&Cs.
These are controlled EnCodec reconstructions, not commercial GenAI stems produced by today’s popular music Gen-AI platforms, and the experiment does not include all the EQ, compression, mastering and creative edits found in finished releases.
So we are not yet at: “this track is 37.4% AI.”, but we may be getting closer by asking the right questions.
For those who prefer the version with Greek letters
This is the simplified blog version, with most of the equations, loss functions and spectrogram settings politely left in the paper.
We’re also attaching the full research paper for the geeks in the front and the signal-processing romantics in the back.
Latest articles
March 13, 2026
First, video killed the radio star, and now AI is going after video: A study on detecting GenAI music in broadcast audio
For the past 20 years, we’ve refined how audio technologies serve the music industry, making them faster, more precise, and more reliable. We invest over €3 million annually in our in [...]
March 8, 2026
The ABCs of Women in Music Technology
Human nature has a long memory for machines and a surprisingly short one for the people behind them. We remember the technology, but less often the researchers, engineers, and visionaries, [...]
March 4, 2026
BMAT Expands Collaboration with MCT Thailand to Scale VOD Royalty Processing in APAC
BMAT, a global leader in music technology and rights and royalties data, and MCT, Thailand’s collective management organisation for composers, authors, and music publishers, are expan [...]