Mtukudzi, Macheso songs appear in AI training datasets, Atlantic database shows
Mtukudzi, Macheso songs appear in AI training datasets, Atlantic database shows
For all the News from Mashonaland Central, Join One of Our Groups
BINDURA — Recordings by some of Zimbabwe’s best-known musicians, including more than 150 songs linked to the late Oliver “Tuku” Mtukudzi, appear in large music datasets circulated for training music-generating artificial intelligence, a Bindura Eye search of a database published by the American magazine The Atlantic has found.
The database, compiled from four collections identified by The Atlantic staff writer Alex Reisner, lists millions of recordings and links that can be used by developers building generative AI tools. Among the Zimbabwean names that appear in the collections are sungura star Alick Macheso, Dendera musician Tryson Chimbetu, Zimdancehall artist Killer T and contemporary hitmaker Jah Prayzah.
According to the database, Mtukudzi’s catalogue is heavily represented, with as many as 153 songs appearing in the data, while other Zimbabwean artists appear with only one or a handful of tracks. The four collections identified by Reisner include two very large datasets of roughly 12 million and nine million tracks, as well as two smaller datasets of about 100,000 recordings each — more than 21 million recordings in total.
Music-generating AI systems learn patterns in melody, rhythm, instrumentation and style by being fed huge volumes of existing music. The datasets identified by The Atlantic function as the catalogues for that process: lists of songs, or links to songs hosted on services such as YouTube and Spotify, that can be downloaded and then used for training.
Reisner reported that three of the four datasets are not packaged audio libraries but lists of links, with recordings then pulled down using automated tools that can bypass features such as logins and adverts. The fourth dataset draws on the Free Music Archive, a library founded by the American radio station WFMU.
Appearing in one of these datasets does not, by itself, prove that any particular AI company trained a product using a specific Zimbabwean song. However, it does show that the recordings were compiled, shared and made downloadable for the purpose of training AI models, without the artists being asked.
There is no indication in the dataset listings that the Zimbabwean artists named were approached for permission or paid. The concern raised by artists and rights organisations internationally is that harvesting music through automated tools can avoid the normal routes through which royalties and subscriber revenue are generated.
Global industry estimates suggest the financial stakes could be significant. A study commissioned by CISAC, the international body for authors’ societies, has estimated that generative AI could take almost a quarter of music creators’ revenue by 2028. Separately, the streaming service Deezer has reported that AI-generated tracks now account for close to half of the new music uploaded to the platform each day.
For artists in Zimbabwe, including those with audiences in Mashonaland Central, the risk is widely viewed as twofold: their existing catalogues may be used to train AI systems, while those same systems then enable large volumes of synthetic music to compete for attention on streaming platforms. Bindura Eye cannot, on the available information, establish a direct loss of earnings for any individual Zimbabwean musician linked to the datasets, and no court has yet ruled on liability relating to these specific collections.
On what artists can do, legal options are currently limited for individuals trying to challenge technology companies based overseas. Observers say the more realistic routes are collective, through institutions such as the Zimbabwe Music Rights Association (ZIMURA) and its links to international rights bodies. Zimbabwe’s Copyright and Neighbouring Rights Act protects the reproduction of musical works, but how that law applies to cross-border AI training remains untested locally.
Major legal battles are under way internationally. Suno and Udio, two prominent AI music generators, face multiple lawsuits over alleged use of copyrighted music in training. Collecting societies including Germany’s GEMA and Denmark’s Koda are among those pursuing cases, and the outcomes may influence how smaller markets such as Zimbabwe seek remedies in future.
Responsibility for scraping and reuse is difficult to trace because training data is often kept secret. The Atlantic reported that the datasets have been downloaded thousands of times, but it is not publicly known who has used most of them. Only Google and Stability AI have publicly acknowledged, in research papers, using material from the Free Music Archive dataset; this does not, on its own, establish use of the specific Zimbabwean recordings identified by Bindura Eye.
In a 2024 court filing, Suno stated it trained its models on “essentially all music files of reasonable quality” it could find online. However, there has been no court ruling linking Suno, Udio or any other company to the specific Zimbabwean tracks identified in these dataset listings, and it remains unclear who scraped Mtukudzi, Macheso or Killer T and where the recordings may ultimately have been used.
Bindura Eye has approached ZIMURA for comment. This remains a developing story.
Mundubile seeks new party route as Hichilema eyes second term
Mundubile seeks new party route as Hichilema eyes second term For all the News from Mashon…








