Their findings, shared exclusively with MIT Technology Review, present a worrying development: AI’s information practices threat concentrating energy overwhelmingly within the arms of some dominant know-how corporations.
Within the early 2010s, information units got here from a wide range of sources, says Shayne Longpre, a researcher at MIT who’s a part of the mission.
It got here not simply from encyclopedias and the net, but in addition from sources corresponding to parliamentary transcripts, incomes calls, and climate experiences. Again then, AI information units have been particularly curated and picked up from completely different sources to swimsuit particular person duties, Longpre says.
Then transformers, the structure underpinning language fashions, have been invented in 2017, and the AI sector began seeing efficiency get higher the larger the fashions and information units have been. At the moment, most AI information units are constructed by indiscriminately hoovering materials from the web. Since 2018, the net has been the dominant supply for information units utilized in all media, corresponding to audio, photos, and video, and a spot between scraped information and extra curated information units has emerged and widened.
“In basis mannequin improvement, nothing appears to matter extra for the capabilities than the dimensions and heterogeneity of the information and the net,” says Longpre. The necessity for scale has additionally boosted the usage of artificial information massively.
The previous few years have additionally seen the rise of multimodal generative AI fashions, which may generate movies and pictures. Like giant language fashions, they want as a lot information as attainable, and one of the best supply for that has develop into YouTube.
For video fashions, as you’ll be able to see on this chart, over 70% of information for each speech and picture information units comes from one supply.
This might be a boon for Alphabet, Google’s mother or father firm, which owns YouTube. Whereas textual content is distributed throughout the net and managed by many alternative web sites and platforms, video information is extraordinarily concentrated in a single platform.