Headlines 24.nl verzamelt actueel nieuws via de rss feeds van online kranten. Op elk moment geven wij al het laatste nieuws overzichtelijk weer.

Tevens kunt u inloggen om uw eigen nieuws pagina samen te stellen en zo alleen het nieuws te zien dat u interesseert.


 
 

AI’s attribution problem gets worse as models scale

21/08 02:15 - AI’s attribution problem gets worse as models scale
Diffusion models are becoming sophisticated enough that they can reproduce an image even when they don’t have access to the original. In a series of ‘what if’ scenarios, researchers associated with MIT’s Computer Science & Artificial Intelligence Laboratory (CSAIL) swapped out different training datasets to test the impact on image outputs when original image data was completely removed. It turns out that, at sufficient scale, nothing changed. The researchers call the phenomenon “attribution decay”: The more data a diffusion model is trained on, and the larger it gets, the less individual inputs matter. “If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output,” Zheng Dai, lead author on the work, explained in an MIT blog post. These findings could have significant ramifications when it comes to resolving growing concerns about intellectual property (IP) and copyright infringement. Models can recreate images even if they’ve never ‘seen’ them Modern generative diffusion models essentially replicate statistical patterns in large training datasets to create realistic reproductions. These powerful tools have achieved “remarkable results” in a wide array of applications, the researchers noted, notably image, video, and audio generation. But they are increasingly under scrutiny by creatives, companies, and policymakers, who all want a way to assign responsibility for generated outputs. Models sit at the center of lawsuits, licensing deals, and proposed regulations around the world. For instance, Stability AI (maker of Stable Diffusion) and Midjourney are embroiled in an ongoing class action lawsuit filed by several artists in federal court in California. The claimants argue that the popular image, video, and audio-creating models are scraping billions of their copyrighted images without their consent. Getty Images also brought claims against Stability AI, but they were struck down by the High Court of Justice Business and Property Courts of England and Wales, although Getty did partly win trademark claims because some AI-generated images closely resembled its work. Attributability, the MIT CSAIL researchers noted, would increase understanding of “machine unlearning,” data poisoning, model interoperability, fairness, and privacy, while also addressing ethical, legal, financial, and regulatory issues. “Developing a method to attribute generated outputs to influential training data would greatly advance our understanding of and ability to regulate these models,” the researchers wrote. In their experiments, they used ablation, which is essentially testing what happens when certain elements are removed by looking at what a model might have produced if it had never “seen” a particular image. Typically, ablation is difficult because models need to be retrained after data is pulled out. But the MIT CSAIL researchers applied the method to a “diffusion ensemble” architecture of many different components trained on different pieces of data. These components could be swapped out to determine how much of an impact, if any, each one had. “Our analysis is based on observing changes in model behavior, or lack thereof, upon omitting a part of the training set,” the researchers explained. To do so, they trained 24 ensembles on datasets containing anywhere from 256 to 160,000-plus images. These were pulled from seven publicly accessible image datasets, including ArtBench (artwork), CIFAR-10 (generic colored images), Fashion-MNIST (clothing and accessories), CelebA (celebrity faces), and MetFaces (human faces). In one example, they presented an image of a famous oil painting generated by a model trained on public domain artwork from 744 artists. It was shown side-by-side with hundreds of seemingly identical images that the model had generated, even when specific artists had been removed from training data. The original was re-imagined in every possible variation, and the researchers quantified attributability by measuring the largest change they could induce by omitting training data. The radius became smaller as datasets became bigger, holding true across different measurements including pixel-by-pixel or semantic meaning. In other words, single artworks by specific artists, or photographs of certain people, could be entirely removed from datasets, and the model could still reproduce that image or style. Essentially, tangible connections are lost, and linking to specific data points responsible for generated samples is “practically impossible,” or can even vanish, the researchers explained. Their method is novel, they said, because prior work has focused on removing large swathes of data rather than targeting smaller pieces, what they called “leave-one-out style attribution.” The impact on attributability Because the experiment shows that, as Dai put it, it “doesn’t make much sense” to attribute a given output to a given piece of data, creatives and others may not be able to provide an audit trail tracing back to their original work. Co-author David Gifford, an MIT professor and CSAIL principal investigator, said the findings have a direct bearing on legal questions around whether model outputs are actually derivative works. “One way to think about this is that these models are creative,” he said. “They are not simply copying what they are fed, but creating brand new outputs.” So if outputs can’t be correlated to individual pieces of training data, questions can be raised around fair use and whether, in fact, model-generated outputs are themselves copyrightable as “novel works,” Gifford said. It could also shift the conversation about how original creators are compensated when what comes out of a model seems a direct recreation of their work, but can’t be traced back to anything on the internet. Ultimately, producing outputs that are guaranteed to be unattributable is an “obligation for the industry, rather than a loophole,” he said. AI builders “need to revise their models to take advantage of the advances in this work, so they can show they’re not creating derivatives of individual people or items.” ...


 
 

Meer over computer

21/08 02:30 Haal meer uit Claude met deze 3 functies in de gratis versie

21/08 02:30 De nieuwe waarom zou AI jouw merk aanbevelen?

21/08 02:30 Wat kan je doen tegen die sterk stijgende AI-kosten?

21/08 02:30 De psychologie achter succesvolle LinkedIn-content

21/08 02:30 Wat automatiseer je wél en wat juist niet met AI?

21/08 02:30 Lessen van de loyaliteitsprogramma’s van KLM, Adidas & Rituals [benchmark]

21/08 02:30 Waarom de stift het wint van de AI-prompt

21/08 02:30 De snelste winst voor je website? Pak deze 2 fouten als eerste aan

21/08 02:30 De meetbaarheidsparadox van social media

21/08 02:30 Van iDEAL naar 4 vragen die elke webshop moet stellen

21/08 02:30 Het gevaarlijkste excuus? Eentje dat gewoon klopt

21/08 02:30 Minder losse content, meer zo bouw je een herkenbaar social kanaal

21/08 02:30 Je klantenbestand kan een bedrijfsgeheim zijn

21/08 02:30 Stop met alleen je concurrenten kijk naar de alternatieven

21/08 02:30 AI bespaart tijd, maar wie krijgt het dividend?

21/08 02:30 De digital native wordt jongeren zijn kritischer op hun digitale leven

21/08 02:30 Aantrekkelijk werkgeverschap begint niet bij een gelikte campagne

21/08 02:30 Durf jezelf overbodig te maken in je huidige rol

21/08 02:30 Schrijft AI steeds meer zoals jij of schrijf jij steeds meer zoals AI?

21/08 02:30 De Custom GPT is lang leve skills

21/08 02:30 Bedrijfsbezoek NSSG in Enschede

21/08 02:30 Tafelstandaard voor UNV/TPV intercom-binnenposten

21/08 02:30 Duitsland verlengt grenscontroles opnieuw met zes maanden

21/08 02:30 Criminele bende vermoedelijk achter inbraakgolf bij brandweerkazernes

21/08 02:30 Amsterdammers voelen zich steeds onveiliger in openbaar vervoer

21/08 02:30 Secusoft stroomlijnt documentbeheer binnen beveiligingsbedrijven

21/08 02:30 Sergio Römer Category Manager bij SmartSD

21/08 02:30 Schiphol laat ultimatum verlopen, acties beveiligers volgen

21/08 02:30 Retailers zetten fysieke beveiliging breder in dan diefstalpreventie

21/08 02:30 Snelle alarmopvolging Multiwacht leidt tot aanhouding na inbraak bouwterrein

21/08 02:30 'Geheugentekorten treffen nu ook de populaire MacBook Air'

21/08 02:30 Avatar The Last Airbender vanaf 22 augustus in Nederland te zien

21/08 02:30 MOVA S70 Roller zo bevalt deze robotstofzuiger in de praktijk

21/08 02:30 De specificaties van Google's Pixel 11-smartphones zijn gelekt

21/08 02:30 Nepkorting of echte aanbieding? Zo herken je het verschil

21/08 02:30 Baas over eigen alle functionaliteit, maar zonder Big Tech

21/08 02:30 Grote techbedrijven praten vandaag met president Trump over AI-veiligheid

21/08 02:30 Snaps slimme 'Specs'-bril arriveert in september, maar Nederland moet wachten

21/08 02:30 iPhone 18- dit weten we nu al over de nieuwe toestellen

21/08 02:30 Review Ring Floodlight Cam (2e gen) – Schijnwerpercamera met verhoogde resolutie

21/08 02:30  Dave Bautista gaat mogelijk Kratos spelen in God of War-serie

21/08 02:30  WhatsApp brengt oplossing uit voor bug die accounts blokkeert

21/08 02:30 Kopiëren en plakken tussen iPhone en Apple werkt aan een gedeeld klembord

21/08 02:30 Pop!_OS met Linux met een kosmische twist

21/08 02:30 Nothing-budgetmerk CMF komt in september met eigen open-ear-oortjes

21/08 02:30 Kerstmis wordt weer gewelddadig dankzij eerste Violent Night 2-trailer

21/08 02:30 Google Chrome krijgt mogelijkheid om Netflix-content in 4k af te spelen

21/08 02:30 Smartphone nat, vuil of oververhit na festival of dagje strand? Zo voorkom je extra schade

21/08 02:30 Versiegeschiedenis in de redder in nood

21/08 02:30 Dit zijn de kleuren van de lichtgevende 'rand' rondom de Google Pixel 11

 

login Member login

Emailadres

Wachtwoord