Wikipedia Asks AI Companies to Stop Scraping Data and to Start Paying Up

Lee Duna@lemmy.nz · edit-2 6 hours ago

Wikipedia Asks AI Companies to Stop Scraping Data and to Start Paying Up

theunknownmuncher@lemmy.world · 9 days ago

This makes no sense, the snapshots are updated regularly and Wikipedia isn’t even that big. Like 25GB.

F/15/Cali@threads.net@sh.itjust.works · edit-2 9 days ago

The answer is simpler than you could ever conceive. Companies piloted by incompetent, selfish pricks are just scraping the entire internet in order to grab every niblet of data they can. Writing code to do what they’re doing in a less destructive fashion would require effort that they are entirely unwilling to put in. If that weren’t the case, the overwhelming majority of scrapers wouldn’t ignore robot.txt files. I hate AI companies so fucking much.

pivot_root@lemmy.world · 9 days ago

“robots.txt files? You mean those things we use as part of the site index when scraping it?”

— AI companies, probably