Google Researchers’ Attack Prompts ChatGPT to Reveal Its Training Data

btp · edit-2 10 months ago

donuts · 10 months ago

I can think of two ways it’s significantly different:

Legally (in the United States specifically) the courts have previously ruled that search engines collecting links to other people’s data is fair use, as it’s a mutually beneficial thing for all parties: users find the info that they’re looking for, search helps drive traffic to providers of info and services, and the search engine profits off connecting them to each other.

Unlike Wikipedia, for example, info that’s chewed up, processed, and regurgitated by “AI” chat bots and the like is totally unsourced, unaccountable, and passed off as original, authentic knowledge. ChatGPT is collecting various data from all of the net and forming it into something that appears to be presentable and correct, but it’s merely recycling ideas from other people’s work without any first-hand knowledge, thought, or attribution. Even the people who create “AI” can’t even connect the dots about why it says what it says, let alone have it properly source where the information came from.

@GarytheSnail · 10 months ago

Thank you for the links!

Do you think the same could be argued: that models collecting links to other people’s data is fair use?