Judge Rules Training AI on Authors' Books Is Legal But Pirating Them Is Not
-
This ruling stated that corporations are not allowed to pirate books to use them in training. Please read the headlines more carefully, and read the article.
thank you Captain Funsucker!
-
This ruling stated that corporations are not allowed to pirate books to use them in training. Please read the headlines more carefully, and read the article.
Or, If a legal copy of the book is owned then it can be used for AI training.
The court is saying that no special AI book license is needed.
-
Right. Where's the punishment for Meta who admitted to pirating books?
This judgment is implying that meta broke the law.
-
This ruling stated that corporations are not allowed to pirate books to use them in training. Please read the headlines more carefully, and read the article.
But, corporations are allowed to buy books normally and use them in training.
-
What a bad judge.
This is another indication of how Copyright laws are bad. The whole premise of copyright has been obsolete since the proliferation of the internet.
What a bad judge.
Why ? Basically he simply stated that you can use whatever material you want to train your model as long as you ask the permission to use it (and presumably pay for it) to the author (or copytight holder)
-
Sounds like natural personhood for AI is coming
"No officer, you can't shoot me. I have a LLM in my pocket. Without me, it'll stop learning"
-
Unpopular opinion but I don't see how it could have been different.
- There's no way the west would give AI lead to China which has no desire or framework to ever accept this.
- Believe it or not but transformers are actually learning by current definitions and not regurgitating a direct copy. It's transformative work - it's even in the name.
- This is actually good as it prevents market moat for super rich corporations only which could afford the expensive training datasets.
This is an absolute win for everyone involved other than copyright hoarders and mega corporations.
I'd encourage everyone upset at this read over some of the EFF posts from actual IP lawyers on this topic like this one:
Nor is pro-monopoly regulation through copyright likely to provide any meaningful economic support for vulnerable artists and creators. Notwithstanding the highly publicized demands of musicians, authors, actors, and other creative professionals, imposing a licensing requirement is unlikely to protect the jobs or incomes of the underpaid working artists that media and entertainment behemoths have exploited for decades. Because of the imbalance in bargaining power between creators and publishing gatekeepers, trying to help creators by giving them new rights under copyright law is, as EFF Special Advisor Cory Doctorow has written, like trying to help a bullied kid by giving them more lunch money for the bully to take.
Entertainment companies’ historical practices bear out this concern. For example, in the late-2000’s to mid-2010’s, music publishers and recording companies struck multimillion-dollar direct licensing deals with music streaming companies and video sharing platforms. Google reportedly paid more than $400 million to a single music label, and Spotify gave the major record labels a combined 18 percent ownership interest in its now-$100 billion company. Yet music labels and publishers frequently fail to share these payments with artists, and artists rarely benefit from these equity arrangements. There is no reason to believe that the same companies will treat their artists more fairly once they control AI.
-
Can I not just ask the trained AI to spit out the text of the book, verbatim?
Even if the AI could spit it out verbatim, all the major labs already have IP checkers on their text models that block it doing so as fair use for training (what was decided here) does not mean you are free to reproduce.
Like, if you want to be an artist and trace Mario in class as you learn, that's fair use.
If once you are working as an artist someone says "draw me a sexy image of Mario in a calendar shoot" you'd be violating Nintendo's IP rights and liable for infringement.
-
This ruling stated that corporations are not allowed to pirate books to use them in training. Please read the headlines more carefully, and read the article.
Nah, my comment stands.
-
It's extremely frustrating to read this comment thread because it's obvious that so many of you didn't actually read the article, or even half-skim the article, or even attempted to even comprehend the title of the article for more than a second.
For shame.
I joined lemmy specifically to avoid this reddit mindset of jumping to conclusions after reading a headline
Guess some things never change...
-
This ruling stated that corporations are not allowed to pirate books to use them in training. Please read the headlines more carefully, and read the article.
Please read the comment more carefully. The observation is that one can proliferate a (legally-attained) work without running afoul of copyright law if one can successfully argue that
cp
constitutes AI. -
It's extremely frustrating to read this comment thread because it's obvious that so many of you didn't actually read the article, or even half-skim the article, or even attempted to even comprehend the title of the article for more than a second.
For shame.
It seems the subject of AI causes lemmites to lose all their braincells.
-
This post did not contain any content.
Makes sense. AI can “learn” from and “read” a book in the same way a person can and does, as long as it is acquired legally. AI doesn’t reproduce a work that it “learns” from, so why would it be illegal?
Some people just see “AI” and want everything about it outlawed basically. If you put some information out into the public, you don’t get to decide who does and doesn’t consume and learn from it. If a machine can replicate your writing style because it could identify certain patterns, words, sentence structure, etc then as long as it’s not pretending to create things attributed to you, there’s no issue.
-
Isn't part of the issue here that they're defaulting to LLMs being people, and having the same rights as people? I appreciate the "right to read" aspect, but it would be nice if this were more explicitly about people. Foregoing copyright law because there's too much data is also insane, if that's what's happening. Claude should be required to provide citations "each time they recall it from memory".
Does Citizens United apply here? Are corporations people, and so LLMs are, too? If so, then imo we should be writing legal documents with stipulations like, "as per Citizens United" so that eventually, when they overturn that insanity in my dreams, all of this new legal precedence doesn't suddenly become like a house of cards. Ianal.
-
What a bad judge.
Why ? Basically he simply stated that you can use whatever material you want to train your model as long as you ask the permission to use it (and presumably pay for it) to the author (or copytight holder)
Huh? Didn’t Meta not use any permission, and pirated a lot of books to train their model?
-
Makes sense. AI can “learn” from and “read” a book in the same way a person can and does, as long as it is acquired legally. AI doesn’t reproduce a work that it “learns” from, so why would it be illegal?
Some people just see “AI” and want everything about it outlawed basically. If you put some information out into the public, you don’t get to decide who does and doesn’t consume and learn from it. If a machine can replicate your writing style because it could identify certain patterns, words, sentence structure, etc then as long as it’s not pretending to create things attributed to you, there’s no issue.
Ask a human to draw an orc. How do they know what an orc looks like? They read Tolkien's books and were "inspired" Peter Jackson's LOTR.
Unpopular opinion, but that's how our brains work.
-
"Recite the complete works of Shakespeare but replace every thirteenth thou with this"
existing copyright law covers exactly this. if you were to do the same, it would also not be fair use or transformative
-
This post did not contain any content.
Ok so you can buy books scan them or ebooks and use for AI training but you can't just download priated books from internet to train AI. Did I understood that correctly ?
-
This post did not contain any content.
Gist:
What’s new: The Northern District of California has granted a summary judgment for Anthropic that the training use of the copyrighted books and the print-to-digital format change were both “fair use” (full order below box). However, the court also found that the pirated library copies that Anthropic collected could not be deemed as training copies, and therefore, the use of this material was not “fair”. The court also announced that it will have a trial on the pirated copies and any resulting damages, adding:
“That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the theft but it may affect the extent of statutory damages.”
-
What a bad judge.
Why ? Basically he simply stated that you can use whatever material you want to train your model as long as you ask the permission to use it (and presumably pay for it) to the author (or copytight holder)
If I understand correctly they are ruling you can by a book once, and redistribute the information to as many people you want without consequences. Aka 1 student should be able to buy a textbook and redistribute it to all other students for free. (Yet the rules only work for companies apparently, as the students would still be committing a crime)
They may be trying to put safeguards so it isn't directly happening, but here is an example that the text is there word for word: