• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

We aren't talking about a library so a big difference, these weren't being preserved in any way, most of them would have simply disappeared from the world without anyone ever opening them again, at least this way the contents are being preserved.
In an AI's probabilities matrix? Again, how about using those digitised pages in an ebook for the kindle to sell?
 
In an AI's probabilities matrix? Again, how about using those digitised pages in an ebook for the kindle to sell?
They could use AI to check which ones were out of copyright!

Genuinely though. I'm not sure people understand how many legit unwanted books are warehoused. I'm mad about a lot of things but not this.

I imagine the OCR'd texts were kept anyways since compressing text for storage is incredibly efficient and cheap. Even just as a rainy day resource.
 
Last edited:
That would for many of the works be illegal. That you own a copy of a book does not mean you have the rights to sell copies of that book.
Most of those rare books are out of copyright. Otherwise the copyright owner could just print a couple thousand more or just offer them online themselves.
 
Last edited:
Are they? In the UK it is the lifetime of the author plus 70 years after their death.
Same as in the USA (where Amazon is based) and most of the western world. But the reason they're rare is that there are few copies still in existence. That takes some time to happen.
 
Same as in the USA (where Amazon is based) and most of the western world. But the reason they're rare is that there are few copies still in existence. That takes some time to happen.
Some topics are pretty niche. One of the books on my bookshelf is a grammar of old Irish, first published in the 1940s. I expect it's pretty rare, and a cursory search seems to confirm that. But I also expect it pretty much started out life that way. If you're trying to scan the entire corpus of human knowledge, and that process is destructive, it's more or less inevitable that you'll be destroying rare books.
 
Last edited:
Some topics are pretty niche. One of the books on my bookshelf is a grammar of old Irish, first published in the 1940s. I expect it's pretty rare, and a cursory search seems to confirm that. But I also expect it pretty much started out life that way. If you're trying to scan the entire corpus of human knowledge, and that process is destructive, it's more or less inevitable that you'll be destroying rare books.
And my point is that it's not inevitable. You've scanned them. You can make an eBook out of those page scans and sell it for however much you think it can sell for. Like even a buck if nobody buys it for more is better than just destroying it. And a book published in the 1940's has expired its 70 years copyright in the 2010's.
 
And my point is that it's not inevitable. You've scanned them. You can make an eBook out of those page scans and sell it for however much you think it can sell for. Like even a buck if nobody buys it for more is better than just destroying it.
No you can't, at least not legally. You don't hold the publication rights.

But the book would be destroyed irrespective of whether the scanned pages are sold on, and presumably the scans are preserved.

And a book published in the 1940's has expired its 70 years copyright in the 2010's.
Lifetime of author plus 70 years. It's also translated from German, so the translators hold IP rights for the English language edition.
 
Before the 1976 change in copyright law in the U.S. and the other more recent changes extending the length of copyright protection, copyrighted works had an initial 28 year copyright that could be extended another 28 years only if proper paperwork was filed. So, there is a possibility that something copyrighted in the 1940s (before around 1948) might not have been renewed and thus could be in the public domain.
 
Before the 1976 change in copyright law in the U.S. and the other more recent changes extending the length of copyright protection, copyrighted works had an initial 28 year copyright that could be extended another 28 years only if proper paperwork was filed. So, there is a possibility that something copyrighted in the 1940s (before around 1948) might not have been renewed and thus could be in the public domain.
Sure, if it wasn't renewed it would be. But the tech giants scanning these books make no claim that they are only seeking public domain books, and I see no evidence that they are making robust efforts to identify whether a book is in the public domain or not--they are attempting to create an enormous corpus of all books published prior to 2022 or so.
 
Last edited:
(skipped a lot)
I made a post elsewhere earlier today and slipped on the Shift key, so it came out "Ai". I think I will be posting it that way from now on, due to he possible confusion (depending on the font) with the lowercase "L". Because I always have to do a double-think when I see "AI" has done something bad, I have to consider if they're talking about me.
 
(skipped a lot)
I made a post elsewhere earlier today and slipped on the Shift key, so it came out "Ai". I think I will be posting it that way from now on, due to he possible confusion (depending on the font) with the lowercase "L". Because I always have to do a double-think when I see "AI" has done something bad, I have to consider if they're talking about me.
weird al music.jpg
 
(skipped a lot)
I made a post elsewhere earlier today and slipped on the Shift key, so it came out "Ai". I think I will be posting it that way from now on, due to he possible confusion (depending on the font) with the lowercase "L". Because I always have to do a double-think when I see "AI" has done something bad, I have to consider if they're talking about me.
I asked Grok about this, the reply? : “You can call me Al”
 
It's too late for the thousands of rare books that Amazon alone has bought, scanned, then destroyed. Not even offer to sell them on Amazon afterwards, but destroyed them.
Destroyed individual books, not every copy that exists. They couldn't sell them again because to be efficient the scanning process requires that the books be disassembled. But even if it didn't, most of them are probably not worth reselling - otherwise they would have been already.

Millions of books are destroyed every year by people who don't want them and can't be bothered trying to resell them, including rare books. Second-hand bookshops are overflowing with books that people gave them which nobody wants. Libraries discard old books all the time too.

Amazon has the right to buy books and use them as they see fit, just like anyone does. If some truly are rare and should be kept then the owners should hold them back, but they don't because they are in it for the money too.

This doesn't just affect books. Almost every product goes through a period where it's just old and worthless, before it's value increases again as nostalgia kicks in and the few remaining examples (if any) skyrocket in value. You may think that's sad, but are you willing to open your wallet and/or use up storage space to save stuff you aren't that interested in? I have a 'few' examples which I saved from being thrown away, but I've reached my limit and some of it will have to go. The rest will probably be dumped when I die.

It's actually worse than Alexandria. If nothing else, based on number of pages alone.
A silly metric.

Library of Alexandria
A single piece of writing might occupy several scrolls, and this division into self-contained "books" was a major aspect of editorial work... At its height, the library was said to possess nearly half a million scrolls.
What we must remember here is that these scroll were all hand-written. That makes every one extremely rare, as well as having vastly higher historical significance than the majority of 'modern' books.

It's a pity that Amazon are only using this data for AI training - assuming they are - because it's a valuable resource by itself. However copyright law makes it tricky to use as eg. an online resource. They would have to set up a lending library, or sell each scanned book and then destroy their own copy. Who knows, once the AI bubble collapses they might do that - or sell the entire dataset to someone else who does.

I don't think it's the big deal being made out. OTOH anything that puts the boot into AI excesses is good. :wink:
 
Last edited:

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom