Before I get started: I'm maintaining a page of links relevant to
the AI/LLM field. It's in loose (very loose) chronological order
and it's usually updated every couple of days. I've tried to avoid
paywalls, but: no guarantees. It's a mix of articles, reports, papers,
opinions, and occasional rants. Here's the page:
Topic-AI
https://www.firemountain.net/~rsk/topic-ai.html
There are currently about 2350 links there. One recent entry of particular
relevance is this:
Generative AI Is an Engineering Disaster
https://www.trueobserver.com/news/6a566b1afb4011d7ecfcb1d1
Now for a comment or three:
I've studied AI/expert systems/statistical pattern recognition/etc.
and I'm very dubious about what the big tech companies are doing.
It's a combination of massive hype and sloppy, careless, haphazard
development that requires exponentially-increasing amounts of power,
water, CPU, memory, disk, and money...and imposes severe costs on the
electric grid, the environment, and anyone unfortunate enough to live
in the vicinity of massive data centers -- as well as on anyone who's
trying to run a web-facing resource, thanks to the abusive behavior of
the web crawlers being used by these companies.
On that last point: I've been thinking about something Grady Booch wrote:
"These deals have absolutely nothing to do with increasing the
accuracy of any large language model.
Such architectures are inherently just next word predictors
and any correlation their output has with truth is only due to
statistical coincidence.
What these deals do have to do with is locking up exclusive
access to sources of information, thereby concentrating power
in the hands of just a few organizations who can afford to do so."
And about something that Sam Altman (OpenAI) said:
"We see a future where intelligence is a utility, like electricity
or water, and people buy it from us on a meter."
I think these explain why AI/LLM web crawlers are so very badly behaved:
it's deliberate. Consider: there are all kinds of open source crawlers
readily available. There's Common Crawl (https://commoncrawl.org/).
There's a wealth of knowledge/experience in this area that makes it
possible for anyone to efficiently and responsibly get their paws on
pretty much anything on the open web; it's just not that hard to do
things the right way, even if you've never done it before.
Yet the AI/LLM web crawlers are ignoring robots.txt, evading firewalls,
hijacking IOT devices, and conducting D/DOS attacks against millions
of web sites. They've forced web site operators to take all kinds
of countermeasures: there are entire software projects (e.g., Anubis)
devoted to stopping them. There are two IETF working groups. There are
all kinds of bot tests that have been deployed (and some of them are
rather badly broken). What was the open web is now a hot mess.
For a while, I thought these were mistakes, that the AI/LLM programmers
were just being sloppy -- but I don't think that any more. I think
it's deliberate strategy. I'm going to refer to those quotes above
and make this assertion: it's the intention of these companies to drink
the wells dry -- and then salt them.
After all, they can't make themselves the exclusive sources of knowledge
(for profit) if there are other people giving it away (for free).
Thus every web site, every library, every archive, every repository must
not only be crawled with no respect for copyright/terms-of-use/etc.,
it must also be rendered unusable so that primary sources are unavailable
and everyone is forced to rely on AI slop.
---rsk
Received on Thu Jul 23 2026 - 07:45:10 EDT