Re: [Ext]Re: [External] Re: llm technology is especially useful

From: Eileen Lopez <0000023cf74d49ef-dmarc-request_at_nyob>
Date: Thu, 9 Jul 2026 11:11:09 +0000
To: CODE4LIB_at_LISTS.CLIR.ORG
[heart]         Eileen Lopez reacted to your message:
________________________________
From: Code for Libraries <CODE4LIB_at_LISTS.CLIR.ORG> on behalf of Charlow, Aurora <000001f5c9675ef3-dmarc-request_at_LISTS.CLIR.ORG>
Sent: Thursday, 09 July 2026 00:59:29
To: CODE4LIB_at_LISTS.CLIR.ORG <CODE4LIB_at_LISTS.CLIR.ORG>
Subject: [Ext]Re: [CODE4LIB] [External] Re: [CODE4LIB] llm technology is especially useful

I've been watching this debate back and forth for several months now, and I'm weighing in to say I sympathize with Alex's point and understand why he left. I won't be doing so, as I value this community, but let me take a moment to elucidate why all of this LLM hype is so upsetting -


  1.
It seems widely documented by this point that LLMs know about everything except what the user knows. They excel at shallow, surface-based analysis while making egregious errors discernible mostly to subject experts. They sound excessively confident while doing this. Even LLM based coding is now being rejected by some experienced software programmers: https://theseniordev.com/blog/why-i-stopped-using-ai-as-a-senior-developer-after-150-000-lines-of-ai-generated-code/.


  2.
Use of LLMs demonstrably reduces cognitive activity: https://www.media.mit.edu/publications/your-brain-on-chatgpt/. Educators have been talking for the last several years about declines in the performance of undergraduate students over-relying on this technology and outsourcing critical thinking tasks. Ignoring this reality is burying your head in the sand in the name of progress. Frankly, I think I function better when my ears aren't full of sand.

  3.
Specifically, to your example, Vishal - how much review did you perform on the output of your LLM data analysis? Were you able to confirm that it was completely free of errors? I hope so. It would be unfortunate for a seasoned professional to craft strategic directions off of any hallucinations, especially because LLMs, mathematically, cannot be prevented from hallucinating: https://www.computerworld.com/article/4059383/openai-admits-ai-hallucinations-are-mathematically-inevitable-not-just-engineering-flaws.html


Maybe you are an AI-whisperer able to review 16 years of data analysis for hallucinations in 72 hours. Okay, sure. Let's take a look at what the impact of that has been, practically speaking.


  1.
Many organizational leaders are reportedly outsourcing decisions to AI and treating LLM agents more and more as so-called 'spiritual advisors.' This is creating widespread concerns among the labor market as administrators grow increasingly divorced from reality and cease managing effectively. https://futurism.com/artificial-intelligence/bosses-obsessed-with-ai


  2.
Development of AI is built on labor exploitation of impoverished workers. It is exacerbating existing inequalities, particularly class and racial inequalities.  https://www.cjr.org/tow_center/qa-uncovering-the-labor-exploitation-that-powers-ai.php.


  3.
AI is powering the continued dominance of fossil fuels. xAI's huge, gas-powered data center in memphis is currently poisoning the local population: https://tennesseelookout.com/2025/07/07/a-billionaire-an-ai-supercomputer-toxic-emissions-and-a-memphis-community-that-did-nothing-wrong/ (and lest you think that only Elon is responsible, Anthropic is currently leasing energy from this location to power Claude).

  4.
Excessive AI use is causing psychosis in some individuals. There have been several suicides in the last few years.  https://nam.edu/news-and-insights/what-is-ai-psychosis/.



Maybe you don't care about negative externalities either. Okay. As librarians, do any of you have opinions on corporations using inherently biased training data to influence image description, article analysis, article writing? Bias that can't be detected by the average person! https://www.psu.edu/news/bellisario-college-communications/story/most-users-cannot-identify-ai-bias-even-training-data


How about this as a microcosm of the entire GLAM community, increasingly forced to outsource our skills to vendors that benefit from de-skilling us so they can charge our underfunded institutions more and more money?

Speaking frankly and directly to the data analytics example: 1. It takes an afternoon (maybe two, if you don't know anything about databases) to learn SQL querying. I have done this. It is one of the simplest coding languages there is. If I can learn how to manipulate XSLT in 3 days, anyone on this mailing list can do the same with SQL. And probably with XSLT too, I'm not special. 2. Deciding on the strategic direction of your department in 2 or 3 days based on the output of an LLM seems like a terrible idea, to me. I sincerely hope it goes well, because I would not wish ill on any library. Frankly, I think that taking several weeks to consider this sort of output is a minimum expectation. Better doesn't always mean faster, and what this AI boom has taught me above all is an appreciation for slow, careful work.

My career, brief as it has been so far, has been spent advocating against the use of LLM technology in various roles. Anyone who watches the economics of these things should know by now that they are functioning at current scale on borrowed time, and when they retreat to their little corner of the tech industry (where they belong) they'll be taking our retirement accounts with them. When that happens, libraries who continue to pay for this technology will be forced to cut access to knowledge, or more realistically, lay off employees to compensate. Claude can't adequately replace your catalogers.

Being an information professional, being an academic, means remaining critical and using some common sense when the fraudsters come knocking at your door.

Thanks to anyone who took the time to read this far.

Sincerely,

Aurora Charlow


[Ohio University Logo]
Aurora Charlow, she/they
DIGITAL ARCHIVIST
UNIVERSITY LIBRARIES
MAHN CENTER FOR ARCHIVES AND SPECIAL COLLECTIONS, PRESERVATION & DIGITAL INITIATIVES
EMAIL
charlowa_at_ohio.edu<mailto:%20charlowa_at_ohio.edu>
PHONE
740.593.0055

________________________________
From: Code for Libraries <CODE4LIB_at_LISTS.CLIR.ORG> on behalf of Vishal Patel <00000293fee5ef0c-dmarc-request_at_LISTS.CLIR.ORG>
Sent: Wednesday, July 8, 2026 05:53 PM
To: CODE4LIB_at_LISTS.CLIR.ORG <CODE4LIB_at_LISTS.CLIR.ORG>
Subject: [External] Re: [CODE4LIB] llm technology is especially useful

Use caution with links and attachments.

Given Alex’s strong response, I feel compelled to weigh in.

  1.
It’s clear that LLMs are divisive. We (namely, Mackenzie Salisbury) started one dedicated to LLMs at library-llm_at_listserv.it.northwestern.edu - it leans more LLM-curious.
  2.
I don’t detect any “LLM slop” in this thread. This appears to me be a constructive evaluation of an LLM case study.
  3.
I would advise against snubbing or canceling a forum or a person if you (even by mistake) catch a whiff of LLM generated content.
     *
Language, ultimately, is just a collection of symbols to which we have assigned meaning, and, when we find ourselves have strong emotional reactions to those symbols, the empirical evidence indicates that it will serve us - individually and collectively - better if we learn to be curious about those emotions, rather than checking out or leaving the group entirely.
     *
It is true that today’s symbols are rapidly commingling with machine-generated symbols (linguistically, visually, & audibly), but librarians have always existed to serve as stewards of collections of symbols for the community. If we opt out when we see or hear a new dialect (e.g. AI-infused English) being written or spoken, not only is that antithetical to the professional values of librarianship, but it fractures our professional body.

Going back to the purpose of Eric’s thought starter, I am greatly appreciating the value of LLMs. As an example, within the paste 24 hrs, I was able to:

  1.
Write SQL queries in BigQuery to analyze 6 months of traffic & search patterns on Lane library’s website
  2.
Analyze 6 years of historical data from our LibAnswers & LibCal systems
  3.
Analyze 16 years worth of proxy server data
  4.
Synthesize all analyzes into a strategic plan
  5.
Summarize the findings as an infographic to share with my executive team

As a former data scientist, I can attest that this type of analysis would have previously taken me 2-3 weeks to assemble, and the polished data visualization would have taken another week itself. LLMs make easy work of data analysis & coding.

Sincerely,

Vishal Patel, MD, PhD
Director, Digital Strategy & Library Technology

Lane Medical Library

Stanford School of Medicine


300 Pasteur Drive, L109, Stanford, CA 94305-5123

vishal.patel_at_stanford.edu<mailto:email_at_stanford.edu>

(650) 723-7196 *41425

Pronouns:  he, him, his

[cid:image001.png_at_01D4F60B.B9209360]

From: Code for Libraries <CODE4LIB_at_LISTS.CLIR.ORG> on behalf of Alex Dunn <adunn_at_UCSB.EDU>
Date: Wednesday, July 8, 2026 at 2:18 PM
To: CODE4LIB_at_LISTS.CLIR.ORG <CODE4LIB_at_LISTS.CLIR.ORG>
Subject: Re: [CODE4LIB] llm technology is especially useful

I think it's high time for this mailing list to enact a ban on LLM
slop.  In the meantime, I'm unsubscribing.  See you all around.

On Wed, Jul 8, 2026 at 8:39 AM Eric Lease Morgan
<00000107b9c961ae-dmarc-request_at_lists.clir.org> wrote:
>
> On Jul 4, 2026, at 7:44 PM, Karen Coyle <lists_at_kcoyle.net> wrote:
>
> > I would feel better about this if these results didn't sound like the platitudes of marketing speak. ("Collaborative refinement of library services" is something I don't think any of us would say about C4L.) Where did the model get this tripe?
>
>
> TL;DNR - I advocate the use of LLM technology in libraries. It can be a supplement to our existing processes, not a replacement.
>
>
> Thank you for the reply because I really feel our community good benefit from discussion on these topics.
>
> Tripe? I had to look up the definition of that word. Where did the result get such a response? Many places, but one of the more significant is my locally configured "system prompt". What's that? A system prompt is an extra little bit, behind the scenes configuration sent to a large-language model (LLM). My current system prompt follows:
>
>   Return results as if they were written by a student
>   attending a liberal arts college. Ask questions, sometimes,
>   but not always. The model is working within a generative-AI
>   system called a RAG, and therefore results are intended to
>   be primarily drawn from the underlying MCP system; results
>   drawn from outside the system are to be kept to a bare
>   minimum. The model is intended to be used as analysis tool
>   not an oracle. Do not voice results in the first person!!!
>   When citing sentences, include item and index values.
>
> If my system prompt said something like "Voice replies as if written by an eighth grader", then the results would be expressed differently. Differences in system prompts make a significant difference in results. You should see the sort of things I get back when I specify second graders or erudite college professors.  :-D
>
>
> > But what really concerns me is that the system is a black box (others have noted this), which means that there is no way to evaluate the result other than ones' gut feeling that it's "right." Your gut feeling returns "true and accurate" while mine concludes that honestly some of the "facts" in the statements below could be wrong. For all that C4L has great discussions, I don't see "code sharing" as taking place often in the body of the emails. (Maybe a snippet or two, but not as stated here.) The last paragraph on MARC doesn't convince me much. The phrase "annoying data format" comes up only once when I search on the C4L archive, and it's in one of your posts, Eric. I wouldn't include this in a summary of C4L list users' statements on MARC even though we are pretty critical of it. I also have doubts that one can conclude from the list that the group is "focusing heavily on MARC records—the standard format for library catalog data." There is a fair amount of discussion about MARC but have posters actually said here that it's the standard format for library catalog data? (We probably assume that everyone here knows that.) Could the software have gotten that from elsewhere, given that there is a fair amount of documentation online?
> >
> > --
> > Karen Coyle
> > kcoyle@kcoyle.net https://urldefense.com/v3/__http://kcoyle.net__;!!G92We9drHetJ8EofZw!emYRrTgEe2gQVprd47jFZCCIprO0hxT1USRlq7XMgipz5vbj2Jf6kauD7i8DN_ZAehOjotejg5xnn7JIIk41Ww$<https://urldefense.com/v3/__http://kcoyle.net__;!!G92We9drHetJ8EofZw!emYRrTgEe2gQVprd47jFZCCIprO0hxT1USRlq7XMgipz5vbj2Jf6kauD7i8DN_ZAehOjotejg5xnn7JIIk41Ww$>
>
>
> Granted, above is more difficult; I more or less (mostly more) agree, but...
>
> The box is more gray than black. The whole process is rooted in the computation of geometric distances between words mapped in a VERY large n-dimentional space. To elaborate, it works something like this. A HUGE pile of text is accumulated. The text is parsed into tokens (think "words") to create a vocabulary. This results in a matrix where each row (millions or billions of them) is a document, and each column is a token from the vocabulary. At the intersection of each row & column is a measurement, and the measurement might be the number of times the given token is found in the given row. This results an a GREAT BIG set of vectors "pointing" to locations in the space. Given some input (at least a word or better yet a large set of words), the input is vectorized in the same manner as the original LLM. The vector is compared to all of the other vectors in the set to identify most similar vectors, where similarity is denoted to something like cosine distance. This is a process of linear algebra, and it is a kind of find or search process. Works in the manner similar to auto-completion or auto-correct but on a REALLY big scale.
>
> For example, in English, the word "the" is very frequently followed by an adjective but ultimately by a noun. These patterns are manifested in the LLM.
>
> Using the sort of process outlined above, given words are mapped with similar words, and the similar word are output. It does not necessarily identify nor extract exact phrases from text unless specifically asked to do so. Also, in my example, I only read one month's of Code4Lib archives, and in that month there may have been more discussion about MARC than not. In any event, the process does work (for the most part), and it is a real-world application of something first articulated by a man named John Firth around 1957 who said, "You shall know a word by the company it keeps" -- context.
>
> It is a black box? Yes, mostly. It black in the same way Google's search algorithms are black. It is back in the same way our bibliographic indexes rank relevance. It is as black as relational database implementations perform join queries through many-to-many relationships.
>
> Very important: I do not advocate the use of LLMs sans context. In other words, I do not advocate asking very general questions, like "What is the best Shakespeare play?" to LLMs because the only context included in the response is what is in the model. On the other hand, I very much advocate the use of retrieval-augmented generation (RAG) and/or model context protocol (MCP) servers. These tools get input from known and verifiable collections of text. The results of RAG and/or MCP queries are sent to an LLM for interpretation, and well-implemented results will include ways to backtrack results to the source. I think the use of RAG and/or MCP technologies applied to library collection can be very useful. LLM are not panaceas though.
>
> For example, the other day I queried various bibliographic indexes for the phrase 'big science'. This resulted in 1,200 abstracts from scholarly journals for a total of 350,000 words (which is bigger than Moby Dick). I then RAG-ed against it to extract definitions of 'big science', learned who helped define it, and how that definition changed over time. It worked. It worked well. It saved me a whole lot of time. It supplemented my reading process, not replaced it.
>
> I have learned two additional things. First and foremost, I take the results as plausible, not truth. Thus, I like to believe I practice information literacy along the way. Second, I am always dubious of the adjectives returned by LLM, especially the superlatives.
>
> What I would really like to see is the creation of one or more LLM built by the library profession. This way we would know whence the model came, how it was implemented, and remove the blackness. Such an effort would be akin to collaboration we have seen in the past when it comes to collection building or metadata sharing. Yes, it would be very expensive, but if we were to pool our resources, then I think it could be done.
>
> Lastly, those were a lot of words. Thank you for listening. I hope the discussion continues.
>
> --
> Eric Lease Morgan, Librarian Emeritus
> University of Notre Dame
Received on Thu Jul 09 2026 - 07:08:38 EDT