Re: llm technology is especially useful

From: Karen Coyle <lists_at_nyob>
Date: Thu, 9 Jul 2026 10:28:09 -0700
To: CODE4LIB_at_LISTS.CLIR.ORG
On 7/8/26 8:38 AM, Eric Lease Morgan wrote:
> It is a black box? Yes, mostly. It black in the same way Google's search algorithms are black. It is back in the same way our bibliographic indexes rank relevance. It is as black as relational database implementations perform join queries through many-to-many relationships.
> 

I agree that Google's search results come from a black box. Actually a 
sinister, manipulative black box. I don't know what ranking is done in 
library systems (I assume that's what "bib indexes" means) but I do know 
that (hope that!) the results aren't based on advertising revenue. And 
if you have access to the RDB tables your results may be complex but not 
black.

I absolutely want to either avoid black boxes or have some way to 
evaluate the truth of the result. One thing we have promoted is that 
libraries can be trusted. Every bit of information we cede to commercial 
entities is a chink in that trust.

> For example, the other day I queried various bibliographic indexes for the phrase 'big science'. This resulted in 1,200 abstracts from scholarly journals for a total of 350,000 words (which is bigger than Moby Dick). I then RAG-ed against it to extract definitions of 'big science', learned who helped define it, and how that definition changed over time. It worked. It worked well. It saved me a whole lot of time. It supplemented my reading process, not replaced it.
> 

Now I think about Wikipedia where every fact much have a citation. Can 
the LLM give you a citation for where it found the fact? I think that 
would be very useful.

> What I would really like to see is the creation of one or more LLM built by the library profession. This way we would know whence the model came, how it was implemented, and remove the blackness. Such an effort would be akin to collaboration we have seen in the past when it comes to collection building or metadata sharing. Yes, it would be very expensive, but if we were to pool our resources, then I think it could be done.

What would the use case for this be? What services are you imagining? I 
would love to see more datamining of bibliographic data - both locally 
and "universally". OCLC had done some very interesting work with 
Identities (like so many projects, now abandoned). Could an LLM help 
provide that kind of information? Or is that just a regular software 
project?

Thanks.
-- 
Karen Coyle
kcoyle@kcoyle.net http://kcoyle.net
Received on Thu Jul 09 2026 - 13:25:34 EDT