Re: Six things that broke while ingesting a layered contract corpus

From: Ross Spencer <exponentialdecay.digipres_at_nyob>
Date: Tue, 8 Sep 2026 03:21:47 -0400
To: CODE4LIB_at_LISTS.CLIR.ORG
Hi Tim,

These are fascinating problems, but some context would be helpful.

What document did you begin with?  
What is its source format? (Your technical language suggests PDF.)  
What are your goals? For example, it sounds like search, but is there a broader perspective?  
Were there no structured text equivalents, such as a website?

One important aspect might be an AI disclaimer and information about any LLM you are using.

You might not have, but there is some terminology and punctuation in your post that suggest its use.

Why is a description of your tooling important?

"Broke" could perhaps benefit from some nuance, and things that are breaking in your tooling might not be issues that break in other tools or approaches.

For example, a more manual approach, using a PDF's text-extraction layer via various tools that support it, such as Tika, and then more manual data wrangling, might help you identify these issues early by eye, rather than performing QA on automated outputs.

Alternative approaches might include finding alternative markup methods before automation to better guide the tools you are using, perhaps drawing inspiration from the digital humanities.

I'm interested to learn more, and it sounds like your work was eventually successful.

Cheers,
Ross
Received on Tue Sep 08 2026 - 03:18:43 EDT