Connecting the Scholastic Corpus


Jeffrey C. Witt (Loyola University Maryland)
https://jeffreycwitt.com | jcwitt@loyola.edu


August 12, 2026, St. Louis University, St. Louis, MO


Slide Deck: http://jeffreycwitt.com/slides/2026-08-12-stlouis

https://creativecommons.org/licenses/by-nc-sa/4.0/

# Introduction
# The Problem There is a contradiction between the means of production and the relations of production

Drawbacks to the Print Publication Paradigm

1. Progress is slow

2. Work is isolated and fragmented

3. Access is tightly controlled

Re-thinking what it means to make an edition "Public"

1. Textual data should be rigorously separated from data presentation

2. Textual data should be continuously revisable

3. Textual data should be freed from the artificial confines of the "codex" or the "website"

## Field Standards * How do we name files * Where do files live * How do we organize files * How do we structure the contents of files --- These banal and boring questions are crucial. It is this organization that allows us to apply automated tools. If they are not answered, then we miss out on the ability to automate workflows and scale our work. --- The specific answer we give is not that important. What is important is that we have the **same** answer.
## Threefold Mission of the Scholastic Commentaries and Texts Archive (SCTA) 1. To develop field standards 2. To aggregate data created according to field standards 3. To publish aggregated data for open and creative re-use by the community.
# Examples 1. Raw Text Similarity Detection 2. Automated Citation Detection, Indices, and Recommendations
## Example 1 ### Text Similarity Detection
"The cat is on the mat" --- 4-grams --- "the cat is on" "cat is on the" "is on the mat"

Paragraph (Document) Feature Vector
as
a Vector recording ngram presence (1) or absence (0)

Paragraph

4gram1

4gram2

4gram3

4gram4

4gram5

4gram6

4gram7

4gram8

4gram9

4gram10

4gram11

4gram12

Doc A

1

1

0

1

0

1

1

1

0

0

1

1

Intersection of shared ngrams in Docs A and B
as
Dot Product of Document Feature Vectors $$ A \cdot B = \sum\limits_{i=1}^{n}{A_i B_i} $$ If Dot Product of A and B >= 6, then Doc Vectors A and B are "similar"

Paragraph

4gram1

4gram2

4gram3

4gram4

4gram5

4gram6

4gram7

4gram8

4gram9

4gram10

4gram11

4gram12

sum

Doc A

1

1

0

1

0

1

1

1

0

0

1

1

8

Doc B

1

1

1

0

1

1

1

1

0

1

1

1

10

A * B (Dot Product Vector)

1x1=1

1x1=1

0x1=0

1x0=0

0x1=0

1x1=1

0x0=0

1x1=1

0x0=0

0x1=0

1x1=1

1x1=1

6

![intersection](https://s3.amazonaws.com/lum-faculty-jcwitt-public/2023-02-01/image5.png) Similarity = X is similar to Y if and only if $$ \\#\\{ a | \forall{ng}\forall{x}\forall{y}(IsFoundIn(ng,x) \land IsFoundIn(ng,y) \land x \neq y \\} >= n $$ where n = 6
“...Only by a **lucky coincidence** [emphasis mine] was I enabled to "unmask" Gracilis' dubious literary honesty. (See Trapp, Damasus, "Augustinian Theology of the 14th Century," Augustiniana 6 (1956): 147-274, p. 254.)
Detecting Successive Text Re-use at Scale with Convolution
||5|6|7|8|9|10|11|12| |---|---|---|---|---|---|---|---|---|---| |2|0|0|0|0|0|0|0|0| |3|0|0|**0**|**0**|**0**|**0**|0|0| |4|0|0|**0**|**1**|**0**|**0**|0|0| |5|0|0|**0**|**0**|**1**|**0**|0|0| |6|0|0|**0**|**0**|**0**|**1**|0|0| |7|0|0|0|0|0|0|0|0| |8|0|0|0|0|0|0|0|0| |9|0|0|0|0|0|0|0|0|
||5|6|7|8|9|10|11|12| |---|---|---|---|---|---|---|---|---|---| |2|0|0|0|0|0|0|0|0| |3|0|0|**0**|**0**|**0**|**0**|0|0| |4|0|0|**0**|**1**|**0**|**0**|0|0| |5|0|0|**0**|**0**|**1**|**0**|0|0| |6|0|0|**0**|**0**|**0**|**1**|0|0| |7|0|0|0|0|0|0|0|0| |8|0|0|0|0|0|0|0|0| |9|0|0|0|0|0|0|0|0|

X

||||| |---|---|---| |1|0|0|0| |0|1|0|0| |0|0|1|0| |0|0|0|1|

=

||||| |---|---|---|---| |0|0|0|0| |0|1|0|0| |0|0|1|0| |0|0|0|1|

=

3

## Example 2 ### Automated Citation Detection, Indices, and Recommendations

Conclusion

1: Machine Learning and AI tools can be positively used in our field when our texts are prepared as data first.

2: The traditional publication paradigm is in conflict with the data-first paradigm required for corpus-wide machine analysis.

C: If we want to reap the benefits of machine analysis, there is an urgent need for scholarly editors and academic institutions to support and demand a new publication workflow.

## Contact me! * Interested in editing a scholastic text as machine actionable data? Contact me! * Interested in using the [Index Biblicus Scholasticus](https://lombardpress.org/index-scholasticus) for research? Contact me! * jcwitt@loyola.edu