Overlap & rarity
How much of the collective collection is widely held versus rare? This is the question behind shared-print, preservation, and "last copy" decisions: a title held by only one institution is a candidate for preservation, while a title held by many is a safe candidate to store or weed.
How records are grouped into titles. Each institution contributes many individual MARC records for the same book. To count titles rather than raw records, records are collapsed using the Gold Rush match key developed by the Colorado Alliance of Research Libraries. The key is a normalized fingerprint derived from a record's title, author, publication year, edition, publisher, pagination, material type, and carrier (print vs. electronic). Records that generate the same key are treated as one title. The key is roughly manifestation-level: because it captures edition and carrier, a print book and its e-book edition count as two distinct titles.
Titles held by N institutions
The rarity curve: how many distinct titles are held by exactly one institution, by two, and so on. The left end is the rare/unique material; the right end is the widely-duplicated core.
Titles held by a single institution
Each institution's uniquely-held titles. This is material that would disappear from the consortium if that copy were lost.
Which collections are most alike
Which pairs of libraries actually resemble each other? How you measure decides the answer, so the control below switches between three measures of the same underlying numbers. They disagree sharply, and the disagreement is the point.
- Shared titles is the raw count of titles both libraries hold. It is mostly a ranking of size — the largest collections top it because they are largest, and Harvard appears in six of the eight highest pairs.
- Jaccard similarity divides that count by the two libraries' combined titles, so scale drops out and what is left is how far their collecting coincides. Harvard leaves the top eight entirely; same-sized peers take over.
- Containment asks a different question in a different direction: what share of this library's titles the other one also holds. It is not symmetric, and it favours the small library in a lopsided pair — which is exactly what a shared-print or last-copy conversation needs to know.
The fifteen closest pairs on the measure selected above.
Every pair, as a matrix. With Containment selected the matrix is asymmetric — read a row across for what share of that library's titles each other library also holds. The diagonal is blank throughout: a collection contains all of itself, and plotting that would flatten every other cell.
Two things limit how far "alike" can be pushed. Jaccard peaks at about 20% here, because 29.7M titles are held by exactly one institution — the collective collection is mostly long tail, so even the closest pair overlaps on a fifth of their combined holdings. And because the Gold Rush key is roughly manifestation-level, this measures edition-level co-holding: two libraries buying the same works in different editions or formats read as dissimilar. Treat it as a floor on how alike two collections are.