podlake — consortial collection analytics

Archives & manuscripts

Archives and manuscript materials are a small but distinctive slice of these catalogs — roughly 0.5–2.3% of each institution's records, but hundreds of thousands of records in absolute terms. They are identified here from the MARC leader: type of record (position 06) t manuscript text, d/f manuscript music/maps, p mixed materials; or bibliographic level (position 08) c collection / d subunit. That is a deliberately broad net — it includes collection-level description of some printed material, not only manuscripts.

Unlike the general collection, these materials are rarely shelf-classified, so the views below lean on what the material is (genre/form) and how it is made findable rather than LC class. As with electronic resources, several of these signals reflect cataloging practice as much as what an institution holds.

Scale & material type

How much archival material each institution holds, and of what kind (leader/06), as record counts — cells are on a square-root color scale so both large (Harvard) and small (Brown) holdings stay legible. Harvard and Penn dominate; "mixed material" is the classic multi-format archival collection.

Genre & form of material

What kinds of things the archives contain, from the MARC 655 genre/form heading — the vocabulary that actually characterizes special collections. We use this instead of LC classification because archives are seldom shelf-classified (LC-class coverage of this subset swings from 8% at Harvard to 71% at Penn).

Genre/form is a long tail — ~5,800 distinct terms, 65% of them used by a single institution — so rather than force a shared axis (which would be sparse and Harvard-skewed), each panel below shows that institution's own top dozen forms. The vocabularies aren't reconciled across libraries and subdivisions are kept intact, because that's exactly where the distinctive strengths show — Arabic and Sanskrit manuscripts, posters, broadsides, scrapbooks.

Vintage: when the material dates from

The decade each archival record's material begins (MARC 008 date1), as a share of the institution's dated archival records, so collection size drops out and the shapes can be compared. The axis runs from ~100 AD to the present: most material is modern (so the recent end dominates), but the deep tail is real — medieval manuscripts around the 1400s, and Duke's documentary papyri back in the first few centuries AD. Old material is dated to the century, so expect round-number clusters; dates are frequently estimated, so treat this as broad-brush. (The 008 can't record BC dates, so ~100 AD is the practical floor.)

Archival records increasingly point to an online finding aid — a fuller guide to the collection than the catalog record itself. First, how many of each institution's archival records carry any online link (856):

Because "finding aid" is not reliably flagged in the record, we don't trust a label — instead we classify each link's host into a fixed set of destination types. This keeps the chart stable as POD adds institutions (each brings its own finding-aid host, which still lands in the same bucket). Persistent-ID resolvers (nrs, arks, purl, hdl) are kept separate because they hide whether the target is a finding aid or a digitized object.