podlake — consortial collection analytics

Source of cataloging

Every other view here asks what is in the collections. This one asks who made the metadata, from MARC 040 — the field that records which agency did the original cataloging ($a) and which agencies have modified the record since ($d).

Read the numbers as attribution practice, not as cataloging labor. Three things matter before anything below means much:

Where each library's records come from

Share of every record the institution holds, by the agency credited in 040 $a. "No 040 field" is kept visible rather than dropped, so each bar is the whole collection. The dominant band almost everywhere is some other agency — mostly commercial suppliers and other libraries' original cataloging, broken out in the next chart.

How that mix has changed

The bar above is the whole of each collection at once. Cutting it by the year each record entered the catalog (008/00-05) shows the mix moving. Panels keep their own vertical scale, so a panel's height is that library's intake curve and the bands are its composition.

The Library of Congress band is not a picture of LC's cataloging. The horizontal axis dates the record in the holding library's system, so an LC record cataloged in 1975 and loaded by Penn in 1986 sits in 1986. Read the blue band as "LC copy arriving here", never as "LC's output that year" — this data has no clock for that. (The one that would is 010 $a: an LCCN encodes the year LC assigned it. Different field, different chart.) The same caveat applies to the "another POD member" band.

The 035 page covers where records travelled; this is who is credited with making them, cut by year. The bands are read as a share of that year's intake, so a band can narrow either because a library took in more of other kinds of record or because it was credited with less.

Which agencies those are

There are tens of thousands of distinct agency codes in 040 $a, so this has to be a selection. It is the union of each institution's own twelve most common codes, not a consortium-wide top-N — because ranking globally drops a small library's principal agencies, and buries a large library whose work is spread thinly across many symbols. Unioning also means the axis keeps working as POD grows: each new member brings its own rows rather than competing for shared ones.

Cells are that institution's share of its records carrying an $aof all of them, not just the rows shown — so the columns do not sum to 100%: the remainder is the long tail of codes no institution ranks highly. A code that is one library's mainstay and absent elsewhere reads as a single bright cell in a row of real zeros.

Codes are shown uppercased because that is the normal form the extract compares on; unrecognized codes pass through raw rather than being dropped.

Cataloging done inside the consortium

Restricting 040 $a to POD members' own codes gives the flow of cataloging between these libraries. First, how much of each library's collection is credited to itself. Read each bar as a floor: it counts only codes somebody has confirmed against a registry, so a library whose retired or unrecorded codes are missing reads lower than its real output.

Then whose cataloging each library holds. This matrix is not symmetric — the number of Harvard-cataloged records at Princeton is a different figure from the number of Princeton-cataloged records at Harvard, and that asymmetry is the whole point. Read a column down to see who supplies a library, and a row across to see where a library's cataloging travels. The diagonal (self-cataloging, shown above) is left out because it is an order of magnitude larger and would flatten everything else. Like the bar above, it can only count codes the map knows, so a thin row may mean unrecorded codes rather than little sharing.

When that cataloging happened

Crossing 040 $a with the record's creation date (008/00-05) dates each library's own cataloging, and shows how much of it is really a conversion project stamped with one year. That has its own page: Original cataloging over time.

How many hands have touched each record

Count of distinct agencies in 040 $d, as a share of all the institution's records. "No 040 field" is kept separate from "None" — a record with no 040 at all is a different thing from one whose 040 simply credits no modifying agency, and the gap between them is substantial at every institution.

Read this as a practice signal first: some systems append a modifying agency on every save and others never do, so the spread says more about ILS habits than about how heavily records have been edited.