← Reference Index

Primary Record

The canonical corpus:
every published document

All 2,731,548 documents are represented and linked. Enriched records show the full archive card; unenriched records show a minimal source-linked row that upgrades in place as enrichment reaches them.

Source documents
2,731,548
Coverage
100%
OCR attempted
60,806
Rich cards
25,676
Two layers, one record

The enriched Reference Archive is unchanged. Its 25,676 rich cards keep their titles, document types, email subjects, summaries, name associations, topics, OCR excerpts and all existing filtering and sorting. Underneath sits complete source coverage: every one of the 2,731,548 published documents is represented and linked, and an unenriched record shows a minimal source-linked row that becomes the full rich card in place once enrichment reaches it. One record, one representation — never a duplicate.

Corpus coverage
100%
2,731,548 of 2,731,548 documents represented and linked
OCR attempted
60,806
2.23% of the corpus
OCR sufficient
56,110
92.3% of what has been read
Meaningfully enriched
25,676
full rich cards
Blocked by OCR
4,696
retained, linked, queued

Section 01

Look Up Any Document

All 2,731,548 identifiers resolve. Enriched records return the full card; unenriched return the minimal representation.

Try: EFTA01998137 (enriched — full card) · EFTA00000002 (OCR insufficient) · EFTA00500000 (pending, never read)

Section 02

Browse by Data Set

Twelve published data sets. Bars show enriched, read-but-unenriched, and low-quality OCR as a proportion of each.

The enriched archive
Browse the 25,676 enriched documents by topic, type and person.

93 topics with full rich cards, type filters and people filters — unchanged and complete. This corpus layer sits underneath it, not in place of it.

Open the reference index →