Skip to content

Processor

After a querier returns results, two Python packages clean and normalize them so entries from different libraries can sit in one catalog.

They map each source's schema onto a shared structure, extract fields, generate UUIDs, parse dates and authors, detect languages, and check URLs. Each data source has its own processing rules. The output is the unified dataset.