Regional Groundwater Database for Arequipa and Surroundings in Southern Peru from Unstructured Sources Using an Integrated Optical Character Recognition and Large Language Model Workflow
Abstract. Groundwater management in the arid Andes is constrained less by an absence of measurements than by their inaccessibility: hydraulic parameters measured over six decades remain in unpublished theses and consultancy reports that are not machine-readable and are indexed nowhere. We present a regional groundwater database for Arequipa, southern Peru, containing 4775 georeferenced records recovered from 3675 source documents together with the national monitoring portal, spanning 1966 to 2025. Records carry depth to water, well geometry, saturated thickness, hydraulic conductivity(K), transmissivity(T), storage coefficient(S), lithology and test type, each linked to the page of the document it came from. Extraction combined optical character recognition with a large language model, followed by a three-phase human-in-the-loop protocol in which every recovered value from the principal thesis corpus was checked against its source page; that audit corrected 4.1% of records, predominantly unit misassignments. Because verifying transcription establishes fidelity but not validity, we additionally evaluate the data against evidence independent of the extraction workflow. Reported transmissivity agrees with the product of reported conductivity and saturated thickness within a factor of two for 94% of records carrying all three quantities, with a median ratio of 1.00. Transmissivities re-derived independently from primary pumping-test reports released by the national water authority are located in the database, and the extreme values in two basins reproduce figures published in the corresponding official bulletins exactly. Conductivities from saturated-zone pumping tests fall inside published material envelopes in 96% of cases. The evaluation also exposes the dominant hazard for reuse: the conductivity column mixes seven measurement supports, and pooling them without filtering biases a regional estimate downward by a factor of about 48. The dataset is therefore distributed with explicit unit and support vocabularies and with guidance to filter on test type before any aggregate statistic is computed. Data are archived on HydroShare under CC BY 4.0 (Venegas-Quiñones et al., 2026).