Data Note: Geocoded Coordinates for a Lower Palaeolithic Site Compilation from India
Pardeep Kumar — Noosphere Engine
This data note documents the geocoding of 1,685 Lower Palaeolithic site records from India and their classification into four location-confidence tiers, for inclusion in the Noosphere Engine’s Paleolithic Sites layer. The underlying site names, districts, and states originate from a compilation published by Vidyarthi and Chauhan (2025); the geocoding procedure described below is independent of that study’s methodology. The two contributions — source data and geocoding method — are attributed separately throughout.
1. Provenance of the Source Data
The site name, district, and state fields used in this release derive from the compilation published in Vidyarthi, V. and Chauhan, P. (2025), “Mapping the Indian Palaeolithic,” Journal of Computer Applications in Archaeology 8(1): 78–93 (DOI: 10.5334/jcaa.148), itself compiled from Indian Archaeology – A Review (Archaeological Survey of India, 1953–1976). The source compilation is also archived on Zenodo. Credit for the identification, compilation, and publication of these site records rests with the original authors. The working file used for this release contained 1,685 records, exceeding the 1,535 Lower Palaeolithic sites reported in the published study; the source of this discrepancy has not been independently established and may reflect an extended or subsequent version of the underlying compilation.
2. Data Cleaning
Prior to geocoding, the source records required the following corrections:
- Of 1,685 records, 202 lacked a usable site name, 258 lacked a district, and 212 lacked a state. Placeholder text (e.g. “not given”, “NA”) was treated as a true blank rather than as data.
- State designations were inconsistently coded (e.g. “Mh” versus “MH”; “HP/PN”; “J and K” versus “JK”) and were normalised to a single code per state. One record, with state recorded as “HP/PN”, remained genuinely ambiguous between Himachal Pradesh and Punjab and was left flagged rather than resolved by assumption.
- Eighty-five site names recur across different district/state combinations in the source data. These were treated as distinct, independently reported occurrences rather than duplicate entries, and processed individually.
- Extraneous whitespace and inconsistent character spacing in name fields were normalised.
3. Geocoding Method
The geocoding procedure applied here is independent of the ESRI ArcGIS World Geocoding Service used in the source publication. It proceeds as follows:
- Reference gazetteer: GeoNames’ India export (approximately 660,000 named features, CC BY 4.0), comprising administrative divisions and populated places with associated coordinates.
- Tiered matching: for each record, a match was first attempted for the site name within its stated district and state; where unsuccessful, the district centroid was used; where the district could not be matched, the state centroid was used.
- String matching: fuzzy matching was performed using a Levenshtein ratio scorer, in preference to partial/substring scoring, to reduce false positives arising from short or generic place-name elements common in South Asian toponymy.
- Manual disambiguation was not performed. Every match reported here was produced by the automated procedure described above, without recourse to secondary literature or remote-sensing imagery. This is the principal methodological divergence from the source study, which resolved ambiguous candidates manually using supplementary sources.
4. Results
| Confidence tier | Records | Proportion | Definition |
|---|---|---|---|
| Precise | 260 | 15.4% | Site name matched and confirmed within the stated district |
| Precise (state-wide) | 584 | 34.7% | Site name matched; district not independently confirmed |
| District | 346 | 20.5% | District matched; district centroid used |
| State | 283 | 16.8% | State matched; state centroid used |
| Unmatched | 212 | 12.6% | Insufficient source information to locate at any level |
5. Comparison with the Source Study’s Results
The source study reports 57% of its 1,535 records geocoded with confidence on first pass, 39.6% requiring manual disambiguation among multiple candidates (“tied” results), and 3.5% unmatched. The corresponding figures obtained here — 50.1% classified Precise, 37.3% resolved only to a district or state centroid, and 12.6% unmatched — occupy a broadly comparable range at the point of automated matching, but the two sets of figures are not directly commensurable. The 37.3% District/State-tier proportion is the automated analogue of the source study’s 39.6% “tied” category; in the source study, tied candidates were subsequently resolved to precise coordinates through manual cross-referencing, a step not performed here. This accounts for the greater part of the difference between the present unmatched rate (12.6%) and that reported in the source study (3.5%). A further portion of the present unmatched category reflects missing information in the source records themselves (absence of any usable name, district, or state), which no geocoding procedure could resolve without recourse to the original archival volumes.
6. Limitations and Intended Use
Consistent with the source study’s own characterisation of its results, the coordinates presented here indicate a probable general location rather than a verified precise findspot; this applies to records in the Precise tier and more markedly to those in the District and State tiers. The dataset is suited to regional-scale pattern analysis (distribution, density, and clustering by area) rather than to findspot-level inference. Because a substantial number of District- and State-tier records share an identical fallback coordinate, raw point density on a map will understate site counts in areas dominated by coarser-tier matches relative to areas with a higher proportion of Precise matches; density comparisons across regions should account for this rather than relying on point counts alone.
7. Data Availability and Licensing
Site names, districts, and states: derived from Vidyarthi and Chauhan (2025), cited in Section 1. Reference gazetteer: GeoNames.org (CC BY 4.0). The coordinates, confidence classification, and processing documented in this note are made available for reuse with attribution to both sources cited above.