|
<< Click to Display Table of Contents >> Navigation: Geospatial Project Design and Research Practice > Research-ready data > Creation and maintenance of geospatial RRD |
Smart geospatial data are typically recorded as point events with high apparent levels of precision. As with statistical and administrative data, this brings disclosure risks which can most easily be addressed in RRD by geographic aggregation. This approach inevitably brings loss of within-zone geographic detail and, where RRD creation involves multiple variables, creates uncertainty of possible ecological fallacies. Yet the effects of aggregation can be documented by data services if the level of the individual has been retained in the underpinning data infrastructure used in RRD creation. Disclosure risks can be managed by use of secure TREs (Trusted Research Environments) to maintain the underpinning individual level data infrastructure. Data services using such facilities typically have additional up-front costs of ensuring safe researcher training, output checking and other data governance procedures.
The ubiquity of georeferencing technology in many smart data sources makes it possible to retain the locations of individuals in underpinning smart data infrastructure. The reliability of positioning measures is one important aspect of internal and external validation of the quality of smart data infrastructure and should be an important element of derivative RRD metadata. Wide dissemination outside secure data environments means that RRD are unlikely to retain identifiable individual observations: however, individual level georeferencing in underpinning data infrastructure makes it possible to create metadata on the effects of aggregation across a full range of scales, consistent with uncompromised disclosure control. This is important in standalone applications such as residential segregation studies and is of value in addressing challenges of ecological fallacy in multivariate studies.
Georeferenced assemblages of raw smart data rarely link a comprehensive range of individual subject characteristics, leading to what Goodchild (2022) has referred to as ‘Balkanisation of the quantitative self’ if smart data sources are siloed. Different sources may nevertheless be linked at individual level in TREs if the necessary data licensing and data subject consents are in place. Augmentation of multiple datasets with each other is also an important aspect of smart data validation, which is otherwise not possible using pre-aggregated statistical or administrative sources. Retention of the level of the individual can also facilitate greater clarity in representing dynamics through updates over time.
The outcome of triangulation of smart data and their augmentation with other sources is spatial data infrastructure that enables geospatial analysis of linked individual characteristics for any convenient spatial aggregation. This infrastructure can be spliced and diced within a secure research environment into RRD products of known provenance and in response to user demands. Successful delivery of RRD derived from multiple sources requires devising data licensing agreements that are sufficiently flexible to support individual level data linkage and modeling. Innovations in service delivery for RRD thus require extended legal and ethical support alongside improved scientific and technical capacity.