|
<< Click to Display Table of Contents >> Navigation: Geospatial Project Design and Research Practice > Research-ready data (RRD) |
Overview
The best geospatial data are often costly to produce - despite huge innovations in sensor technologies and ubiquitous use of digital devices to record the characteristics and activities of citizens. But once created, they are very cheap to reproduce. The onus is on new innovation-led data services to share up-front costs of database development and maintenance between ever-increasing numbers of research users. Just as the sharing of GIS software development costs a generation ago enabled the GIS industry, so generic ‘off-the-peg’ Research Ready Data (RRD: sometimes called ‘analysis ready data’) are today enabling wide dissemination of exciting new data sources that can enable efficient, effective and safe geospatial analysis.
In this section we describe these developments from the standpoint of socioeconomic data, but the core idea is perhaps best illustrated by remote sensing in environmental science. Classic remotely sensed satellite imagery records the reflectance characteristics of the Earth’s surface, rather than the land cover (and by inference, land use) characteristics that most users require. Surface reflectance turns out to be a direct indicator of land cover, and established and widely accepted procedures exist to assign different reflectance characteristics to land cover classes. These procedures are well-documented, as are caveats to their use that have become evident through particular research case studies. It therefore makes little sense for multiple users of the same data source to repeat the same steps in image preparation: instead, a single ‘research ready’ processed dataset can be made available for wide use, and new research can focus upon innovative analysis and interpretation.
This straightforward example raises several important issues that merit consideration when thinking about RRD for socioeconomic applications:
(a)It is understood that the satellite scene is ‘all that there is’ – that it offers complete coverage of a scene at a level of detail that is understood. If any areas are obscured by cloud, this will be evident and all or part of the image will be replaced by another taken on a cloud free day. Comprehensive coverage means that the data can be construed as data infrastructure.
(b)Cloud cover is not randomly scattered above the Earth’s surface, so it may be necessary to pay more attention to scenes for areas that are systematically prone to cloud cover – such as mountainous areas or locales with convectional microclimates.
(c)Inference of land cover from surface reflectance is widely understood, generally accepted and well-documented. It is a geographically invariant procedure that can be applied to any cloud free dataset.
(d)This said, image classification is a probabilistic process that is inherently uncertain. Moreover, assignments to some land use categories are more uncertain than others, and so the property of uncertainty is not geographically invariant.
(e)Once classified, provision of RRD is ‘vertically integrated’ – that is, the organization that collected the satellite data either pre-processes them itself, or has formal data licensing arrangements with a partner organization to process them, render them fit for a prescribed range of purposes, and make them available to an agreed user base.
Here we develop these observations in the context of RRD for socioeconomic GIS. More data are collected about citizens today than at any point in human history, and more of these data can be described as ‘smart’ – that is, created through everyday interactions between humans and digital devices. As yet, only tiny fractions of these data are consolidated into usable datasets, or are made available to researchers through FAIR (Findable, Accessible, Interoperable and Reusable) channels. Even where smart data do become available, uncertainty about their inherent incompleteness and bias can limit their usefulness in socially inclusive research, although researchers may not be aware of deficiencies in data that they have not themselves collected. This chapter: (a) sets out differences between use of smart data in social investigation and conventional survey-based research designs; (b) assesses the prospects for designing RRD infrastructure that manages the risk of bias in geospatial analysis; and (c) sets out how innovation-led data services can serve smart data to today’s mass research culture.