|
<< Click to Display Table of Contents >> Navigation: Big Data, AI and Geospatial Analysis > Conclusions |
Geospatial analysis is fundamentally changing in the wake of the exaflood of new sources and forms of Big Data, and the accompanying development of machine learning and other AI methods to analyze them. From a geospatial perspective, new data sources and their appropriate governance offer finer spatial and temporal resolution – and hence data volume. But the absence of research design in their creation means that new kinds of data may disrupt established approaches to representation of social and physical systems. For example, data-based weather forecasting may already be faster, cheaper, and more accurate than physics-based forecasting. Similarly, cities exist because of interactions, and cellphone information about connectivity between individuals and locations is changing many of our concepts of how cities function.
Widespread re-purposing of new data sources challenges the implied view that GIS software can remain agnostic with respect to purpose. Failure to undertake data engineering and due diligence in assessing the quality and provenance of Big Data creates dangers of unreliable and unethical use. Analysts must anticipate the principle of caveat emptor when using data of undocumented provenance acquired from third parties. Prospectively it may be possible to render GIS software more ethical – for example by flagging workflows that open the possibility of ecological fallacy, for example, or by enhancing the use of metadata to flag uses for which the data are known to be unfit. In the absence of Big Data that are engineered to predefined research ready data standards, this raises much broader questions about assigning the responsibility for assembling functions into workflows largely to the user. It also speaks to the impossibility of educating all GIS users to be aware of every ethical issue.
Machine learning is becoming integral to data driven geospatial analysis in an era of science which is dominated by complexity, since there are seemingly no simple scientific truths left to be discovered. Yet the results of machine learning have been shown to be non-replicable and unreproducible, black box rather than transparent, incapable of dealing with uncertainty, and not at all consistent with traditional science. Machine learning is in important respects akin to curve-fitting, with the scale of its adoption tantamount to abandonment of many of the core organizing principles of science. Explanation, discovery and understanding are all increasingly problematic in a world dominated by machine learning.