Establishing the Data Framework

Phase one — Auditing Existing Data and the Sensoric Manifold. Sensoric is used here in its Kantian sense: not the output of hardware instrumentation, but the sensoric manifold (Mannigfaltigkeit der Anschauung)—the raw, unstructured extent of given spatial reality prior to any conceptual synthesis. This phase asks two distinct, sequential questions of the Universe of Discourse identified in Phase One: first, for each concept, does the required data for representing the concept already exist, and if so, how faithfully does it correspond to the ideal concept? Second, where it does not exist, can it be created—through field observation, reclassification, or other data harvesting—and at what cost? Existing data is never ontologically neutral—it was collected, classified, and bounded under someone else’s categories, for someone else’s purpose, and those categories frequently diverge from the workflow’s ideal ontology in ways that are easy to overlook. A land-use dataset built for taxation purposes, for instance, encodes a different implicit ontology than one built for ecological management, even where the two appear, superficially, to describe the same parcels. This phase therefore demands not merely an inventory of what data exists, but an interrogation of the ontology that produced it.

Phase Three — Formalising the Domain of Discourse. Boole himself narrows the unconstrained universe to “a less spacious field” once a discourse’s practical limits are made explicit; this manuscript directly adopts that idea of narrowing the “universe of discourse.” We use the term Domain of Discourse to denote this narrower, more strictly defined version of the Universe of Discourse. Where the Universe of Discourse holds concepts in their unconstrained, ideal form, the Domain of Discourse holds them formalised: we call a concept, once bounded by the five commitments below, a Cognised Existence (die erkennbare Existenz)—a mass noun, denoting a kind rather than a countable instance of it. This formalisation requires five commitments the Universe of Discourse does not yet make.

First, whether the Cognised Existence itself is discrete or continuous—a bounded, individually identifiable phenomenon, or a sampled instance of an unbounded field. This is the most consequential of the five commitments, and it is made first because every commitment after it depends on which answer is given. A single building is unambiguously discrete; a temperature reading, a pollution measurement, or an elevation value is typically a sampled point drawn from a continuous surface that has no natural boundary of its own. This is not a question of measurement, but of ontology: the field/object split is one of the most basic distinctions in spatial thinking itself, prior to and independent of anything this manuscript adds to it, and it must be settled before any of the commitments below are even askable. A continuous Cognised Existence has no boundary to specify in the way a discrete one does, no minimum mapping unit to fall below, and no natural discrete extent for a topological containment or adjacency test to apply to.

Second, an explicit thematic specification for every Cognised Existence: the necessary and sufficient conditions that settle what belongs to it, given the discrete or continuous character already settled above. What those conditions must answer, however, is not the same question in both cases.

For a discrete Cognised Existence, thematic specification is a boundary test: a practitioner, confronted with a specific phenomenon in the world, must be able to decide whether or not it constitutes a legitimate instance of that Cognised Existence. It is not enough to name a category—“building,” “venue,” “road”—without also specifying its boundary conditions. Consider a houseboat, moored permanently at a fixed jetty and used for public events: does it satisfy the Domain of Discourse’s definition of building, or does that Cognised Existence presuppose a fixed connection to land? The answer is not self-evident, and a Domain of Discourse that has not stated it in advance will be answered inconsistently and silently by whichever practitioner happens to digitise it. The consequence compounds where categories depend on one another: if venue is itself defined, elsewhere in this Domain of Discourse, as a building used for public assembly, then the unresolved status of the houseboat as a building is inherited directly by the question of whether it is a venue—the ambiguity does not stay contained to a single phenomenon, but propagates through every Cognised Existence built upon it.

For a continuous Cognised Existence, there is no individual phenomenon to test against a boundary — a field has no discrete instances to admit or exclude. Thematic specification asks a different question instead: what varies within the field, and along which dimensions, before it stops being the same Cognised Existence and becomes a different one? Noise is a useful case. A single sound-pressure reading, taken at one place and one moment, is a discrete measurement of a continuous field — but is “noise,” as a Cognised Existence, one field that simply varies with time of day, or two: a daytime/evening Noise and a night-time Noise, evaluated against different thresholds because the same decibel reading carries a different meaning at 2pm and at 2am? Neither answer is self-evident, and the Domain of Discourse must commit to one before the field can be evaluated at all — an unstated commitment here does not stay contained to noise, any more than the houseboat’s did to building: whatever Evaluation Matrix or Analytical Schema later treats “acceptable noise” as a single threshold is silently assuming the one-field reading, whether or not that assumption was ever made explicitly.

Thematic specification, in either form, is an intensional definition of each Cognised Existence — a statement of what a phenomenon must be, or what a field must remain, rather than merely a list of phenomena or readings it happens to include — and, together with the discrete/continuous commitment above, it is the precondition for every commitment that follows: MMU, typing, and topology can only be meaningfully applied to a Cognised Existence whose membership, or whose identity as a single field, has already been settled.

Third, a defined spatial and temporal resolution—though what this resolution consists of now depends on the discrete/continuous character settled at the first commitment. For a discrete Cognised Existence, the key element of spatial resolution is a Minimum Mapping Unit (MMU)—below which distinctions defined in the thematic specification are not, in practice, honoured. A continuous Cognised Existence has no such minimum unit to fall below; its analogous commitment is a minimum sampling density or interpolation resolution—how finely the underlying field must be sampled before the surface built from it can be trusted—treated in full alongside the discrete/continuous distinction in the chapter Operationalising Confronting the Sensoric Manifold: data specification. The temporal resolution plays the same role in time, i.e. a minimum temporal precision below which distinctions defined in the thematic specification are not, in practice, honoured, and it applies to both the discrete and continuous cases alike. Fourth, an explicit typing, in two parts. The first part is measurement typing for every attribute: nominal, ordinal, interval, or ratio (NOIR), following Stevens (1946). The second part of typing this manuscript must add to Stevens, since spatial representation was no part of his problem in 1946: an explicit Geometric Type for every Cognised Existence—Point, Line, Polygon, or Field—stating which spatial operations are legitimate upon it, exactly as NOIR states which statistical operations are legitimate upon an attribute. A road may be typed as a Line for a network analysis or as a Polygon for a land-use calculation; either is a legitimate Geometric Type, but the choice is not free, and an area operation is no more askable of a Line than a mean is askable of a nominal attribute. A continuous Cognised Existence is typed, once and for good, as a Field—this follows directly from the first commitment, and is not a fresh decision made here. Fifth, an explicit specification of the topological and spatial-relationship constraints governing how entities of the Domain may legitimately relate to one another—not merely what each entity is, but what configurations between entities are permissible or impossible. A road network and a hydrological layer may each be perfectly well-typed in isolation, yet the Domain of Discourse remains incomplete until it states, for instance, that a road may not intersect a lake except at a bridge, or that administrative boundaries must not leave gaps or overlaps where full coverage is assumed. These are not properties of any single entity, but constraints on the relations between entities, and without them the formalised model cannot detect a large and common class of error—one that neither MMU nor NOIR typing, applied independently to each attribute, is capable of catching.

These five commitments are not bureaucratic housekeeping; together they are much of what allows the Analytical Schema to be validated before execution rather than failing silently within it. A schema step cannot legitimately compute a mean across nominal land-cover codes, apply a Euclidean buffer to an attribute that describes a continuous field rather than discrete, bounded geometry, or run a network analysis over a road layer whose topology has not been verified as connected—and none of these checks are even askable of an entity whose thematic membership in the Domain was never settled in the first place, or whose discrete or continuous character was never settled before that. The formalised Domain of Discourse thus functions as more than a description of the data—it is a gate the Analytical Schema must pass through.

It is not, however, the only gate. The five commitments—discrete/continuous character, thematic specification, MMU, typing (NOIR and Geometric Type), and topological/spatial-relationship constraints—formalise what a Cognised Existence is and which operations are legitimately askable of it at all. They do not, however, guarantee that its digital representation is fit for the operations the Analytical Schema will actually perform. A failure already introduced in this manuscript’s Introduction illustrates the difference: a buffer computed on data still held in a geographic, unprojected coordinate system does not correspond to any real-world distance, however correctly that data was typed, geometrically typed, and bounded above. This is not an inadmissible operation in the sense the five commitments police—a buffer is exactly what one legitimately does to a Polygon or Line—it is a computable operation silently returning a wrong answer. None of the five commitments is capable of catching it, because none of them concerns the fidelity of the coordinate space the geometry is recorded in. Securing a coordinate reference system fit for the Schema’s intended operations is Recording’s distinct responsibility (Phase Five, below), not a further test the five commitments already perform. Until Recording has been done correctly, the Domain of Discourse’s formal correctness is necessary, but not sufficient, for the Analytical Schema to execute safely.

Formalising the Domain of Discourse is, however, not the final step of this activity but the trigger for one more: a Backward Glance to Phase Two, re-evaluating any candidate data identified there against the now-rigorous criteria this phase has established. A dataset may satisfy the loose Concept of Phase One while failing the more rigorous Domain of Discourse of Phase Three—most commonly where its scale of observation does not honour the MMU, its attributes are typed too coarsely, its thematic boundaries were drawn differently than this Domain of Discourse now requires (a secondary venues dataset, for instance, may silently include or exclude houseboat venues the Domain deliberately excludes or includes), or its discrete or continuous character does not match what this Domain of Discourse now requires (a dataset built from discrete station readings, for instance, may not yet be the continuous Field the Domain of Discourse commits to). It is this failure, discovered only once formalisation makes it visible, that the Epistemic Choice below exists to resolve. Phase Four — Confronting the Sensoric Manifold. Only once the Domain of Discourse has been formalised does the practitioner turn to the sensoric manifold proper—the raw, unstructured given of spatial reality prior to any conceptual synthesis, in the strict Kantian sense (Mannigfaltigkeit der Anschauung). This is not a further audit of data already collected by others, but direct confrontation with the field itself, and it is always performed, even where the expected outcome is simple confirmation. This confrontation has a single function: validation—does the formalised Domain of Discourse survive contact with the manifold, or does field observation reveal that the discrete/continuous character, categories, MMU, or topological assumptions formalised in Phase Three do not, in practice, hold? Where an existing dataset survived Phase Two, validation can test it against both its own proclaimed domain of discourse and the efficient domain of interest its actual content reveals; where no dataset survives, it can only test the Domain of Discourse’s own proclaimed domain of interest against the manifold directly. This is not collection, and the phase performs no second, parallel act of harvesting—but where no dataset survives, the physical activity of validating is largely indistinguishable from the physical activity of collecting data; the difference is purpose, not motion. What results, incidentally, is a set of validated Realisations, which it is Phase Five’s task to give schema shape. (Data collection methodology proper is treated in full in Chapter [X]; this phase establishes only its epistemic position within the Data Framework.)

The Epistemic Choice: The confrontation between Phases One and Two routinely surfaces a gap the formalised model of Phase Three cannot simply paper over. Should the ideal data model be compromised to match existing, flawed, or coarsely-typed data? Or must the workflow absorb the labour-intensive burden of harvesting new data—collecting field observations, reclassifying existing categories, refining the MMU—to satisfy the model as originally conceived? Neither answer is free, and the choice made here is itself a piece of Design Rationale, not a technical footnote.

Phase Five — Recording the Domain of Discourse. The four preceding phases are conducted, deliberately, without committing to any particular software or storage format—formalising the Domain of Discourse is not a technical task of creating database tables, but the act of operationalising the Universe of Discourse into rules a Realisation can be tested against. Those rules must nonetheless, at some point, be given a physical, digital shape: a spatial table whose geometry type, field names, and field types faithfully carry the Domain of Discourse’s discrete/continuous character, thematic specification, MMU, typing (NOIR and Geometric Type), and topological/spatial-relationship constraints — and whose coordinate reference system is, in addition, fit for whatever operations the Analytical Schema will actually perform. A Realisation recorded in a geographic rather than a projected coordinate system where the Schema requires a metrically accurate buffer will satisfy the Domain of Discourse’s five commitments perfectly and still cause the Schema to fail silently downstream. Closing this second gate is Recording’s distinct responsibility, not a restatement of Phase Three’s. Recording is this instantiation, and it happens twice over, not once: once for the Realisations re-evaluated from Auditing Existing Data (Phase Two, revisited), and again for the Realisations validated at Confronting the Sensoric Manifold (Phase Four). Both provenances converge on the same schema, because both are Realisations of the same formalised Domain of Discourse—the practitioner should not be able to tell, from the finished spatial table alone, whether a given row was transformed from an existing register or digitised from a field observation.

A brief terminological note is owed here, since the next activity uses “GIS” in a way that could otherwise seem to contradict this phase. Recording typically does involve GIS software—a spatial table is, after all, usually built and populated inside a desktop GIS or a spatial database. But it uses that software as an instrument of record: creating a schema and entering or digitising data into it commits nothing about how that data will later be analysed. This is a different use of the technology stack from the one Activity 5 describes below, where GIS is used as an instrument of analysis—executing the mathematical verbs the Analytical Schema specifies. The distinction, not the timing of software use, is what “GIS enters the workflow” refers to at Activity 5.

Authoring the Analytical Schema

Design the abstract algorithmic logic (the mathematical “verbs”) required to bridge the gap between the chaotic data and the scientific intent. This must maintain strict tool-agnosticism. The schema dictates what must happen mathematically, not how the software will achieve it.

The Schema’s only legitimate starting point is the Recorded Domain of Discourse—the Realisations given physical, digital shape at Phase Five of Establishing the Data Framework, whether by adopting and transforming existing data or by digitising newly collected field observations. Neither the unconstrained Concepts of the Universe of Discourse (Phase One) nor the raw Sensoric Manifold itself (Phase Four) is an admissible input here: a mathematical verb cannot be applied to a phenomenon that has not yet been typed, bounded, and given a geometry, any more than a statistical operation can be applied to a measurement that has not yet been assigned a NOIR level. It is precisely Recording’s work—settling geometry type, field types, and coordinate reference system—that makes the Domain of Discourse admissible as the Schema’s raw material. Authoring the Analytical Schema is, in this sense, entirely dependent on Establishing the Data Framework having already closed both of its gates.

The most useful way to conceive of the Analytical Schema itself is as a Directed Acyclic Graph (DAG): a network of nodes connected by one-directional edges, containing no path that loops back on itself. Each node in the Schema’s DAG is a single spatial operation—a buffer, an intersection, a spatial join, a reclassification—and each edge represents a Realisation flowing from the output of one operation into the input of the next. The graph is directed because data only ever flows forward, from Recorded input toward final result, never backward into an operation that has already run; it is acyclic because no operation may legitimately depend, even indirectly, on its own output—a constraint that is not a technicality but the formal guarantee that the Schema terminates in a finite, well-ordered sequence of steps rather than an undefined circular dependency. An operation’s output is not itself a return to the Domain of Discourse—it is a new, derived Realisation, typically of a Cognised Existence the Domain of Discourse never needed to define, because it exists only as an intermediate artefact of the analysis (a distance-to-nearest-venue value, a reclassified buffer zone, a filtered subset of noise incidents). The Schema must nonetheless remain as explicit about the thematic meaning and typing of these derived, intermediate Realisations as the Domain of Discourse was about the Recorded ones feeding it, since an operation three nodes downstream is just as capable of failing silently on a poorly-typed intermediate output as the first operation was on poorly-Recorded input.

Conceived this way, the DAG is also what secures the Schema’s tool-agnosticism. The graph of nodes and edges specifies only which mathematical verb transforms which input into which output, and in what order—it says nothing about whether that verb is executed by a QGIS processing algorithm, an ArcGIS geoprocessing tool, or a line of Python calling a spatial library. The same DAG can be translated into more than one tool-bound Analytical Recipe, and it is only at that translation—not at authoring—that a specific Stack enters the picture.

4. Designing Spatial Results Communication

While a Geospatial Workflow can yield many deliverables, the brief will often specify the communication of results via reports, maps, and graphs. It is worth being precise, at the outset, about what design means in this context. This manuscript uses the term not in its colloquial sense of layout or aesthetic composition, but in the sense of suitability: design, here, denotes the fitness of a representational choice to the data it represents and the audience it must reach.

Two distinct suitability judgements govern this Vantage Point. The first concerns what kind of deliverable the result should take. MacEachren’s cartography cube (MacEachren and Taylor, 1994) is useful here: it positions any map or visualisation along three axes—from public to private audience, from presenting knowns to revealing unknowns, and from low to high interaction. A static report map intended for a public planning consultation occupies a very different position in this cube than an interactive digital twin built for a client’s internal exploratory analysis, and the practitioner’s first task is to judge, from the brief established at Scoping, where along these three axes the deliverable ought to sit.

The second judgement concerns how the chosen deliverable represents the underlying data, and here the practitioner’s decisions carry the same epistemic weight as any made at the Analytical Schema. Consider classification method alone: the same continuous surface, classified by equal interval, equal count (quantile), or natural breaks (Jenks), will produce visibly different maps from identical underlying data. Equal interval emphasises the extremes of a distribution, at the cost of collapsing distinctions among values clustered near the mean; equal count guarantees a balanced visual weighting across classes, at the cost of potentially separating near-identical values across a class boundary and merging quite different ones within it. Neither is more “correct” in the abstract—each is more or less suitable depending on what the map is meant to reveal, and to whom.

Because these representational choices depend directly on the statistical character and uncertainty of the analytical output, arriving at this Vantage Point requires a Backward Glance to the Analytical Schema: does the chosen representation honestly convey the uncertainty and distributional shape of the result, or does it imply a precision or evenness the underlying analysis does not support? Where it does not, the practitioner returns—not to redesign the map alone, but to establish, in dialogue with the Schema, a representation that is suitable to what was actually computed. Only once both judgements—deliverable and representation—are secured is the mathematical result legible to the Lifeworld it was produced to serve.

5. Executing the Analytical Schema

Only once the semantic workflow is fully documented is the logic delegated to the computational engine—the operational GIS—for syntactic execution. It is worth pausing here to restate, at the point where it matters most, the terminological convention this manuscript has held throughout: GIS, in this text, denotes strictly the technology stack—the software or code through which spatial logic is carried out, whether a desktop application such as QGIS or ArcGIS, or a heterogeneous ecosystem of Python or R libraries. Software of this kind may already have been used earlier, at Recording (Establishing the Data Framework, Phase Five), to build and populate a spatial table—but there it served only as an instrument of record. It is precisely at this fifth activity, and nowhere before it, that GIS in this narrow sense enters the workflow as an instrument of analysis.

This delayed entry is deliberate. The preceding four activities were conducted in deliberate tool-agnosticism: the Analytical Schema authored in Activity 3 specifies the mathematical verbs—buffer, overlay, interpolate, aggregate—required to bridge data and intent, without committing to how any particular software will carry them out. Execution is the activity in which that abstraction is finally surrendered. The Schema does not survive this activity unchanged; it is translated, and the translation is neither trivial nor value-neutral.

From Schema to Recipe. The product of this translation is what we termed, among the Semantic Assets, the Analytical Recipe: a decoupled, reproducible, and now tool-bound script—a Directed Acyclic Graph of operations—capable of being executed, audited, or handed to another practitioner (human or computational agent) without requiring their presence in the room. Where the Schema answers what must happen mathematically, the Recipe answers how, in this specific software environment, it was made to happen. The same Schema may legitimately yield several different Recipes: a chained sequence of native tools within a desktop GIS, a scripted pipeline in Python using GeoPandas and rasterio, or a hybrid of both. None of these Recipes is more “correct” than another with respect to the Schema; they diverge only in their operational form, not in their semantic intent.

The translation is not free. Practitioners should resist the assumption that this step is merely mechanical. Software-specific defaults—interpolation algorithms, topology tolerances, coordinate transformation methods, null-value handling—frequently reintroduce exactly the kind of tacit, unexamined decision-making that the earlier activities were designed to surface and document. A Schema that specifies “interpolate a continuous surface from point observations” is genuinely software-agnostic; the moment a specific interpolation method, kernel, or tolerance is chosen, the workflow has re-entered the operational domain. Where such choices meaningfully affect the analytical outcome, they must be returned to Workflow Stewardship as an addendum to the Design Rationale—not silently absorbed into the Recipe as an implementation detail. The boundary between Schema and Recipe is porous in practice, even where it is sharp in principle.

Delegation and its limits. It is at this activity, and only this activity, that delegation to an autonomous computational agent becomes a legitimate option rather than a pedagogical shortcut. Because the Schema was authored independently of any execution environment, and because the Design Rationale documents the reasoning behind it, a well-formed Schema can, in principle, be handed to an AI-supported coding environment for execution with the human intent still fully auditable. This is precisely the capability that renders the procedural knowledge of button-pressing dispensable, as argued in the Introduction—but it is only trustworthy to the extent that Activities 1 through 4 were rigorously performed. An agent asked to execute an ill-formed or under-specified Schema will not surface that failure; it will simply produce an output, syntactically valid and semantically wrong. Execution, whether performed by hand or delegated, is therefore never a check on the quality of the preceding four activities—it is a test that only a well-stewarded workflow can pass, and a poorly-stewarded one will fail silently.

The Reality of Iteration (Workflow Stewardship)

The five activities of the Geospatial Workflow are best understood not as sequential steps but as five Vantage Points the practitioner occupies and revisits over the course of a workflow. Each Vantage Point discloses an aspect of the workflow that no other vantage can: Scoping reveals human intent and constraint; the Data Framework reveals ontological fit between the ideal and the given; the Analytical Schema reveals the mathematical logic required to bridge them; Communication reveals legibility to the Lifeworld the workflow ultimately serves; and Execution reveals the operational reality of a specific computational engine. No single Vantage Point offers a complete view of the workflow—this is precisely why the practitioner must be able to move between them. Two are fixed in position: Scoping is always the first Vantage Point occupied, and Execution is always the last. Between these two fixed points, the practitioner may return to any Vantage Point as often as the Domain Project’s complexity demands—authoring the Schema may reveal a flaw visible only from the Data Framework; designing Communication may surface a gap only Scoping can resolve. What matters is not the order in which these Vantage Points are visited, but that the practitioner always knows which one they currently occupy, what it does and does not show them, and can account—through Workflow Stewardship—for why they moved.

Movement between Vantage Points is not merely spatial repositioning; it is accompanied by the Backwards Glance—the phase, within Workflow Stewardship, in which the practitioner identifies what backflow is required and documents its cause. A Backwards Glance is not confined to the immediately preceding Vantage Point. Because each Vantage Point depends on the integrity of every Vantage Point before it, a single unresolved question can cascade backwards through the entire chain. Discovering, while authoring the Analytical Schema, that private-hire and taxi services were never included in a transport dataset is, on its surface, a Data Framework problem—but it frequently is not, at root: it may reveal that the scope of “transport” was never explicitly agreed with the client at all. The Backwards Glance, in such cases, does not stop at the nearest plausible vantage; it travels back to wherever the unresolved question actually originates—here, to Scoping itself, where the boundary of the Domain Project’s intent must be renegotiated before the Schema can proceed. This cascading property is not a failure of the framework; it is the mechanism by which Workflow Stewardship keeps the entire chain of Vantage Points honest to one another.

The Epistemological Nature: Heuristics vs. Teleology

In the introduction, we presented the activities of a geospatial workflow as a series of vantage points from which different aspects of the workflow can be observed. It is important to understand that the path connecting these vantage points is not one-way, and there is almost always some going back and forth between them, especially if the project is a bit more complex. Another factor that determines how we move between the vantage points is the type of workflow. It is here relevant to distungish betwee two types of workflow Exploratory (Heuristic) and Prescriptive (Teleological) workflows

Exploratory (Heuristic) Workflow: The Path of Discovery

The brief poses an open-ended question where the outcome is unknown (e.g., “Investigate the spatial correlation between urban canopy cover and respiratory health”).

  • The Intent: To uncover patterns, relationships, or anomalies hidden within the data’s extent.

  • The Process: This follows the numerical order of the five activities. However, it is rarely linear; it is characterised by Heuristic Backflow, in which findings from Authoring the Analytical Schema may require a return to Establishing the Data Framework or even a refinement of the initial Intent in the project scoaping.