The Data Framework Protocol

The Universe of Discourse:

  • Inventory the concepts identified in the “semantic analysis” of the brief and the “Addressing Spatial Omission and Bias” from the “workflow scoping”

Auditing existing data

For each concept in the universe of discourse, list potential sources (CVR, DAR, Field Notes). Do not list all datasets in the world; only those relevant enough to be evaluated. For each dataset/combination of datasets, note the status of the evaluation:

  • Accept datasets that align with the concept in the Universe of Discourse
  • Transform: Datasets/combination of datasets that can be brought to align with the concept in the universe of discourse through different processing steps.
  • Reject: datasets that have been considered but cannot be brought to align with the concept in the universe of discourse. Importantly, note why this dataset, which was initially considered relevant, has been rejected. If a concept cannot be represented in a dataset or a transformation of one or more datasets, note whether the concept is so central that primary data collection is needed, or whether it is better to initiate a backflow to the workflow scoping to remove or redefine the concept.

Domain of Discourse.

For each Cognised Existence defined in the domain of discourse, specify:

  • Thematic specification: including whether the realisations of the Cognised Existence are discrete entities or a continuous gradient.
  • The spatial and temporal resolution
  • Typing: the thematic scale for each attribute (NOIR), and the Geometric Type of the Cognised Existence itself (point/line/polygon/field)
  • Topological and Spatial-Relationship Constraints

Confronting the Sensoric Manifold

Stress test the “domain of discourse” by confronting it with the “Sensoric Manifold”

Re-auditing existing data

Based on the protocol of the phase 2 “Auditing existing data” perform a re-audit of existing data, but this time based not on the “concepts” from the “universe of discourse” but on the “Cognised Existence” of the “domain of discourse”

Recording the Domain of Discourse

For each Cognised Existence, define a schema and populate it:

  • Schema: one spatial table per Cognised Existence, with a geometry matching its declared Geometric Type (point/line/polygon; where the Cognised Existence is a continuous gradient, its Geometric Type is Field — record it as such, e.g. a raster or other field-native structure, rather than leaving it ungeometried), and one field per attribute in its NOIR specification, typed to match (nominal/ordinal as text or coded values, interval/ratio as numeric). Where the Recorded evidence for a Field is itself discrete (e.g. Lidar point returns), interpolating it into the Field is still Recording’s own work, not the Analytical Schema’s.
  • Population — existing data: load the datasets re-evaluated above (Accept/Transform) into the matching schema, applying whatever join, filter, or clip operations were identified as necessary.
  • Population — field data: digitise the Realisations collected at Confronting the Sensoric Manifold into the same schema, reconciling the field record against the schema’s thematic specification, MMU, and topological/spatial-relationship constraints as you go; log any reconciliation judgement call in the Design Rationale.
  • Both populations must land in the same schema: do not give existing-data and field-collected Realisations of the same Cognised Existence different tables or different field sets.