Primary Data Collection: Field Work
Related: 2.2 Establishing the Data Framework (Part 2) — this exercise performs the fieldwork that chapter narrates.
This note operationalises the fifth phase of Establishing the Data Framework Recording the Domain of Discourse — into a concrete field exercise. Part 2’s Hermetic Example Workflow showed you what this phase looks like on paper: Public Social Space, having no register to Accept or Transform, was sent to Direct Procurement, and the fieldwork was described rather than performed. Here, when you perform it. This note assumes the schema-design work is already done — see Operationalising Confronting the Sensoric Manifold: data specification if it isn’t. You will already have a Domain of Discourse containing at least one Cognised Existence of your own choosing, and a schema built from it and stress-tested against the coincidence problem that note describes. What follows is concerned only with populating that schema: taking it into the field, testing whether it survives contact with reality, and bringing back a Realisation.
From Observation to Recording
According to Kant, we can never observe reality as it is in itself — the Ding an sich, the noumenon. What we have access to is phenomena: the objects, events, and signals that appear to our senses or instruments, before we have assigned them any formal meaning. We can observe phenomena; however, in order to record them, we need the detailed definitions associated with our Cognised Existence of the outer Domain of Discourse.
Our Domain of Discourse is the rulebook that lets us move from unstructured Observation to structured Recording: the act of creating data. “Unstructured” and “structured” are the two ends of a continuum, and where a given fieldwork session sits on it depends on how tightly the Domain of Discourse has been specified. A large, multi-participant project running over months needs a far more tightly defined Domain of Discourse than a single afternoon’s solo survey — without it, individual recorders will silently diverge from one another.
Even a carefully formalised Domain of Discourse rarely survives contact with the complexity of Sensoric Manifold unchanged. This is why, before doing any recording, we always confront our domain of discourse with the “Sensoric Manifold” as described in the section Confronting the Sensoric Manifold; In larger projects, this is done as part of a formal pilot study where we not only stress-test our “Domain of Discourse” but also calibrate our understanding of it within the group doing the fieldwork. For more ad hoc projects, the confrontation with the “Sensoric Manifold” is a less formal walk around, finding one or two realisations of each Cognised Existence to ensure a match.
Mapping Is a Sample, Not a Census
Whatever you record in the field is a sample of an underlying population, not a complete census of it — whether the population is every user of a park over a year, or every building in a municipality. This raises the question of whether the sample is representative of the population and how to ensure the best possible representation given the resources. A classic illustration of this problem comes from the statistician Abraham Wald’s work for the U.S. military’s Statistical Research Group during the Second World War (Wald, 1943/1980; see Mangel and Samaniego, 1984, for an accessible account of the method). Engineers wanted to know where to add armour to aircraft, and had assembled damage records like the one below
the observed distribution of bullet and flak holes across aircraft that returned from combat missions. Their first instinct was to reinforce wherever the damage clustered most densely. Wald’s insight was the reverse. Every aircraft in that dataset, however badly marked, had made it home — its damage pattern was evidence of where a plane could be hit and still survive, not of where it was vulnerable. The aircraft that should have driven the decision, those hit in the engines, the cockpit, the fuel lines, were missing from the sample precisely because they did not return to be counted. The correct inference was to armour the areas with the fewest observed hits, not the most, because the absence of data there was itself the data.
The same structure recurs in every field mapping exercise. What gets recorded is not a neutral cross-section of the underlying population but whatever survived the process by which the practitioner came to observe it — whatever was accessible on foot, visible during the hours surveyed, safe enough to approach, or simply lay along the route the surveyor happened to walk. A map of observed instances is, in this sense, always a map of the sampling process as much as it is a map of the phenomenon itself. The temptation, exactly as it was for the wartime engineers, is to read areas of dense observation as areas of real interest, while the blank areas on the map are quietly read as absence of the phenomenon rather than as absence of access to it. The corrective is the same as Wald’s: before drawing conclusions from where the data is, the practitioner must deliberately ask what could not have made it into the sample, and why — and treat that question as being just as informative as the pattern actually observed.
has two immediate consequences for how you plan a session.
Spatial and temporal autocorrelation. People do not distribute themselves randomly across space, and events do not distribute themselves randomly across time. The people using a park have each chosen to be there — their positions cluster, rather than scatter evenly, which is itself a form of positive spatial autocorrelation. Counts of any recurring activity — cyclists at an intersection, cars on a street — vary systematically by time of day and day of week rather than holding constant. A sampling protocol that ignores this (a single fifteen-minute count taken at an arbitrary hour) will not generalise to the population you actually care about, and this is a design decision to make explicit in your Design Rationale, not an afterthought to note once the data looks strange.
What We Register: Object Classes and Properties
Two kinds of concept get registered in the field, and the distinction maps directly onto the thematic specification work you have already done in your Domain of Discourse.
Object classes describe discrete, countable phenomena understood to have a well-defined boundary — a building, a lake, a parcel of land, a bench. This is the same discrete Cognised Existence that thematic specification defines.
Properties describe the measurable characteristics of a phenomenon, and — echoing the discrete/continuous distinction already made under thematic specification — a property can be conceptualised in two different ways. As an attribute of an object class, a property is a single value attached to a discrete object: room_temperature is an attribute of a specific room. As a field, the same kind of property varies continuously across space, independent of any single bounded object: outdoor temperature across a landscape is a field, not an attribute of anything in particular. Deciding which of the two applies to a given property of your Cognised Existence is not a recording detail — it determines which observation mode (below) is even capable of capturing it.
The Phases of Data Recording
A field campaign moves through five phases, and each maps onto a Vantage Point you already know:
- Design the preliminary ontology — desk work, building a first-pass Domain of Discourse from prior knowledge and literature, before any fieldwork.
- Test the ontology (Pilot Fieldwork) — the purpose is to test the recording rules the preliminary Domain of Discourse implies against reality in the Area of Interest. Unstructured tools (paper, notebooks) are used deliberately, to maximise the flexibility needed to notice where the model breaks.
- Refine the final ontology — the Backward Glance triggered by the pilot: modifying the Domain of Discourse to match what the pilot actually revealed.
- Define the data schema — specifying how the refined Domain of Discourse will be stored: this is where the Attribute Framework (the NOIR-typed table) is committed to.
- Execute the final registration — a new, structured recording pass using apps, pre-printed forms, or other structured tools, applying the now-refined ontology consistently.
Pilot Fieldwork versus Final Registration. Whichever of the four observation modes below you use, the pilot and the final registration differ in three things: your goal, your tools, and your required level of precision. In the pilot, the goal is exploration — you use unstructured tools (notebooks, paper maps) specifically to test the ontology and surface its weaknesses. In the final registration, the goal is production and accountability — you use structured tools (survey apps, pre-printed forms) to apply the now-finalised ontology consistently and produce clean data.
The Four Modes of Observation
Each mode answers a different core question about the phenomena you are observing, and each produces a different kind of geospatial output — which is exactly why the choice of mode is not incidental to your Domain of Discourse, but a direct consequence of it: a discrete, point-like Cognised Existence calls for a different mode than a continuous field.
| Mode | Core Question | Pilot Approach | Final Registration | Geospatial Translation |
|---|---|---|---|---|
| Inventorying | What and where are the things? | Notebook and map: list what you find (“found a strange concrete bench — not in my schema” is a successful pilot finding, not a failure). | Survey app with a pre-defined form: a new feature per object, attributes filled in (material: 'wood', condition: 'good'). | Point or polygon features, one per discrete entity (e.g. trees with species/height, building footprints with land_use). |
| Delineating | What are the zones, and where are their boundaries? | Sketch a boundary on a paper map — the goal is not a perfect line, but finding where your classification rules get fuzzy (e.g. the border between a “picnic area” and a “sports area”). | GPS-enabled app, or desktop digitisation over aerial photos: no longer questioning the rules, only applying them consistently. | Polygon or raster features (land cover, soil type, image-classified surfaces). |
| Tracing | Where does it go, and what path does it take? | Sketch the path of a bicycle, pedestrian, or animal by hand — a rough record of the nature of the path itself. | Walk the path or attach a GPS tracker to the object being tracked, producing an accurate LineString. Tracing can also be elicited through interview (ask the informant to draw their route) or by observing desire lines. | Line features — a trajectory (a GPS track, a pedestrian route, a stream’s path). |
| Monitoring | How does the frequency of an event change over time at this location? | Stand at a fixed point with a tally counter for a short interval, and note anything that biases the count (“a delivery truck blocked the intersection, stopping all cyclists” is a pilot finding about your protocol, not noise to discard). | A permanent sensor records counts automatically, producing a clean, structured time series (location_id, timestamp, count). | A numerical attribute, typically a time series attached to a fixed location. |
A fifth, indirect mode: reading traces. Not every observation is of the phenomenon directly. Human activity leaves traces — litter, worn dirt patches on grass, desire lines across a lawn — that are themselves evidence of use, and can be registered through counting, photographing, or mapping even where the activity that produced them was never directly witnessed. Treat this as a variant of Inventorying or Monitoring depending on whether you are recording a static trace or its accumulation over time, but flag it explicitly in your Design Rationale: a trace is a proxy for a phenomenon, not the phenomenon itself, and inherits all the caveats a Proxy Entity carries elsewhere in this manuscript.
Tools for Recording in the Field
No single tool suits every mode or every project. A few, in rough order of formality:
- Hand-drawn interview or observation maps — flexible and cheap, but time-consuming to digitise afterwards. Aggregated hand-drawn maps can be digitised more efficiently using printed ArUco markers (print-get-map.org).
- Mobile mapping apps (ArcGIS Field Maps, Survey123) — structured, GPS-located recording, well suited to Final Registration. Can feel alienating or slow in an interview setting.
- Web maps (e.g. Esri Survey123 forms) — remote or self-service data entry.
- Mobile phone tracking / fitness apps — needs either the informant to install an app, or researcher access to phone-tracking data, which can be hard to obtain and raises consent questions; collects only the physical location of the informant.
- Dedicated tracking devices — can feel intrusive, produce easy-to-use data, but are relatively expensive and must be physically retrieved.
- For Danish student projects: kort.plandata.dk/spatialmap and combining QGIS with Dataforsyningen data are established local starting points.
Gehl’s qualitative tools are a useful complement to the four modes above, particularly where the phenomenon resists clean quantification: photographing (documenting where urban life and form interact, or fail to), keeping a diary (recording nuance that can be categorised or quantified later), and test walks (a semi-systematic walk along a route, aimed at noticing problems and potentials rather than counting anything).
Recording through interviews. Letting an informant draw their own map is a distinct way of eliciting spatial knowledge, and it requires deliberately balancing cartographic precision against expressive freedom — a sketch map or mental map is not a survey, and should not be treated as one. Variants include using physical tokens to mark locations on a base map, “air-brushing” intensity onto a digital map, the walkalong interview (walking a route together with the informant, recording their commentary in place), and the sensor-enriched post-walk interview (reviewing GPS or video footage of a walk together with the informant afterwards, to elicit commentary on what it shows).
Exercise: Your Own Primary Data Collection
You already hold a formalised Domain of Discourse for at least one Cognised Existence from an earlier exercise. This exercise takes it through the phases above, in order:
- Choose one Cognised Existence from your Domain of Discourse, and identify whether it is discrete or continuous — this determines which observation mode(s) below apply to it.
- Select the observation mode(s) that fit: Inventorying or Delineating for a discrete entity or a zone; Tracing or Monitoring if temporal dynamics are part of its thematic specification.
- Conduct a Pilot Fieldwork session using unstructured tools only (notebook, paper map, or a hand sketch). Your goal is exploration, not clean data — actively look for phenomena your Domain of Discourse did not anticipate, and write them down as they occur, however awkward the fit.
- Perform a Backward Glance. Where the pilot surfaced an ambiguity — a boundary condition your thematic specification did not settle, an MMU that excluded something that mattered, a topological or spatial-relationship constraint you had not stated — record it explicitly in your Design Rationale, and refine the Domain of Discourse accordingly. An ambiguity discovered here is a successful outcome of the exercise, not a failure of your original model.
- Define the Attribute Framework for the refined Domain of Discourse as a spatial table, exactly as worked through above.
- Execute a short Final Registration pass using a structured tool (a survey app, or a pre-printed paper form built from your Attribute Framework), and produce a small, genuinely populated spatial table — your own Realisation of a Cognised Existence, carried through Confronting the Sensoric Manifold from formal specification to recorded data.