Joins each row of dat to the nearest feature of source within the same
time period, and returns dat with source's columns added.
Usage
matchData(
dat,
source,
temporal_resolution = c("auto", "hour", "day", "month", "year"),
record_source = TRUE,
speciesDat = NULL,
envDat = NULL
)Arguments
- dat
the points to add columns to: observations, stations, tag positions, anything with coordinates and time. Needs year and month columns, plus a day column when matching at daily resolution and an hour column at hourly. Columns whose names begin with those words are recognised, whatever their case, so Yearandmonth_utcwork butobs_monthdoes not — rename it, or pass it asMONTH. Day-of-year names (yearday,dayofyear,jday,doyand the like) are never used, even though some of them do begin withyearorday: they hold an ordinal day rather than a year or a day of the month. An unrecognised name is named in the error, so nothing has to be guessed at.- source
the points to take values from: a grid from any access function — accessCopernicus(),accessFVCOM(),accessHYCOM(),accessCCMP()oraccessERDDAP(). Must carryYEAR/MONTH/DAY, andHOURas well when matching hourly.- temporal_resolution
one of "auto"(default),"hour","day","month", or"year"."auto"uses the step the access function recorded onsource, or infers it fromsource's time steps.- record_source
add a <var>_sourcecolumn for each column joined, naming which source and archive produced it. On by default, and only has an effect whensourcecarries the stamp an access function leaves — seesource_of(). SetFALSEfor the narrower table.- speciesDat, envDat
deprecated names for
datandsource. Still accepted, with a warning.
Details
Neither side has to be species observations or environmental data. It is a
spatiotemporal nearest-feature join between two sf point objects that carry
YEAR/MONTH/DAY columns, so it works equally for stations against a
covariate grid, tag positions against a model field, moorings against
satellite retrievals, or one gridded product against another.
Which geometries can be matched
dat may hold points, lines or polygons — survey stations, tow tracks,
transects, statistical areas. The join is nearest-feature against the whole
geometry, so a tow matches the nearest grid cell to the track rather than to
any one end of it, and an area matches the nearest cell to the area.
LON/LAT are then a representative point rather than the geometry: for a
line or polygon they come from sf::st_point_on_surface(), which is
guaranteed to lie on the feature where a centroid need not. The geometry
column itself is untouched, so nothing is lost — but do not read LON/LAT
as the position of an area.
A caution about extended geometries: a long tow or a large area may lie nearer one cell while spanning several, and nearest-feature returns exactly one. Where a track crosses a front, consider splitting it, or matching its vertices as points and summarising afterwards.
source should be points, as every access function returns.
Matching in time
The time period is source's own resolution: hourly data matches on
year/month/day/hour, daily data on year/month/day, monthly data (Copernicus
...P1M-m means, say) on year/month, and annual data on year alone.
Matching hourly needs dat to say which hour each row belongs to, in a column
recognised the same way as the others — HOUR, hour, hour_utc, but not
obs_hour, which does not begin with the word. On UTC, as the environmental
data is: an observation timestamped in local time will match the wrong hour,
and nothing in the join can detect that.
That matters because a day-exact join against monthly data matches nothing. A
monthly product carries one time step per month, while observations fall on
arbitrary days. temporal_resolution overrides the inference when the data
cannot speak for itself.
What is preserved
One row out per row of dat, in the same order, whatever happens. A period
source does not cover gives NA for its columns and a warning naming the
periods, rather than dropping those rows — a silent change in row count is a
worse outcome than a visible gap.
dat keeps its own columns. One of source's that collides with a name
already in dat is suffixed .matched, so nothing of dat's is overwritten
or renamed.
Which source a column came from
The access functions share variable names on purpose, so SST from
Copernicus, FVCOM, HYCOM and MUR all arrives in a column called SST and
everything downstream works unchanged. The cost is that a table with several sources chained onto it
has no record of which produced what.
So each joined column gets a companion <var>_source naming the source and
archive — "hycom:GLBv53X", "fvcom:GOM3" — in the same spirit as the
<var>_source column fill_satellite_gaps() writes. Pass
record_source = FALSE to omit them.
These are provenance rather than measurements, but they travel with the
variable they describe: covariate_columns() reports them, and resampling
carries them as a categorical - the commonest value when aggregating, the
nearest when interpolating - so a regridded or retimed object still says
which source each value came from. plot_series() skips them, since a source
tag has no mean.
See also
accessCopernicus(), accessFVCOM(), accessHYCOM(), accessCCMP()
and accessERDDAP() for the usual source; attach_bathymetry() and
attach_climate_index() for covariates that are not matched this way
Examples
if (FALSE) { # \dontrun{
env <- accessCopernicus(vars = "SST", years = 2010, months = 1:12, bounding_box = bb)
matched <- matchData(observations, env)
# Chains, so several sources land on one table
matched <- matchData(matched, chlorophyll)
} # }