Skip to contents

datamatch 0.3.0

New features

  • Two new sources: accessCEFI() and accessOBDAAC(). Seven sources now sit behind one interface, and each of the two is here for something the other five do not do.

    NOAA CEFI (accessCEFI()) reads the Changing Ecosystems and Fisheries Initiative regional model - MOM6-COBALT run for the Northwest Atlantic at a twelfth of a degree, 1993-2023, over OPeNDAP from the NOAA Physical Sciences Laboratory, with no account. It is the only source here carrying coupled physics and biogeochemistry on one grid at one resolution: NO3, PO4, O2, PH, PCO2, PHYC, MESOZOO and BOTO2 beside SST, SSS, BOTT, BOTS, SSH, MLD, SIC, UO and VO. The Copernicus biogeochemical reanalysis is a quarter degree and a separate product, so pairing them otherwise means regridding across a factor of three.

    It also publishes sea-floor salinity outright, so BOTS from CEFI needs none of the deepest-wet-level derivation GLORYS forces, and returns no BOTS_depth. cefi_archive() reaches the domains that do not ship - Northeast Pacific, Arctic, Pacific Islands, Great Lakes - the way fvcom_archive() does for FVCOM.

    NASA OB.DAAC (accessOBDAAC()) reads Level-3 mapped ocean colour from NASA’s Ocean Biology DAAC. What it adds over accessERDDAP() is length: SeaWiFS from September 1997, then both MODIS instruments and all three VIIRS, where the ERDDAP chlorophyll entries begin in 2012. Beyond CHL it carries KD490, PAR, POC, PIC, SST, SST_NIGHT and NFLH.

    It is the one source in this package needing an account. obdaac_credentials() documents both routes - an OB.DAAC appkey, or a ~/.netrc entry for urs.earthdata.nasa.gov - and where each is looked for.

  • A failed Earthdata login is reported as one. NASA answers an unauthenticated request with HTTP 200 and the login page, not a 401, so a download that trusts the status code writes nine kilobytes of HTML into a file named .nc and the mistake surfaces much later as a corrupt netCDF. accessOBDAAC() checks the first bytes for the netCDF signature instead, names the credential as the cause, and leaves nothing on disk.

  • CEFI forecasts are supported, and say that they are experimental. experiment = "decadal_forecast" warns on every call, and requires init and member rather than defaulting them. Neither has a defensible default: a decadal file is ten years from one January and there are sixty of them, and each holds ten ensemble members whose average is a modelling decision this package will not make silently. source_of() records both, as cefi:NWA12-decadal_forecast-r20250925-i198001-m01.

  • A CEFI source tag names the release, not the alias. release = "latest" is the default because a pinned release goes stale, but it is a poor thing to record - two runs a year apart would carry the same tag and different numbers. The tag is taken from the filename actually read, so a fetch made with "latest" records cefi:NWA12-hindcast-r20250715.

  • Both new sources are described in the EML write_eml() writes. known_variables() gathers labels and units from every source catalog, and a catalog missing from it is silent: the source still fetches, still joins, and then produces a document in which its columns have no definition and no units. CEFI’s and OB.DAAC’s catalogs are now listed, and a test checks every catalog reaches it rather than only checking the units.

    Seven units came with them. Only mol/m3 had a standard EML spelling - molePerCubicMeter, singular on both words where milligramsPerCubicMeter is plural on both, which is the sort of thing only validation tells you. The other six (umol/kg, uatm, mol/m2, 1/m, einstein/m2/day, W/m2/um/sr) are emitted as declared custom units, and a test writes a table carrying all of them and validates it against the schema.

  • Both new sources are cited in the EML methods section, and the data is cited as well as the paper. source_reference() switches on the source family, and a family with no branch returns NA - at which point the methods section names the source and cites nothing, which is the failure that section exists to prevent, arriving silently. A test now walks the families instead of trusting the list to stay complete.

    For CEFI that is the model paper plus the acknowledgement NOAA PSL asks for, with the release named, since releases revise the record. For OB.DAAC it is the mission paper plus a pointer to the dataset DOI, which is a separate obligation: OB.DAAC mints one per mission, suite and reprocessing, and the version token is not predictable from the suite - Aqua MODIS PAR/2022 resolves where SST/2019 does not. A hard-coded table of them would be wrong at the next reprocessing and wrong silently, so the citation names the dataset and says where its DOI lives, as the Copernicus one already did.

Bug fixes

  • ccmp_dictionary() and erddap_dictionary() no longer print as FVCOM. print.datamatch_dictionary() told the dictionaries apart by the absence of the Copernicus product column, which is true of four of them, so every non-Copernicus dictionary printed FVCOM’s header and FVCOM’s footer - telling a CCMP user about sigma layers and pointing them at ?accessFVCOM. Each dictionary now carries a tag saying which source it describes, and prints its own header, its own explanation of the last column, and its own Full descriptions: line.

Refusals worth knowing about

  • CEFI’s seasonal forecast and reforecast cannot be read, and are refused with the reason rather than attempted. Those files carry a 64-bit integer coordinate that DAP2 has no type for, so the PSL server returns HTTP 500 to every OPeNDAP request for one; the DAP4 endpoint describes them and then reads them wrongly, transposing start against count and crashing R with a bus error. A reader that segfaults the session is worse than no reader. The files download whole from the THREDDS fileServer endpoint and open correctly once local, at roughly 260 MB per variable per initialisation.

  • CEFI’s long-term projection and multi-decadal outlook have directories and no files. The portal announces an experiment before its output is posted. cefi_archive() will read them when they appear.

  • OB.DAAC eight-day composites are refused. They are the obvious answer to cloud gaps, but matchData() joins on an hour, a day, a month or a year, and an eight-day bin is none of those: stamped as a day it would demand an observation fall on the bin’s first date, and nearly every row would go unmatched. Use frequency = "monthly", or fill_satellite_gaps() on the daily field.

  • PACE is not among the OB.DAAC sensors. Its standard Level-3 mapped catalog publishes no chlorophyll suite, and its filenames carry a processing version that changes with each reprocessing and cannot be constructed - only looked up, through a file search that accepts POST only. Hard-coding a version would work until the next reprocessing and then fail silently.

  • CEFI has no PP. It writes intpp, production integrated over the water column in mol/m2/s, which is not the satellite PP’s volumetric rate in mg/m2/day. Giving them one name would make a nonsense of any comparison.

datamatch 0.2.0

Breaking changes

  • covariate_columns() now reports <var>_source columns. They are provenance, but they travel with the variable they describe rather than being left behind by it, and both resamplers were already written on that assumption: upscale_grid() and upscale_time() carry a non-numeric column as a categorical - the commonest value when aggregating, the nearest or a step when interpolating - and resolve_methods() says in as many words that “a source column riding along should not force the caller to enumerate every column”.

    Because covariate_columns() is the default vars for both, a source column never actually rode along. Regridding or retiming a gap-filled object dropped the record of which source each value came from, at the point it matters most.

    <var>_depth is still excluded, since the mean of two depths is not the depth any value came from, as is the internal per-row source tag.

    Callers relying on covariate_columns() to return only measurements should filter on is.numeric(), as plot_series() does.

New features

  • A date that has not happened yet is refused before anything is fetched. All five access functions now check the requested days against today, and stop if any are past it. Previously a typo in a year, or a projection window that ran off the end of the record, cost a download attempt per day and then an error from the server about exceeding the dataset coordinates - which named neither the date nor the mistake.

    mode = "forecast" keeps its horizon: the analysis-and-forecast products run about ten days past today, and dates inside that are still fetched. The check is on the calendar, not on what a given product has published yet - a reanalysis running months behind the present is a separate problem, and one the lag varies too much between products to guess at.

  • accessHYCOM(archive = "continuous") reads across the archives. HYCOM reaches 1994 to September 2024, but only as a reanalysis followed by a chain of shorter operational experiments, and until now that meant one call per archive. "continuous" picks an archive per day — the reanalysis wherever it reaches, the earliest-starting operational archive after that — so a single call spans the record.

    It is opt-in, and the default is unchanged, because the seam is real: crossing out of the reanalysis leaves one internally consistent hindcast for the model as it was running at the time, so a step in a series across that date can be the change of run rather than a change in the ocean. The call warns once, naming the day it happens and the archives it used.

    Provenance is recorded per row for such a fetch, rather than once for the object, so a value from the operational model never claims to be the reanalysis. source_of() reports every run present, and matchData() carries the per-row tag into <var>_source. A continuous read that never leaves one archive is stamped once, as before.

Breaking changes

  • accessCopernicus() takes its arguments in the same order as the other four access functions. It led with product_id and dataset_id, where accessFVCOM(), accessHYCOM(), accessCCMP() and accessERDDAP() all lead with vars. All five now open vars, years, months, bounding_box, dates, put frequency sixth where the source has more than one step, and end with overwrite; the source-specific arguments sit between. Swapping one access function for another is now a rename rather than a rewrite.

    Calls that name their arguments — every example in the documentation, and every call in datamatch, taupatch, derivoce and msomgom — are unaffected. A call written positionally against the old order would now put a product identifier into vars, which would fetch nothing rather than fail, so that case is caught at the call with a message saying what changed.

    product_id and dataset_id can usually be dropped entirely; they are inferred from the variable names.

    A new test file pins the shared order, so the five cannot drift apart again.

Documentation

  • The README is now a front door rather than a manual. It had grown to 1,618 lines, much of it duplicating the four vignettes. The per-source deep dives, the full reference list and the error-message guide now live in vignette("sources"), and the README keeps orientation, installation, a quick start, and a short section per source — 535 lines.

    The documentation guards move with the content: every export is still named in the README, and the citation guards now read the vignettes as well, so a source is still required to be cited somewhere a reader will meet it.

New features

  • The FVCOM mesh itself, through fvcom_mesh(). accessFVCOM() returns values at points — nodes for scalars, element centroids for velocities — which is what matching needs but is not the grid. Drawing it shows dots where the model has triangles.

    fvcom_mesh() reads the connectivity array a point fetch never touches and returns the triangles as POLYGONs, with DEPTH per cell. It carries no time dimension: it is the grid. what = "nodes" and "elements" return the two point sets on their own.

    plot_mesh() draws it — bare, or shaded by a covariate. Pass a fetch as values and the spatial join is done for you, which matters more than it sounds: a fetch is subset to the bounding box, so its row order says nothing about the mesh’s element numbering, and joining by position would silently mislabel every cell.

    Node values sit at triangle corners and element values at centroids, so the join tests intersection rather than containment — a corner lies on the boundary between cells, and nothing contains it. Joining by containment returned no node values at all.

    A triangle is kept when any vertex is inside the box, so the mesh covers what was asked for rather than stopping short of it, and the edge is ragged by design. GOM7 remeshes between months, so date says which month’s mesh.

  • EML metadata, through write_eml(). Writes an Ecological Metadata Language document for a matched table — the standard EDI, LTER and DataONE expect alongside a deposited dataset.

    Most of it comes from the data: the bounding box and date range from the object, and an attribute for every column with its definition, units and measurement scale taken from whichever source catalog defines that name. This package’s own conventions are described too, so the time columns, LON/LAT and the _source and _depth columns are not left as bare names.

    The methods section is the part worth having. Because matchData() records <var>_source on every join, a table with several sources chained onto it produces a methods statement naming each and a citation for each — assembled from the source catalogs rather than by hand, which is the most tedious part of depositing a derived dataset and the easiest to get wrong.

    One thing that would otherwise fail at submission rather than at write time: EML validates units against a fixed vocabulary, and PSU and N/m2 are not in it — salinity and wind stress, two core variables here. Both are written as custom units and declared in the document’s own unitList. Which names are valid was settled by validating documents against the schema, not by reading a list: milligramsPerCubicMeter is accepted and milligramPerCubicMeter is not. write_eml() validates before returning, so an invalid document is an error rather than a rejection later.

    Needs emld, a new Suggests.

  • MUR and VIIRS, through accessERDDAP(). Satellite SST and chlorophyll from NOAA’s ERDDAP servers: MUR (0.01°, daily, 2002-06 onward) for SST, and VIIRSCHL / VIIRSCHL2018 for chlorophyll. MUR is by a wide margin the finest field the package can reach — about 1 km, against 9 km for the physics reanalysis.

    No account is needed. The same products at PO.DAAC require an Earthdata login; ERDDAP serves them openly and subsets server-side, so there are no credentials for this package to handle or for the user to configure. That is why this route was chosen.

    Three things the documentation is explicit about. A satellite SST is the foundation temperature, below the daily warming layer, where a model SST is its topmost level — they differ by a degree or more on a calm sunny day, and both arrive in a column called SST. MUR is gap-free by construction because it is an analysis, so SST_ERROR is where the interpolation shows up. And there is no global VIIRS SST: the VIIRS SST on these servers covers the US West Coast only, so MUR — a blend taking VIIRS among its inputs — is the SST to use instead.

    As with FVCOM, the backend is general rather than a fixed list: erddap_dataset(server, dataset_id) describes any of the thousands of other griddap datasets for accessERDDAP().

  • HYCOM now reaches 2024, and FVCOM 2025. Both had stopped short — HYCOM at 2015 and FVCOM at 2013 — because each shipped a single archive. Neither was a limit of the data.

    hycom_archives() now carries the 1994–2015 reanalysis plus the seven operational experiments that follow it, to September 2024, and hycom_covering(date) says which span a given day. They overlap, so where two cover a date the choice between them is a judgement — the reanalysis is more internally consistent, the operational run more recent — and nothing picks for you. One archive is read per call, and a request falling outside the one named is told which others hold it rather than being stitched to them silently.

    The seam that matters is the run, not the grid: GLBv0.08 and GLBy0.08 hold the same 126 latitudes between 40 and 45 °N identically, so on a mid-latitude shelf the cells line up across it. GLBy0.08 stores longitude 0–360 where GLBv0.08 stores −180…180; the convention is read off the coordinates rather than recorded, so a box given negative west works against either.

    fvcom_archives() gains GOM7, the archived operational forecast: hourly on a 207,081-node mesh from 2025. It is deliberately not presented as a continuation of the GOM3 hindcast — the mesh, the step and whether the output is retrospective all change at once, and GOM7 saves no wind stress. Between 2014 and 2024 NECOFS publishes neither, which is a gap in the source.

    accessFVCOM() gains frequency and hour for sub-daily archives, defaulting to one snapshot a day because a month of hourly GOM7 is 720 reads of a 207,081-node field. A snapshot is an instant, not a daily mean; upscale_time() makes a real one.

  • CCMP winds, through accessCCMP(). The Cross-Calibrated Multi-Platform ocean surface wind analysis from Remote Sensing Systems: WSPD, UWND, VWND and NOBS, six-hourly from January 1993 to within days of the present. No account is needed — the registration RSS asks for covers their FTP service, and the HTTPS archive is open.

    It is the longest and finest-in-time wind record the package can reach. The Copernicus L4 wind is monthly from mid-1994 or hourly only from 2007, with nothing daily between; CCMP is six-hourly throughout.

    Two costs. It carries no wind stress, which the Copernicus product does, and stress cannot be recovered from these winds without choosing a drag coefficient — so where stress is the covariate, use Copernicus. And there is no server-side subsetting: RSS publishes static files, so a day is one 33 MB global file however small the bounding box, and a year is roughly 12 GB of transfer. Subsets are cached, and a request for more than 30 days says what it is about to download first.

    CCMP is stored on a 0–360 longitude grid, alone among the sources here. bounding_box is given negative west as everywhere else, converted on the way in, and the returned coordinates are negative west too — so a CCMP result overlays the others without adjustment.

  • HYCOM, through accessHYCOM(). Reads the HYCOM + NCODA GOFS 3.1 reanalysis (GLBv0.08 expt_53.X, 1994–2015) from the Naval Research Laboratory’s THREDDS server, in the same sf shape as everything else.

    It is worth having for two things GLORYS cannot give. HYCOM publishes salinity_bottom and water_temp_bottom as fields, so BOTS is an ordinary request rather than a derivation from the full depth column. And it is an independent model, so agreement with Copernicus is evidence in a way either alone is not.

    HYCOM publishes instantaneous fields every three hours and no mean at all, so none is invented: frequency = "daily" takes one snapshot per day at hour, and frequency = "3hourly" returns every step with an HOUR column. A real mean is upscale_time()’s job. The archive also has gaps — some steps are simply absent — so a daily request at a missing hour skips that day and says so.

  • Any FVCOM endpoint, through fvcom_archive(). FVCOM is a model rather than a data product, so there is no global archive of it: groups run it for their own coastlines and publish on their own servers, and fvcom_archives() ships only what one server publishes for the Northeast US.

    fvcom_archive(url) describes any other FVCOM endpoint — opening it once to read the mesh size, period, and which fields that run actually saved — and the result is passed to accessFVCOM(archive = ). This works because the reader depends on FVCOM’s structure rather than on the region. A URL that is wrong, blocked, or not FVCOM fails there with the reason instead of part-way through a fetch.

  • FVCOM, through accessFVCOM(). Reads NECOFS — the Northeast Coastal Ocean Forecast System — from the UMass Dartmouth THREDDS server over OPeNDAP, and returns the same sf shape accessEnvDat() does, so matchData() and everything downstream work unchanged. fvcom_variables() and fvcom_archives() say what can be read; currently the 30-year GOM3 hindcast, 1978–2013, monthly means.

    Variables reuse the Copernicus names for the same quantities, so a covariate lands in a column of the same name from either source. They are not interchangeable for that reason — one is a regional model on a triangular mesh, the other a global reanalysis. Which was used is worth recording.

    Two things behave unlike the Copernicus path, both because FVCOM is an unstructured mesh:

    • Scalars are on mesh nodes and velocities on element centroids, which are two different sets of points (48,451 and 90,415 on GOM3). Fetching both in one call is an error rather than a silent interpolation of one onto the other, in the same spirit as refusing to mix two Copernicus grids.
    • BOTS costs nothing. FVCOM’s sigma coordinate makes the deepest layer the sea floor at every node, so bottom salinity is a layer index rather than the derivation it needs against GLORYS. A sigma layer is not a depth, though: the deepest sits at 98.9% of the local column.

    Only the monthly means are offered. The hourly hindcast carries 342,348 time steps, and nc_open() reads the whole time coordinate before returning, so the request exceeds the server’s DAP timeout — the failure takes over ten minutes to arrive and says only NetCDF: DAP failure. Sub-monthly FVCOM would mean reading the per-file datasets behind the aggregation, a month at a time. Reading FVCOM needs ncdf4, a Suggests, and a route to a plain-HTTP service on port 8080.

  • Surface wind and stress. Six new variables from the Copernicus L4 wind product — WSPD (speed), UWND/VWND (components), TAUX/TAUY (stress components) and TAU (stress magnitude). Wind is its own product on its own 0.25° grid, running from June 1994, so it is a separate accessEnvDat() call from SST as usual.

    Stress rather than speed is what sets mixing and Ekman pumping, and the two are not interchangeable: stress is roughly quadratic in speed. Note also that the monthly WSPD is averaged as a speed, so it exceeds the magnitude of the mean vector wherever direction varied within the month.

  • frequency = "hourly". Copernicus publishes its L4 wind hourly or monthly and nothing between, so there is no daily wind to fetch. Hourly is the way to a sub-monthly wind field, and upscale_time(to = "day") aggregates it — which keeps the choice of summary, mean or maximum, with the caller.

    Hourly results carry an HOUR column, and matchData() joins on it, so observations matched against hourly data need an hour of their own, on UTC. A day is one download at either step, so hourly costs no extra requests — it costs 24 times the rows.

    frequency = "daily" is refused for the wind variables, and "hourly" for everything else, both before anything is downloaded. WSPD and TAU are magnitudes the hourly product does not carry, so they are monthly only.

  • upscale_time(to = "day"). The time axis now goes hour → day → month → year rather than starting at day.

  • Bottom salinity, as BOTS. GLORYS12V1 publishes potential temperature at the sea floor but no salinity counterpart, so in reanalysis mode this is derived: the full salinity column is fetched and the deepest wet level in each cell kept. The depth that value came from is returned as BOTS_depth rather than left to be assumed — the deepest wet model level is not the sea floor, and in deep water can sit well above it.

    Because it needs the whole depth column, BOTS must be fetched on its own; mixing it with a surface variable is an error rather than a quiet second download. It is also much more expensive than a surface variable, downloading roughly fifty levels to keep one, and accessEnvDat() says so when it starts.

    In forecast mode none of this applies. The analysis-and-forecast product publishes sob outright, so BOTS is an ordinary variable there, fetches alongside BOTT, and returns no BOTS_depth. The two are the same quantity but not the same number, so a record spanning both modes has a seam in it.

  • variable_dictionary("wind") filters the catalog to the wind variables, and variable_dataset(frequency = "hourly") reports which have an hourly dataset.

Documentation

  • The README is orientation now, and the depth is in vignettes. It had grown to 11,600 words — a forty-five minute read — because every capability added a section to it. It is now 9,000, and what left went into two new articles rather than being deleted:

    • Working with what comes back — resampling between grids and time steps, filling satellite gaps, the four plots, and what n_workers actually buys.
    • Static and basin-scale covariates — seafloor terrain and the climate indices, including what TPI measures and why the indices are not interchangeable.

    The deeper parts of “Satellite or model?”, “Bottom salinity” and “Downloads run in parallel” moved into the existing articles, each leaving a short summary and a link behind. Four articles now: getting started, choosing a source, working with the result, and the other covariates.

  • A function reference table. Every export, grouped by what it is for, in one place in the README. That is what the depth moving out required: the README says what each function is and where to read more, rather than explaining every one at length.

  • The contents list was rebuilt to match, and every internal link checked — two pointed at sections that had moved to a vignette.

Breaking changes

  • accessEnvDat() is now accessCopernicus(). The old name dates from when Copernicus was the only source. With five, “access environmental data” reads as though it fetches from all of them, sitting beside accessFVCOM(), accessHYCOM(), accessCCMP() and accessERDDAP(), which each say what they read.

    accessEnvDat() still works and warns, on the same reasoning as the speciesDat/envDat arguments in matchData(): taupatch and any script written against the old name keep working, and their authors find out that the name moved rather than discovering it when the alias is eventually removed. Nothing else changes — the arguments, the result, and the copernicus: source tag are all as they were.

Bug fixes

  • upscale_time() and downscale_time() did not record the step they produced. The result’s resolution was left to be inferred from its own time stamps, which fails when there are too few periods to tell: one day of hourly data aggregated to daily is a single day in a single month, indistinguishable by inspection from monthly data. detect_temporal_resolution() fell back to monthly, and matchData() would then join by month and quietly ignore the day. Both now record the target step, which they know exactly.

  • A yearday column could be matched on as though it were the year. matchData() finds a table’s time columns by prefix, and yearday begins with year. With no plain year column beside it to win the tiebreak, it was renamed to YEAR without complaint, and every row went into a period no source covers — so the join came back all NA behind a warning about uncovered periods, which reads like a gap in the data rather than a mistake.

    Day-of-year names — yearday, dayofyear, jday, doy and the usual variants — are now never matched on, for any time column. A table carrying one and no real year or day column gets the missing-column error instead.

  • The missing-column error now names the columns it passed over. The lookup stays a prefix match, so obs_month and survey_year are still not recognised. Rather than only saying what it looked for, the error lists the similar names that are present and asks for the intended one to be renamed. Suggesting is not selecting: widening the rule to a contains match would have swept up jday and yearday alongside obs_month, trading a clear error for a wrong answer.

    matchData()’s own documentation had claimed obs_month would be recognised, which was never true. It now describes the prefix rule the code implements.

datamatch 0.1.0

Bug fixes

  • Variables could be silently mislabelled in multi-variable downloads. This can affect results already in hand, so it is worth checking rather than just reading.

    accessEnvDat() named result columns positionally from vars. Copernicus returns layers in the NetCDF’s own order, which is alphabetical by variable code, so the two agreed only by coincidence. Requesting c("SST", "SSS") sends c("thetao", "so") and gets back so first, so the columns were labelled the wrong way round:

    # before the fix, Gulf of Maine, January 2010
    SST = 32.69     # 32.69 PSU of salinity, in the temperature column
    SSS = 6.18      # 6.18 C of temperature, in the salinity column

    There was no warning. Any download of two or more variables whose requested order differed from the alphabetical order of their Copernicus codes was affected. Single-variable downloads never were, and neither were requests that happened to be in alphabetical code order.

    Checking existing results: look for a column whose values sit in the wrong range for its name, such as salinity near 32 in an SST column. Re-running the fetch fixes it. The cached NetCDF files were always correct; only the naming was wrong.

    The fix: order_layers() strips the depth suffix that 3D variables carry (thetao_depth=0.494025) and selects layers by matching the variable code, so naming cannot drift from content.

  • Expected N variable column(s) blamed the depth range for every cause. The check that caught the problem above could only count columns, so it guessed:

    Error: Expected 4 variable column(s) but the download returned 2. This usually
    means the depth range spans several model levels...

    For a request whose real problem was two variables the dataset never returned, that pointed in the wrong direction. The two causes are now distinguished: a variable the download omitted is named along with what did arrive, and a code appearing on several layers is reported as the depth-range problem it is, with the levels it returned.

  • Daily datasets whose identifier ends in P1D fetched one day per month. accessEnvDat() decided whether a dataset was daily by reading a single character three from the end of dataset_id. That is D in ..._P1D-m, which the model products use, but P in ..._P1D, which the satellite ocean colour products use — so an explicitly passed ocean colour daily dataset was treated as monthly and only day 1 of each month was downloaded. The frequency token is now matched as such.

  • matchData() returned rows grouped by period, not in the order they went in. Rows are processed a period at a time, so a table whose periods were interleaved came back reordered. Anyone aligning the result against the input by position - cbind(), or assigning a column straight across - would have got silently mismatched rows. The order is now restored before returning.

  • Sparse daily data could be matched as though it were monthly. detect_temporal_resolution() infers daily data from more than one day within a month, so a set of survey dates — one per month — was indistinguishable from monthly data by inspection, and matchData() would join by month and ignore the day.

    accessEnvDat() knows which dataset it fetched, so it now records the step on the result and detect_temporal_resolution() trusts that over the heuristics. Passing temporal_resolution explicitly still overrides both.

  • Failed downloads surfaced as an unreadable file. A run of failures warned and returned a status code the caller ignored, so the error appeared later from terra::rast() on a file that was never written. Failures now raise where they happen, carrying the Copernicus client’s output.

Breaking changes

  • matchData()’s arguments are now dat and source, replacing speciesDat and envDat.

    The function was never specific to species observations or environmental data. It is a spatiotemporal nearest-feature join between two sf point objects, and works as well for tag positions against a model field, moorings against satellite retrievals, or one gridded product against another. The old names described one use of it as though it were the only one.

    The old names still work, with a warning, so existing scripts and taupatch keep running. They will be removed in a later version.

    matchData(observations, env)                 # positional, unchanged
    matchData(dat = observations, source = env)  # new names
    matchData(speciesDat = obs, envDat = env)    # still works, warns

    One related change: a source column whose name collides with one already in dat is now suffixed .matched rather than .env, for the same reason.

  • The BigelowLab/copernicus dependency is gone. It was used in two places, both in accessEnvDat(), and both are now internal. This removes the Remotes: field and the hand-created ~/.copernicusdata file that a new machine previously needed.

    The download cache has moved to tools::R_user_dir("datamatch", "cache"). Files downloaded under the old scheme are not found there and would be re-downloaded. To keep using them:

    options(datamatch.cache = readLines("~/.copernicusdata"))   # in .Rprofile

    getOption("datamatch.cache") and the DATAMATCH_CACHE environment variable both override the default. The Copernicus client is located via getOption("datamatch.copernicusmarine") or the PATH, and a missing client now says so with install instructions.

New features

  • Daily data. accessEnvDat(frequency = "daily") fetches the daily datasets rather than the monthly means, expanding each requested month into its days.

    Daily is not simply the monthly identifier at a finer step, and the catalog records what Copernicus actually publishes. PH, PP, DIATO and DINO are monthly composites only, and requesting them daily is refused before anything is downloaded. Daily CHL comes from the gap-free interpolated ocean colour dataset, whose cloud gaps Copernicus has already filled — so fill_satellite_gaps() has nothing to do on it, and accessEnvDat() says so when it makes the substitution.

    variable_dataset() takes frequency too, and returns NA where no daily dataset exists.

    dates names the exact dates to fetch, which is the argument to use when matching daily data to observations:

    accessEnvDat(vars = "SST", dates = unique(observations$date), bounding_box = bb)

    Survey dates differ from month to month, and a day-of-month rule does not describe them. YYYYMMDD strings, YYYY-MM-DD strings, and Date objects are all accepted. Dates are sorted and deduplicated, and a date the calendar does not have is an error naming it rather than a silently dropped request.

    dates replaces years and months, which must not be given alongside it, and implies frequency = "daily". Passing it with an explicit frequency = "monthly" is a contradiction and an error.

    Since it fetches only the dates named, it is also how a long record is sampled rather than fetched whole.

    Monthly remains the default. A call naming neither frequency nor days fetches monthly means exactly as before.

  • Downloads run in parallel. Only the days not already cached are fetched, and they go out n_workers at a time (default 4). A month of daily SST and SSS takes about 40 seconds rather than about 160.

    A failed day no longer abandons the rest: every day is attempted, successful ones stay in the cache, and the error names each day that failed, so re-running the same call retries only those. Reading and converting the files stays in the calling session — those are local reads, and returning each day’s data frame from a worker would cost more than the read saves.

  • Resampling. upscale_grid(), downscale_grid(), upscale_time(), and downscale_time() move data between grids and time steps, each with a choice of method. Targets may be a resolution or another object’s grid. Aggregation guards partial coverage with min_coverage; interpolation defaults to the methods that do not invent structure.

  • Plotting. plot_env(), plot_coverage(), plot_series(), and plot_matched(), in base graphics, each returning the data it drew.

  • TPI joins the bathymetry covariates. A cell’s depth relative to its neighbours, which separates banks from basins where plain depth cannot.

  • LCR, the Labrador Current retroflection index, joins the climate index catalog. Published with Jutras et al. (2023) and cited in index_dictionary().

  • Every climate index now declares its units, surfaced in index_dictionary(). Most are standardized anomalies, in standard deviations rather than anything physical; only AMO (degrees C) and AMOC (Sverdrups) carry real units. A coefficient fitted to one of those is not comparable with one fitted to NAO, and a table of model output gave no way to tell.

  • AMOC, the overturning transport measured by the RAPID array at 26.5°N, joins it too. Unlike the other indices this is a direct measurement rather than a pressure or SST pattern — which is also its limitation: it begins in April 2004 and is recorded far south of the shelf, so it describes the basin-scale circulation the Labrador and slope currents sit within rather than local conditions.

    RAPID publishes only NetCDF at a stable URL, so this one is downloaded as bytes and read with ncdf4 (a Suggests) rather than parsed as text. The twelve-hourly series is averaged to monthly. The file is cached like the Copernicus downloads, since it is over a megabyte and RAPID times out on repeated requests, and the time origin is read from the file rather than hard-coded — RAPID has re-based it across releases.

  • Cached indices now expire on their provider’s publishing cadence, so a living dataset does not quietly stop being current. Each index records how its source updates — monthly for the NOAA indices, roughly yearly for AMOC, never for LCR, which was published with a paper and ends at 2014 — and that sets how long a cached copy is reused: a week, a month, or forever. Living indices re-download on their own; finished ones are not re-fetched pointlessly.

    Staleness is judged from the data rather than the cache, because a fresh download of a file the provider stopped updating is still stale. A series ending further behind the present than its source’s usual lag is reported, with the command to force a refresh.

    A failed download with a usable cached copy on disk returns that copy and warns, rather than erroring: a provider being briefly unreachable should not become an outage here, and the warning is what stops the old copy being mistaken for a current one.

    climate_index_status() reports what is cached and what is due without downloading anything, so it is safe offline. refresh_climate_index() forces a re-fetch and skips the indices that cannot change.

    The text indices were previously re-downloaded on every call and are now cached too, since downloading moved out of the NetCDF reader into a step every format shares.

  • A monthly workflow checks the variable catalog against the live Copernicus catalogue and opens an issue when dataset identifiers drift.