Skip to contents

datamatch fetches from seven sources and joins any of them to your observations the same way. This is about which one to reach for, and what changes when you do.

Nothing here is downloaded when the vignette is built — the fetches are shown but not run, because each needs a network and some need several gigabytes.

The seven, side by side

The four models first:

accessCopernicus() accessFVCOM() accessHYCOM() accessCEFI()
Source Copernicus Marine NECOFS / any FVCOM HYCOM GOFS 3.1 NOAA CEFI, MOM6-COBALT
Kind global reanalysis, forecast, satellite regional coastal model global model regional coupled model
Grid 0.083°–4 km, regular unstructured mesh 0.08° regular 0.083° regular (regridded)
Extent global Gulf of Maine global NW Atlantic only
Steps monthly, daily, hourly monthly (GOM3), hourly (GOM7) 3-hourly monthly, daily
Record 1993– 1978–2013, then 2025– 1994–2024 across archives 1993–2023
Biogeochemistry 0.25°, separate product — — same grid as physics
Bottom salinity derived free free free
Wind stress yes GOM3 only — —
Subsetting server-side client-side server-side server-side
Account needed yes no no no

Then the three observational products:

accessCCMP() accessERDDAP() accessOBDAAC()
Source RSS CCMP v3.1 NOAA ERDDAP NASA OB.DAAC
Kind wind analysis satellite analysis satellite Level-3 composites
Grid 0.25° regular 0.01°–0.04° regular 4 km or 9 km regular
Steps 6-hourly daily daily, monthly
Record 1993–present 2002– (SST), 2012– (chl) 1997–
Gaps under cloud n/a none (analysed) yes, single-sensor
Subsetting whole globe server-side whole globe
Account needed no no yes, Earthdata

Two of those records are chains rather than one run, and the gaps matter more than the endpoints. FVCOM has nothing between 2014 and 2024 — the GOM3 hindcast stops in 2013 and the GOM7 forecast archive starts in 2025, on a different mesh. HYCOM reaches 2024 only by crossing from the reanalysis into a series of operational experiments, which is a seam in how the values were produced. hycom_covering(date) says which archives hold a given day, and accessHYCOM(archive = "continuous") reads across them in one call — warning once at the crossing, and recording on every row which run it came from, since a step in a series across that date can be the change of run rather than the ocean.

variable_dictionary()   # Copernicus
fvcom_dictionary()      # FVCOM
hycom_dictionary()      # HYCOM
cefi_dictionary()       # CEFI
ccmp_dictionary()       # CCMP
erddap_dictionary()     # MUR and VIIRS
obdaac_dictionary()     # SeaWiFS, MODIS, VIIRS

CEFI and OB.DAAC are the two most recently added, and each is here for one thing the other six do not do. CEFI is the only source carrying physics and biogeochemistry on one grid at one resolution, which is what a question about nutrients or oxygen alongside temperature actually needs. OB.DAAC is the only one reaching back to 1997, which is what a question about trend rather than state needs.

They share variable names on purpose

A covariate arrives in a column of the same name whichever source supplied it, so everything downstream — matchData(), resampling, plotting, and whatever model you fit — works unchanged:

bb <- list(xmin = -70, xmax = -66, ymin = 41, ymax = 44)

copernicus <- accessCopernicus(vars = "SST", years = 2010, months = 1:12, bounding_box = bb)
fvcom      <- accessFVCOM(vars = "SST", years = 2010, months = 1:12, bounding_box = bb)
hycom      <- accessHYCOM(vars = "SST", years = 2010, months = 1:12, bounding_box = bb)

All three produce an SST column. They are three different models and three different numbers. Sharing the name makes them mechanically interchangeable, which is the point; it does not make them scientifically interchangeable, which is why it is worth recording which you used.

That property is useful deliberately. Fetching the same variable from two sources and comparing them is a real check on a result:

matched <- matchData(observations, copernicus)          # SST
matched <- matchData(matched, hycom)                    # SST.matched

plot(matched$SST, matched$SST.matched)

A colliding name is suffixed .matched rather than overwriting yours, so both survive the join and the disagreement is visible.

Which to reach for

Start with Copernicus. It has the widest variable list — physics, biogeochemistry, satellite ocean colour, wind and stress — the longest record with a forecast, and server-side subsetting. Everything else here is for something Copernicus does not do well.

Reach for FVCOM when the coast is the point. A Gulf of Maine box holding 1,742 GLORYS cells holds 6,579 GOM3 nodes, concentrated where the bathymetry is complicated. If your question is about a shelf, a bank, or a channel rather than a basin, that resolution is the reason. The hindcast ends in 2013; the GOM7 forecast archive picks up in 2025 on a different mesh, with nothing in between.

Reach for HYCOM for bottom fields, or for a second opinion. It publishes salinity_bottom and water_temp_bottom as real fields where Copernicus has no bottom salinity at all, and being an independent model it makes agreement meaningful. The reanalysis covers 1994–2015; operational archives carry it to September 2024.

Reach for CCMP for winds over a long record. Six-hourly from 1993 to within days of the present, where the Copernicus wind is monthly from mid-1994 or hourly only from 2007. But it carries no stress.

Reach for ERDDAP for resolution. MUR is 0.01° — about a kilometre, against nine for the physics reanalysis — and needs no account, where the same product at PO.DAAC needs an Earthdata login. It is a satellite analysis of the foundation temperature rather than a model level, which is a different quantity from a model SST, not a better measurement of the same one.

Reach for CEFI for biogeochemistry on the shelf. It is the only source here where nitrate, oxygen, pH, pCO2, phytoplankton carbon and mesozooplankton biomass sit on the same twelfth-degree grid as the temperature and salinity — the Copernicus biogeochemical reanalysis is a quarter degree and a separate product, so pairing them means regridding across a factor of three. CEFI also publishes sea-floor salinity outright rather than by derivation. Its cost is extent: the Northwest Atlantic and nothing else, so a box outside roughly 98°W–36°W and 5°N–58°N is refused. Its daily output is biogeochemistry only.

Reach for OB.DAAC for the length of the ocean colour record. SeaWiFS begins in September 1997, where the ERDDAP chlorophyll entries begin in 2012 — the difference between a decade of chlorophyll and nearly three. It carries what ocean colour measures besides chlorophyll (KD490, PAR, POC, PIC) from the same retrieval, and both MODIS instruments alongside all three VIIRS so there is a sensor either side of any gap you need to bridge. Its costs are an Earthdata Login, whole-globe downloads, and cloud: these are single-sensor composites, so a daily field outside the tropics is mostly NA. Composite to monthly, or use fill_satellite_gaps().

Two warnings that belong with it. The sensors are not interchangeable — where they overlap they disagree, and a series stitched across a mission boundary has an instrumental step in it; where a long consistent record matters more than any one sensor, the Copernicus-GlobColour entries are multi-sensor and built for that. And PACE is not among the sensors, deliberately: its standard Level-3 mapped catalog publishes no chlorophyll suite, and its filenames carry a processing version that changes with each reprocessing and cannot be constructed. ?obdaac_sensors gives the detail.

Bottom salinity, five ways

BOTS is the clearest case of the same name costing different amounts:

# Copernicus reanalysis: derived. Fetches ~50 depth levels to keep the deepest
# wet one, so it must be fetched alone, and reports the depth it used.
bots <- accessCopernicus(vars = "BOTS", years = 2010, months = 1:12, bounding_box = bb)
bots$BOTS_depth

# Copernicus forecast: published outright as `sob`, no derivation.
accessCopernicus(vars = c("BOTT", "BOTS"), years = 2026, months = 7,
             bounding_box = bb, mode = "forecast")

# FVCOM: the deepest sigma layer is the sea floor everywhere. Free.
accessFVCOM(vars = c("BOTT", "BOTS"), years = 2010, months = 1:12, bounding_box = bb)

# HYCOM: `salinity_bottom` is its own field. Free.
accessHYCOM(vars = c("BOTT", "BOTS"), years = 2010, months = 1:12, bounding_box = bb)

# CEFI: `sob` is the model's own sea-floor diagnostic. Free, and NW Atlantic only.
accessCEFI(vars = c("BOTT", "BOTS"), years = 2010, months = 1:12, bounding_box = bb)

Only the first is expensive, and only the first returns BOTS_depth — because only the first had to choose a level. The deepest wet model level is not the sea floor, and in deep water sits well above it, so the depth is reported rather than left to assume.

Sub-daily data, and means that are not means

Four of the seven publish below daily somewhere, and none of them publishes a daily mean:

Source Native step frequency
Copernicus wind hourly "hourly"
FVCOM GOM7 hourly "hourly", or "daily" for a snapshot
HYCOM 3-hourly "3hourly", or "daily" for a snapshot
CCMP 6-hourly "6hourly", or "daily" for a snapshot

The other three stop at daily. CEFI is monthly or daily and nothing finer; ERDDAP is daily; OB.DAAC is daily or a monthly composite, and its eight-day composite is refused because matchData() has no step that means eight days.

Where a source has no mean, none is invented. frequency = "daily" on HYCOM and CCMP takes one snapshot at a chosen hour — an instant, not an average. A real mean is upscale_time()’s job, which keeps the aggregation visible and the choice of summary yours:

steps <- accessCCMP(vars = c("UWND", "VWND"), frequency = "6hourly",
                    dates = "2010-06-15", bounding_box = bb)

mean_wind <- upscale_time(steps, to = "day")
peak_wind <- upscale_time(steps, to = "day", method = "max")

The coverage check knows each source’s own step, so a complete CCMP day scores 4/4 rather than 4/24.

One caution that is easy to miss: a mean of the components is not a mean speed. Averaging UWND and VWND over a day and taking the magnitude gives the net displacement of air; averaging WSPD gives how hard it blew. On a day the wind reversed, the first is near zero and the second is not.

Longitude, and one trap handled for you

Pass bounding_box with longitudes negative west to every function here. CCMP is stored on a 0–360 grid and is converted on the way in and back on the way out, so its results overlay the others without adjustment. You should never have to think about it — but if you fetch CCMP by hand elsewhere, that is the difference between the Gulf of Maine and the Indian Ocean.

Putting several together

The pattern is the same as chaining two Copernicus products: each call adds columns and leaves the row count alone.

matched <- matchData(observations, accessCopernicus(vars = c("SST", "MLD"),
                                                years = 2010, months = 1:12,
                                                bounding_box = bb))
matched <- matchData(matched, accessHYCOM(vars = "BOTS", years = 2010,
                                          months = 1:12, bounding_box = bb))
matched <- matchData(matched, accessCCMP(vars = "WSPD", years = 2010,
                                         months = 1:12, bounding_box = bb))

matched <- attach_bathymetry(matched, fetch_bathymetry(bounding_box = bb),
                             c("DEPTH", "SLOPE"))
matched <- attach_climate_index(matched, "NAO")

colSums(is.na(sf::st_drop_geometry(matched)))

Each source fails differently at the edges of its record — FVCOM after 2013, HYCOM after 2015, LCR after 2014 — so that last line is worth running before fitting anything.

What to cite

Every source here is somebody else’s work, and the obligation to cite travels with the data rather than with this package. Cite whichever products you actually fetched from. Several catalogs carry their references at runtime:

fvcom_archives()$GOM3$reference
hycom_archives()$GLBv53X$reference
ccmp_versions()$`v03.1`$reference
index_dictionary()          # carries the climate index references

Copernicus Marine Service

Copernicus asks for a specific form, including the access date:

Product Title. E.U. Copernicus Marine Service Information (CMEMS). Marine Data Store (MDS). DOI: 10.48670/moi-xxxxx (Accessed on DD MMM YYYY)

Reanalysis products, used by default:

Product Supplies DOI
Global Ocean Physics Reanalysis (GLORYS12V1) SST, SSS, BOTT, BOTS, UO, VO, SSH, MLD, SIC 10.48670/moi-00021
Global Ocean Biogeochemistry Hindcast CHL_MODEL, NPP_MODEL, NO3, PO4, O2, PH 10.48670/moi-00019
Global Ocean Colour (Copernicus-GlobColour) satellite CHL, PP, DIATO, DINO 10.48670/moi-00281
Global Ocean Monthly Mean Sea Surface Wind and Stress WSPD, UWND, VWND, TAUX, TAUY, TAU 10.48670/moi-00181
Global Ocean Hourly Reprocessed Sea Surface Wind and Stress the same, with frequency = "hourly" 10.48670/moi-00185

BOTS is derived from GLORYS’s salinity field rather than published by it, so it carries that product’s citation like any other variable taken from it.

The two wind products are separate records with separate DOIs. A study using monthly wind cites the first; one using hourly wind, or a daily field aggregated from it, cites the second.

Analysis-and-forecast products, used with mode = "forecast":

Product DOI
Global Ocean Physics Analysis and Forecast 10.48670/moi-00016
Global Ocean Biogeochemistry Analysis and Forecast 10.48670/moi-00015

variable_dataset() says which product a variable came from, so only the ones you used need citing. Downloads go through the Copernicus Marine Toolbox, which publishes no DOI of its own — cite the products.

Ocean and wind models

  • FVCOM / NECOFS — Chen C, Beardsley RC, Cowles G (2006). An unstructured grid, finite-volume coastal ocean model (FVCOM) system. Oceanography 19(1):78–89. https://doi.org/10.5670/oceanog.2006.92

  • HYCOM / GOFS 3.1 — Chassignet EP, Hurlburt HE, Smedstad OM, Halliwell GR, Hogan PJ, Wallcraft AJ, Baraille R, Bleck R (2007). The HYCOM (HYbrid Coordinate Ocean Model) data assimilative system. Journal of Marine Systems 65:60–83. https://doi.org/10.1016/j.jmarsys.2005.09.016

  • NOAA CEFI / MOM6-COBALT-NWA12 — Ross AC, Stock CA, Adcroft A, Curchitser E, Hallberg R, Harrison MJ, Hedstrom K, Zadeh N, Alexander M, Chen W, Drenkard EJ, du Pontavice H, Dussin R, Gomez F, John JG, Kang D, Lavoie D, Resplandy L, Roobaert A, Saba V, Shin S-I, Siedlecki S, Simkins J (2023). A high-resolution physical–biogeochemical model for marine resource applications in the northwest Atlantic (MOM6-COBALT-NWA12 v1.0). Geoscientific Model Development 16:6943–6985. https://doi.org/10.5194/gmd-16-6943-2023

    The paper is the model; the data is separately owed an acknowledgement, in the form PSL asks for: Data provided by the NOAA Physical Sciences Laboratory, Boulder, Colorado, USA, from https://psl.noaa.gov/cefi_portal/. There is no separate data DOI.

    Say which release — source_of() records it, as cefi:NWA12-hindcast-r20250715. Releases extend and revise the record, so two results from different releases are not the same numbers. For a forecast, say which initialisation and ensemble member as well.

  • CCMP — Mears C, Lee T, Ricciardulli L, Wang X, Wentz F (2022). Improving the Accuracy of the Cross-Calibrated Multi-Platform (CCMP) Ocean Vector Winds. Remote Sensing 14(17):4230. https://doi.org/10.3390/rs14174230

  • MUR — Chin TM, Vazquez-Cuervo J, Armstrong EM (2017). A multi-scale high-resolution analysis of global sea surface temperature. Remote Sensing of Environment 200:154–169. https://doi.org/10.1016/j.rse.2017.07.029

  • VIIRS gap-filled chlorophyll — Liu X, Wang M (2018). Gap filling of missing data for VIIRS global ocean color products using the DINEOF method. IEEE Transactions on Geoscience and Remote Sensing 56:4464–4476. https://doi.org/10.1109/TGRS.2018.2820423

  • NASA OB.DAAC — two citations are owed, not one. The mission paper below describes the instrument; the data has a DOI of its own, under the 10.5067 prefix, minted per mission, per suite and per reprocessing — for example 10.5067/AQUA/MODIS/L3M/CHL/2022 for Aqua MODIS chlorophyll.

    The trailing version is not predictable from the suite: for Aqua MODIS the PAR DOI ending 2022 resolves where the SST one ending 2019 does not, and a reprocessing changes it. So this package points at the DOI rather than shipping a table of them that would go stale silently — look yours up at https://www.earthdata.nasa.gov/centers/ob-daac and cite it alongside the mission. obdaac_sensors() carries the mission papers at runtime, and source_of() records which sensor answered, as obdaac:MODISA-4km.

    • SeaWiFS — McClain CR, Feldman GC, Hooker SB (2004). An overview of the SeaWiFS project and strategies for producing a climate research quality global ocean bio-optical time series. Deep-Sea Research II 51:5–42. https://doi.org/10.1016/j.dsr2.2003.11.001
    • MODIS — Esaias WE, Abbott MR, Barton I, Brown OB, Campbell JW, Carder KL, Clark DK, Evans RH, Hoge FE, Gordon HR, Balch WM, Letelier R, Minnett PJ (1998). An overview of MODIS capabilities for ocean science observations. IEEE Transactions on Geoscience and Remote Sensing 36:1250–1265. https://doi.org/10.1109/36.701076
    • VIIRS — Wang M, Liu X, Tan L, Jiang L, Son S, Shi W, Rausch K, Voss K (2013). Impacts of VIIRS SDR performance on ocean color products. Journal of Geophysical Research: Atmospheres 118:10347–10360. https://doi.org/10.1002/jgrd.50793

Say which archive as well as which model — a value from the HYCOM reanalysis and one from an operational experiment are not the same run, and source_of() records exactly that.

Seafloor terrain

fetch_bathymetry() requests the 60 arc-second bedrock grid (ETOPO_2022_v1_60s_bed) through marmap.

Climate indices

Two are the published output of specific work and should be cited when used:

  • LCR — Jutras M, Dufour CO, Mucci A, Talbot LC (2023). Large-scale control of the retroflection of the Labrador Current. Nature Communications 14:2623. https://doi.org/10.1038/s41467-023-38321-y

  • AMOC — Moat BI, Smeed DA, Rayner D, Johns WE, Smith R, Volkov D, Elipot S, Petit T, Kajtar J, Baringer MO, Collins J (2026). Atlantic meridional overturning circulation observed by the RAPID-MOCHA-WBTS array at 26°N from 2004 to 2024 (v2024.1a). British Oceanographic Data Centre, NERC, UK. https://doi.org/10.5285/48d0bf43-0598-ceb2-e063-7086abc062f1

    BODC mints a new DOI for each release and retires the old one, so this changes when RAPID publishes a new version. index_dictionary() carries the current reference.

The other four are operational products with no single paper behind them. Credit the provider:

Software this is built on

citation("datamatch") gives this package’s own entry, and citation() works on any of the above.

Keeping these current

Citations go stale without anyone touching them. Data centres reissue a DOI when a record is superseded and retire the old one, so a reference that was correct when written stops resolving on its own. The AMOC entry here has already been through that once, when BODC published a newer RAPID release.

Two scheduled workflows watch for it, and open an issue rather than editing anything, since choosing a replacement is a judgement about which version to track:

  • Citation check, quarterly, resolves every DOI cited in the README, the catalogs, and NEWS.md.
  • Copernicus catalog check, monthly, compares the variable catalog against the live Copernicus catalogue, since dataset identifiers are revised too.

Both run on demand from the Actions tab, and locally:

# from the package root
system("Rscript inst/scripts/check_citations.R")
system("Rscript inst/scripts/check_catalog.R")

Satellite or model?

Copernicus serves chlorophyll and primary production from two very different sources, and both are available:

Name Source Resolution Trade-off
CHL, PP Copernicus-GlobColour, satellite 4 km Observed, finer — but surface-only and gappy under persistent cloud
CHL_MODEL, NPP_MODEL Biogeochemistry reanalysis 0.25° Gap-free and depth-resolved — but simulated, and coarser

The plain names default to satellite, since the values are observed rather than simulated. Switch to the model versions where cloud gaps would matter more than resolution.

One caution: satellite PP and model NPP_MODEL are not the same quantity. PP is depth-integrated (mg/m2/day), NPP_MODEL volumetric (mg/m3/day). Substituting one for the other is a units error, not a resolution difference.

Phytoplankton functional types come from the same satellite plankton dataset as CHL, so they can be fetched together:

env <- accessCopernicus(
  vars = c("CHL", "DIATO", "DINO"),   # one request, one dataset
  years = 2003:2017, months = 1:12,
  bounding_box = list(xmin = -76, xmax = -65, ymin = 35, ymax = 45)
)

DIATO and DINO are diatom and dinophyte chlorophyll — the spring-bloom species large copepods prefer, and the later stratified-water group respectively.

Bottom salinity

BOTS pairs with BOTT, but it is not fetched the same way, because GLORYS12V1 does not publish it. The reanalysis has temperature at the sea floor and no salinity counterpart. So it is derived: the full salinity column is fetched and the deepest wet level in each cell kept.

bots <- accessCopernicus(vars = "BOTS", years = 2010:2014, months = 1:12,
                     bounding_box = bb)
#> BOTS is not published by this product. Deriving it from the full 'so' column
#> and keeping the deepest wet level in each cell; the depth used comes back as
#> BOTS_depth.

The depth each value came from is returned as BOTS_depth rather than left to be assumed. That matters because the deepest wet model level is not the sea floor: level spacing coarsens with depth, so in deep water the value can sit a long way above the bottom. In shelf water it is within a few metres.

Two consequences, both deliberate:

  • It must be fetched on its own. The whole depth column is a different request from the single level SST wants, so mixing them is an error rather than a quiet second download. Call twice and chain matchData().
  • It costs far more. Roughly fifty levels are downloaded over the same box to keep one, so a large box is much slower than the same box of SST.

In forecast mode none of this applies. The analysis-and-forecast product publishes sea-floor salinity outright, so BOTS is an ordinary variable there, fetches alongside BOTT, and returns no BOTS_depth:

accessCopernicus(vars = c("BOTT", "BOTS"), years = 2026, months = 8,
             bounding_box = bb, mode = "forecast")

The two are the same quantity by construction but not the same number — one is the deepest level of a 50-level grid, the other Copernicus’s own diagnostic — so a record spanning both modes has a seam in it.

variable_dictionary("biogeochemical")        # filter by product
variable_dictionary("wind")                  # or just the winds
fvcom_dictionary()                           # the FVCOM catalog, for accessFVCOM()
fvcom_archives()                             # which FVCOM archives are built in
hycom_dictionary()                           # the HYCOM catalog, for accessHYCOM()
hycom_archives()                             # which HYCOM archives can be read
hycom_covering("2019-06-15")                 # which of them span a given date
ccmp_dictionary()                            # the CCMP catalog, for accessCCMP()
ccmp_versions()                              # which CCMP versions can be read
erddap_dictionary()                          # MUR and VIIRS, for accessERDDAP()
erddap_datasets()                            # which ERDDAP datasets ship
variable_dataset(c("SST", "CHL"))            # which dataset each comes from
as.data.frame(variable_dictionary())$description  # full descriptions

Error messages

Most of what this package refuses to do, it refuses deliberately. These are the messages you are most likely to meet, and what each one means.

Could not find the Copernicus Marine client

R cannot see copernicusmarine on its PATH. Common when it lives in a conda environment that RStudio does not inherit. Point at it directly:

options(datamatch.copernicusmarine = "~/miniconda3/bin/copernicusmarine")

These variables come from different Copernicus datasets

Expected, not a fault. SST is physics and CHL is biogeochemistry, on different grids. Fetch them separately and chain matchData(), as in Putting several together. variable_dataset() shows which dataset each variable comes from.

In forecast mode this happens more often, because the forecast products split variables across more datasets than the reanalysis does. SST and UO share a dataset as reanalysis but not as forecast.

Copernicus publishes no daily dataset for: UWND, VWND

Expected. This wind is published hourly or monthly and nothing between. Fetch frequency = "hourly" and aggregate with upscale_time(to = "day"). The same message for PH, PP, DIATO or DINO means the opposite — those are monthly composites, so fetch them monthly.

Copernicus publishes no hourly dataset for: WSPD

WSPD and TAU are magnitudes the hourly wind product does not carry. Fetch UWND and VWND hourly and compute the magnitude, or take these monthly.

BOTS must be fetched on its own

Bottom salinity is derived from the whole depth column, which is a different request from the single level a surface variable wants. Call accessCopernicus() once for BOTS and once for the rest, then chain matchData(). See Bottom salinity.

These variables sit on different parts of the FVCOM mesh

Scalars sit on mesh nodes and velocities on element centroids, so the two kinds cannot be read together. Call accessFVCOM() once for each and chain matchData().

The CEFI hindcast saves no daily SST

Expected. CEFI’s daily output is biogeochemistry only — CHL, NO3, PH, PCO2, PHYC, MESOZOO and BOTO2. Everything else is monthly, and the message lists what is available daily.

The bounding box selects no CEFI cells

CEFI is a regional model, not a global one. The Northwest Atlantic domain runs roughly 98°W–36°W and 5°N–58°N, and the message gives the actual extent. Longitudes are negative west, and the regridded files are already on that convention, so no conversion is needed in either direction.

CEFI's 'seasonal_forecast' cannot be read through this package

Not a fault in your call. Those files carry a 64-bit integer coordinate that DAP2 has no type for, so the server returns HTTP 500 to every OPeNDAP request; the DAP4 endpoint describes them and then reads them wrongly, crashing R. The files download whole from the THREDDS fileServer endpoint and open correctly once local, at roughly 260 MB per variable per initialisation.

init / member is required for CEFI's decadal forecast

Deliberate. A decadal file is ten years from one January and there are sixty of them, so which initialisation you mean is a question about your analysis. Each holds ten ensemble members, and averaging them is a modelling decision this package will not make silently — read the members you want and combine them yourself, so the combination is visible in your code.

Reading OB.DAAC needs an Earthdata Login

The one account this package cannot work around. Register at https://urs.earthdata.nasa.gov/users/new, generate an appkey at https://oceandata.sci.gsfc.nasa.gov/appkey/, and put it in ~/.Renviron as EARTHDATA_APPKEY=. A ~/.netrc entry for urs.earthdata.nasa.gov works too.

the server returned a login page rather than the file

The credential is missing, wrong, or revoked. NASA answers an unauthenticated request with HTTP 200 and the login page rather than a 401, so this is caught by checking the bytes; nothing is written to disk. A date the sensor did not return lands here too, since the two are indistinguishable from the response.

MODISA does not publish: NFLH (or similar)

The sensor does not carry that suite. SeaWiFS measures no temperature at all; only MODIS carries NFLH; VIIRSJ2 has no SST yet. The message lists what that sensor does carry, and obdaac_sensors() has the full table.

The download did not return: uo, vo

Those variables were requested but are not in the returned file. Either the dataset does not serve them, or they are unavailable at the requested depth or date. The message lists what did arrive.

The depth range returned several model levels

depth spanned more than one model level, so a variable came back on several layers and there is no way to know which was wanted. Request a single level with depth = c(0, 1).

Copernicus download failed after 3 attempt(s)

The request was retried and kept failing. The message carries the client’s own output. Rapid repeated downloads can also be rate-limited, in which case waiting and re-running works.

Expected N variable column(s) but the download returned M

From a version before the layer-matching fix. Update the package: this message blamed the depth range for causes it could not distinguish, and the two real ones now report themselves. See NEWS.md.

Values in the wrong range for their column name

Salinity around 32 in an SST column means you are on a version predating the layer-ordering fix, where multi-variable downloads could be mislabelled silently. Update and re-fetch. NEWS.md describes what was affected and how to check.