Stacks the components of g(0) — availability, perception, anything else
measured separately — multiplies them, and propagates their variance. The
result is an object to hand to the abundance step, not a number applied here.
Arguments
- ...
One or more component tables, each with
key,component,valueandsecolumns, and optionallysource.availability()returns this shape.
Value
An object of class dsfit_g0: a list with table (one row per key,
carrying value, se and cv) and components (the rows it was built
from, source included, kept so the correction can always be taken apart
again).
Why this is an object and not an argument
dsfit fits detection functions; it does not compute abundance. The
correction is applied one layer out, where density is calculated, and that
layer is a targets repository rather than a package.
What this can do from here is make the correction impossible to get wrong
quietly. A dsfit_g0 cannot be built without naming its components, cannot
be built without standard errors, and cannot be built at all by accident —
so whatever consumes it downstream is handed a documented object rather than
a bare number of unknown provenance.
The three rules
From docs/01-plan.md, and each one is enforced rather than documented:
- Never silently 1
There is no default. A component that is absent is warned about by name, because assuming perception is 1 while believing you have corrected for it is how these estimates go wrong by a factor.
- Propagate the variance or refuse
A component without a standard error is an error, not a point estimate. The CV of
g(0)routinely dominates the CV of abundance, so dropping it inverts which uncertainty matters.- Name the components separately
Availability, perception and the geometric blind spot are three different things measured three different ways. They are kept as rows, not collapsed, so it stays visible which have been applied.
How the components combine
Multiplicatively, and independently — availability comes from dive data and perception from a double-observer trial, so they are estimated from separate studies and the covariance between them is not merely unknown but usually undefined.
Under independence the delta method gives a result worth remembering: the squared CVs add.
$$CV(g_0)^2 = \sum_i CV(x_i)^2$$
So the least precise component sets the floor. A perception estimate with a 30% CV cannot be rescued by an availability estimate with a 2% one.
Keys, and the trap in them
Components are matched on key. A component with a single NA key applies
to every key — a perception estimate that does not vary by month, say,
combined with availability that does.
Components keyed on different things are refused rather than joined. Availability by month and perception by year is the common case, and it has no correct join: the answer is a value per month-year, which means building the cross product yourself and keying on it. Silently pairing January with 1998 because both are first in their vectors is exactly the failure this refuses to commit.
Recording where a component came from
Components may carry an optional source column, and it is worth using.
Perception is the reason. Estimating it needs two independent observer teams, and whether a survey programme ran them is a property of that programme rather than of the archive its data ends up in. A NARWC extract may or may not carry the structure, so for many datasets the only available perception estimate is one borrowed from a different programme. Roberts et al. (2024) did exactly this — perception was estimable only from NOAA's AMAPPS surveys, and they applied those corrections to the other ten institutions' data, cautioning in print that it could have biased their density estimates.
A borrowed correction that looks local is the failure this guards against. So
source travels with the component into the object and is printed every
time, and a component with none is shown as source not recorded rather than
silently blank.
References
Laake, J.L. and Borchers, D.L. (2004) Methods for incomplete detection at distance zero. In Advanced Distance Sampling, pp. 108-189. Oxford University Press. Why the components are separate quantities.
See also
availability() for the availability component. sweep_models()
for why none of this enters the detection function fit.
Examples
avail <- availability(surface = 60, dive = 240, window = 24,
se_surface = 8, se_dive = 25)
avail$source <- "focal follows, this survey"
# A perception estimate that is not this survey's, said so
perception <- data.frame(
key = NA_character_, component = "perception", value = 0.68, se = 0.09,
source = "AMAPPS double-observer, borrowed - not measured on this survey"
)
g0(avail, perception)
#> <dsfit_g0>
#> components:
#> availability focal follows, this survey
#> perception AMAPPS double-observer, borrowed - not measured on this survey
#>
#> key g0 se cv
#> - 0.1878 0.032 0.171
#>
#> Divide abundance by g0, and propagate cv. This package does not
#> apply it: correction happens at the abundance step.
# Without perception at all, which is said too
suppressWarnings(g0(avail))
#> <dsfit_g0>
#> assumed 1: perception
#> components:
#> availability focal follows, this survey
#>
#> key g0 se cv
#> - 0.2761 0.0297 0.108
#>
#> Divide abundance by g0, and propagate cv. This package does not
#> apply it: correction happens at the abundance step.