Skip to contents

Stacks the components of g(0) — availability, perception, anything else measured separately — multiplies them, and propagates their variance. The result is an object to hand to the abundance step, not a number applied here.

Usage

g0(...)

Arguments

...

One or more component tables, each with key, component, value and se columns, and optionally source. availability() returns this shape.

Value

An object of class dsfit_g0: a list with table (one row per key, carrying value, se and cv) and components (the rows it was built from, source included, kept so the correction can always be taken apart again).

Why this is an object and not an argument

dsfit fits detection functions; it does not compute abundance. The correction is applied one layer out, where density is calculated, and that layer is a targets repository rather than a package.

What this can do from here is make the correction impossible to get wrong quietly. A dsfit_g0 cannot be built without naming its components, cannot be built without standard errors, and cannot be built at all by accident — so whatever consumes it downstream is handed a documented object rather than a bare number of unknown provenance.

The three rules

From docs/01-plan.md, and each one is enforced rather than documented:

Never silently 1

There is no default. A component that is absent is warned about by name, because assuming perception is 1 while believing you have corrected for it is how these estimates go wrong by a factor.

Propagate the variance or refuse

A component without a standard error is an error, not a point estimate. The CV of g(0) routinely dominates the CV of abundance, so dropping it inverts which uncertainty matters.

Name the components separately

Availability, perception and the geometric blind spot are three different things measured three different ways. They are kept as rows, not collapsed, so it stays visible which have been applied.

How the components combine

Multiplicatively, and independently — availability comes from dive data and perception from a double-observer trial, so they are estimated from separate studies and the covariance between them is not merely unknown but usually undefined.

Under independence the delta method gives a result worth remembering: the squared CVs add.

$$CV(g_0)^2 = \sum_i CV(x_i)^2$$

So the least precise component sets the floor. A perception estimate with a 30% CV cannot be rescued by an availability estimate with a 2% one.

Keys, and the trap in them

Components are matched on key. A component with a single NA key applies to every key — a perception estimate that does not vary by month, say, combined with availability that does.

Components keyed on different things are refused rather than joined. Availability by month and perception by year is the common case, and it has no correct join: the answer is a value per month-year, which means building the cross product yourself and keying on it. Silently pairing January with 1998 because both are first in their vectors is exactly the failure this refuses to commit.

Recording where a component came from

Components may carry an optional source column, and it is worth using.

Perception is the reason. Estimating it needs two independent observer teams, and whether a survey programme ran them is a property of that programme rather than of the archive its data ends up in. A NARWC extract may or may not carry the structure, so for many datasets the only available perception estimate is one borrowed from a different programme. Roberts et al. (2024) did exactly this — perception was estimable only from NOAA's AMAPPS surveys, and they applied those corrections to the other ten institutions' data, cautioning in print that it could have biased their density estimates.

A borrowed correction that looks local is the failure this guards against. So source travels with the component into the object and is printed every time, and a component with none is shown as source not recorded rather than silently blank.

References

Laake, J.L. and Borchers, D.L. (2004) Methods for incomplete detection at distance zero. In Advanced Distance Sampling, pp. 108-189. Oxford University Press. Why the components are separate quantities.

See also

availability() for the availability component. sweep_models() for why none of this enters the detection function fit.

Examples

avail <- availability(surface = 60, dive = 240, window = 24,
                      se_surface = 8, se_dive = 25)
avail$source <- "focal follows, this survey"

# A perception estimate that is not this survey's, said so
perception <- data.frame(
  key = NA_character_, component = "perception", value = 0.68, se = 0.09,
  source = "AMAPPS double-observer, borrowed - not measured on this survey"
)

g0(avail, perception)
#> <dsfit_g0>
#>   components:
#>     availability  focal follows, this survey
#>     perception    AMAPPS double-observer, borrowed - not measured on this survey
#> 
#>  key     g0    se    cv
#>    - 0.1878 0.032 0.171
#> 
#>   Divide abundance by g0, and propagate cv. This package does not
#>   apply it: correction happens at the abundance step.

# Without perception at all, which is said too
suppressWarnings(g0(avail))
#> <dsfit_g0>
#>   assumed 1:   perception
#>   components:
#>     availability  focal follows, this survey
#> 
#>  key     g0     se    cv
#>    - 0.2761 0.0297 0.108
#> 
#>   Divide abundance by g0, and propagate cv. This package does not
#>   apply it: correction happens at the abundance step.