Writes a CSV matching the schema original/create_database.R produces, so the
pipeline can be exercised end-to-end without the real survey data, which stays
on its owner's machine.
Arguments
- config
a config list, as returned by
load_config()- n_stations
number of station visits to generate per year
- seed
random seed
Details
Two details are copied from the real database's structure rather than invented, because the pipeline's behavior depends on them:
Stage columns that a survey does not resolve are filled with zeros, not NA (
create_database.Rlines 32-55 for ECOMON), so<species>_totalcollapses to the single stage ECOMON measures.datasettakes the valuesECOMON,ECOMON_STAGED,MBON,CPR, so the config'sdataset_filteris actually exercised rather than matching every row.
Abundance is given real spatial and seasonal structure — a temperature-like gradient plus a summer peak — so a model fit to it should score meaningfully better than chance. A smoke test that passed on pure noise would not be testing anything.