Method
This page describes how a release is built. The first release will publish the build code, and this page will link to it.
Sources
A release combines official constants published by agencies (NOAA CO-OPS, Kartverket), constants OTC fits from GESLA gauge records, and, where they pass the checks, other open sources. TICON-4, the XTide harmonics files and the Slackwater database are used only as cross-checks. Each source, its licence, attribution, update cadence, coverage and rank is described on the Sources page.
Fitting
- Basis: least squares on the reference predictor's own equation, h(t) = Z0 + c·(t − t̄) + Σ fk·(ak cos Xk + bk sin Xk), where Xk = speedk·t + (V0+u)k. Speeds, V0+u and f come from the same tables the predictor uses. The fitted constants therefore plug into the predictor exactly, and the convention mismatches on the Why page cannot arise.
- Span: the last 19 years of good data (one nodal cycle is 18.61 years). A shorter record is fitted whole; records under 1 year are not used, and records of 1–3 years are flagged as short. A linear trend is fitted alongside, so sea-level rise does not leak into the annual constituent.
- Quality control before fitting: null values, values flagged as not usable, and duplicate time stamps are dropped. Sub-hourly records are reduced to one value per hour. Spikes are removed against a rolling median of the residual, and the fit is repeated.
- Uncertainty: for each constituent, the larger of the residual noise in its frequency band and half the difference between fits to the two halves of the span.
- Determinism: fixed random seeds, sorted inputs, fixed output precision. Two builds from the same inputs must give byte-identical files.
Time-base checks
Every record that a build might use is checked for a wrong or undeclared clock. The metadata cannot be trusted for this: in GESLA-4.1, 6,681 of 6,682 records declare a time zone offset of 0, including records that are not UTC.
- Comparators, in order of preference: official predictions near the gauge; NOAA harmonic constants for NOAA records; and other records at the same gauge from a different contributor, after their own correction.
- Measurements: the monthly time lag at 1-min resolution, the lag per constituent, amplitude ratios, a direct test of instantaneous against averaged sampling, and the missing hour on daylight-saving change days.
- Verdict classes, each with a fixed correction: UTC instantaneous; UTC hourly mean stamped at the start of the hour (shift +30 min and correct the amplitudes); UTC plus a whole number of hours; local time with daylight saving; a step at a date; isolated bad months.
- Pass rule: after the correction, the residual M2 lag must be within ±3 min in every month used. A record that matches no class is not used as it is.
- Positive controls on every build: pinned GESLA-4.0 WSV files must come out as local time with daylight saving; a record shifted by +60 min must be detected; an hourly-averaged copy must be detected. If a control fails, the build stops.
Broken-record checks
A record is excluded, and listed with its numbers, when its main constituents are not stable from year to year or between the two halves of the span, when it does not correlate with a nearby gauge that correlates with its own neighbours, or when its values are quantised or out of range. The Hirtshals record (no tide in the input) is the reference case. The thresholds are proposals that the first build calibrates and publishes.
Record selection
When several records exist at one gauge, the build picks one by fixed rules, in this order: the most good data in the 19-year span; recent records before records that ended long ago; the gauge operator's own record before a re-distribution of it; instantaneous before averaged; sub-hourly before hourly; then the lowest file id. When two eligible records disagree beyond their uncertainty, the record with the lower error against an official reference is chosen. Every choice and its reason is recorded.
Constituent selection
- Candidates: every constituent the predictor's tables define: NOAA's 37 plus EPS2, MKS2, SIG1, N4, S3, 2MK5, 2MO5 and 2MS6, in a fixed priority order.
- Record length (Rayleigh rule): a candidate is kept only when the record is long enough to separate it from every higher-priority constituent already kept.
- Signal to noise: after a first fit, a constituent is kept only when its amplitude is at least twice the residual noise in its frequency band. Then the fit is repeated.
- Long-period constituents (SA, SSA, MM, MF, MSF) are left out of the published set at river and lake gauges, and where SA is 0.25 m or more, because official predictions do not contain river and seasonal discharge signals.
- Non-tidal daily cycles (S1, S3, S4) are left out at river and lake gauges, and where they are larger than both M2 and K1.
In the proof of concept this kept 41–45 constituents for 19-year records and 23–28 for records of 1.5–2.5 years. Every constituent in a release records why it was kept or dropped.
Per-source convention check
Every source is compared on every build with an independent comparator at the same gauges: our own GESLA fit, predicted with the reference predictor. The proposed pass rule is that the median M2, S2, K1 and O1 differences are within 1 cm and 2° over at least 3 overlapping gauges. A source that fails is not published in that release until its declared convention is fixed. A source with no overlapping gauge cannot be added.
Automation
- Watchers check every source weekly and produce a list of changed records.
- Only changed sets are rebuilt. A new GESLA release, a change in the predictor's tables, or a format change triggers a full rebuild.
- Builds run on GitHub Actions, so the build history is public.
- After a release, the webcaltides app receives the new release through an automated, pinned update that merges only when its own checks pass.
The release gate (no person in the loop)
Every judgement in a build is made in three tiers, and nothing waits for a person:
- Fixed rules decide every case they can classify: pass, fail or fall back.
- One batched automated review per build handles the cases the rules cannot classify. Each answer is either accept or fall back, with a reason that cites the evidence. Answers are stored in a public table keyed by the evidence, and reused while the evidence is unchanged, so a rebuild is reproducible.
- Fallback: anything still undecided keeps the station's data from the last release, and the reason is listed in the changelog.
A release is published only when every whole-release gate passes, including the round-trip checks of the converters and the uploads. If any gate fails, nothing is published, the last good release stays live, and an alert opens a public issue. Licence and publication policy stay with the project owner and never block a build.