How far can two cameras bound a wave?
A stereo-video rig recovers the sea surface from the pixel offset between two images. Every pixel is a box — a disparity known to a fraction of a pixel, a vertical coordinate to half of one, a shutter that fired a few milliseconds late — and the set of elevations the rig cannot tell apart at a given range is a cell whose height is computed here exactly, in outward-rounded interval arithmetic, over inputs that may themselves be boxes. Turn the baseline, the height, the lens, the lag, the sea; read the bound at any range and the range at which the bound first exceeds a tolerance. Preset: the two-smartphone rig of Vieira, Guimarães, Violante-Carvalho, Benetazzo, Bergamasco and Pereira (2020) at Leme, with the paper's numbers as literals and the ones it does not state as boxes.
Four curves, each the UPPER end of an enclosure: the quantization cell (disparity to ±δd pixels, vertical coordinate to ±½ pixel; tight, because each measured quantity enters once), the cell with the synchronisation lag's texture shift added to the disparity, the slope term (a horizontal misplacement read on the steepest slope of the wave), and their total with the surface's own motion over the lag. The shaded band is the preset's imaged range; the thin vertical is the range you chose; the vertical marker is the reach at the tolerance. Solid where every input is a literal or a box from a source; dotted where any input is a choice.
The Leme rig, re-read
Two Samsung Galaxy J5 Pro phones 0.98 m apart, 3.5 m above the water, imaging 10–35 m of surf with a pressure gauge at ~34 m; 12.8-second swell of 0.35 m. The paper gives the baseline, the height, the focal length and the sea; it does not give the pixel pitch of the 1080p video or the precision of its matches, so those are boxes, and the audio synchronisation is taken to the frame.
Read together: every observed RMSE lies between the best and worst corners of the budget, so the geometry with the paper's own numbers accounts for the deviation from the gauge without invoking anything else — and the disparity at the gauge is only 43–95 pixels, which is why a pixel of matching error is 32.7 cm of elevation there. The quoted quantization of 1.1 mm is below what this model gives for any corner of the box at any range in the imaged area; the paper takes it from a formula in its reference [42] that we do not hold, so this is a number we could not reproduce, not one we refute.
What is decided, and what is assumed
The model is the standard rectified pinhole pair with both cameras aimed at the surface point: Z = fB/d, Y = Zv/f, and a rotation by the pitch. Nothing in it is new; what is new is that the cell is an enclosure — each measured quantity appears once in the elevation and once in the range, so the interval evaluation over the pixel box is tight, and it holds over the whole of a box input such as an unstated pixel pitch. The three added terms are bounds, not cells: the texture the matcher correlates is taken to move with the water (its orbital speed) plus a texture speed you choose; the surface rises or falls by at most ωH/2 over the lag; a horizontal misplacement is read on the steepest slope, kH/2. Deep-water dispersion gives ω and k from the period. A total is an interval and its upper end is what the verdict uses.
It bounds geometry, not matching. A matcher can fail outright on glare, foam or a textureless surface, and no error budget bounds a failed match; the disparity precision you enter is the precision of the matches that succeed. The pitch of the optical axis is set by aiming at the point — a point at the edge of a wide field of view sits off-axis and its cell differs. Lens distortion, calibration error and refraction are not in the budget. The Nazaré preset is a scenario with every number chosen, in the shape of the Big Wave Tracker, to be replaced by the rig's own.
node instruments/stereo/battery.js
node playground/build.js