Running only the throughput benchmark#
notebooks/throughput_scaling_benchmark.ipynb measures operator-construction
and solve throughput as a function of batch size, under three regimes: one
operator shared across a batch of repeated acquisitions of the same
footprint, the same footprint with the operator rebuilt every item, and a
batch of distinct geographic patches. It is the paper’s batch-scaling
benchmark (supplement), and it can be run on its own — none of the other
experiment notebooks are prerequisites.
The commands below are documentation only; they do not run during the documentation build (the benchmark needs a CUDA GPU to be meaningful).
Setup#
git clone https://github.com/GRID4EARTH/healpix-resample.git
cd healpix-resample
pixi install -e notebooks
pixi run -e notebooks jupyter lab
Choosing the patch pool#
The benchmark selects its patch pool automatically, in order of preference:
Archived Sentinel-2 regions (the paper’s configuration): the 40 real region patches of Zenodo record 22210945. Install them by running
notebooks/load_data_in_zenodo.ipynbfirst.Esri-derived cache: the 360 regenerated multi-region patches, if present (see the full reproduction tutorial).
Synthetic fallback: deterministic synthetic textures on the versioned site-manifest geometries — no download needed at all.
Operator-construction cost depends on patch geometry only, so the scaling
shape is meaningful in every tier; to reproduce the paper’s numbers, use
tier 1. The chosen pool is recorded as patch_pool_source in
notebooks/tables/throughput_scaling_environment.csv.
Running#
Open throughput_scaling_benchmark.ipynb and run it top to bottom. Two
cells deserve attention before the sweep:
the cost-estimator cell prints a runtime estimate extrapolated over
BATCH_SIZESandN_BATCH_REPEATS, so you can decide whether to shortenBATCH_SIZESfirst (the full 1–2048 sweep takes a few hours on a datacenter GPU);N_BATCH_REPEATS = 1reproduces the paper’s single descriptive sweep; raise it for repeat-to-repeat uncertainty at the cost of runtime.
Outputs and what to expect#
The run writes, under notebooks/:
tables/throughput_scaling_by_batch.csv— one row per regime × batch size,tables/throughput_scaling_by_item.csv— per-item build and solve timings,tables/throughput_scaling_environment.csv— hardware and configuration snapshot,figures/throughput_scaling_curve.pdfandfigures/throughput_scaling_memory.pdf.
Timing values are hardware-dependent: expect the ratios, not the absolute times, to reproduce. On the paper’s reference system (NVIDIA L4, Sentinel-2 pool), sharing one operator across a batch of 2048 repeated acquisitions reaches 0.075 s/patch against ~0.86 s/patch for both rebuild regimes — an ~11.5× amortized speedup — while batches of distinct geographic patches gain nothing, since each needs its own operator.
Two reading rules from the notebook’s reporting guidance: the
different_pointing regime samples the pool with replacement once the
batch exceeds the pool size (check unique_patches_used before citing such
a point), and this is a single-GPU sequential benchmark — do not
extrapolate to multi-GPU or parallelized deployments.