From 8cdfad5e06769fb7955c8f7a76ec8efe68826be2 Mon Sep 17 00:00:00 2001 From: Claude Date: Sat, 22 Aug 2026 21:19:03 +0000 Subject: [PATCH] Document the DFT/NUFFT trade-off on the sparse-operator setup path MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both transformers work with `apply_sparse_operator` and agree to ~3e-13 relative, but which is *faster* was undocumented, and the answer is not the one the codebase's existing 10,000-visibility guard implies. Measured on CPU (`use_jax=False`), the crossover is governed by the product `N_vis * N_pix`, not by visibility count alone — which follows from the DFT setup being `O(N_vis * N_pix)` against the NUFFT's `O((N_vis + N_pix) log N)` plus a fixed ~2s overhead. Seven measurements agree on a crossover near 1e7: N_vis x N_pix DFT/NUFFT time ratio 1.3e6 0.21x (DFT faster) 2.6e6 0.27x 5.2e6 0.70x 1.0e7 1.16x (crossover) 2.1e7 1.50x 2.3e7 1.41x 8.3e7 1.92x (NUFFT faster) So at a typical 64x64 mask the crossover sits near 5,000 visibilities, but on a 32x32 mask the DFT still wins at 4,000. Quoting a visibility count alone would be wrong on half the grids. Memory is the more decisive axis and is what makes this more than a micro-optimisation. The DFT's allocation grows with `N_vis * N_pix` (measured 60 -> 239 -> 293 -> 446 MB as the grid grows at fixed N_vis=4000) while the NUFFT path allocates nothing measurable beyond its working buffers. At 10.7 bytes per element that extrapolates to ~109 GB at 1M visibilities, which independently corroborates the ~123 GB figure recorded in bd18a769 and confirms the NUFFT is the only feasible path at ALMA scale. Adds the guidance as an INFO on the existing setup log, emitted only when a NUFFT-backed transformer is used below the crossover — i.e. only on an explicit, expensive, opt-in call, and never when the advice would be moot. Deliberately NOT a warning on `Interferometer` construction: that class defaults to `TransformerNUFFT`, so a small-dataset warning would fire on the default path for every tutorial, test and workspace example. The dangerous direction is already a hard `DatasetException` above 10,000 visibilities with an opt-out, so nothing is left unguarded by keeping this advisory quiet. Verified the advisory fires for small+NUFFT and stays silent for both small+DFT and large+NUFFT. Suite green: 1179 passed. --- autoarray/dataset/interferometer/dataset.py | 29 +++++++++++++++++++++ 1 file changed, 29 insertions(+) diff --git a/autoarray/dataset/interferometer/dataset.py b/autoarray/dataset/interferometer/dataset.py index 3738230d..69ac03ec 100644 --- a/autoarray/dataset/interferometer/dataset.py +++ b/autoarray/dataset/interferometer/dataset.py @@ -232,6 +232,21 @@ def apply_sparse_operator( the number of visibilities and the real-space mask resolution — potentially hours for large datasets). The result can be cached to disk and reloaded to avoid recomputation. + Both `TransformerDFT` and `TransformerNUFFT` are supported here and agree to ~3e-13 relative. + Which is faster is governed by the product `N_vis * N_pix`, because the DFT setup cost scales + as `O(N_vis * N_pix)` whereas the NUFFT scales as `O((N_vis + N_pix) log N)` on top of a fixed + ~2s overhead: + + - below `N_vis * N_pix ~ 1e7`, `TransformerDFT` is faster (measured 0.2-0.7x the NUFFT time) + - above it, `TransformerNUFFT` is faster (1.2-1.9x at 1e7-1e8) + - above `~1e8`, the DFT's `O(N_vis * N_pix)` allocation makes it infeasible rather than merely + slow — extrapolating a measured 446 MB at N_vis=4e3 / N_pix=1e4 gives ~109 GB at 1M + visibilities, whereas the NUFFT path allocates nothing beyond its working buffers. + + As a rule of thumb at a typical 64x64 mask the crossover sits near 5,000 visibilities, but on a + coarser mask the DFT stays ahead well beyond that — it is the product that matters, not the + visibility count alone. + Parameters ---------- nufft_precision_operator @@ -264,6 +279,20 @@ def apply_sparse_operator( "INTERFEROMETER - Computing NUFFT Precision Operator; runtime scales with visibility count and mask resolution, CPU run times may exceed hours." ) + n_vis = self.uv_wavelengths.shape[0] + n_pix = self.real_space_mask.pixels_in_mask + + if n_vis * n_pix < 10**7 and not isinstance( + self.transformer, TransformerDFT + ): + logger.info( + f"INTERFEROMETER - This dataset is small for the NUFFT setup path " + f"(N_vis x N_pix = {n_vis * n_pix:.1e}, below the ~1e7 crossover). " + f"`TransformerDFT` computes this operator faster below that point; " + f"`TransformerNUFFT` wins above it, and above ~1e8 it is the only " + f"option that fits in memory. Both are supported here." + ) + nufft_precision_operator = self.psf_precision_operator_from( chunk_k=chunk_k, show_progress=show_progress,