Filtering
CoarseGrainingEnergyFluxes.Filtering.DEFAULT_CACHE_BYTE_BUDGET — Constant
DEFAULT_CACHE_BYTE_BUDGETDefault byte budget for AutoCache's size check on a real-space footprint cache (256 MiB).
CoarseGrainingEnergyFluxes.Filtering.AbstractCacheStrategy — Type
AbstractCacheStrategyWhether a real-space footprint over a genuinely nonuniform axis (StructuredGrid with a Vector axis, CurvilinearGrid, or ND with a non-Range axis) stores its full per-point neighbour list, or recomputes it on the fly at apply time. The per-point neighbour/weight computation itself is always the same either way (there is no shared translation-invariant table to exploit on a nonuniform axis, unlike the FilterFootprint fast path) — this only controls whether that computation's RESULT is kept in memory for reuse across separate future filter_apply! calls, or re-derived each time.
CoarseGrainingEnergyFluxes.Filtering.AbstractFilterMethod — Type
AbstractFilterMethodHow the convolution is evaluated: RealSpace (physical-space footprint, any grid/mask) or Spectral (transform-space multiply — FFT for uniform periodic Cartesian, spherical harmonics for the uniform sphere, NUFFT/NUFSHT for scattered points; provided by extensions).
CoarseGrainingEnergyFluxes.Filtering.AbstractFilterPlan — Type
AbstractFilterPlanA prebuilt filter (grid + kernel + scale + mask strategy + backend) that can be applied to many fields without redoing setup. Physical-space backends precompute a FilterFootprint; the spectral extensions (FFTW/FINUFFT/SHT) hold cached transform plans. Declared here rather than alongside PhysicalFilterPlan further down so that filter_field!'s filter_plan::Union{Nothing, AbstractFilterPlan} keyword annotation, just below, can name it.
CoarseGrainingEnergyFluxes.Filtering.AbstractFilterScratch — Type
AbstractFilterScratchTransient per-apply buffers, sized by the grid and the applied array's rank. This is the only part of a plan mutated during filter_apply!, so one scratch may not be shared between concurrent applies: a driver that runs applies in parallel (filter_slices!, the batch pipeline drivers, Distributed/MPI ranks) must give each worker its own. Scales within one sweep run sequentially and therefore share.
CoarseGrainingEnergyFluxes.Filtering.AbstractGridPlan — Type
AbstractGridPlanPrecomputation determined by the grid, the kernel family, the mask strategy and the method — never by the filter scale. One instance serves every scale of a sweep; see plan_filter_sweep.
CoarseGrainingEnergyFluxes.Filtering.AbstractMaskStrategy — Type
AbstractMaskStrategyHow masked (inactive) cells enter the filter normalization.
CoarseGrainingEnergyFluxes.Filtering.AlwaysCache — Type
AlwaysCache <: AbstractCacheStrategyForce building the full per-point neighbour-list cache regardless of cache_byte_budget — for a caller who knows more memory is available than the conservative default budget assumes.
CoarseGrainingEnergyFluxes.Filtering.AutoCache — Type
AutoCache <: AbstractCacheStrategyBuild and store the full per-point neighbour-list cache only if its estimated size is under cache_byte_budget (default DEFAULT_CACHE_BYTE_BUDGET); otherwise fall back to recomputing neighbours on the fly at apply time. This is the only cache-strategy knob most callers ever need — it caches whenever doing so is affordable, which is strictly better than not caching whenever a plan will be reused across more than one filter_apply! call.
CoarseGrainingEnergyFluxes.Filtering.AutoMethod — Type
AutoMethod <: AbstractFilterMethodPick the engine from real capability, the same contract AutoCache and AutoBackend follow: Spectral only where a transform is available AND exact for this grid, otherwise RealSpace.
Selects Spectral only where every axis is periodic and uniform and a transform backend is loaded — the case where a periodic transform is the exact filter. Otherwise RealSpace.
CoarseGrainingEnergyFluxes.Filtering.BandedScratch — Type
BandedScratch{T,MT,VM} <: AbstractFilterScratchPer-apply buffers for the banded engine: the masked input mask · field, one per field in flight.
Empty for an unmasked grid — there the engine convolves the caller's array directly, since mask · field == field and the copy would be pure cost. Slots are added on demand for a batched apply and then reused, so repeat applies allocate nothing.
CoarseGrainingEnergyFluxes.Filtering.Deformable — Type
Deformable <: AbstractMaskStrategyMasked cells are excluded from BOTH numerator and denominator, so the kernel is renormalized over the locally-included area only ("deformable kernel"). Excluded cells are genuinely dropped, but the kernel becomes inhomogeneous near a mask boundary (breaks the strict commutation theorems).
CoarseGrainingEnergyFluxes.Filtering.FilterFootprint — Type
FilterFootprint{T,VI,VT,MT,SC,MS}Precomputed convolution footprint for a structured grid + kernel + scale. The in-support neighbour offsets (di, dj) and their geometric weights w = kernel_weight(distance) * cell_area are stored in a flat CSR-like layout, grouped into axis-2 (y) bands (ptr[b]:ptr[b+1]-1). For Cartesian grids the footprint is translation-invariant → a single band; for a spherical grid it is invariant in x (longitude) → one band per y (latitude) value.
The offsets and weights are geometry only. The normalization is not: invden is the reciprocal window mass, which depends on the mask, the mask strategy and where the window is truncated by a domain edge. It is accumulated once at plan build by running the SAME tap loop the numerator uses over a mass field — 1 everywhere for ZeroFill, whose denominator counts every in-support tap regardless of the mask, and the mask itself for Deformable, which drops inactive cells from both sums.
Precomputing it is what lets the apply hold the tap index in the outer loop and convolve by contiguous axpy along x, at every column of every grid. A denominator accumulated per point would force the tap index inside, turning the inner loop into a floating-point reduction that cannot be reassociated and so does not vectorize. Masked grids, and the columns within one filter radius of a domain edge, therefore run the same vectorized path as the unmasked interior.
CoarseGrainingEnergyFluxes.Filtering.FilterFootprintND — Type
FilterFootprintND{N, T}General N-dimensional footprint: in-support neighbour offsets (NTuple{N,Int}) and their geometric weights w = kernel_weight(distance) · cell_measure. Used for 1D and 3D (Cartesian, translation-invariant ⇒ a single offset set); the 2D path uses the optimized per-row FilterFootprint.
CoarseGrainingEnergyFluxes.Filtering.FilterPlanFamily — Type
FilterPlanFamily(grid_plan, plans, scratch)The plans for a whole sweep of scales over one grid, together with the two pieces they share: the scale-independent AbstractGridPlan and the per-apply AbstractFilterScratch.
Indexing and iteration go straight to plans, so anywhere a Vector of per-scale plans was expected a family works unchanged.
grid_plan/scratch are nothing for an engine that has no shared state to hoist; the family is then just the vector of plans and costs nothing extra.
Because the scales share one scratch, the plans of one family may not be applied concurrently with each other. Give each concurrent worker its own family.
CoarseGrainingEnergyFluxes.Filtering.NDScatteredCache — Type
NDScatteredCache{N, T}The full per-target-point neighbour list for an NDScatteredFilterPlan (absolute neighbour multi-indices + weights), built only when the plan's AbstractCacheStrategy calls for it.
CoarseGrainingEnergyFluxes.Filtering.NDScatteredFilterPlan — Type
NDScatteredFilterPlan{N, T, K}N-D (1D or 3D) analog of ScatteredFilterPlan: compact scalar metadata (per-axis window limits, periodicity/period, geometry flag) for when at least one of the N axes is a plain AbstractVector (no type-level uniformity proof) — no translation-invariance assumption, correct for any spacing pattern. cache holds the materialized NDScatteredCache only when the plan's cache strategy decided to build it, nothing otherwise (apply-time recomputation).
CoarseGrainingEnergyFluxes.Filtering.NeverCache — Type
NeverCache <: AbstractCacheStrategyForce recomputing neighbours on the fly at every filter_apply!/filter_apply_batch! call, never storing the cache — for a genuine, known memory ceiling AutoCache's budget check doesn't already account for (e.g. a GPU's device memory budget, or deliberately running many large plans concurrently). Not a speed/memory preference: for any workflow that reuses a plan across more than one call, caching is strictly better whenever it fits in memory, so this should be reached for only when a specific external memory constraint is known, not by default.
CoarseGrainingEnergyFluxes.Filtering.NodeFilterPlan — Type
NodeFilterPlan{T}Real-space filter footprint for a node set: per-node CSR neighbour blocks with their geometric weights w = kernel_weight(d) · control volume.
A node set has no axes, so there is no index window to bound a search with and the neighbourhood cannot be re-derived per apply the way a structured grid's can. It is found once, at plan time, and stored — which also means the search is paid once per plan rather than once per field, and a single compute_Π! makes six to nine applies against one plan.
The node itself is included. Connectivity.neighbors_within excludes it, matching stencil semantics where a cell is not its own neighbour, but a filter's zero offset is a genuine contribution.
CoarseGrainingEnergyFluxes.Filtering.PhysicalFilterPlan — Type
Physical-space plan: a precomputed footprint reused across all longitudes, fields, and layers — for EVERY backend, not just serial. kernel/scale are retained only so the cached-footprint path can still call each backend's row-parallel hook (which takes them positionally); they're not used to rebuild the footprint once footprint is already built.
CoarseGrainingEnergyFluxes.Filtering.PrefixSum3DScratch — Type
PrefixSum3DScratch{T,A} <: AbstractFilterScratchThe 3-D top-hat engine's per-apply buffer: the axis-1 cumulative scan of mask · field, sized (Nx+1, Ny, Nz). Refilled at the start of every apply, so one copy serves a whole sweep.
CoarseGrainingEnergyFluxes.Filtering.PrefixSumGridPlan — Type
PrefixSumGridPlan{T,VT,DT,BT,WX,WY,MS} <: AbstractGridPlanThe half of the prefix-sum top-hat engine that the filter scale does not reach: everything fixed by the grid, its mask and the mask strategy. One instance serves every scale of a sweep.
Holds the separable measure factors (wx,wy), the extended axis-1 coordinate array the two-pointer sweep walks (tripled when axis 1 is periodic, so a wrapped support interval is still one contiguous run), and the two mask/measure prefix scans.
prefix_den (Deformable masking) is the prefix sum of mask·wx; ZeroFill's denominator is mask-independent and needs only the 1-D prefix_wx. Neither depends on ℓ, which is why an S-scale sweep builds them once rather than S times.
CoarseGrainingEnergyFluxes.Filtering.PrefixSumScratch — Type
PrefixSumScratch{T,MT,VM} <: AbstractFilterScratchThe prefix-sum engine's per-apply buffers: one cumulative mask·field·wx scan per field in flight. They are refilled at the start of every apply, so their contents never outlive one call and one set serves the whole sweep. Row-disjoint, so the threaded numerator fill writes them concurrently by row.
A single-field apply uses slot 1. A batched apply needs all K live at once so that one walk of the support interval can feed every field — see apply_prefixsum_tophat_batch_row! — so slots are added on demand and then reused, which keeps repeat applies allocation-free.
CoarseGrainingEnergyFluxes.Filtering.PrefixSumTopHat3DPlan — Type
PrefixSumTopHat3DPlan{T,SC,A,WT,MS}Exact O(N·w_y·w_z) top-hat footprint for a uniform 3-D Cartesian StructuredGrid.
wcell[dj, dk] is the axis-1 half-width of the ball's slice at that offset, or -1 where the slice is empty. It is a function of the offsets alone — a uniform grid makes the ball translation-invariant — so it is a (2·dj_lim+1) × (2·dk_lim+1) table, not a per-point one.
invden is the reciprocal window mass, built once per scale under the plan's mask strategy; see the section comment above.
CoarseGrainingEnergyFluxes.Filtering.PrefixSumTopHatPlan — Type
PrefixSumTopHatPlan{T,GP,SC,MT,WT}Exact O(N·dj_lim) top-hat footprint for a rectilinear 2D StructuredGrid — see the section comment above for the derivation.
Only the ℓ-dependent state lives here. The grid half is reached through grid_plan and the per-apply buffer through scratch, both shared with every other scale of the same sweep, so S scales hold one copy of each rather than S.
Two field-independent quantities are built here, once per scale, so that the apply is numerator-only:
invden— the reciprocal window mass, a function of the grid, the mask, the strategy and ℓ alone. Accumulating it alongside the numerator would double the arithmetic of the apply's inner loop and repeat that for each of the five to nine fields a flux calculation filters.hw/wcell— the per-(band, row)support half-width and, on a uniform ascending axis, the constant cell half-width it implies.hwcalls the geometry'smetric_band, a transcendental on the sphere, andwcellis found by a linear walk, so both are worth tabulating.
invden is built for one mask strategy, so the apply asserts the strategy it is handed matches grid_plan.strategy. Build the plan with the strategy you intend to apply with.
CoarseGrainingEnergyFluxes.Filtering.RealSpace — Type
Physical-space direct-sum convolution (works on any grid, mask, and geometry).
CoarseGrainingEnergyFluxes.Filtering.ScatteredCache — Type
ScatteredCache{T}The full per-target-point neighbour list for a ScatteredFilterPlan (absolute neighbour indices + weights), built only when the plan's AbstractCacheStrategy calls for it. Same CSR-style layout as before: target t = i + (j-1)*Nx, entries ptr[t]:ptr[t+1]-1.
CoarseGrainingEnergyFluxes.Filtering.ScatteredFilterPlan — Type
ScatteredFilterPlan{T,K}Real-space footprint for a genuinely nonuniform 2D StructuredGrid axis, or a CurvilinearGrid (exactly the periodic_x = periodic_y = false case of the same candidate-window/distance-gate computation). FilterFootprint is a translation-invariant cache — the SAME index offset (and its weight) is reused for every target point — which is only valid on a uniform axis; here that assumption is false (offset +3 means a different physical displacement depending on where you start), so there is no way to share one offset/weight set across points. That does NOT mean the result must be stored, though: since the search window (di_lim/dj_lim) is already a global scalar bound (not per-point), the exact same candidate enumeration + distance/kernel_weight gate that determines a point's neighbours can be re-run identically at apply time from these few scalars alone — cache holds the materialized ScatteredCache only when the plan's cache strategy decided to build it (see AbstractCacheStrategy), nothing otherwise (apply-time recomputation).
CoarseGrainingEnergyFluxes.Filtering.SeparableFootprint — Type
SeparableFootprint{T}Precomputed 1D Gaussian weight vectors (gx,gy) plus preallocated scratch buffers for the row-pass/column-pass separable convolution. invrenorm (Deformable masking only) is the precomputed reciprocal local kernel mass over active cells — the SAME separable machinery run once, at plan-build time, on Float.(mask) instead of field, mirroring FFTWFilterPlan/SHTFilterPlan's established invrenorm pattern. Nx_profile/Ny_profile (ZeroFill masking) are the mask-INDEPENDENT denominator profiles: Σ w over valid (in-bounds/periodic) offsets is itself separable into one Nx-length and one Ny-length vector, since which offsets are valid depends only on i (resp. j) and periodicity, never on the mask.
Note: unlike the disk-truncated (d <= rad) non-separable footprint, this truncates each axis independently at the SAME per-axis rad (the Gaussian's own 1D marginal decays at the identical rate kernel_radius was derived from) — a square window, not a disk. The Gaussian has no true hard support (only a numerical truncation tolerance), so this is an equally valid truncation, just a different shape — matched against RealSpace within a measured tolerance, not asserted bit-identical.
CoarseGrainingEnergyFluxes.Filtering.SeparableFootprintND — Type
SeparableFootprintND{N,T}A separable kernel in N dimensions: one weight table per axis, applied as N successive 1-D passes.
This is the same factorization the 2-D path uses, and the reason it matters grows with N. A FilterFootprintND enumerates the whole ∏(2wᵈ+1) box per point; N passes cost ∑(2wᵈ+1). At w = 20 in 3-D that is 68,921 multiply-adds per point against 123.
Weight tables follow _sepw: a vector where the axis is uniform, an Nᵈ × (2wᵈ+1) matrix where it is stretched.
CoarseGrainingEnergyFluxes.Filtering.SeparableScratch — Type
SeparableScratch{T,MT} <: AbstractFilterScratchThe separable engine's two pass buffers: the masked input, and the intermediate the row pass writes and the column pass reads.
Every table the engine holds — the per-axis tap weights, the denominator profiles, invrenorm — is a function of the filter scale, so these buffers are the whole of what the scales of a sweep can share. They are sized by the grid alone, hence one set per sweep rather than one per scale.
One set serves one apply at a time, so a driver running applies concurrently needs one per worker.
CoarseGrainingEnergyFluxes.Filtering.Spectral — Type
Spectral <: AbstractFilterMethodTransform-space filtering (kernel applied as a multiply on the transformed field). Requires a spectral extension and a compatible grid (e.g. using FFTW for a uniform, periodic Cartesian grid).
CoarseGrainingEnergyFluxes.Filtering.ZeroFill — Type
ZeroFill <: AbstractMaskStrategyExcluded cells are treated as zero-valued: they contribute to the denominator (kernel weight) but zero to the numerator. The kernel is homogeneous (same shape everywhere), which preserves domain averages and commutation with derivatives (the Storer 2022 / Aluie 2019 "fixed kernel" mode).
CoarseGrainingEnergyFluxes.Filtering._banded_build_invden — Method
_banded_build_invden(grid, di, dj, w, ptr, nbands, strategy, periodic_x, periodic_y) -> invdenThe reciprocal window mass, accumulated once per plan by running _banded_row_accumulate! over a mass field:
ZeroFillcounts every in-support tap whether or not its source cell is active, so its mass field is1everywhere and the result varies only where a domain edge truncates the window.Deformabledrops inactive cells from both sums, so its mass field is the mask.
The target-activity test and the degeneracy floor are folded in here too, so the apply's epilogue is one multiply with no branch.
CoarseGrainingEnergyFluxes.Filtering._banded_row_accumulate! — Method
_banded_row_accumulate!(oc, src, di, dj, w, lo, hi, periodic_x, periodic_y, Nx, Ny, j) -> ocAccumulate one output row's weighted tap sum, contiguously along x.
This is the whole banded engine. It is used twice with different inputs: on mask · field to build the numerator at apply time, and on the mass field to build invden at plan build. Holding the tap index in the OUTER loop is what makes the inner loop a unit-stride axpy — accumulating one output point at a time would make it a floating-point reduction, which cannot be reassociated and so never vectorizes.
On a bounded axis a tap contributes only where its source index is in range. Clamping the RANGE, rather than branching inside the loop, keeps that run contiguous — which is what lets the columns within one filter radius of an edge take the same vectorized path as the interior. At a wide kernel those columns are most of every row, so they are not an edge case worth handling separately.
CoarseGrainingEnergyFluxes.Filtering._banded_source — Function
_banded_source(fp, field) -> srcThe array the tap loop actually convolves: mask · field on a masked grid, and the caller's array itself when there is nothing masked out, since the copy would then be pure cost.
CoarseGrainingEnergyFluxes.Filtering._build_footprint_curvilinear — Method
_build_footprint_curvilinear(grid, kernel, scale; cache_strategy=AutoCache(), cache_byte_budget=DEFAULT_CACHE_BYTE_BUDGET) -> ScatteredFilterPlanCompact plan (and, if cache_strategy calls for it, the full per-point cache) for a FlowGeometries.Grids.CurvilinearGrid. Enumeration goes through the grid's own ball query, as it does for a StructuredGrid; what differs is the SIZE ESTIMATE the cache budget is checked against. Connectivity.metric_window bounds a window from per-axis spacing, and a curvilinear mesh has no separable axes to bound with, so the estimate comes from the smallest adjacent-node spacing in each index direction instead. It feeds no computation — only the AutoCache decision — so being loose costs a cache that would have fit, never a wrong answer.
CoarseGrainingEnergyFluxes.Filtering._build_footprint_scattered — Method
_build_footprint_scattered(grid, kernel, scale; cache_strategy=AutoCache(), cache_byte_budget=DEFAULT_CACHE_BYTE_BUDGET) -> ScatteredFilterPlanBuild the compact nonuniform-axis plan (O(1) scalar metadata) and, only if cache_strategy calls for it, the full per-point ScatteredCache — correct for any spacing pattern (Cartesian or spherical, uniform or not), since it never assumes translation invariance. The search window comes from _scattered_window_bounds, i.e. the grid's own Connectivity.metric_window; the exact d <= rad check still gates inclusion, so a loose bound only costs iterations, never a missed cell.
CoarseGrainingEnergyFluxes.Filtering._build_prefixsum_grid_plan — Method
_build_prefixsum_grid_plan(grid; mask_strategy = ZeroFill()) -> PrefixSumGridPlanBuild the scale-independent half of the prefix-sum top-hat engine. Nothing here reads the filter scale, so one call serves an entire sweep — see plan_filter_sweep.
CoarseGrainingEnergyFluxes.Filtering._build_prefixsum_tophat — Method
_build_prefixsum_tophat(grid, kernel::TopHatKernel, scale; mask_strategy=ZeroFill(), kwargs...)Build the exact O(N·dj_lim) prefix-sum top-hat plan for one scale. grid_plan/scratch let a sweep hand in the shared pieces; omitted, they are built here for a standalone single-scale plan. kwargs... absorbs (and ignores) cache_strategy/cache_byte_budget — there is no per-point neighbour list here to cache at all.
CoarseGrainingEnergyFluxes.Filtering._build_separable_footprint — Method
_build_separable_footprint(grid, kernel::GaussianKernel, scale; mask_strategy=ZeroFill(), kwargs...) -> SeparableFootprintBuild the per-axis weight tables, the mask-dependent normalization data for mask_strategy, and the preallocated scratch buffers for the separable path. kwargs... absorbs (and ignores) cache_strategy/cache_byte_budget — there is no per-point neighbour list to cache here at all, so those knobs (which only govern the scattered path) don't apply.
CoarseGrainingEnergyFluxes.Filtering._default_method — Method
_default_method(grid) -> AbstractFilterMethodWhich method plan_filter takes on grid when the caller names none. The sweep resolves it the same way before asking whether the engine has a shared grid plan.
RealSpace() wherever an engine for it exists, on every grid architecture alike. Coarse graining is defined as a convolution against a compact kernel (Aluie 2019), so the filter at a point reads only its own neighbourhood, which survives a regional domain and a coastline; ZeroFill keeps that kernel position-independent, so filtering commutes with ∇ and the Π budget closes.
A transform reproduces the same convolution only on a global, unmasked domain and only up to its bandlimit, so it is reached by an explicit method = Spectral(), or by AutoMethod() choosing an evaluator on capability. The grid architecture never changes which operator ran.
A layout with no real-space engine gets Spectral(), whose builder then names the backends that exist.
CoarseGrainingEnergyFluxes.Filtering._is_flat_cell — Method
_is_flat_cell(grid) -> BoolWhether grid names a cell by a single integer — FlowGeometries.Grids.FlatCells().
The node CSR engine and everything built on it are written against exactly that: fold_within(grid, t) with one index, measure(grid, j), isactive(grid, t). A node set carries the trait, and so does every spherical pixelization whose cells carry one id: ring, cubed-sphere, healpix, icosahedral and Yin–Yang layouts. The trait is asked at dispatch, so a layout upstream adds is served the day it declares itself flat-celled.
The counterpart is CartesianCells(), whose cells are an index tuple; those grids take the rectilinear and curvilinear engines above.
CoarseGrainingEnergyFluxes.Filtering._nd_scattered_window_bounds — Method
_nd_scattered_window_bounds(grid, rad) -> (lim, periodic, period, is_cartesian)Compact per-axis scalar window-bound derivation for the nonuniform N-D (1D/3D) case — the same values an NDScatteredFilterPlan stores. The window itself is Connectivity.metric_window, read from the grid's cached axis statistics in O(1); nothing here scans the grid.
CoarseGrainingEnergyFluxes.Filtering._prefixsum3d_fill_plane! — Method
_prefixsum3d_fill_plane!(P, src, grid, masked, j, k) -> nothingCumulative sum along axis 1 of src for one (j, k) plane, with P[1, j, k] = 0 as the empty prefix. Writes only that plane's column, so planes are mutually independent and the threaded driver fills them concurrently.
CoarseGrainingEnergyFluxes.Filtering._prefixsum_batch_fuses — Method
_prefixsum_batch_fuses(fp) -> BoolWhether a batched apply should share one support walk across the fields, or simply run them one at a time.
On a uniform axis the support interval comes from wcell in O(1), so there is no walk to amortize and fusing is pure cost: the batch must hold K numerator scans live where the per-field form works through one at a time, and the larger working set is what decides it.
On a nonuniform axis the two-pointer walk dominates the apply and is identical for every field, so one walk feeding K accumulators is strictly less work.
CoarseGrainingEnergyFluxes.Filtering._prefixsum_build_invden — Method
_prefixsum_build_invden(grid, gp, dj_lim, hw, wcell) -> invdenAccumulate the window mass once per scale and store its reciprocal.
The mass depends on the grid, the mask, the mask strategy and ℓ — never on the field — so it belongs here rather than in the apply, where accumulating it would double the inner loop's arithmetic and be repeated for every field. Folding isactive and the degeneracy floor into the same table turns the apply's epilogue from a branch and a division into one multiply.
CoarseGrainingEnergyFluxes.Filtering._prefixsum_numerators! — Method
_prefixsum_numerators!(sc, K) -> the first `K` numerator slotsAllocates any slot that does not exist yet, by similar on slot 1 so a new buffer inherits whatever array type the engine is actually running on rather than assuming a host Matrix. Row 1 of each scan is the empty-prefix zero and is never written by the fill, so a fresh slot must be zeroed.
Growth happens once per (scratch, batch width) — compute_Π! asks for 2 and then 3 — so it is off the repeated-apply path and the allocation gates still see zero.
CoarseGrainingEnergyFluxes.Filtering._prefixsum_scratch — Method
_prefixsum_scratch(gp::PrefixSumGridPlan, grid) -> PrefixSumScratchAllocate the engine's per-apply numerator scan. Sized by the grid alone, so one scratch serves every scale — but only one apply at a time; see AbstractFilterScratch.
CoarseGrainingEnergyFluxes.Filtering._rectilinear_measure_factors — Method
_rectilinear_measure_factors(grid) -> (wx, wy)The grid's own per-axis measure factors, so that measure[i, j] == wx[i] * wy[j] exactly. Read from the stored SeparableMeasure rather than rebuilt here: rebuilding has to reproduce every convention the measure was constructed under, and a degenerate direction is where that fails — a single-latitude grid measures arc length R·Δλ and a single-longitude one R·Δφ, neither of which is the R²cosφ area form.
CoarseGrainingEnergyFluxes.Filtering._scattered_window_bounds — Method
_scattered_window_bounds(grid::StructuredGrid{...,2}, rad) -> (di_lim, dj_lim, periodic_x, periodic_y, x_period, y_period, is_cartesian)The widest window the grid's own ball query will scan, taken from Connectivity.metric_window rather than re-derived here. On a rectilinear grid the window depends on the row, not the column, so one evaluation per row covers the grid.
CoarseGrainingEnergyFluxes.Filtering._sep_serial — Method
_sep_serial(f, indices)The default pass driver: apply f to every index in order. Every output point of a pass is independent, so a backend can substitute a parallel driver of the same shape and get a bit-identical result — only the barrier BETWEEN passes is required.
CoarseGrainingEnergyFluxes.Filtering._sepw — Method
_sepw(g, i, k) -> TWeight of stencil slot k at axis position i.
The Gaussian factorizes on ANY rectilinear grid — exp(-α(Δx²+Δy²)/ℓ²) = Gx(Δx)·Gy(Δy) needs no constant spacing — but on a uniform axis Gx depends on the OFFSET alone, while on a stretched one it depends on the position too. Both are the same convolution with a different weight table, so the two are one code path distinguished by the table's rank: a vector is shared across positions and a matrix is (2·lim+1) × N, column-major so each position's stencil is contiguous.
CoarseGrainingEnergyFluxes.Filtering.analyze_buffer — Function
analyze_buffer(plan, field) -> F̂ or nothing
filter_analyze!(F̂, field, plan) -> F̂
filter_synthesize!(out, F̂, plan) -> outSplit of a spectral apply into its two halves. A spectral filter is forward transform → multiply by Ĝ(|k|, ℓ) → inverse transform, and only the multiply depends on the scale, so a sweep over S scales needs the forward ONCE per field rather than once per (field, scale): 5 + 5S transforms instead of 10S.
analyze_buffer returns nothing for an engine with no transform to share — a real-space filter does all its work per scale — and callers fall back to filter_apply!.
Plans for different scales over the same grid share a forward transform, so F̂ produced with any one of them may be synthesized with any other.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint! — Method
apply_footprint!(out, field, grid, fp::PrefixSumTopHatPlan, strategy, periodic_x, periodic_y) -> outapply_footprint!-shaped entry point for the prefix-sum plan, so the generic whole-grid convolve name works uniformly across every footprint type. periodic_x/periodic_y are accepted for interface compatibility but IGNORED: unlike the offset-based footprints, this plan captured its periodicity (and the matching extended-coordinate layout) from the grid at build time, and cannot honour a different choice at apply time.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint! — Method
apply_footprint!(out, field, grid, fp::ScatteredFilterPlan, strategy, periodic_x, periodic_y)Whole-grid convolve using a ScatteredFilterPlan (the nonuniform-axis/curvilinear fallback). periodic_x/periodic_y are accepted only for a uniform call signature with the FilterFootprint method above — periodicity for this footprint kind lives in fp itself, not these arguments.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint! — Method
apply_footprint!(out, field, grid, fp, strategy, periodic_x, periodic_y)Convolve field with a precomputed fp into out, applying the mask strategy. out and field are 2D (a single layer). The masking branch specializes on the strategy type.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint! — Method
apply_footprint!(out, field, grid, fp::NodeFilterPlan, strategy) -> outWeighted mean over each node's stored neighbourhood, with the same two mask conventions the structured engines use: ZeroFill keeps a masked neighbour in the denominator and contributes nothing for it, Deformable drops it from both.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint_nd_batch! — Method
apply_footprint_nd_batch!(outs, fields, grid, fp, strategy) -> outsBatched point-indexed apply over the whole grid. The batch shares each point's neighbour enumeration across all K fields, so the enumeration is paid once rather than K times.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint_nd_batch_over! — Method
apply_footprint_nd_batch_over!(outs, fields, grid, fp, strategy, indices) -> outsThe batched apply restricted to indices. Each output point depends only on its own neighbourhood, so a parallel backend can hand disjoint index blocks to different tasks and get the serial answer. outs must already be zeroed — the caller owns that, since a block only writes its own points.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint_row! — Method
apply_footprint_row!(out, field, grid, fp::ScatteredFilterPlan, strategy, periodic_x, periodic_y, j)Fill output row j from a ScatteredFilterPlan: if fp.cache !== nothing, read the precomputed per-point neighbour list (absolute (ii,jj) indices, periodic wrap already resolved at build time); otherwise recompute each point's neighbours/weights on the fly from fp's compact scalar metadata. Both branches enumerate candidates through _scattered_foldl, so they are bit-identical by construction rather than by convention. The accumulator is threaded through the fold's return value rather than captured and mutated, which is what keeps the streaming branch free of per-iteration allocation (verified by @allocated tests) without a second copy of the loop.
CoarseGrainingEnergyFluxes.Filtering.apply_footprint_row_batch! — Method
apply_footprint_row_batch!(outs, fields, grid, fp, strategy, periodic_x, periodic_y, j) -> outsBatched banded row apply: one contiguous axpy per field per tap, every field sharing the plan's invden.
The split is exactly along what depends on the field. The numerator does, so it is computed once per field. The window mass does not — it is a function of the geometry, the mask and the strategy — so it is precomputed per scale and read here, which is also what keeps each field's inner loop a vectorizable axpy rather than a per-point reduction.
CoarseGrainingEnergyFluxes.Filtering.apply_prefixsum_tophat! — Method
apply_prefixsum_tophat!(out, field, grid, fp, strategy) -> outWhole-grid exact top-hat convolve: one O(N) prefix-sum pass over the field, then one O(N·dj_lim) two-pointer sweep.
CoarseGrainingEnergyFluxes.Filtering.apply_prefixsum_tophat_batch! — Method
apply_prefixsum_tophat_batch!(outs, fields, grid, fp, strategy) -> outsWhole-grid batched top-hat convolve: one prefix scan per field, then a single sweep that walks each support interval once and feeds every field from it.
CoarseGrainingEnergyFluxes.Filtering.apply_prefixsum_tophat_batch_row! — Method
apply_prefixsum_tophat_batch_row!(outs, grid, fp, strategy, j) -> outsFill row j of every output in a batch from the (already current) per-field prefix scans, walking the support interval once for the whole batch.
That walk is the only thing a batch here can share, and it is shared only where it exists — see _prefixsum_batch_fuses:
- Uniform axis: every interval comes from
wcellinO(1), so there is nothing to amortize, and hoisting the band loop above the field loop would re-stream each output row once per band instead of keeping it resident across all of them. Those fields run the ordinary per-field row body. - Nonuniform axis: the two-pointer walk dominates, so the band loop goes outermost and every field is fed from one walk.
lo/histay in registers — no interval table is materialized.
CoarseGrainingEnergyFluxes.Filtering.apply_prefixsum_tophat_row! — Method
apply_prefixsum_tophat_row!(out, grid, fp, strategy, j) -> outFill output row j from the (already current) prefix sums. For each row offset in the band, sweeps axis 1 with two MONOTONE pointers — O(1) amortized per point, not a per-point binary search — so the whole row costs O(Nx·dj_lim). Touches only row j of out, so rows may run concurrently.
This sums the NUMERATOR only. The window mass and its reciprocal are tabulated per scale in fp.invden (they depend on the grid, mask, strategy and ℓ, never on the field), which halves the inner loop and turns the epilogue into a multiply — and the saving repeats for every one of the five to nine fields a flux calculation filters at each scale.
CoarseGrainingEnergyFluxes.Filtering.apply_separable! — Method
apply_separable!(out, field, grid, fp::SeparableFootprint, strategy) -> outApply the separable fast path: masked_input = mask .* field (the SAME numerator input for both mask strategies — see the struct docstring), one shared row-pass/column-pass convolution, then divide by whichever mask-strategy-specific denominator fp holds.
CoarseGrainingEnergyFluxes.Filtering.apply_slice_serial! — Method
apply_slice_serial!(out, field, plan) -> outOne slice, forced down the serial engine regardless of the backend recorded in plan. The slice loop owns the parallelism; see filter_slices!.
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid, kernel, scale; kwargs...) -> NodeFilterPlanReal-space footprint for a flat-cell grid — see _is_flat_cell — from Connectivity.fold_within, the grid's own metric ball query. The neighbourhood therefore honours the geometry's distance and any periodic wrap as the grid defines them, and the fold hands back each neighbour's distance for the weight to use directly.
A node set and a spherical pixelization present the same interface here: one integer per cell, a measure per cell, and a ball query. So one engine serves both, and nothing in it reads a coordinate directly.
The sweep visits every cell, which is where a spatial index pays for itself: Connectivity.default_sweep_topology builds one when NearestNeighbors is loaded, taking the build from O(n²) to O(n log n), and returns the unindexed topology (same rows, linear scan) when it is not. Either way it is paid once per plan and reused by every filter_apply!, including the six to nine a single compute_Π! makes.
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid, kernel, scale; kwargs...) -> ScatteredFilterPlanGeneral path: at least one axis is a plain (non-Range) AbstractVector, which makes no type-level uniformity guarantee — its values might happen to be evenly spaced, but nothing proves it, so no assumption is made and the always-correct per-point plan is built instead. (Less specific than the method above, so Julia only reaches this one when the fast method's constraint doesn't match.)
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid::StructuredGrid{T,<:AbstractGeometry,2}, kernel::TopHatKernel, scale; kwargs...) -> PrefixSumTopHatPlanExact prefix-sum top-hat path for any rectilinear 2D StructuredGrid (see the section comment above). More specific than the generic 2D methods (constrained on kernel type), so Julia selects it whenever kernel isa TopHatKernel.
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid::StructuredGrid{T,Cartesian,2}, kernel::GaussianKernel, scale; kwargs...) -> SeparableFootprintSeparability does not require constant spacing: exp(-α(Δx²+Δy²)/ℓ²) factorizes on any rectilinear grid, and a stretched axis only makes the per-axis weight depend on position as well as offset — see _separable_axis_weights. So a stretched Cartesian grid gets the same two-pass O(N·(wx+wy)) convolution rather than falling to the O(N·wx·wy) scattered engine, which for a Gaussian at w = 20 is a factor (2w+1)/2 in operations and a much larger one in per-operation cost.
The Range-axis method above is strictly more specific and resolves what would otherwise be an ambiguity with the generic Range-axis method; both build the same footprint.
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid, kernel, scale) -> FilterFootprintFast path — real multiple dispatch, not a runtime check: both axes are AbstractRange, a compile-time proof of constant spacing, so the footprint is genuinely translation-invariant and can be shared via a single (Cartesian) or per-latitude-band (spherical) offset/weight cache. Spacing is read via step(...) directly from the axis that's already proven uniform by its type — not from the geometry's separately-stored dx/dy scalar, so there's no possibility of the two disagreeing.
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid::StructuredGrid{T,Cartesian,2,S,TP,<:Tuple{AbstractRange,AbstractRange}}, kernel::GaussianKernel, scale; kwargs...) -> SeparableFootprintFast path for a GaussianKernel on a uniform (Range-axis) Cartesian grid — see the "Separable Gaussian fast path" section above. More specific than the generic Range-axis method (constrained on kernel type too), so Julia picks this one whenever kernel isa GaussianKernel.
CoarseGrainingEnergyFluxes.Filtering.build_footprint — Method
build_footprint(grid::CurvilinearGrid, kernel, scale; kwargs...) -> ScatteredFilterPlanReal-space direct-sum footprint for a curvilinear grid (see _build_footprint_curvilinear).
CoarseGrainingEnergyFluxes.Filtering.filter_apply! — Method
filter_apply!(out, field, plan) -> outApply a prebuilt plan_filter to a single 2D field, dispatching to whichever backend the plan was built for — the footprint is ALWAYS the one cached in plan, never rebuilt here, for every backend (serial, threaded, distributed, GPU, MPI).
CoarseGrainingEnergyFluxes.Filtering.filter_apply_batch! — Method
filter_apply_batch!(outs, fields, plan) -> outsApply one prebuilt plan to SEVERAL separate fields sharing its grid, writing outs[i] from fields[i]. Equivalent to calling filter_apply! per field and asserted bit-identical to it, but the engines that derive per-point geometry — the scattered/node footprints — derive each target point's neighbourhood once for the whole batch instead of once per field, which is where the saving is.
compute_Π! filters five to nine fields per scale, so this is the shape its inner loop uses.
Distinct from filter_slices!, which takes independent fields on DIFFERENT grids, one plan each; here there is one grid and one plan. For a single array carrying a trailing batch axis, this dispatches on to the fused-launch path.
CoarseGrainingEnergyFluxes.Filtering.filter_apply_batched! — Method
filter_apply_batched!(out, field, plan) -> outApply plan across a trailing batch axis: out and field are (spatial..., Nb) over the plan's grid. The filter gathers only along spatial axes, so slices are independent and the batch index is carried through untouched.
On a device this issues ONE launch of prod(spatial) * Nb work items instead of Nb launches of prod(spatial), which is what matters when a single slice does not fill the device — a 64² slice is 4k work items. On the host it is the same work as slicing and looping, and exists so callers have one shape-generic entry point rather than reimplementing the loop.
Differs from filter_apply_batch!, which takes several separate fields sharing a grid; here the batch is one contiguous array, which is what allows the fused launch.
CoarseGrainingEnergyFluxes.Filtering.filter_field! — Method
filter_field!(out, field, grid, kernel, scale; mask_strategy=ZeroFill(), filter_plan=nothing, backend=AutoBackend())Filter a field on a grid using kernel at characteristic full width scale (ℓ), writing the result to out (returned).
Keyword Arguments
mask_strategy::AbstractMaskStrategy=ZeroFill(): masking strategy —ZeroFill()(excluded cells count in the denominator as zero; the kernel stays homogeneous) orDeformable()(excluded cells dropped from numerator and denominator; renormalized over the locally-included area).Near a boundary the footprint is truncated, and both strategies inherit the same shape distortion from that: measured on a straight coast, Gaussian at
ℓ = 16cells, one cell inshore the footprint's centroid sits 0.21ℓ offshore and its width is 62% of the interior value (75% atℓ/4inshore, 90% atℓ/2, 100% atℓ). Points within≈ℓof a boundary are contaminated either way.They differ on the footprint's mass:
ZeroFillleaves it at the truncated value, so the kernel is position-independent and filtering commutes with spatial derivatives — the step the flux budget is derived by. A uniform field ≡ 1 then filters to 0.543 one cell from the coast, 0.948 atℓ/2, 0.9996 atℓ.Deformabledivides it out, so a constant filters to 1.000 everywhere, at the cost of a position-dependent kernel, which does not commute with derivatives.
ZeroFillis the default because the flux budget is a statement about commuting operators; preferDeformablewhen a locally unbiased amplitude near a coast matters more than a budget that closes.Neither conserves the ACTIVE-cell integral on a masked domain, and they fail differently.
ZeroFillconserves it over the whole domain exactly (7e-18 relative, unmasked periodic) but smears part of it onto masked cells, which report zero;Deformablerenormalizes that away and tracks the active-cell integral better — 4.8e-4 relative drift againstZeroFill's 1.1e-2 on a masked periodic grid. On a bounded grid the domain edge costs both about 1e-2 atℓ = 6Δx. A closed energy budget wants a periodic unmasked domain; otherwise expect anO(ℓ/L)boundary residual.filter_plan::Union{Nothing,AbstractFilterPlan}=nothing: a prebuiltplan_filterresult to reuse instead of building one from scratch — the zero-(re)allocation entry point for a repeated sweep (many timesteps/fields over the same grid/kernel/scale). When supplied,mask_strategy/backend/methodare ignored (already baked into the plan); build it once withplan_filterand pass it here on every subsequent call.backend::AbstractExecutionBackend=AutoBackend(): execution backend (SerialBackend, ThreadedBackend, GPUBackend, …). Ignored whenfilter_planis supplied.
For spherical grids the longitude footprint wraps only when the grid is periodic (isperiodic); distances use the great-circle (Haversine) metric.
Examples
geom = CartesianGeometry()
grid = StructuredGrid(geom, 0.0:1000.0:99_000.0, 0.0:1000.0:99_000.0, mask)
out = zeros(100, 100)
filter_field!(out, field, grid, TopHatKernel(), 5000.0; mask_strategy = Deformable())CoarseGrainingEnergyFluxes.Filtering.filter_fields! — Method
filter_fields!(outs, fields, grid, kernel, scale; mask_strategy=ZeroFill(), backend=AutoBackend())Filter several fields that share the same grid/kernel/scale, building the footprint/plan ONCE and applying it through filter_apply_batch!, so each target point's neighbours are enumerated once for the whole batch rather than once per field. outs and fields are indexable collections of matching arrays — a tuple of velocity components, or a vector of them.
CoarseGrainingEnergyFluxes.Filtering.filter_slices! — Method
filter_slices!(outs, fields, plans; backend = AutoBackend()) -> outsApply plans[t] to fields[t], writing outs[t], over a collection of independent slices.
This is a different parallel axis from filter_apply_batch!, which shares one grid across several fields: here each slice has its own grid, plan and point count, and slices share nothing, so there is no synchronization at all. Where a workload has many slices, this is the outermost race-free axis and the one that converts thread count into throughput — threading within one slice saturates once the slice is small enough that per-task overhead dominates its work.
Each slice runs serially inside, whatever backend its own plan carries: nesting a threaded apply under a threaded slice loop would have both levels claim the whole thread pool.
CoarseGrainingEnergyFluxes.Filtering.plan_filter — Method
plan_filter(grid, kernel, scale; mask_strategy=ZeroFill(), backend=AutoBackend()) -> AbstractFilterPlanBuild a reusable filter plan: the footprint is precomputed ONCE regardless of backend (serial, threaded, distributed, GPU, or MPI) and reused across every subsequent filter_apply! call — no backend rebuilds it per call. Apply with filter_apply!(out, field, plan).
CoarseGrainingEnergyFluxes.Filtering.plan_filter — Method
plan_filter(grid, kernel, scale; method = <the grid's own default>, …)Entry point for every grid architecture with no rectilinear or curvilinear engine of its own.
A flat-cell grid — see _is_flat_cell — has the node CSR real-space engine over its own ball query, and a spectral one wherever a transform targets it: the nonuniform-FFT backend for a Cartesian node set, the nonuniform spherical-harmonic one for a spherical node set. _default_method chooses between them per grid, and both plan in O(n log n).
RealSpace() applies the kernel as written, with compact support. A transform is exact for a band-limited field and its per-apply cost does not grow with the filter scale, at the price of global support. build_footprint raises for a layout with neither engine.
CoarseGrainingEnergyFluxes.Filtering.plan_filter_sweep — Method
plan_filter_sweep(grid, kernel, scales; kwargs...) -> FilterPlanFamilyPlan a whole sweep of scales at once, building the grid-determined half of the engine and the per-apply scratch once instead of once per scale.
This is what plan_filter cannot do on its own: given only one scale it has no way to know another is coming, so it must own its own copies. For the prefix-sum top-hat engine, the shared half is the measure prefix scans, the extended axis and the mask scan; a sweep of S scales over a 1024² grid holds one copy of each rather than S.
Accepts and forwards every plan_filter keyword.
family = plan_filter_sweep(grid, TopHatKernel(), scales)
for (s, plan) in zip(scales, family)
filter_apply!(out, field, plan)
endCoarseGrainingEnergyFluxes.Filtering.prefixsum_fill_numerator! — Method
prefixsum_fill_numerator!(fp, field, grid)The single O(N) per-apply pass: refill the scratch numerator scan (cumulative mask·field·wx per row) from the current field. Must run before any apply_prefixsum_tophat_row! call for that field — both the serial and threaded drivers guarantee this.
CoarseGrainingEnergyFluxes.Filtering.prefixsum_fill_numerator_batch_row! — Method
prefixsum_fill_numerator_batch_row!(fp, fields, grid, j)Fill row j's numerator scan for every field of a batch, one table each. This is the whole of the per-field work in a batched apply — measured at ~1% of it — which is why the batch shares the support walk rather than the scan.
CoarseGrainingEnergyFluxes.Filtering.prefixsum_fill_numerator_row! — Method
prefixsum_fill_numerator_row!(fp, field, grid, j)Fill row j's numerator prefix scan. Writes only column j of the scratch, so rows are mutually independent and this may be run concurrently across j (the threaded backend does exactly that).
CoarseGrainingEnergyFluxes.Filtering.prepare_row_apply! — Method
prepare_row_apply!(fp, field, grid) -> nothing
prepare_row_apply!(fp, fields, grid) -> nothingWhole-grid work an engine needs done ONCE before any of its per-row applies run.
Every parallel backend decomposes over rows and calls apply_footprint_row! / apply_footprint_row_batch! directly rather than going through the whole-grid entry point, so anything that is not per-row has to be hoisted here or it is silently skipped on those paths. The banded engine needs its mask · field source materialized; the other engines need nothing, and get the no-op fallback.
Call it before opening a parallel region, never inside one: it writes the whole buffer.
CoarseGrainingEnergyFluxes.Filtering.prepare_workspace — Method
prepare_workspace(backend, grid, footprint) -> workspaceBackend hook run ONCE by plan_filter, whose result becomes the plan's stored workspace. The default returns the footprint unchanged; a backend that needs its own residency — the GPU's device buffers — returns something its apply step consumes directly, so no transfer is repeated per call.
CoarseGrainingEnergyFluxes.Filtering.slice_costs — Method
slice_costs(plans) -> Vector{Int}Per-slice work proxy: the number of target points each plan writes. A slice's cost grows at least linearly in this, so it is what a longest-first schedule should sort on.
CoarseGrainingEnergyFluxes.Filtering.spectral_grid_plan — Method
spectral_grid_plan(spectral_backend, grid, kernel; mask_strategy, batch) -> AbstractGridPlan or nothingThe scale-independent half of a spectral plan: the transform objects themselves, the wavenumber grids, and the mask. Only the transfer function Ĝ(|k|, ℓ) and the Deformable renormalization depend on the filter scale, so a sweep builds this once and each scale keeps only those two.
Planning a transform is not cheap — FFTW measures, and a nonuniform transform additionally sorts its points — so paying it once per grid rather than once per scale is the whole reason this hook exists.
Returns nothing for a backend that has not been given one; the sweep then falls back to a full plan per scale, which is correct, just not shared.
CoarseGrainingEnergyFluxes.Filtering.spectral_scratch — Method
spectral_scratch(grid_plan) -> AbstractFilterScratch or nothingThe transient half of a spectral plan — the complex spectrum buffer, the mask · field staging array, and whatever else the apply overwrites — sized from the grid plan the transforms were built against.
Splitting it out is what makes spectral_grid_plan's result shareable: transforms and wavenumber grids are read-only during an apply, these buffers are not. A sweep builds one and hands it to every scale; a driver that applies concurrently builds one per worker.
Returns nothing for a backend that keeps no apply-time buffers.