Execution

FlowGeometries.Execution._reduce_chunks — Method
_reduce_chunks(f, op, n, backend)

f(range) over a partition of 1:n, combined left to right with op.

Distinct from map_chunks because the serial path has nothing to collect: it hands f the whole range and returns that value directly. map_chunks builds a one-element Vector there, which the compiler usually erases; this path keeps a serial reduction allocation-free unconditionally.

source
FlowGeometries.Execution.allocate — Method
allocate(backend, T, n) -> AbstractVector{T}

An uninitialised length-n vector the loops running under backend can write.

A CSR build allocates its degree, offset and neighbour arrays through this, so the buffers land where the passes that fill them run: an ordinary Vector under nothing and the host backends, device memory under a backend that launches kernels.

source
FlowGeometries.Execution.chunk_ranges — Method
chunk_ranges(n, k) -> Vector{UnitRange{Int}}

Partition 1:n into at most k contiguous ranges of near-equal length. Contiguity matters: each chunk then touches one span of every array the loop indexes.

source
FlowGeometries.Execution.exclusive_scan! — Function
exclusive_scan!(out, counts, backend = nothing; init = 1) -> out

Write the exclusive prefix sums of counts into out, which is one element longer: out[1] = init and out[i+1] = out[i] + counts[i].

This is a CSR offset array — out[k] is where row k starts and out[end] - init is the total — and every connectivity builder here calls it between its counting pass and its filling pass. Under a threaded backend it is two passes over counts plus a serial scan of one sum per chunk.

source
FlowGeometries.Execution.map_chunks — Method
map_chunks(f, n, backend) -> Vector

Apply f(range) over a partition of 1:n and collect one result per chunk, for a bulk operation that reduces. The caller combines them, so f needs no lock and the combining order is the caller's to fix, which keeps a threaded reduction independent of scheduling for an operation that is associative but not commutative.

f is called at least once, on an empty range when n == 0, since a reduction has to produce a value; run_chunks has nothing to write there and does not call f. Being callable on an empty range is part of the contract, and it is how the result's element type is known before the chunks run.

source
FlowGeometries.Execution.reduce_indices — Method
reduce_indices(f, op, init, n, backend)

Reduce f(i) over i in 1:n with op, under the execution policy backend names.

The per-index counterpart of _reduce_chunks, and the reduction a device can run: f reads index i and nothing else, so the work splits without a chunk body. Every integral, norm and count over a grid goes through this.

op must be associative and init its identity. init seeds each partial as well as the whole, so the partials combine in any grouping.

The grouping follows the partition, never the scheduling, so a threaded result is deterministic. It matches the serial left fold wherever op is exactly associative; floating-point + is not, so a threaded sum can differ from the serial one in its last bits.

source
FlowGeometries.Execution.run_chunks — Method
run_chunks(f, n, backend)

Apply f(range) over a partition of 1:n, under the execution policy backend names.

f must be safe to run on disjoint index ranges concurrently — every write it makes has to be determined by the index, never accumulated across chunks. nothing means serial and hands f the whole range in one call, so the serial path adds no partitioning at all.

source
FlowGeometries.Execution.run_indices — Method
run_indices(f, n, backend)

Apply f(i) for each i in 1:n, under the execution policy backend names. f must write only what i determines, with no accumulation across indices — the same contract as run_chunks, stated per index, the form a device launch expresses.

This is the form a kernel maps onto: with KernelAbstractions loaded and a device backend, f becomes the body of a launch over 1:n.

source