Skip to main content

GPUFlagOffsets

Overview

GPUFlagOffsets turns one packed binary flag stream into exclusive dense indices and a GPU-resident count. It is the smallest useful materialization step after a classifier.

At a glance

QuestionAnswer
ProblemTurn one packed binary flag stream into stable dense indices and a count.
Reads / writesReads uint32 zero-or-one flags; writes exclusive uint32 offsets and one scalar count.
OwnershipFlags, offsets, and count are caller-owned; hierarchical scan scratch is graph-owned.
Output contractOne source-aligned exclusive offset per flag and an exact count modulo uint32.
Expected workOne hierarchical exclusive scan plus one scalar publication pass.
ChunksAtomic and vector views may have independent boundaries; only the flag-length destination prefix is written.
Conditions / budgetsContributes ordinary graph nodes and never compiles, submits, maps, or reads back.
Neighborhoodformat classifier → GPUFlagOffsets → compaction destinations, counts, or GPUSegmentOffsets.

When to use it

Use it when later GPU work needs a stable destination for every accepted slot, or when an indirect consumer needs the number of accepted slots. Typical uses include nullable-value layout, stable scatter, variable-length value offsets, and the logical-element stream at one nesting depth.

Use GPUScan directly if no total count is needed. Use GPUCompaction if the operation should also move uint32 values. Use GPUSegmentedLayout if value, element, and group layout are all needed at one depth. For nested layouts, compose one shared GPUFlagOffsets for leaf validity with one element GPUFlagOffsets and GPUSegmentOffsets per depth.

Contract

ViewLengthMeaning
flagsslot countPacked values that must be exactly zero or one
offsetsat least slot countExclusive prefix sum; the dense destination for each slot
countat least 1Total set flags in element zero

flags and offsets may independently be atomic views or GraphVectorViews with different partitions. The scan carries across chunks without repacking their buffers, and count covers the complete flag sequence. Extra offset capacity is left untouched. An empty input writes count[0] = 0. No offset element exists for an empty input. Counts and offsets use uint32 arithmetic, so callers must retain page or batch boundaries before overflow.

Usage

import {GPUCommandGraph, GPUFlagOffsets} from '@luma.gl/gpgpu/gpu-core';

const graph = new GPUCommandGraph(device, {id: 'nullable-column'});

graph.add(new GPUFlagOffsets({
id: 'present-values',
flags: validity,
offsets: valueOffsets,
count: nonNullValueCount
}));

The class contributes graph nodes only. It does not allocate public outputs, compile the graph, submit work, or map the count. A compiled graph can be reused when its view sizes and topology stay fixed; imported buffers may be rebound for each encoding.

Costs and mistakes

The complete flag stream is scanned even when few values are present. Keep flags GPU-resident when several consumers reuse the result; a CPU round trip solely to obtain the count defeats the main benefit. Values greater than one are summed, not clamped, so normalize arbitrary predicates before calling the operation.