Module: Thin Instances
Package path:
packages/babylon-lite/src/mesh/(data + GPU sync + GPU culling),packages/babylon-lite/src/material/standard/,packages/babylon-lite/src/material/pbr/,packages/babylon-lite/src/material/node/, andpackages/babylon-lite/src/material/shader/(rendering)
Purpose
Thin instances allow a single mesh to be drawn thousands of times with unique per-instance world matrices and optional per-instance RGBA colors, using a single instanced draw call. This is the primary mechanism for rendering large crowds, particle-like effects, and procedural grids. The system is split into three layers — CPU data model, GPU buffer sync, and material integration — designed so that scenes that don't use thin instances pay zero bundle-size cost.
Thin instances are supported by all four mesh material families: Standard, PBR, NodeMaterial, and ShaderMaterial (custom user-WGSL). For ShaderMaterial integration specifics (auto-injected world0..world3 / instanceColor vertex attributes and the user-shader contract) see 24-shader-material.md.
A ShaderMaterial draw may explicitly ignore an existing instance-color stream with
useThinInstanceColors: false. This is draw-local: the mesh keeps its colors for other materials, while the
selected material binds only the instance matrices. The canonical use is a color-independent depth override
sharing the visible mesh's matrices without paying for unused color synchronization or vertex fetch.
Public API Surface
Interfaces
/** Per-mesh thin-instance state. Stored as mesh.thinInstances. */export interface ThinInstanceData { matrices: Float32Array; // 16 floats per instance (column-major 4x4 world matrix) count: number; // active instance count _capacity: number; // allocated capacity (≥ count) _version: number; // bumped by every mutating helper; checked by render system _gpuBuffer: GPUBuffer | null; // matrix GPU buffer, managed by render system _gpuVersion: number; // last _version uploaded to GPU colors?: Float32Array | null; // optional RGBA per instance (4 floats each) _colorVersion: number; // independent of _version; bumped by setThinInstanceColors _colorGpuBuffer: GPUBuffer | null; _colorGpuBufferStorage: boolean; _colorGpuVersion: number; _gpuCullingEnabled: boolean; // opt-in GPU frustum culling + indirect draw _lodPartner?: Mesh | null; // lower-detail mesh receiving the far cull bucket _lodSource?: Mesh | null; // source mesh when this mesh is the LOD partner}Functions — CPU Data Model (thin-instance.ts)
/** Bulk-set all instance matrices. Creates ThinInstanceData if absent. */export function setThinInstances(mesh: Mesh, matrices: Float32Array, count: number): void;
/** Change the active count and mark the complete active matrix range dirty. */export function setThinInstanceCount(mesh: Mesh, count: number): void;
/** Change the active draw count without marking matrix or color data dirty. */export function setThinInstanceDrawCount(mesh: Mesh, count: number): void;
/** Pre-create stable indirect draw arguments during scene warm-up. */export function enableThinInstanceDynamicDrawCount(mesh: Mesh): void;
/** Add one instance. Returns the new instance index. Grows capacity 2× when full. */export function addThinInstance(mesh: Mesh, matrix: Mat4): number;
/** Overwrite the matrix at a specific index. */export function setThinInstanceMatrix(mesh: Mesh, index: number, matrix: Mat4): void;
/** Remove instance at index via swap-remove (last instance moves into the gap). */export function removeThinInstance(mesh: Mesh, index: number): void;
/** Bump _version after direct manipulation of the matrices Float32Array. */export function flushThinInstances(mesh: Mesh): void;
/** Include exact thin-instance placements in default-camera and environment auto-sizing. */export function enableThinInstanceWorldBounds(mesh: Mesh): void;
/** Set per-instance RGBA colors. Bumps _colorVersion. */export function setThinInstanceColors(mesh: Mesh, colors: Float32Array): void;
/** Enable/disable per-pass GPU frustum culling. Must be called before registerScene(). */export function enableThinInstanceGpuCulling(mesh: Mesh, enabled?: boolean): void;
interface ThinInstanceLodPartnerOptions { distance: number; // camera-space switch distance in world units band?: number; // deterministic per-instance dither width, default 0}
/** Split one GPU-culling result into near/full-detail and far/LOD buckets. */export function setThinInstanceLodPartner(fullMesh: Mesh, lodMesh: Mesh, options: ThinInstanceLodPartnerOptions): void;
/** Restore the two meshes to independent rendering. */export function clearThinInstanceLodPartner(fullMesh: Mesh): void;setThinInstanceDrawCount() is the count-only path for fixed-capacity pools whose CPU array and GPU
buffer are already populated. It accepts an integer in [0, _capacity], updates only count, and leaves
matrix/color versions and dirty ranges untouched. Cached draws observe the count through their stable
indirect argument buffer, so changing the active prefix neither replaces instance buffers nor invalidates
render bundles. A caller exposing newly written slots must mark those exact matrix/color ranges dirty with
the corresponding update API; the count-only setter deliberately does not upload vertex data. The pool
must complete one full-capacity GPU synchronization before this setter is used.
setThinInstanceCount() remains the convenience path that changes the count and marks the complete active
matrix range [0, count) dirty for upload.
Call enableThinInstanceDynamicDrawCount() before registerScene() when a synchronized pool will change
counts interactively. Its next normal GPU sync creates the stable indirect argument buffer during warm-up,
so the first later count change does not invalidate cached render bundles.
All thin-instance material families, including Node Material main, transparent, and override/shadow bindings, synchronize the active count in their per-frame update. A binding initially recorded with a direct draw promotes to one stable indirect argument buffer when its count first changes; the visibility epoch then invalidates that one stale bundle. Subsequent transitions, including nonzero-to-zero and zero-to-nonzero changes within capacity, update only the indirect arguments and keep the matrix buffer and cached bundle stable.
The CPU indirect-argument array is also the index-count/base-vertex snapshot; these values
are not duplicated in parallel fields. The stored instance count acknowledges a successful
upload and is invalidated before writing, so a failed upload retries without allocating new
argument buffers. Base vertices are compared as signed 32-bit words.
Matrix buffers always carry STORAGE for GPU picking, so switching culling on/off does
not recreate or re-upload them. Color buffers still track their optional storage usage.
enableThinInstanceWorldBounds() is a setup-time, tree-shakable opt-in for hand-built thin-instance meshes
that will be consumed by createDefaultCamera() or automatic environment sizing. It expands the prototype's
object-local box through mesh.worldMatrix × instanceMatrix for every active instance and ignores parked
instances whose linear transform is effectively zero. Call it after setThinInstances() and before camera or
environment creation. The glTF EXT_mesh_gpu_instancing feature enables it automatically.
Functions — Hierarchy Instance Pools (hierarchy-instance-pool.ts)
For the old Babylon.js parentNode.instantiateHierarchy() prop workflow, Lite exposes a small opt-in helper that keeps the rendering path thin-instance based while preserving child mesh offsets/rotations/scales:
/** Build a fixed-capacity pool from a template hierarchy. Call before registerScene(). */export function createHierarchyInstancePool(root: SceneNode, capacity: number): HierarchyInstancePool;
/** Add one logical hierarchy instance and return its slot index. */export function addHierarchyInstance(pool: HierarchyInstancePool, matrix: Mat4): number;
/** Update one logical hierarchy instance root matrix. */export function setHierarchyInstanceMatrix(pool: HierarchyInstancePool, index: number, matrix: Mat4): void;
/** Remove one logical hierarchy instance via swap-remove. */export function removeHierarchyInstance(pool: HierarchyInstancePool, index: number): void;
/** Change active logical count without reallocating buffers. */export function setHierarchyInstanceCount(pool: HierarchyInstancePool, count: number): void;The instance matrix composes with the template hierarchy — it behaves like a parent transform node added above the root, so the final world of each descendant is matrix * meshWorld (stored per mesh as meshWorld⁻¹ * matrix * meshWorld, since the shader draws at finalWorld = mesh.world * instanceMatrix). The identity matrix therefore reproduces the template unchanged. The pool must not divide out the root's world matrix: a loadGltf() root carries the RH→LH conversion as scaling (-1, 1, 1), and cancelling it would mirror every instance and invert its winding.
createHierarchyInstancePool() walks all descendant meshes and assigns each one its own thin-instance matrix buffer at the requested capacity, then sets active count to zero. The source meshes therefore become render carriers for the pool: they do not draw the template by themselves, but they must stay visible !== false so their thin instances can draw. Do not hide a hierarchy pool with setSubtreeVisible(root, false); clear it with setHierarchyInstanceCount(pool, 0) instead. When growing, prefer addHierarchyInstance(pool, matrix) so the newly visible slot has a defined matrix before the next frame.
Functions — GPU Sync (thin-instance-gpu.ts)
/** * Sync CPU thin-instance data to GPU vertex buffers and bind them to the render pass. * Returns the next free vertex buffer slot. */export function syncThinInstanceBuffers( engine: EngineContextInternal, ti: ThinInstanceData, pass: GPURenderPassEncoder | GPURenderBundleEncoder, slot: number, hasColor: boolean, drawBuffers?: ThinInstanceDrawBuffers | null): number;Feature Flag Constants
// standard-pipeline.tsexport const THIN_INSTANCES = 1 << 15; // matrix instancingexport const THIN_INSTANCE_COLOR = 1 << 16; // per-instance color bufferexport const DISABLE_LIGHTING = 1 << 17; // skip light loop, output emissive × diffuse × baseColor
// mesh-features.tsexport const MSH_HAS_THIN_INSTANCES = 1 << 4;export const MSH_HAS_INSTANCE_COLOR = 1 << 5;Material Property
// StandardMaterialProps — extended with:disableLighting: boolean; // default false. When true, skip all lighting; output emissive * diffuse * baseColor.Internal Architecture
Data Flow
User code │ setThinInstances(mesh, matrices, count) │ setThinInstanceColors(mesh, colors) ▼thin-instance.ts → mesh.thinInstances: ThinInstanceData │ _version / _colorVersion bumped on every mutation ▼standardGroupBuilder (standard-material.ts) │ detects meshes.some(m => !!m.thinInstances) │ dynamic import('./thin-instance-gpu.js') ← lazy-loaded chunk │ passes syncThinInstanceBuffers as tiSync callback ▼buildStandardMeshRenderables (standard-renderable.ts) │ stores the version-gated GPU sync in the per-frame update closure │ binds the stable matrix/color buffers in the draw closure ▼Per-frame update + draw │ update: if mesh.thinInstances, upload dirty matrix/color ranges │ draw: if mesh.thinInstances and tiSync: │ slot = tiSync(device, ti, pass, slot, hasInstanceColor) │ pass.drawIndexed(indexCount, ti.count) ← instanced draw ▼GPU vertex shader │ world0..world3 → instanceWorld → finalWorld = mesh.world * instanceWorld │ instanceColor → vInstanceColor varying ▼GPU fragment shader │ lighting or disableLighting path │ final color.rgba *= vInstanceColor.rgba (if instance color)Capacity Growth & Swap-Remove
Capacity growth — addThinInstance starts at capacity 16. When count === _capacity, a new Float32Array is allocated at _capacity * 2 and the old data is copied:
const newCap = ti._capacity * 2;const newData = new Float32Array(newCap * 16);newData.set(ti.matrices);ti.matrices = newData;ti._capacity = newCap;Swap-remove — removeThinInstance copies the last instance matrix into the removed slot using copyWithin, then decrements count:
ti.matrices.copyWithin(index * 16, last * 16, last * 16 + 16);ti.count--;ti._version++;This avoids shifting the entire array, keeping removal O(1). Callers must be aware that the last instance's index changes.
Version Tracking
| Version field | Bumped by | Checked by |
|---|---|---|
_version | Matrix/count helpers plus color helpers (color-only changes leave the matrix dirty range empty but still dirty static shadows) | syncThinInstanceBuffers (matrix sync/version) |
_colorVersion | setThinInstanceColors | syncThinInstanceBuffers (color upload) |
_gpuVersion | syncThinInstanceBuffers (after matrix upload) | — |
_colorGpuVersion | syncThinInstanceBuffers (after color upload) | — |
GPU upload is skipped when _version === _gpuVersion (or _colorVersion === _colorGpuVersion), avoiding redundant writeBuffer calls for static instances.
setThinInstanceDrawCount does not mark either data stream dirty; draw-argument synchronization observes
count independently. It advances the existing thin-instance version without marking matrix/color ranges
dirty, so shadow caches observe the change while GPU synchronization performs no attribute upload. Static
ESM, PCF, and CSM maps therefore redraw only when one of their actual casters changes.
Opaque main-color and shadow draws are normally recorded into cached render bundles. Recording-time
setVertexBuffer calls retain the GPU buffer handle, not a snapshot of its contents, so runtime matrix
mutation must upload through that same buffer from the renderable's per-frame update closure before either
bundle executes. Uploading only from the draw closure is incorrect: that closure does not run again while a
cached bundle is reused. The update remains version-gated, and a same-capacity mutation must not recreate the
buffer or invalidate/re-record the bundle. Buffer growth or usage changes still follow the existing
visibility-epoch invalidation path because those operations replace the captured handle.
GPU Buffer Sync (thin-instance-gpu.ts)
Matrix Buffer
- Compare
ti._version !== ti._gpuVersion— skip if equal. - Compute
byteSize = ti.count * 64(16 floats × 4 bytes). - If
ti._gpuBufferis null orti._gpuBuffer.size < byteSize:- Destroy old buffer (if any).
- Create new buffer:
size = ti._capacity * 64,usage = VERTEX | COPY_DST, plusSTORAGEwhen GPU culling is enabled.
device.queue.writeBuffer(ti._gpuBuffer, 0, ti.matrices.buffer, ti.matrices.byteOffset, byteSize).- Set
ti._gpuVersion = ti._version. - Bind:
pass.setVertexBuffer(slot++, ti._gpuBuffer).
Color Buffer (conditional: hasColor && ti.colors)
- Compare
ti._colorVersion !== ti._colorGpuVersion— skip if equal. - Compute
byteSize = ti.count * 16(4 floats × 4 bytes). - If
ti._colorGpuBufferis null orti._colorGpuBuffer.size < byteSize:- Destroy old buffer (if any).
- Create new buffer:
size = ti._capacity * 16,usage = VERTEX | COPY_DST, plusSTORAGEwhen GPU culling is enabled.
device.queue.writeBuffer(ti._colorGpuBuffer, 0, ti.colors.buffer, ti.colors.byteOffset, byteSize).- Set
ti._colorGpuVersion = ti._colorVersion. - Bind:
pass.setVertexBuffer(slot++, ti._colorGpuBuffer).
Return Value
Returns the updated slot number — the next free vertex buffer slot after all thin-instance buffers have been bound.
Optional GPU Frustum Culling (thin-instance-gpu-culling.ts)
GPU culling is opt-in through enableThinInstanceGpuCulling(mesh). The helper only flips state on existing ThinInstanceData; the compute module is dynamically imported (via the shared thin-instance-cull-binding.ts lifecycle helper) by the Standard, PBR, and ShaderMaterial group builders only when at least one mesh in that material family has _gpuCullingEnabled === true.
The per-binding cull lifecycle is factored into one shared module, packages/babylon-lite/src/mesh/thin-instance-cull-binding.ts, used identically by all three material families. Its tryBind() seam is called from each renderable's bind(): it gates on opt-in + opaque-only, creates the per-binding ThinInstanceGpuCullState, registers its disposal, and returns a TiCullBinding whose update() dispatches the compute cull pass and whose draw() issues drawIndexedIndirect (or falls back to a normal instanced draw when culling did not run).
Scope
- Supported: opaque Standard, opaque PBR, and opaque ShaderMaterial thin instances.
- Excluded: transparent thin instances, transmissive PBR surfaces (
needsTaskRefraction), and arbitrary non-instanced meshes. - The helper must be called before
registerScene()so the material group builder can import the culling module and mark the renderable as direct-drawn.
Per-Binding State
Cull state is owned by each DrawBinding, not by ThinInstanceData, because the same mesh can be rendered by multiple render tasks/cameras. Each binding owns:
| Resource | Usage | Purpose |
| --------------------- | --------- | --------- | ---------------------------------------------------------------------------- | --------------------------------------------------------------------------- |
| source matrix buffer | VERTEX | COPY_DST | STORAGE | Full CPU-authored instance matrix list |
| source color buffer | VERTEX | COPY_DST | STORAGE | Full CPU-authored color list, when present |
| visible matrix buffer | VERTEX | STORAGE | Compacted visible matrices written by compute and read by the vertex shader |
| visible color buffer | VERTEX | STORAGE | Compacted visible colors, when present |
| indirect args buffer | INDIRECT | STORAGE | COPY_DST | [indexCount, visibleInstanceCount, firstIndex, baseVertex, firstInstance] |
| params uniform | UNIFORM | COPY_DST | Six frustum planes, mesh world matrix, local bounding sphere, instance count |
Per-Frame Flow
RenderTask.prepareRenderTaskPass() exposes the active camera and render-target size in DrawUpdateContext; the culling module uses the engine's current command encoder for its compute pass. The material binding update then:
- Syncs source matrices/colors to STORAGE-capable GPU buffers.
- Writes the indirect args buffer with
instanceCount = 0and constant draw fields. - Extracts six normalized world-space frustum planes from the active camera view-projection matrix.
- Uploads culling params, including
mesh.worldMatrixand a conservative local bounding sphere computed from_cpuPositions. - Dispatches one compute invocation per source instance.
- The compute shader transforms the local sphere by
mesh.world * instanceWorld, tests it against all planes, atomically appends visible instances to compacted buffers, and atomically incrementsargs[1]. - The draw closure binds compacted buffers and calls
drawIndexedIndirect(argsBuffer, 0).
Cull-enabled opaque renderables remain bundle-compatible because their compacted buffers and indirect argument buffers are stable. Any buffer replacement bumps the global visibility epoch, forcing the opaque render bundle to record the new handles before it executes.
Compute Shader Outline
let world = params.meshWorld * srcMatrices[i];let center = (world * vec4<f32>(params.localSphere.xyz, 1.0)).xyz;let radius = params.localSphere.w * max(length(world[0].xyz), length(world[1].xyz), length(world[2].xyz));
if (sphereIntersectsFrustum(center, radius)) { let outIndex = atomicAdd(&args[1], 1u); dstMatrices[outIndex] = srcMatrices[i]; dstColors[outIndex] = srcColors[i]; // color variant only}The test is conservative: spheres touching a plane stay visible. This preserves parity with non-culled rendering.
Distance LOD Pairing
setThinInstanceLodPartner(fullMesh, lodMesh, { distance, band }) extends the opt-in GPU culler with a second compacted output. Both meshes must already have thin-instance state and must be distinct, opaque, non-transmissive meshes. The full-detail mesh must be the only source for the partner, and neither mesh may simultaneously participate in another LOD chain. LOD-paired thin-instance state cannot be shared through cloneTransformNode; pair meshes only before cloning or use independent thin-instance data. distance and band are finite non-negative world-space values.
The source mesh retains in-frustum instances closer than the threshold. Farther instances are compacted into stable matrix/color buffers and an indirect argument buffer consumed by the partner. band applies a deterministic instance-index hash in [-band/2, +band/2], spreading transitions without frame-to-frame noise. The partner's geometry and material pipeline are its own, but its drawn matrices and optional instance colors come from the source. A partner that renders instance colors therefore requires the source draw to provide them too.
The far bucket is published per RenderTargetSignature, so main, shadow, and auxiliary camera passes never share compacted state. Partner renderables use the direct-draw phase and resolve the current bucket after all binding updates and the shared compute batch have completed; scene insertion order therefore cannot introduce a one-frame lag or capture stale cross-renderable handles in an opaque bundle. Source renderables retain their normal stable-buffer bundle path. If source culling is unavailable, disabled, hidden, or empty, the source falls back to drawing all active instances and the paired partner draws nothing.
Pairing is configured after setThinInstances() and before registerScene(). Repeating the call with the same pair updates distance/band live. clearThinInstanceLodPartner() makes the partner fall back to its independent instance draw immediately; pairing links are also detached automatically when either mesh is disposed so no draw retains destroyed far-bucket buffers.
Pipeline Configuration
Feature Flag Logic
computeFeatures(material, receiveShadows) sets:
if (material.disableLighting) f |= DISABLE_LIGHTING;The renderable builder sets:
if (mesh.thinInstances) features |= THIN_INSTANCES;if (mesh.thinInstances.colors) features |= THIN_INSTANCE_COLOR;Vertex Buffer Layouts
Base per-vertex attributes use sequential shaderLocation values starting at 0:
| Location | Attribute | Format | Stride | Step Mode |
|---|---|---|---|---|
| 0 | position | float32x3 | varies | vertex |
| 1 | normal | float32x3 | varies | vertex |
| 2 (opt) | uv | float32x2 | varies | vertex |
| next (opt) | uv2 | float32x2 | varies | vertex |
When THIN_INSTANCES is set, an instanced buffer layout is appended:
{ arrayStride: 64, // 4 × vec4<f32> = 4 × 16 bytes stepMode: 'instance', attributes: [ { shaderLocation: nextAttr++, offset: 0, format: 'float32x4' }, // world0 (row 0) { shaderLocation: nextAttr++, offset: 16, format: 'float32x4' }, // world1 (row 1) { shaderLocation: nextAttr++, offset: 32, format: 'float32x4' }, // world2 (row 2) { shaderLocation: nextAttr++, offset: 48, format: 'float32x4' }, // world3 (row 3) ],}When THIN_INSTANCE_COLOR is set, another instanced buffer layout is appended:
{ arrayStride: 16, // 1 × vec4<f32> stepMode: 'instance', attributes: [ { shaderLocation: nextAttr++, offset: 0, format: 'float32x4' }, // instanceColor (RGBA) ],}Shader Logic
Vertex Shader
Instance matrix attributes (when THIN_INSTANCES):
@location(N) world0: vec4<f32>,@location(N+1) world1: vec4<f32>,@location(N+2) world2: vec4<f32>,@location(N+3) world3: vec4<f32>,Instance color attribute (when THIN_INSTANCE_COLOR):
@location(M) instanceColor: vec4<f32>,World matrix composition:
let instanceWorld = mat4x4<f32>(world0, world1, world2, world3);let finalWorld = mesh.world * instanceWorld;finalWorld replaces mesh.world in all subsequent vertex transforms (position, normal).
ShaderMaterial exposes this specialization through the opt-in enableShaderMaterialInstanceWorld(material)
helper. A custom vertex shader that declares the "world" system uniform can then call
getFinalWorld(input) for both regular and thin-instanced meshes; the regular variant returns the mesh world
matrix without requiring world0..world3.
ShaderMaterial can similarly opt into enableShaderMaterialFinalColor(material). The generated
getFinalColor(input) returns white, mesh vertex color, thin-instance color, or their product according to
the active regular/instanced pipeline layout and the color attributes it binds.
Instance color passthrough (varying):
// Vertex output struct@location(K) vInstanceColor: vec4<f32>,
// Vertex mainout.vInstanceColor = instanceColor;Fragment Shader
Normal lighting path (DISABLE_LIGHTING not set):
Standard Blinn-Phong lighting loop (diffuse, specular, ambient, emissive, shadows). After composition, if THIN_INSTANCE_COLOR:
color = vec4<f32>( color.rgb * input.vInstanceColor.rgb, color.a * input.vInstanceColor.a);Disabled lighting path (DISABLE_LIGHTING set):
Skips: lighting function definitions, light loop, shadow sampling, ambient, reflection, lightmap. Emits:
var color = vec4<f32>( clamp(emissiveContrib * diffuseColor, vec3<f32>(0.0), vec3<f32>(1.0)) * baseColor.rgb, alpha);Then, if THIN_INSTANCE_COLOR, the same instance-color multiplication is applied:
color = vec4<f32>( color.rgb * input.vInstanceColor.rgb, color.a * input.vInstanceColor.a);Renderable Integration (standard-renderable.ts)
tiSync Callback Type
type ThinInstanceSync = ( engine: EngineContextInternal, ti: ThinInstanceData, pass: GPURenderPassEncoder | GPURenderBundleEncoder, slot: number, hasColor: boolean, drawBuffers?: ThinInstanceDrawBuffers | null) => number;Draw Function
buildStandardMeshRenderables accepts an optional tiSync callback. In the per-mesh draw closure:
const ti = mesh.thinInstances;if (ti && tiSync) { slot = tiSync(engine, ti, pass, slot, hasInstanceColor, cullResult?.drawBuffers ?? null); if (cullResult) { pass.drawIndexedIndirect(cullResult.argsBuffer, 0); } else { pass.drawIndexed(g.indexCount, ti.count); }} else { pass.drawIndexed(g.indexCount);}The regular path uses drawIndexed(indexCount, instanceCount) and draws all instances in a single GPU call. The cull path uses drawIndexedIndirect() and draws only the compute-compacted visible instances.
Dynamic Loading Architecture (standard-material.ts)
Group Builder
The standardGroupBuilder function detects thin instances at build time:
const hasTI = meshes.some((m) => !!m.thinInstances);const hasTICulling = meshes.some((m) => m.thinInstances?._gpuCullingEnabled === true);let tiSync;if (hasTI) { const mod = await import("../../mesh/thin-instance-gpu.js"); tiSync = mod.syncThinInstanceBuffers;}if (hasTICulling) { tiCull = await import("../../mesh/thin-instance-gpu-culling.js");}const { buildStandardMeshRenderables } = await import("./standard-renderable.js");return buildStandardMeshRenderables(scene, meshes, { tiSync, tiFragment, tiCull });This ensures thin-instance-gpu.ts is only fetched when a scene actually uses thin instances, and thin-instance-gpu-culling.ts is only fetched when a scene explicitly opts in.
Bundle Size Impact
The thin instance feature is designed for zero bundle-size impact on scenes that don't use it:
| Layer | Cost | When Loaded |
|---|---|---|
thin-instance.ts (CPU data model) | ~1 KB | Only if user imports setThinInstances() etc. |
thin-instance-gpu.ts (GPU sync) | ~0.9 KB | Dynamic import, only when standardGroupBuilder detects thin instances |
thin-instance-gpu-culling.ts (compute culling) | opt-in chunk | Dynamic import, only when _gpuCullingEnabled is true |
| Shader/pipeline feature flag checks | ~400 bytes | Always present in standard shader composer (unavoidable — feature flags are checked in shared composition functions) |
Scene 16 chunk breakdown: scene16.js (18.1 KB) + standard-renderable (22.5 KB) + thin-instance-gpu (0.9 KB) = 41.5 KB total.
State Machine / Lifecycle
Initialization
- User calls
setThinInstances(mesh, matrices, count)/addThinInstance(mesh, matrix)for a single mesh, orcreateHierarchyInstancePool(root, capacity)for a prop hierarchy. ThinInstanceDatais created onmesh.thinInstanceswith initial capacity.- Optionally, user calls
setThinInstanceColors(mesh, colors)for per-instance RGBA. - Optionally, user calls
enableThinInstanceGpuCulling(mesh)beforeregisterScene().
Per-Frame Render
1. standardGroupBuilder detects mesh.thinInstances2. Dynamically imports thin-instance-gpu.ts (cached after first load)3. Retains the shared sync helper for per-frame upload and draw-time binding4. For each mesh with thinInstances, before cached main/shadow bundles execute: a. the renderable update checks _version vs _gpuVersion b. Creates / resizes GPU buffer if needed (capacity x 64 bytes for matrices) c. writeBuffer from CPU Float32Array → GPU d. Bumps _gpuVersion = _version5. During bundle recording or direct transparent drawing: a. setVertexBuffer(slot, stable matrixBuffer); slot++ b. If hasColor: bind the stable color buffer; slot++ c. drawIndexed(indexCount, ti.count), or drawIndexedIndirect(argsBuffer, 0) after GPU culling6. On steady opaque frames, the cached main and shadow bundles reuse those buffer handles and consume the contents uploaded by step 4 without re-recording.Mutation (Runtime)
addThinInstance→ grows capacity 2× if full, copies old data, bumps_version.removeThinInstance→ swap-removes (O(1)), bumps_version.setThinInstanceMatrix→ overwrites 16 floats in-place, bumps_version.flushThinInstances→ bumps_versiononly (for direct array manipulation).setThinInstanceColors→ replaces colors array, bumps_colorVersion.addHierarchyInstance→ writes one logical root matrix into every descendant mesh buffer and bumps all counts.removeHierarchyInstance→ swap-removes the same logical slot from every descendant mesh buffer.setHierarchyInstanceCount→ changes the active logical count on every descendant mesh without reallocating.
PBR Material Integration
PBR thin instances are fully implemented. The system mirrors the Standard material path but uses the ShaderFragment composition system:
Fragment-Based Integration
pbr-renderable.ts detects thin instances at build time and dynamically imports the thin-instance fragment:
if (meshes.some((m) => !!m.thinInstances)) { const { createThinInstanceFragment } = await import("../../shader/fragments/thin-instance-fragment.js"); fragments.push(createThinInstanceFragment(hasInstanceColor)); const { syncThinInstanceBuffers } = await import("../../mesh/thin-instance-gpu.js"); tiSync = syncThinInstanceBuffers;}Thin Instance Fragment (shader/fragments/thin-instance-fragment.ts)
The fragment contributes:
- Vertex attributes:
world0..world3(instance matrix rows) + optionalinstanceColor - Vertex slot: Composes
finalWorld = mesh.world * instanceWorld— replaces the mesh world matrix in all subsequent transforms - Fragment slot: Multiplies base color/alpha by instance color (when
MSH_HAS_INSTANCE_COLORis set)
Draw Path
if (ti && tiSync) { slot = tiSync(engine, ti, pass, slot, hasInstanceColor, cullResult?.drawBuffers ?? null); if (cullResult) pass.drawIndexedIndirect(cullResult.argsBuffer, 0); else pass.drawIndexed(indexCount, ti.count);}The same GPU sync function (syncThinInstanceBuffers) is shared between Standard and PBR paths.
Scene 17
Scene 17 (scene17-pbr-std-thin-instances) validates PBR thin instances: a PBR box with 2 thin instances + per-instance colors, alongside Standard material thin instances in the same scene.
Node Material Integration
Node materials opt into thin instances with
parseNodeMaterialFromSnippet(engine, id, { hasInstances: true }). When the
graph reaches an InstancesBlock, its world output declares four
float32x4, step-mode instance attributes and evaluates:
meshU.world * mat4x4<f32>(in.world0, in.world1, in.world2, in.world3)The block's instanceID output reads @builtin(instance_index) and converts it
to the Babylon-compatible f32 graph value. If instance support is disabled,
the block returns meshU.world and 0.0; neither the attributes nor the
builtin are emitted.
The renderable uses the shared thin-instance GPU helper in two distinct phases:
- each opaque packet update and each transparent renderable update performs the version-gated upload before rendering;
- bundle recording/direct drawing binds the resulting stable matrix buffer and
issues
drawIndexed(indexCount, thinInstances.count).
This separation is required because opaque Node Material main and shadow
bundles are cached: their draw closures record once, while their update
closures continue to run every frame. Both passes bind the same matrix buffer,
so one pre-render upload makes the current contents visible to both without
separate dirty flags or duplicate writes. The GPU sync module is dynamically
imported while parsing only when hasInstances is enabled and is retained on
that material. Consequently an ordinary node-material scene does not include
thin-instance upload code. Per-instance colors and GPU culling are not part of
the node-material opt-in.
Node packet ownership follows the same synchronous visibility / deferred GPU
retirement split for ordinary and thin-instance meshes. Main-scene packet
disposers are reachable both from scene._meshDisposables and from the scene
lifetime registry so partially built scenes still clean up safely. Removing or
synchronously rebuilding a mesh unregisters the duplicate scene-lifetime
reference, unlinks its packet immediately, and consumes the group's one-shot
owner-empty callback when the final packet leaves. An empty merged renderable
therefore cannot reach binding, and the lifetime registry cannot retain
retired meshes. Packet callbacks reference-count the shared Node UBO, destroy
it only after the final mesh resource retires, and then unregister its
idempotent scene-lifetime fallback.
Auxiliary renderables instead register every allocation in the caller's
MeshRebuildResources._lifetimeDisposers sink and never acquire main-scene
ownership. Node construction tracks all allocations until publication: if any
later bind-group or feature-binding step throws, every unpublished buffer is
released and any provisional main-mesh registration is removed before the
error propagates.
Babylon.js Equivalence Map
| Babylon Lite | Babylon.js |
|---|---|
setThinInstances(mesh, matrices, count) | mesh.thinInstanceSetBuffer("matrix", data, 16) |
setThinInstanceColors(mesh, colors) | mesh.thinInstanceSetBuffer("color", data, 4) |
enableThinInstanceGpuCulling(mesh) | No direct core equivalent; comparable to engine-level custom GPU culling before thin-instance draw |
addThinInstance(mesh, matrix) | mesh.thinInstanceAdd(matrix) |
removeThinInstance(mesh, index) | mesh.thinInstanceRemove(index) |
setThinInstanceMatrix(mesh, index, matrix) | mesh.thinInstanceSetMatrixAt(index, matrix) |
flushThinInstances(mesh) | mesh.thinInstanceBufferUpdated("matrix") |
material.disableLighting = true | material.disableLighting = true |
THIN_INSTANCES feature flag | Internal #define THIN_INSTANCES |
THIN_INSTANCE_COLOR feature flag | Internal #define THIN_INSTANCE_COLOR |
| Per-instance color via vertex attribute | Per-instance color via vertex attribute |
finalWorld = mesh.world * instanceWorld | finalWorld = world * instanceWorld |
Swap-remove with copyWithin | Swap-remove with buffer manipulation |
| Dynamic import of GPU sync module | Always loaded (no code splitting) |
Same math, minimal code. No class hierarchy — just typed arrays, version counters, and a GPU sync function.
Dependencies
- WebGPU instanced drawing (
drawIndexed(indexCount, instanceCount)) - WebGPU indirect indexed drawing (
drawIndexedIndirect(argsBuffer, 0)) for GPU-culled thin instances - WebGPU compute shaders and storage buffers for opt-in culling
- WebGPU vertex buffer
stepMode: 'instance' device.queue.writeBufferfor CPU → GPU transfer- Standard material shader composition (feature-flag-driven WGSL generation)
- Dynamic
import()for lazy chunk loading
Test Specification
| Test | Description |
|---|---|
| Scene 16 parity | Pixel comparison of 64K colored cubes against Babylon.js reference |
| Live reference | Opens babylon-ref-scene16.html, captures live-ref.png, compares against Lite |
| Golden fallback | Falls back to babylon-ref-golden.png if live capture fails |
| MAD threshold | Full-image Mean Absolute Difference ≤ 1 |
| Exact match ratio | ≥ 95% of pixels must be exact matches |
| Capacity growth | addThinInstance beyond initial capacity → doubles array, preserves existing data |
| Swap-remove correctness | removeThinInstance(i) → last instance moves to slot i, count decrements |
| Version skip | Static instances: GPU upload skipped when _version === _gpuVersion |
| Color independence | Matrix mutation does not trigger color re-upload (separate version counters) |
| Count-only draw update | Active count changes update draw args without dirtying matrix or color buffers |
| Zero-cost loading | Scenes without thin instances never fetch thin-instance-gpu.js chunk |
File Manifest
| File | Purpose |
|---|---|
src/mesh/thin-instance.ts | CPU-side data model + public API (ThinInstanceData, setThinInstances, etc.) |
src/mesh/thin-instance-gpu.ts | GPU buffer sync — lazy-loaded chunk (syncThinInstanceBuffers) |
src/mesh/thin-instance-gpu-culling.ts | Opt-in compute frustum culling + compacted visible buffers + indirect args |
src/material/standard/standard-material.ts | disableLighting property + standardGroupBuilder with dynamic sync loading |
src/material/standard/standard-pipeline.ts | THIN_INSTANCES, THIN_INSTANCE_COLOR, DISABLE_LIGHTING flags + pipeline vertex buffer layouts |
src/material/standard/standard-template.ts | instanceColor varying + disableLighting fragment path + instance world matrix composition |
src/material/standard/standard-renderable.ts | tiSync callback integration + instanced drawIndexed / drawIndexedIndirect |
src/material/mesh-features.ts | MSH_HAS_THIN_INSTANCES, MSH_HAS_INSTANCE_COLOR feature flag constants |
src/material/pbr/pbr-renderable.ts | PBR thin-instance detection, fragment/culling loading, instanced draw |
src/material/pbr/pbr-pipeline.ts | PBR pipeline vertex buffer layouts for thin instances |
src/material/node/blocks/instances-block.ts | Node graph instance-world attributes and instanceID builtin |
src/material/node/node-material.ts | Node-material thin-instance opt-in and lazy GPU-sync loading |
src/material/node/node-renderable.ts | Node-material matrix binding and instanced draw |
src/shader/fragments/thin-instance-fragment.ts | ShaderFragment for instance matrix/color — shared by PBR and Standard |
lab/lite/src/lite/scene16.ts | Reference/check scene: 40x40x40 = 64K colored cubes with opt-in GPU culling |
lab/lite/src/lite/scene17.ts | Reference/check scene: PBR + Standard thin instances in one scene, both culling-enabled |
lab/lite/src/lite/scene35.ts | Reference/check scene: glTF EXT_mesh_gpu_instancing with opt-in GPU culling |
tests/lite/parity/scene16-thin-instances.spec.ts | Parity test for Standard thin instances |
tests/lite/parity/scene17-pbr-std-thin-instances.spec.ts | Parity test for PBR + Standard thin instances |
Limitations
- No per-instance custom data — only world matrix and RGBA color are supported as instance attributes.
- Swap-remove reorders instances — removing an instance changes the index of the last instance. Callers managing external index mappings must account for this.
- Max 4 floats per color — RGBA only, no HDR or extended per-instance data.
- GPU culling is opt-in — call
enableThinInstanceGpuCulling(mesh)beforeregisterScene(). - GPU culling is opaque-only in v1 — transparent and transmissive thin instances use the regular draw path.
- GPU culling compacts instance order nondeterministically — correct for opaque rendering, but not suitable for transparent sorting.