VoxelBlox · engine notes

Voxels in Python, on the GPU

An endless, fully editable block world, drawn at a few milliseconds a frame from Python. Here's how it's built: where the work runs, how the terrain reaches the GPU in one call, how the distant land fills in, why there's no CPU occlusion culling, and the measuring habits that decided each step.

1
draw call for all visible solid terrain (GL 4.3 multi-draw indirect; water is a second)
2.69 ms
median frame in the reference smoke world (was 3.88)
228 → 1
draw calls for distant terrain at 4 km
0 of 57,600
pixels differ between the fast path and the fallback

00What it looks like

Screenshots from the game itself, taken through the real GPU path at 1280×720 (the on-screen guide and HUD left in). Click any one for full size.

01The stack, and why Python

VoxelBlox is Python 3.13 with pygame for the window and input, moderngl for OpenGL (3.3 core, 4.3 where the card has it), and numba compiling the heavy loops over numpy arrays. There are no asset files: every texture, model, sound and piece of music is generated in code at load.

Python is slow per call, so the whole design follows one rule: Python touches single blocks; numba touches many. Generation, lighting, meshing, raycasts over many rays and bulk queries are compiled kernels over arrays. Python orchestrates, handles the player's own edit, and issues as few GL calls as possible. A browser port was researched and declined; Windows + Python + numba was a deliberate choice.

02The world's data

The world is a grid of columns, 16×16 blocks wide and 384 tall (y from −128 to 255: under y 0 lie 128 blocks of deep rock, magma and mantle). A column is keyed by integer cells, (x >> 4, z >> 4), never by Python object identity (Python reuses ids, and a sister project once collapsed 8 missiles into 2 that way).

At render distance 32 the raw columns would be 1,114 MB. A ColumnStore keeps columns beyond the near ring zlib-packed (blocks about 30×, meshes about 3×) and unpacks them on read: 181 MB.

03Generation and streaming off the main thread

The terrain kernel, the light flood and the mesher are all compiled nogil, so they run truly in parallel on a thread pool (cores − 2, at most 8) while the main thread gathers inputs and integrates results. Output is byte-identical to a single thread. Loading at render distance 16 went from 1.97 s to 0.78 s.

Worker threads (nogil) terrain kernel light flood column mesher nearest ground first Main thread feature passes the player's own edit: re-meshed the same frame upload only what changed cull columns (numpy) build the command list GPU one arena buffer MultiDraw ElementsIndirect depth test does the occlusion
Where the work runs. The main thread never generates: the compiled kernels do it on the workers.

Two lessons shaped the scheduling:

04Meshing: 12 bytes a vertex

The mesher takes a column padded by one cell on every side (258 × 18 × 18), so faces on the border cull against the neighbours and ambient occlusion and smooth light see across the seam. It emits 4 vertices a visible face, each six int16:

// 6 x int16 = 12 bytes a vertex
x, y, z      column-local corner, in 1/16 block (shaped blocks need the sixteenths)
layer        texture layer in a sampler2DArray (every texture generated at load)
info         face:3 | ao:2 | sky:4 | flow:4 | glow:1 | fluid-surface:1
block light  0-15, flooded by the light kernel
Not built yetGreedy merging. The shader's UVs are already world-space and tile with fract(), so merged quads would drop in unchanged, but so far vertex count hasn't been what limits a frame.

05The terrain in one draw call

In Python, every GL call costs real CPU time, so the aim is to keep terrain to a handful of calls. All column meshes live in one arena buffer, sub-allocated first-fit in spans of 64 quads and doubled on the GPU when full. Each frame, numpy culls the columns against the view (six planes against every column's box at once), sorts them near-first, and writes one indirect command per visible column:

# one row per visible column: (count, instances, first index, base vertex, base instance)
cmd = np.zeros((n, 5), dtype=np.uint32)
cmd[:, 0] = (quads - wet) * 6        # the dry quads' indices
cmd[:, 1] = 1
cmd[:, 3] = base                     # where this column's vertices start in the arena
cmd[:, 4] = np.arange(n)             # baseInstance -> this column's camera-relative offset
vao.render_indirect(commands, moderngl.TRIANGLES, count=len(cmd))

Each column's camera-relative offset is a per-draw attribute read through baseInstance, so the floating origin survives without one uniform per column. The whole terrain is one glMultiDrawElementsIndirect; water is a second, sorted far-first for blending, from the wet quads kept at the end of each column's mesh.

One arena buffer free free Indirect commands (visible columns, near first) cmd 0 → col A cmd 1 → col D cmd 2 → col F one call draws them all
Column meshes sub-allocated in one buffer; culled columns simply get no command.

The per-column path (GL 3.3: one buffer and one draw per column) stays as the fallback, and a check requires the two to render identical pixels (0 of 57,600 differ). Paired smoke runs: terrain draw 2.01 / 2.17 ms per column → 1.78 / 1.77 ms multi-draw, and the p99 frame 11.0 / 8.8 ms → 5.7 / 5.9 ms.

Upload only what changed. A sister project once re-uploaded a 112 MB terrain every frame while streaming and fell from 40–50 FPS to 1–4. Here a new mesh is written into its span with a byte offset, and nothing else moves.

06The far land, and the holes a jet outruns

Beyond the meshed columns, a far land of coarse height tiles (a region each) runs out to the horizon, built on the workers from the same generation rule the columns follow. Towns, cities and megastructures appear on it: each tile samples the heights of its structures' own voxel programs, coloured by their top blocks.

columns meshed not meshed yet: filled by the holes pass far land tiles (one indirect call) flying fast, the ring ahead isn't meshed yet: the far tiles cover it, so there's no sky hole.
The far land fills exactly where the near columns aren't ready, and nowhere else. A check requires every pixel on a drawn column to be identical with the fill on and off.

07Occlusion: letting the depth buffer do it

Eniko Fox's excellent software-rendered occlusion culling in Block Game rasterises nearby opaque blocks into a small CPU depth buffer and skips hidden chunks, cutting a lot of chunks (50–60% above ground, up to 95% in caves). It's a great fit for an engine that pays per chunk it draws.

VoxelBlox made the opposite call, for two reasons:

  1. The draw calls are already gone. All visible terrain is one indirect call, so skipping a column saves only GPU vertex work, not CPU time per chunk. The CPU cost of drawing doesn't grow with the number of columns in view.
  2. A CPU occlusion test where the GPU writes depth was a bug as well as a cost. In the sister project, VoxelStrike, machines blinked as they crested hills: the CPU's answer lagged or disagreed with what the GPU drew. The rule since: where the GPU writes depth, the CPU never tests occlusion.
Where the article's idea still pointsTwo things would fit the rule if a GPU measurement ever shows terrain vertices are the bound: cave culling by connectivity (flood each section's air at mesh time, record which faces see each other, and walk only open air from the camera: conservative, on the workers, with no depth and so no blinking), and GPU occlusion (a compute pass testing boxes against a depth pyramid of the GPU's own depth, writing the indirect command list directly). Neither is built: measure first.

08Frame pacing and the CPU side

The first windowed profile had p50 16.6 ms but p99 35–57 ms, while the logged work summed to ~6 ms. Every spike was in flip: the driver's windowed vsync stalled 2–4 refreshes. With vsync off, p99 was 5.6 ms. So the game paces itself: sleep, then spin the last 1.5 ms, at the monitor's rate, and the result was 15.15 ms intervals with 0.00 jitter over 1,125 frames.

On the CPU, profiling ranked things nobody would guess:

Net result on the reference smoke world (render distance 8): frame p50 3.88 → 2.69 ms, draw 2.43 → 1.33 ms.

09The method: measure through the player's path

The engine is less interesting than the habits that built it. A few rules that did most of the work:

The tools are part of the codebase: perf_tour.py drives one world through the real loop to the costly places (a metropolis, a spaceport mid-launch, a jet low over mountains) and reports percentiles, the part that dominates slow frames, and CPU by function. Spike reports capture what any frame over 250 ms was doing. Seeded fuzzing plays through the input path, and a crash leaves a replayable recording. perf_tour once took the city's worst frame from 224 ms to 18 ms and the metropolis's from 667 ms to 20 ms by finding one pure-Python search per building every 2 seconds.

10All the numbers

WhatBeforeAfterHow
Frame p50, smoke world3.88 ms2.69 msmulti-draw, batched posing, resting creatures, cached HUD
Frame p99, paired smoke runs11.0 / 8.8 ms5.7 / 5.9 msterrain in one indirect call
Windowed p9935–57 ms5.6 msown frame pacer instead of driver vsync
HUD draw4.2 ms1.1 msflat array, cached layouts
Distant terrain calls, 4 km2281tiles in an arena, one indirect call
Holes in view, jet at 18 chunks4.8%0.01%the holes pass (u_cover)
Ground drawn ahead of the jet (p50)165–178 m1,800 mthe holes pass
A dig, deep underground148 ms8 msre-mesh only what the edit dirtied; deep zone
Load, render distance 161.97 s0.78 snogil kernels on a thread pool
Memory, render distance 321,114 MB181 MBfar columns zlib-packed
City worst frame224 ms18 msperf_tour found a per-building search
Metropolis worst frame667 ms20 msthe same, compiled and spread out

All figures measured on the development machine and recorded in the project's engineering log with the run that produced them.

11Should you copy this?

Honestly: copy the techniques and the habits; think hard before copying the language. Here's each piece graded for someone starting their own voxel game.

PieceVerdictWhy
One buffer + multi-draw indirectKEEPThe standard modern answer to draw-call cost, in any language. The single biggest win here.
Kernels off the main threadKEEPGeneration, light and meshing as pure functions over arrays, run on workers, nearest ground first. It's language-independent and it removes most hitches.
Pure-hash generation, journal saveKEEPAny chunk regenerates identically in any order, and a save stores only what the player changed. That makes streaming, saves and multiplayer simple.
Floating originKEEPCheap, and it removes a whole class of far-from-origin jitter bugs for good.
Your own frame pacerKEEPWindowed vsync stalled 2–4 refreshes here. Measure your flip before you blame your code.
The far land + holes passKEEPA cheap horizon, and the fill trick means fast travel never shows sky through the ground.
Witnessed checks, perf toursKEEPThe habit that paid most: every fix was measured through the player's path and every check shown to fail.
Python + numba as the engineONLY IFIt works because the heavy loops are compiled and the draw calls are gone. But everything left in Python (creatures, UI, feature passes) costs many times what it would in C#, Rust or C++, first-run compiles are slow, and shipping is heavier. Choose it if you love Python or want to learn the ideas. For a commercial game, Godot, Unity, Bevy or a C++ engine will get you there with less effort.
Columns as the draw and cull unitIMPROVEA 16×384 column is simple, but it's coarse. 16³ sections cull caves and empty sky separately and enable connectivity culling.
12-byte vertices, no greedy mergeIMPROVEFine so far, but greedy meshing or vertex pulling with packed 4–8-byte faces would cut memory and vertex work several times.
CPU occlusion cullingSKIP HEREIt's a big win when you pay per chunk. With one indirect call you don't, and a CPU answer that disagrees with the GPU's depth makes things blink. Prefer connectivity culling or GPU-side tests.

12What's missing, and what I'd do differently

13What's next: faster, then finer

Two research tracks, each item with a measurement to take first and a condition that stops it. Nothing here ships until it beats the current path on more than one machine and matches its pixels.

Faster: first inside Python, then out of it

  1. Measure better first. Build a GPU lab tool (per-pass GPU time, quads, overdraw, arena use for every screenshot scene), get a second, weaker baseline machine, and gate frame budgets per system so a regression fails a check instead of being noticed by feel.
  2. Sections and cave culling, then greedy meshing and packed vertices, then GPU-driven culling (a compute pass writing the indirect list), each only if the GPU is shown to be the bound.
  3. The known CPU cost: the per-column feature passes still on the main thread while streaming, and the per-object Python work (creatures, people, vehicle tools) that grows with the scene.
  4. Leaving Python, piece by piece. Replace the numba kernels one at a time with native ones (Rust or C++, called from Python), each behind a setting beside its twin and required to produce byte-identical output: the mesher first, since its reference oracle already exists. Only when the kernels are native, profile again and decide whether the renderer and streaming core should follow. A full port to an engine would be the last resort, because it throws away the tools and checks that made this work.
  5. A second graphics API (wgpu: Vulkan, Metal, DX12) behind the same renderer interface, if macOS or GPU culling ever needs it.

Finer: less blockiness, built on what already works

14Lessons learned

  1. The instrument is wrong more often than the engine. A CPU meter read the game at 0% (a truncated process handle); a stall blamed on the code was the driver's vsync. Check the tool before believing its number.
  2. The cost is almost always redundant work, not a bad algorithm. Parked vehicles re-posed on every column streamed in, the deep re-meshed for the whole view, a whole terrain re-uploaded every frame. Look for work done twice before rewriting anything.
  3. A green test suite proves nothing unless a check asks the observable question. Read back pixels and positions, not internal flags.
  4. A check never shown to fail isn't a check. Revert the line it protects and watch it go red. Never derive the bar from the setting it checks.
  5. Ship a new render path off until an on-device frame rate says it's fine, and keep the old path as a pixel-identical oracle.
  6. Python pays per call. Batch everything that repeats: one indirect draw, one numpy pass for all creatures, a flat array('f') for the HUD instead of 800,000 tuples.
  7. Never key anything by object identity, and never draw from a global random stream during play. Both produced bugs that only showed up much later.
  8. Where you put a player, ask for room and for ground. "Not inside a block" alone once put a driver on top of the hill above the tunnel they'd just dug.