Voxels in Python, on the GPU
An endless, fully editable block world, drawn at a few milliseconds a frame from Python. Here's how it's built: where the work runs, how the terrain reaches the GPU in one call, how the distant land fills in, why there's no CPU occlusion culling, and the measuring habits that decided each step.
00What it looks like
Screenshots from the game itself, taken through the real GPU path at 1280×720 (the on-screen guide and HUD left in). Click any one for full size.










01The stack, and why Python
VoxelBlox is Python 3.13 with pygame for the window and input, moderngl for OpenGL
(3.3 core, 4.3 where the card has it), and numba compiling the heavy loops over numpy arrays.
There are no asset files: every texture, model, sound and piece of music is generated in code at load.
Python is slow per call, so the whole design follows one rule: Python touches single blocks; numba touches many. Generation, lighting, meshing, raycasts over many rays and bulk queries are compiled kernels over arrays. Python orchestrates, handles the player's own edit, and issues as few GL calls as possible. A browser port was researched and declined; Windows + Python + numba was a deliberate choice.
02The world's data
The world is a grid of columns, 16×16 blocks wide and 384 tall (y from −128 to 255: under y 0
lie 128 blocks of deep rock, magma and mantle). A column is keyed by integer cells,
(x >> 4, z >> 4), never by Python object identity (Python reuses ids, and a sister project once
collapsed 8 missiles into 2 that way).
- A cell is two bytes and a byte: a
uint16block id and auint8state (facing, water level, upside-down). Anything that stores or sends blocks carries both: the world, the journal, saves, the network, blueprints. - Light is a byte per cell: sky and block light, flooded by a numba kernel.
- Floating origin: every vertex is column-relative, and the camera offset is subtracted in float64 on the CPU. The view matrix is rotation only, so float32 never sees a large coordinate. A check renders the same scene at the origin and at 30 million blocks out: 93 of 57,600 pixels differ.
- Endless by construction: generation is a pure hash of int64 coordinates, so any column regenerates identically, in any order.
- The journal is the save: per column, only the cells that differ from generation. What you build stays built across unload, reload and save.
At render distance 32 the raw columns would be 1,114 MB. A ColumnStore keeps columns beyond the near
ring zlib-packed (blocks about 30×, meshes about 3×) and unpacks them on read: 181 MB.
03Generation and streaming off the main thread
The terrain kernel, the light flood and the mesher are all compiled nogil, so they run truly in
parallel on a thread pool (cores − 2, at most 8) while the main thread gathers inputs and integrates results. Output
is byte-identical to a single thread. Loading at render distance 16 went from 1.97 s to 0.78 s.
Two lessons shaped the scheduling:
- Ground where the player is comes first. Flying fast at render distance 18, generation once held every worker slot until the whole ring (~1,300 columns) was generated, and the ground within 4 columns fell to 0% drawn. Now never-drawn ground near the player gets half the workers as soon as it's ready: 100% drawn at 40 m/s, and the whole view in 2–3 s when stopped.
- The deep is meshed only round the player. Digging straight down once lowered the mesh floor for the whole view and re-meshed about 1,000 columns on the main thread: a one-second freeze per step. Now the deep is meshed in a zone 4 columns round the player that follows a tunnel, and a dig re-meshes only what that edit dirtied: 148 ms → 8 ms a dig.
04Meshing: 12 bytes a vertex
The mesher takes a column padded by one cell on every side (258 × 18 × 18), so faces on the border cull
against the neighbours and ambient occlusion and smooth light see across the seam. It emits 4 vertices a visible face,
each six int16:
// 6 x int16 = 12 bytes a vertex x, y, z column-local corner, in 1/16 block (shaped blocks need the sixteenths) layer texture layer in a sampler2DArray (every texture generated at load) info face:3 | ao:2 | sky:4 | flow:4 | glow:1 | fluid-surface:1 block light 0-15, flooded by the light kernel
- Per-vertex AO from three neighbours; a quad whose AO is anisotropic is emitted rotated by one vertex so its diagonal follows the light (otherwise it shows as seams).
- Fluids are drawn at their level: each top corner is the mean surface height of the fluid cells sharing it, so neighbours meet exactly and runs slope downhill. The flow direction (8 ways, or falling) rides in 4 spare bits, and the shader scrolls runs, pours falls and churns lava.
- Shaped blocks (stairs, doors, furniture, wedges for 45° slopes) are box models turned by the cell's state, lit with the same corner light as a full cube so a hillside of slopes shades like the blocks round it.
- One shared index buffer (0,1,2, 0,2,3 per quad) serves every mesh.
- A slow reference mesher in plain Python is the oracle: a check requires the numba mesher to match it face for face.
fract(), so merged quads would drop in unchanged, but so far vertex count hasn't been what limits a frame.05The terrain in one draw call
In Python, every GL call costs real CPU time, so the aim is to keep terrain to a handful of calls. All column meshes live in one arena buffer, sub-allocated first-fit in spans of 64 quads and doubled on the GPU when full. Each frame, numpy culls the columns against the view (six planes against every column's box at once), sorts them near-first, and writes one indirect command per visible column:
# one row per visible column: (count, instances, first index, base vertex, base instance) cmd = np.zeros((n, 5), dtype=np.uint32) cmd[:, 0] = (quads - wet) * 6 # the dry quads' indices cmd[:, 1] = 1 cmd[:, 3] = base # where this column's vertices start in the arena cmd[:, 4] = np.arange(n) # baseInstance -> this column's camera-relative offset vao.render_indirect(commands, moderngl.TRIANGLES, count=len(cmd))
Each column's camera-relative offset is a per-draw attribute read through baseInstance, so the
floating origin survives without one uniform per column. The whole terrain is one
glMultiDrawElementsIndirect; water is a second, sorted far-first for blending, from the wet quads kept at
the end of each column's mesh.
The per-column path (GL 3.3: one buffer and one draw per column) stays as the fallback, and a check requires the two to render identical pixels (0 of 57,600 differ). Paired smoke runs: terrain draw 2.01 / 2.17 ms per column → 1.78 / 1.77 ms multi-draw, and the p99 frame 11.0 / 8.8 ms → 5.7 / 5.9 ms.
Upload only what changed. A sister project once re-uploaded a 112 MB terrain every frame while streaming and fell from 40–50 FPS to 1–4. Here a new mesh is written into its span with a byte offset, and nothing else moves.
06The far land, and the holes a jet outruns
Beyond the meshed columns, a far land of coarse height tiles (a region each) runs out to the horizon, built on the workers from the same generation rule the columns follow. Towns, cities and megastructures appear on it: each tile samples the heights of its structures' own voxel programs, coloured by their top blocks.
- One call for all tiles. The tiles live in an arena too, with a shared index buffer (fine grid, coarse
grid, box pattern) and one
render_indirect. At 4 km: 228 calls → 1, CPU 0.5 → 0.2 ms, pixel-identical to the per-tile path. - The holes pass. Flying the jet at 112 m/s at render distance 18, the ground was only drawn 98–114 m
ahead (p5; p50 165–178 m): the far land's shader discards everything inside the render distance because "the columns are drawn
there", and where they weren't yet, there was sky. The fix draws the far tiles again inside the ring, depth-tested
with the columns' own projection, keeping only fragments over a column that isn't on the GPU yet. That's an R8
texture with one texel a column (
u_cover), rebuilt only when a mesh comes or goes. Holes went from 4.8% of the view to 0.01%, and the ground is drawn to the full 1,800 m ahead (p50).
07Occlusion: letting the depth buffer do it
Eniko Fox's excellent software-rendered occlusion culling in Block Game rasterises nearby opaque blocks into a small CPU depth buffer and skips hidden chunks, cutting a lot of chunks (50–60% above ground, up to 95% in caves). It's a great fit for an engine that pays per chunk it draws.
VoxelBlox made the opposite call, for two reasons:
- The draw calls are already gone. All visible terrain is one indirect call, so skipping a column saves only GPU vertex work, not CPU time per chunk. The CPU cost of drawing doesn't grow with the number of columns in view.
- A CPU occlusion test where the GPU writes depth was a bug as well as a cost. In the sister project, VoxelStrike, machines blinked as they crested hills: the CPU's answer lagged or disagreed with what the GPU drew. The rule since: where the GPU writes depth, the CPU never tests occlusion.
08Frame pacing and the CPU side
The first windowed profile had p50 16.6 ms but p99 35–57 ms, while the logged work summed to ~6 ms. Every spike was in flip: the driver's windowed vsync stalled 2–4 refreshes. With vsync off, p99 was 5.6 ms. So the game paces itself: sleep, then spin the last 1.5 ms, at the monitor's rate, and the result was 15.15 ms intervals with 0.00 jitter over 1,125 frames.
On the CPU, profiling ranked things nobody would guess:
- The HUD: 800k tuple appends plus
np.arrayover them was 2.6 ms of a 4.2 ms draw. A flatarray('f')and cached label layouts cut draw from 4.2 to 1.1 ms. - Posing creatures one by one in numpy: batched into one pass (
pose_many), identical rows, 1.2 → 0.5 ms. Parked vehicles are posed once per column and uploaded together. - Resting creatures skip physics until the world changes (41% of creature ticks).
- No full-screen CPU surfaces: UI and effects are GL quads and shader uniforms. A CPU HUD upload once cost 1.1 ms GPU at 1080p and 4.4 ms at 4K.
Net result on the reference smoke world (render distance 8): frame p50 3.88 → 2.69 ms, draw 2.43 → 1.33 ms.
09The method: measure through the player's path
The engine is less interesting than the habits that built it. A few rules that did most of the work:
- The instrument is wrong more often than the engine. Every tool gets checked before its numbers are believed, and a dramatic number is reproduced before anything changes.
- Verify through the path the player takes: keys and mouse through the real event handler, the real frame loop, the GPU's actual pixels. Checks read back pixels, not internal flags.
- A check isn't a check until it has failed. Every assertion is witnessed: revert the one line it protects, run the check, watch it fail, restore. The bar is never computed from the knob it checks.
- Profile to rank, wrap to size, look for redundant work first. The cost is almost always redundant work (the re-posed parked vehicles, the re-meshed deep, the re-uploaded terrain), not a bad algorithm.
- A new render path ships off until an on-device frame rate says it's fine. The far land's one-call path shipped as a setting for exactly that reason.
The tools are part of the codebase: perf_tour.py drives one world through the real loop to the costly
places (a metropolis, a spaceport mid-launch, a jet low over mountains) and reports percentiles, the part that
dominates slow frames, and CPU by function. Spike reports capture what any frame over 250 ms was doing. Seeded fuzzing
plays through the input path, and a crash leaves a replayable recording. perf_tour once took the city's
worst frame from 224 ms to 18 ms and the metropolis's from 667 ms to 20 ms by finding one pure-Python search
per building every 2 seconds.
10All the numbers
| What | Before | After | How |
|---|---|---|---|
| Frame p50, smoke world | 3.88 ms | 2.69 ms | multi-draw, batched posing, resting creatures, cached HUD |
| Frame p99, paired smoke runs | 11.0 / 8.8 ms | 5.7 / 5.9 ms | terrain in one indirect call |
| Windowed p99 | 35–57 ms | 5.6 ms | own frame pacer instead of driver vsync |
| HUD draw | 4.2 ms | 1.1 ms | flat array, cached layouts |
| Distant terrain calls, 4 km | 228 | 1 | tiles in an arena, one indirect call |
| Holes in view, jet at 18 chunks | 4.8% | 0.01% | the holes pass (u_cover) |
| Ground drawn ahead of the jet (p50) | 165–178 m | 1,800 m | the holes pass |
| A dig, deep underground | 148 ms | 8 ms | re-mesh only what the edit dirtied; deep zone |
| Load, render distance 16 | 1.97 s | 0.78 s | nogil kernels on a thread pool |
| Memory, render distance 32 | 1,114 MB | 181 MB | far columns zlib-packed |
| City worst frame | 224 ms | 18 ms | perf_tour found a per-building search |
| Metropolis worst frame | 667 ms | 20 ms | the same, compiled and spread out |
All figures measured on the development machine and recorded in the project's engineering log with the run that produced them.
11Should you copy this?
Honestly: copy the techniques and the habits; think hard before copying the language. Here's each piece graded for someone starting their own voxel game.
| Piece | Verdict | Why |
|---|---|---|
| One buffer + multi-draw indirect | KEEP | The standard modern answer to draw-call cost, in any language. The single biggest win here. |
| Kernels off the main thread | KEEP | Generation, light and meshing as pure functions over arrays, run on workers, nearest ground first. It's language-independent and it removes most hitches. |
| Pure-hash generation, journal save | KEEP | Any chunk regenerates identically in any order, and a save stores only what the player changed. That makes streaming, saves and multiplayer simple. |
| Floating origin | KEEP | Cheap, and it removes a whole class of far-from-origin jitter bugs for good. |
| Your own frame pacer | KEEP | Windowed vsync stalled 2–4 refreshes here. Measure your flip before you blame your code. |
| The far land + holes pass | KEEP | A cheap horizon, and the fill trick means fast travel never shows sky through the ground. |
| Witnessed checks, perf tours | KEEP | The habit that paid most: every fix was measured through the player's path and every check shown to fail. |
| Python + numba as the engine | ONLY IF | It works because the heavy loops are compiled and the draw calls are gone. But everything left in Python (creatures, UI, feature passes) costs many times what it would in C#, Rust or C++, first-run compiles are slow, and shipping is heavier. Choose it if you love Python or want to learn the ideas. For a commercial game, Godot, Unity, Bevy or a C++ engine will get you there with less effort. |
| Columns as the draw and cull unit | IMPROVE | A 16×384 column is simple, but it's coarse. 16³ sections cull caves and empty sky separately and enable connectivity culling. |
| 12-byte vertices, no greedy merge | IMPROVE | Fine so far, but greedy meshing or vertex pulling with packed 4–8-byte faces would cut memory and vertex work several times. |
| CPU occlusion culling | SKIP HERE | It's a big win when you pay per chunk. With one indirect call you don't, and a CPU answer that disagrees with the GPU's depth makes things blink. Prefer connectivity culling or GPU-side tests. |
12What's missing, and what I'd do differently
- Sections and cave culling first. Split columns into 16³ sections, record at mesh time which faces of each section can see each other through air, and walk only open air from the camera. It's conservative, it runs on the workers, and it catches the underground case, where Block Game's article saw up to 95% of chunks hidden.
- Greedy meshing or vertex pulling. The shader already tiles world-space UVs, so merged quads drop in. Vertex pulling (face data in a storage buffer, the vertex shader expanding quads) would shrink the arena further.
- A real sky-light flood. Sky light is exposure from the height map today; a flood would light overhangs and cave mouths properly.
- GPU occlusion, if measured to matter. A compute pass testing boxes against a depth pyramid of last frame's depth, writing the indirect commands itself. The command list is already on the GPU, so it would slot in.
- Measure on more than one machine. Every number here comes from one development PC (a big desktop GPU, played over Remote Desktop). The ratios transfer; the absolute milliseconds won't.
13What's next: faster, then finer
Two research tracks, each item with a measurement to take first and a condition that stops it. Nothing here ships until it beats the current path on more than one machine and matches its pixels.
Faster: first inside Python, then out of it
- Measure better first. Build a GPU lab tool (per-pass GPU time, quads, overdraw, arena use for every screenshot scene), get a second, weaker baseline machine, and gate frame budgets per system so a regression fails a check instead of being noticed by feel.
- Sections and cave culling, then greedy meshing and packed vertices, then GPU-driven culling (a compute pass writing the indirect list), each only if the GPU is shown to be the bound.
- The known CPU cost: the per-column feature passes still on the main thread while streaming, and the per-object Python work (creatures, people, vehicle tools) that grows with the scene.
- Leaving Python, piece by piece. Replace the numba kernels one at a time with native ones (Rust or C++, called from Python), each behind a setting beside its twin and required to produce byte-identical output: the mesher first, since its reference oracle already exists. Only when the kernels are native, profile again and decide whether the renderer and streaming core should follow. A full port to an engine would be the last resort, because it throws away the tools and checks that made this work.
- A second graphics API (wgpu: Vulkan, Metal, DX12) behind the same renderer interface, if macOS or GPU culling ever needs it.
Finer: less blockiness, built on what already works
- Heightfield shapes. The 45° wedges are already a height function within a cell; generalise it to dome caps, quarter-rounds, rounded corners and arches, one primitive for all of them. Add side slopes, posts and pillars.
- Micro-blocks. A cell split into 4³ or 8³ sub-voxels for carved detail: rounded edges, domes, railings, sculpture. Only detailed cells pay, and far away they draw as a plain block. This grows what a cell is from (id, state) to (id, state, detail), so every system that stores or sends blocks has to carry it, with round-trip checks.
- Smooth terrain, the grid way first: slopes and corners placed automatically on generated hills, as a per-region trait so some land stays rugged. Truly organic ground (surface nets) is a research spike for a separate world type, not a change to this one.
- Rounder vehicles and objects. Start with shading that fakes a bevel on every model part, which is a shader-only change, then add rounded primitives (cylinders, rounded boxes, tapers, domes), then prototype a vehicle modelled as a fine voxel grid (meshed once, with AO, and instanced), comparing screenshots and the parked-city frame cost side by side.
14Lessons learned
- The instrument is wrong more often than the engine. A CPU meter read the game at 0% (a truncated process handle); a stall blamed on the code was the driver's vsync. Check the tool before believing its number.
- The cost is almost always redundant work, not a bad algorithm. Parked vehicles re-posed on every column streamed in, the deep re-meshed for the whole view, a whole terrain re-uploaded every frame. Look for work done twice before rewriting anything.
- A green test suite proves nothing unless a check asks the observable question. Read back pixels and positions, not internal flags.
- A check never shown to fail isn't a check. Revert the line it protects and watch it go red. Never derive the bar from the setting it checks.
- Ship a new render path off until an on-device frame rate says it's fine, and keep the old path as a pixel-identical oracle.
- Python pays per call. Batch everything that repeats: one indirect draw, one numpy pass for all creatures, a
flat
array('f')for the HUD instead of 800,000 tuples. - Never key anything by object identity, and never draw from a global random stream during play. Both produced bugs that only showed up much later.
- Where you put a player, ask for room and for ground. "Not inside a block" alone once put a driver on top of the hill above the tunnel they'd just dug.