VoxelBlox · architecture and engineering notes

Engineering a procedural voxel world in Python

A case study in architecture and engineering practice: an open-world voxel sandbox with an endless editable world, no hand-made assets, peer-to-peer multiplayer and a mod API, built and run by one developer. It covers the system's shape, the decisions and their trade-offs, the verification strategy that keeps a large codebase honest, the incidents that shaped it, and how the rendering engine meets its frame budget from Python.

Last updated 30 September 2026

Recent changes (30 September 2026)In-world physics games (bowling, golf, basketball, darts, pool) with discoverable play spots; eighteen screen games with rules pages and casino-accurate blackjack and roulette; a post-round coach in every game; learning by doing (topics taught by playing, every building and person teaching something); procedural audio for every game, venue and job; populated venues; banking and a simulated stock market; smooth rail gradients; a static check class for draw-path-only code; eight incident reviews.
622
automated checks in the gate, each witnessed failing on a revert
~94k
lines of Python across 141 engine and 92 content modules
0
asset files: every texture, model, sound and building is generated
1
draw call for all visible solid terrain (GL 4.3 multi-draw indirect)

00What it looks like

Screenshots from the game itself, taken through the real GPU path at 1280×720 (the on-screen guide and HUD left in). Click any one for full size.

Recent work: games in the world

The latest captures: physics games played in the world itself, and the screen games at the venues that host them. Same GPU path, 1280×720.

01The system at a glance

VoxelBlox is an open-world voxel sandbox: dig, build, drive, fly, travel to orbit. Three constraints shape everything else. The world is endless and fully editable. Content is procedural and code-defined: there are no asset files, so a texture, a sound effect, a town and a cruise ship are all programs. And the codebase is maintained by one developer, so the architecture has to keep a large surface area safe to change without a team's worth of review.

Content as code blocks, items, structures, creatures, games, sounds: registered through a mod API Registries and world options append-only ids, saves store names; generation changes gated per world World model deterministic generation, one edit API, journal = save Simulation (session) fixed tick, host-authoritative, player intents as named actions Rendering numba kernels on workers, GPU-driven terrain, GL quads UI Networking WebRTC peer-to-peer or LAN, world deltas from the edit journal Tooling and verification 622 checks 41 tools 364 GPU scenes
The layers. The verification layer is drawn alongside rather than beneath: it exercises every layer through the same entry points the player uses.

Much of the design is a single-developer answer to problems a studio would solve with process and headcount. Where a team relies on code review to catch a broken invariant, this codebase encodes the invariant in one function or one registry and gates it with a check. Where a team would keep architecture knowledge in people's heads, it is kept in an engineering log and a lessons file that the gate itself keeps in step with the code.

02Architecture decisions and trade-offs

The significant decisions, in the spirit of lightweight architecture decision records: what was chosen, what it costs, and the alternative it was weighed against. Each has a dated entry in the engineering log with the measurements behind it.

DecisionWhyCost / alternative
Python + numba, not an engineFast iteration and one language end to end; the hot loops compile to native code and release the GIL, so they run truly parallel on worker threads.Everything left in Python (creatures, UI, feature passes) costs several times what it would in C#, Rust or C++. Godot, Unity or Bevy would have been the conventional choice. A browser build was researched and declined.
Content as code, through registriesNo asset pipeline; every building is a small program over a builder API, and mods register content the same way core does. Registries are append-only and saves store names, so content can grow and be renamed without breaking a save.Content changes are code changes, with code's review cost. Visual iteration needs tooling (a design preview, a gallery), which had to be built.
Deterministic generation, the journal as the saveGeneration is a pure hash of coordinates, so any region regenerates identically in any order. A save stores only the cells a player changed.Every generation change would silently alter existing worlds, so each one ships behind a per-world option: on for new worlds, off for old saves. In effect, feature flags with a backward-compatibility policy.
One edit API for the worldEvery change goes through one call that journals it, supports undo and produces the multiplayer delta. Networking and persistence can't drift from each other.No shortcut writes, even in tools. Bulk generation is the single sanctioned exception.
Player intents as named actionsAnything that changes state (a purchase, a bank deposit, a game button) is a registered action the host executes. It's the command pattern, and it makes multiplayer safe by construction.Slightly more ceremony for small UI features.
One rule, one functionA rule the UI displays and the model enforces lives in one function: layout, recipes, stats, and now each game's rules page. The drawing and the checks call the same code.More indirection in UI code, in exchange for UI and model never disagreeing.
Explanation as a system, not copyRules, scoring, real-world context and coaching live in registries that every surface reads: the rules page, the in-world rules card, the Almanac, the Learn tab. A game records what happened (notes); a coach function turns that into feedback after the round, stored with the player.Every new game needs its coach and notes written; a gate requires them, and requires the coach to say different things for careful and careless play.
GPU-driven terrain, no CPU occlusionOne multi-draw-indirect call removes per-chunk CPU cost, which is the thing Python can least afford. The depth buffer does occlusion.Vertex work isn't culled early; connectivity or compute-based culling are ready options if a measurement ever shows the GPU is the bound.

03Verification strategy

With one developer and a fast-moving codebase, the test suite has to catch what review would. Its central idea is borrowed from mutation testing and applied to every assertion rather than sampled across the codebase.

Why this matters more than coverageLine coverage says a line ran. A witnessed check says a specific wrong version of that line would have been caught. For a solo codebase that changes daily, the second is the property that lets you refactor without fear.

Designing for learning

The game is meant to teach how things work, so learning is engineered like any other feature: measured, gated and fed by the game's own data.

04Observability and performance engineering

05Delivery model: from request to done

The workflow is deliberately lightweight, but it has the traceability of a larger process:

  1. Requirements in the stakeholder's words. Each request becomes a row in an open-items file, quoted verbatim, with its status, the checks that prove it and how to try it. A row closes when the feature is confirmed in the product, not when tests pass: definition of done as acceptance, not as a build status.
  2. A branch per change, trunk kept releasable. Merges require the targeted checks, the documentation check and the code-map check to pass. The full suite runs in parallel workers on demand.
  3. An engineering log as the decision record. Every change ends with a dated entry holding the evidence (the measurements, what was tried, what was rejected), and a one-line lesson, tagged by area, in a lessons file that's read before touching that area again.
  4. Every job leaves a tool. A question that was hard to answer once becomes a tool that answers it always: 41 so far, all on one result protocol with JSON output, saved baselines and comparisons, and one local web page that runs them.

06Incident reviews

Blameless, short, and focused on the detection gap: why the existing checks didn't catch it, and what now does.

A draw-path-only crash that every check missed

Impact: the game would have crashed on the first frame of play. It was caught by pre-release screenshots before reaching a player. Root cause: new code in the draw path called the settings lookup with a default value; the lookup takes only a name. Detection gap: all checks for the feature exercised its logic directly, and no check draws a frame during play, so that line never ran under test.

Remediation: a static check parses the source and fails on any settings lookup with the wrong arity, and a practice of capturing a GPU screenshot after any change to the draw path.

GPU driver resets during test runs

Impact: repeated graphics driver resets on the development machine, taking down every program using the GPU. Root cause: the Windows event log showed 1,334 driver errors in one day, all during test runs: up to seven concurrent runs, their workers each creating and destroying 30 to 50 OpenGL contexts, beside an inference server holding 7 GB of video memory. The driver's context-switching engine failed. Detection gap: the dying workers looked like crashes in whatever check they happened to be running, so each failure first read as a different bug.

Remediation: one GL context per process, at most four rendering workers, a machine-wide lock, workers bound to their parent's lifetime, and a driver-fault count printed at the end of every run.

Collision tunnelling in the bowling physics

Impact: a perfect pocket roll knocked down two pins. Root cause: at 8 m/s the ball advanced 40 cm per tick against a pin 13 cm wide, so most pins were stepped over. Detection gap: none. The first end-to-end check caught it before merge, which is the point of writing that check first.

Remediation: adaptive substeps capping any fast body's movement at 4 cm (up to 24 per tick). A pocket roll now takes nine pins, and just off the head pin is a strike.

A market with no volatility

Impact: simulated share prices barely moved. Root cause: prices summed four draws from CRC32 of the time; CRC32 is linear, so the draws were correlated and their sum nearly constant. Detection gap: the check asserted determinism (the same prices for everyone), not the distribution.

Remediation: a seeded PRNG per block of 256 ticks, which is deterministic across clients and swings 0.3–2.4× over a long run. The check now covers the range as well.

Content that existed in code and never in a world

Impact: two new building types never generated. Root cause: the planetarium lost every weighted draw because of a bias in the hash on the lot grid, and the garden centre never fit a town lot. Detection gap: unit checks built each building directly; nothing measured the population.

Remediation: adjusted weights and footprints, and a census tool that plans 500+ settlements and reports types that never appear, and why. A type that never generates is treated as a defect.

Eighteen screens nobody had looked at

Impact: screen games shipped with placeholder visuals: the home run derby was two white rectangles and a dot. Root cause: the games were verified by bots and checks, which don't judge appearance. Detection gap: looking required the GPU, which is rationed.

Remediation: a software renderer for the games' view description, which draws any game's screen without a GPU. The first redraw cost 475 shapes a frame; lines and fields became polygons (64 shapes), and a gated budget now caps every game screen at 160.

A learning agent that couldn't express the policy

Impact: after blackjack moved to real casino rules (six-deck shoe, 3:2 blackjack, doubling, dealer peek), the learned player lost 10 chips a round, and the balancing gate failed. Root cause: a linear scorer over raw features can't represent a conjunction such as “12–16 and the dealer shows 7+”. Detection gap: none. The gate did its job.

Remediation: the domain concepts as features. The agent now plays level with textbook basic strategy, which over 300 rounds ends about 3 chips down: the house edge, as in the real game.

A font without the glyphs

Impact: card suits would have rendered as empty boxes. Root cause: the UI font atlas covers ASCII only; the default font's metrics for ♥ ♦ ♣ ♠ were one identical box. Detection gap: caught by measuring before building.

Remediation: suits composed from the view's own primitives (discs, rectangles, one polygon for the diamond, after a GPU screenshot showed stacked bars reading as a plus sign).

07The rendering engine

The engine's problem statement: draw an endless, editable voxel world at a few milliseconds a frame when the host language pays heavily per call. The answer follows the direction the industry has taken since the mid-2010s: compute on workers, GPU-driven submission, and as few API calls per frame as possible.

08The stack, and why Python

VoxelBlox is Python 3.13 with pygame for the window and input, moderngl for OpenGL (3.3 core, 4.3 where the card has it), and numba compiling the heavy loops over numpy arrays. There are no asset files: every texture, model, sound and piece of music is generated in code at load.

Python is slow per call, so the whole design follows one rule: Python touches single blocks; numba touches many. Generation, lighting, meshing, raycasts over many rays and bulk queries are compiled kernels over arrays. Python orchestrates, handles the player's own edit, and issues as few GL calls as possible. A browser port was researched and declined; Windows + Python + numba was a deliberate choice.

09The world's data

The world is a grid of columns, 16×16 blocks wide and 384 tall (y from −128 to 255: under y 0 lie 128 blocks of deep rock, magma and mantle). A column is keyed by integer cells, (x >> 4, z >> 4), never by Python object identity (Python reuses ids, and a sister project once collapsed 8 missiles into 2 that way).

At render distance 32 the raw columns would be 1,114 MB. A ColumnStore keeps columns beyond the near ring zlib-packed (blocks about 30×, meshes about 3×) and unpacks them on read: 181 MB.

10Generation and streaming off the main thread

The terrain kernel, the light flood and the mesher are all compiled nogil, so they run truly in parallel on a thread pool (cores − 2, at most 8) while the main thread gathers inputs and integrates results. Output is byte-identical to a single thread. Loading at render distance 16 went from 1.97 s to 0.78 s.

Worker threads (nogil) terrain kernel light flood column mesher nearest ground first Main thread feature passes the player's own edit: re-meshed the same frame upload only what changed cull columns (numpy) build the command list GPU one arena buffer MultiDraw ElementsIndirect depth test does the occlusion
Where the work runs. The main thread never generates: the compiled kernels do it on the workers.

Two lessons shaped the scheduling:

11Meshing: 12 bytes a vertex

The mesher takes a column padded by one cell on every side (258 × 18 × 18), so faces on the border cull against the neighbours and ambient occlusion and smooth light see across the seam. It emits 4 vertices a visible face, each six int16:

// 6 x int16 = 12 bytes a vertex
x, y, z      column-local corner, in 1/16 block (shaped blocks need the sixteenths)
layer        texture layer in a sampler2DArray (every texture generated at load)
info         face:3 | ao:2 | sky:4 | flow:4 | glow:1 | fluid-surface:1
block light  0-15, flooded by the light kernel
Not built yetGreedy merging. The shader's UVs are already world-space and tile with fract(), so merged quads would drop in unchanged, but so far vertex count hasn't been what limits a frame.

12The terrain in one draw call

In Python, every GL call costs real CPU time, so the aim is to keep terrain to a handful of calls. All column meshes live in one arena buffer, sub-allocated first-fit in spans of 64 quads and doubled on the GPU when full. Each frame, numpy culls the columns against the view (six planes against every column's box at once), sorts them near-first, and writes one indirect command per visible column:

# one row per visible column: (count, instances, first index, base vertex, base instance)
cmd = np.zeros((n, 5), dtype=np.uint32)
cmd[:, 0] = (quads - wet) * 6        # the dry quads' indices
cmd[:, 1] = 1
cmd[:, 3] = base                     # where this column's vertices start in the arena
cmd[:, 4] = np.arange(n)             # baseInstance -> this column's camera-relative offset
vao.render_indirect(commands, moderngl.TRIANGLES, count=len(cmd))

Each column's camera-relative offset is a per-draw attribute read through baseInstance, so the floating origin survives without one uniform per column. The whole terrain is one glMultiDrawElementsIndirect; water is a second, sorted far-first for blending, from the wet quads kept at the end of each column's mesh.

One arena buffer free free Indirect commands (visible columns, near first) cmd 0 → col A cmd 1 → col D cmd 2 → col F one call draws them all
Column meshes sub-allocated in one buffer; culled columns simply get no command.

The per-column path (GL 3.3: one buffer and one draw per column) stays as the fallback, and a check requires the two to render identical pixels (0 of 57,600 differ). Paired smoke runs: terrain draw 2.01 / 2.17 ms per column → 1.78 / 1.77 ms multi-draw, and the p99 frame 11.0 / 8.8 ms → 5.7 / 5.9 ms.

Upload only what changed. A sister project once re-uploaded a 112 MB terrain every frame while streaming and fell from 40–50 FPS to 1–4. Here a new mesh is written into its span with a byte offset, and nothing else moves.

13The far land, and the holes a jet outruns

Beyond the meshed columns, a far land of coarse height tiles (a region each) runs out to the horizon, built on the workers from the same generation rule the columns follow. Towns, cities and megastructures appear on it: each tile samples the heights of its structures' own voxel programs, coloured by their top blocks.

columns meshed not meshed yet: filled by the holes pass far land tiles (one indirect call) flying fast, the ring ahead isn't meshed yet: the far tiles cover it, so there's no sky hole.
The far land fills exactly where the near columns aren't ready, and nowhere else. A check requires every pixel on a drawn column to be identical with the fill on and off.

14Occlusion: letting the depth buffer do it

Eniko Fox's excellent software-rendered occlusion culling in Block Game rasterises nearby opaque blocks into a small CPU depth buffer and skips hidden chunks, cutting a lot of chunks (50–60% above ground, up to 95% in caves). It's a great fit for an engine that pays per chunk it draws.

VoxelBlox made the opposite call, for two reasons:

  1. The draw calls are already gone. All visible terrain is one indirect call, so skipping a column saves only GPU vertex work, not CPU time per chunk. The CPU cost of drawing doesn't grow with the number of columns in view.
  2. A CPU occlusion test where the GPU writes depth was a bug as well as a cost. In the sister project, VoxelStrike, machines blinked as they crested hills: the CPU's answer lagged or disagreed with what the GPU drew. The rule since: where the GPU writes depth, the CPU never tests occlusion.
Where the article's idea still pointsTwo things would fit the rule if a GPU measurement ever shows terrain vertices are the bound: cave culling by connectivity (flood each section's air at mesh time, record which faces see each other, and walk only open air from the camera: conservative, on the workers, with no depth and so no blinking), and GPU occlusion (a compute pass testing boxes against a depth pyramid of the GPU's own depth, writing the indirect command list directly). Neither is built: measure first.

15Frame pacing and the CPU side

The first windowed profile had p50 16.6 ms but p99 35–57 ms, while the logged work summed to ~6 ms. Every spike was in flip: the driver's windowed vsync stalled 2–4 refreshes. With vsync off, p99 was 5.6 ms. So the game paces itself: sleep, then spin the last 1.5 ms, at the monitor's rate, and the result was 15.15 ms intervals with 0.00 jitter over 1,125 frames.

On the CPU, profiling ranked things nobody would guess:

Net result on the reference smoke world (render distance 8): frame p50 3.88 → 2.69 ms, draw 2.43 → 1.33 ms.

16All the numbers

WhatBeforeAfterHow
Frame p50, smoke world3.88 ms2.69 msmulti-draw, batched posing, resting creatures, cached HUD
Frame p99, paired smoke runs11.0 / 8.8 ms5.7 / 5.9 msterrain in one indirect call
Windowed p9935–57 ms5.6 msown frame pacer instead of driver vsync
HUD draw4.2 ms1.1 msflat array, cached layouts
Distant terrain calls, 4 km2281tiles in an arena, one indirect call
Holes in view, jet at 18 chunks4.8%0.01%the holes pass (u_cover)
Ground drawn ahead of the jet (p50)165–178 m1,800 mthe holes pass
A dig, deep underground148 ms8 msre-mesh only what the edit dirtied; deep zone
Load, render distance 161.97 s0.78 snogil kernels on a thread pool
Memory, render distance 321,114 MB181 MBfar columns zlib-packed
City worst frame224 ms18 msperf_tour found a per-building search
Metropolis worst frame667 ms20 msthe same, compiled and spread out

All figures measured on the development machine and recorded in the project's engineering log with the run that produced them.

17Patterns from adjacent projects

Several of the rules here were paid for first in sister projects. Each one generalises well beyond games.

18Should you copy this?

Honestly: copy the techniques and the habits; think hard before copying the language. Here's each piece graded for someone starting their own voxel game.

PieceVerdictWhy
One buffer + multi-draw indirectKEEPThe standard modern answer to draw-call cost, in any language. The single biggest win here.
Kernels off the main threadKEEPGeneration, light and meshing as pure functions over arrays, run on workers, nearest ground first. It's language-independent and it removes most hitches.
Pure-hash generation, journal saveKEEPAny chunk regenerates identically in any order, and a save stores only what the player changed. That makes streaming, saves and multiplayer simple.
Floating originKEEPCheap, and it removes a whole class of far-from-origin jitter bugs for good.
Your own frame pacerKEEPWindowed vsync stalled 2–4 refreshes here. Measure your flip before you blame your code.
The far land + holes passKEEPA cheap horizon, and the fill trick means fast travel never shows sky through the ground.
Witnessed checks, perf toursKEEPThe habit that paid most: every fix was measured through the player's path and every check shown to fail.
Python + numba as the engineONLY IFIt works because the heavy loops are compiled and the draw calls are gone. But everything left in Python (creatures, UI, feature passes) costs many times what it would in C#, Rust or C++, first-run compiles are slow, and shipping is heavier. Choose it if you love Python or want to learn the ideas. For a commercial game, Godot, Unity, Bevy or a C++ engine will get you there with less effort.
Columns as the draw and cull unitIMPROVEA 16×384 column is simple, but it's coarse. 16³ sections cull caves and empty sky separately and enable connectivity culling.
12-byte vertices, no greedy mergeIMPROVEFine so far, but greedy meshing or vertex pulling with packed 4–8-byte faces would cut memory and vertex work several times.
CPU occlusion cullingSKIP HEREIt's a big win when you pay per chunk. With one indirect call you don't, and a CPU answer that disagrees with the GPU's depth makes things blink. Prefer connectivity culling or GPU-side tests.

19What's missing, and what I'd do differently

20What's next: faster, then finer

Two research tracks, each item with a measurement to take first and a condition that stops it. Nothing here ships until it beats the current path on more than one machine and matches its pixels.

Faster: first inside Python, then out of it

  1. Measure better first. Build a GPU lab tool (per-pass GPU time, quads, overdraw, arena use for every screenshot scene), get a second, weaker baseline machine, and gate frame budgets per system so a regression fails a check instead of being noticed by feel.
  2. Sections and cave culling, then greedy meshing and packed vertices, then GPU-driven culling (a compute pass writing the indirect list), each only if the GPU is shown to be the bound.
  3. The known CPU cost: the per-column feature passes still on the main thread while streaming, and the per-object Python work (creatures, people, vehicle tools) that grows with the scene.
  4. Leaving Python, piece by piece. Replace the numba kernels one at a time with native ones (Rust or C++, called from Python), each behind a setting beside its twin and required to produce byte-identical output: the mesher first, since its reference oracle already exists. Only when the kernels are native, profile again and decide whether the renderer and streaming core should follow. A full port to an engine would be the last resort, because it throws away the tools and checks that made this work.
  5. A second graphics API (wgpu: Vulkan, Metal, DX12) behind the same renderer interface, if macOS or GPU culling ever needs it.

Finer: less blockiness, built on what already works

21Principles

  1. The instrument is wrong more often than the engine. A CPU meter read the game at 0% (a truncated process handle); a stall blamed on the code was the driver's vsync. Check the tool before believing its number.
  2. The cost is almost always redundant work, not a bad algorithm. Parked vehicles re-posed on every column streamed in, the deep re-meshed for the whole view, a whole terrain re-uploaded every frame. Look for work done twice before rewriting anything.
  3. A green test suite proves nothing unless a check asks the observable question. Read back pixels and positions, not internal flags.
  4. A check never shown to fail isn't a check. Revert the line it protects and watch it go red. Never derive the bar from the setting it checks.
  5. Ship a new render path off until an on-device frame rate says it's fine, and keep the old path as a pixel-identical oracle.
  6. Python pays per call. Batch everything that repeats: one indirect draw, one numpy pass for all creatures, a flat array('f') for the HUD instead of 800,000 tuples.
  7. Never key anything by object identity, and never draw from a global random stream during play. Both produced bugs that only showed up much later.
  8. Where you put a player, ask for room and for ground. "Not inside a block" alone once put a driver on top of the hill above the tunnel they'd just dug.
  9. If a line only runs while drawing, only a picture tests it. Take a screenshot after touching the draw path, and have a check read the code for the mistakes pictures find.
  10. Look before you polish. A rough software drawing of a screen showed more in a minute than a day of reading its code.
  11. Physics needs small steps. Anything fast moves further in one tick than the thing it should hit is wide. Split the step.
  12. “It exists in the code” isn't “it exists in the world”. Count what actually gets generated.
  13. Give a learner the concepts, not just the numbers. A simple scorer can't discover “this and that” on its own.
  14. Code that only runs while drawing needs its own detection. Either render in the test or read the source for the failure shape. Neither happens by default.
  15. Test the population, not only the unit. A generator whose every part passes its checks can still never produce a given output. Measure what comes out.
  16. Budget what you can't see. A render cost cap in the gate stopped a 7× regression that no functional check would have noticed.
  17. Discrete simulation needs step control. Anything that moves further in a tick than the thing it should hit is wide will tunnel. Substep adaptively.
  18. Give models the domain concepts. Feature engineering beat more training for a simple learned player.