A case study in architecture and engineering practice: an open-world voxel sandbox with an endless
editable world, no hand-made assets, peer-to-peer multiplayer and a mod API, built and run by one developer. It
covers the system's shape, the decisions and their trade-offs, the verification strategy that keeps a large
codebase honest, the incidents that shaped it, and how the rendering engine meets its frame budget from Python.
Last updated 30 September 2026
Recent changes (30 September 2026)In-world physics games (bowling, golf, basketball, darts,
pool) with discoverable play spots; eighteen screen games with rules pages and casino-accurate blackjack and roulette;
a post-round coach in every game; learning by doing (topics taught by playing, every building and person teaching
something); procedural audio for every game, venue and job; populated venues; banking and a simulated stock market;
smooth rail gradients; a static check class for draw-path-only code; eight incident reviews.
622
automated checks in the gate, each witnessed failing on a revert
~94k
lines of Python across 141 engine and 92 content modules
0
asset files: every texture, model, sound and building is generated
1
draw call for all visible solid terrain (GL 4.3 multi-draw indirect)
00What it looks like
Screenshots from the game itself, taken through the real GPU path at 1280×720 (the on-screen guide and HUD
left in). Click any one for full size.
A city's streets: every building a voxel program, streamed in and drawn by the one indirect call.The far land: coarse tiles to the horizon in one draw, the pyramids' shapes sampled from their own programs.Generation is a pure hash of the coordinates: a canyon regenerates identically, in any order.Fluids drawn by level: corners share the mean surface height, so runs slope and meet exactly.Underground: the deep is meshed only in a zone round the player that follows the tunnel.Vehicles and machines are instanced part models, posed in one numpy pass.A rocket launch: one of the places the perf tour measures mid-flight.An old town's cobbled street, with shaped blocks lit like the cubes round them.A theme park, streamed and drawn like everything else.Towers seen from above: the view is culled per column in numpy, then drawn front to back.
Recent work: games in the world
The latest captures: physics games played in the world itself, and the screen games at the venues that host them. Same GPU path, 1280×720.
Bowling in the world: a physics-driven ball on a real lane. The tag over the pins is the play-spot marker, found from loaded chunks once a second.Every game opens on a rules page laid out by the same function the layout checks call: how to play, scoring, thresholds, real-world context.Blackjack to casino rules: six-deck shoe, 3:2 blackjack, doubling, dealer peek. Suits are composed from view primitives.Home run derby: timing maps to exit direction and distance against a 400/330 ft fence; fouls beyond 45 degrees.Hockey shootout: goalie model with glove and blocker sides, goal light, per-shot outcome logging.Hospital triage: vital signs derived per patient and gated never to contradict the correct department.Fire-station hose drill: procedural flames, smoke and steam within a gated per-frame shape budget.
01The system at a glance
VoxelBlox is an open-world voxel sandbox: dig, build, drive, fly, travel to orbit. Three constraints shape
everything else. The world is endless and fully editable. Content is procedural and code-defined: there
are no asset files, so a texture, a sound effect, a town and a cruise ship are all programs. And the codebase is
maintained by one developer, so the architecture has to keep a large surface area safe to change without a
team's worth of review.
The layers. The verification layer is drawn alongside rather than beneath: it exercises every layer
through the same entry points the player uses.
Much of the design is a single-developer answer to problems a studio would solve with process and headcount.
Where a team relies on code review to catch a broken invariant, this codebase encodes the invariant in one function
or one registry and gates it with a check. Where a team would keep architecture knowledge in people's heads, it is
kept in an engineering log and a lessons file that the gate itself keeps in step with the code.
02Architecture decisions and trade-offs
The significant decisions, in the spirit of lightweight architecture decision records: what was chosen, what it
costs, and the alternative it was weighed against. Each has a dated entry in the engineering log with the
measurements behind it.
Decision
Why
Cost / alternative
Python + numba, not an engine
Fast iteration and one language end to end; the hot loops compile to
native code and release the GIL, so they run truly parallel on worker threads.
Everything left in Python
(creatures, UI, feature passes) costs several times what it would in C#, Rust or C++. Godot, Unity or Bevy would
have been the conventional choice. A browser build was researched and declined.
Content as code, through registries
No asset pipeline; every building is a small program over a
builder API, and mods register content the same way core does. Registries are append-only and saves store names,
so content can grow and be renamed without breaking a save.
Content changes are code changes, with code's
review cost. Visual iteration needs tooling (a design preview, a gallery), which had to be built.
Deterministic generation, the journal as the save
Generation is a pure hash of coordinates, so any
region regenerates identically in any order. A save stores only the cells a player changed.
Every
generation change would silently alter existing worlds, so each one ships behind a per-world option: on for new
worlds, off for old saves. In effect, feature flags with a backward-compatibility policy.
One edit API for the world
Every change goes through one call that journals it, supports undo and
produces the multiplayer delta. Networking and persistence can't drift from each other.
No shortcut writes,
even in tools. Bulk generation is the single sanctioned exception.
Player intents as named actions
Anything that changes state (a purchase, a bank deposit, a game
button) is a registered action the host executes. It's the command pattern, and it makes multiplayer safe by
construction.
Slightly more ceremony for small UI features.
One rule, one function
A rule the UI displays and the model enforces lives in one function: layout,
recipes, stats, and now each game's rules page. The drawing and the checks call the same code.
More
indirection in UI code, in exchange for UI and model never disagreeing.
Explanation as a system, not copy
Rules, scoring, real-world context and coaching live in
registries that every surface reads: the rules page, the in-world rules card, the Almanac, the Learn tab. A game
records what happened (notes); a coach function turns that into feedback after the round, stored with the
player.
Every new game needs its coach and notes written; a gate requires them, and requires the coach to
say different things for careful and careless play.
GPU-driven terrain, no CPU occlusion
One multi-draw-indirect call removes per-chunk CPU cost, which
is the thing Python can least afford. The depth buffer does occlusion.
Vertex work isn't culled early;
connectivity or compute-based culling are ready options if a measurement ever shows the GPU is the bound.
03Verification strategy
With one developer and a fast-moving codebase, the test suite has to catch what review would. Its central idea is
borrowed from mutation testing and applied to every assertion rather than sampled across the codebase.
Witnessed checks (mutation testing by hand). Every assertion is shown to fail: a tool reverts the one line
it protects, runs the check, requires it to go red, and restores the line. Tools like mutmut or Stryker
sample mutations to score a suite. Here each check is proven individually when it's written. A check that can't
be witnessed is either redundant or asserting the wrong thing, and both turn up more often than you'd expect.
End-to-end through the player's path. Checks drive keyboard and mouse events through the real event
handler and frame loop, then assert on observable state: positions, pixels, scores. Internal flags are not
evidence.
Oracles and golden images. A slow reference mesher in plain Python is the oracle for the compiled one
(face for face). Every fast render path keeps its fallback, and a check requires identical pixels between them.
Property and robustness testing. Seeded fuzzing plays the game through the input path, and a crash
leaves a replayable recording. The compiled kernels run under bounds checking on fuzzed inputs to prove no
out-of-range write.
Simulation-based balancing. Each of the eighteen screen games has a careful bot, a careless bot and a
learned player. Difficulty is tuned by cross-entropy optimisation against target star ratings, and a gated check
requires the careful player to beat the careless one by a margin and the learned player to stay human-fair.
Static checks for what runtime can't reach. Some code only runs while a frame is drawn, which most checks
never do. The source is parsed and checked for known failure shapes (see the first incident below).
Docs as code. A documentation check fails the gate when a documented command, check name or scene no
longer exists, and a code-map check fails when a module doesn't state its purpose.
Why this matters more than coverageLine coverage says a line ran. A witnessed check
says a specific wrong version of that line would have been caught. For a solo codebase that changes daily, the second
is the property that lets you refactor without fear.
Designing for learning
The game is meant to teach how things work, so learning is engineered like any other feature: measured, gated and
fed by the game's own data.
Feedback from what actually happened. Each game model records the moments a teacher would comment on
(a blackjack decision against basic strategy, a swing 90 ms early, a penalty taker who always goes left). After
the round a coach turns them into one to three specific lines: "9 of 15 decisions went against basic strategy,
e.g. DOUBLE on 18 against a 7; the book says STAND." The same lines return on the next rules page.
Coverage as a gated property. Every building type and every person in the world teaches a topic; every
game has rules, scoring, a real-world note and a coach; every vehicle job has a note on how the real job works.
Each is a check, so new content can't quietly arrive unexplained. That check caught eleven building types and
thirteen staff roles added without anything to teach.
Learning by doing. Topics are taught at the moment they become relevant: a first game of darts teaches
darts, a first bank deposit teaches how banks lend, walking into an aquarium teaches why its windows are 30 cm
thick.
Honest numbers, derived rather than typed. The slot machine uses a real reel strip of 22 stops and a pay
table. Its payback (93.1%) is not a constant someone typed: it is computed by enumerating all 10,648 combinations of
stops, and the machine's glass, the rules page and a 30,000-spin check all read that one function. Because the only
skill in a slot machine is knowing when to stop, the star rating rewards keeping to a loss limit. It plays with
pretend chips the casino hands out once a game day. The bowling scoreboard follows the same pattern: one pure
function turns the balls into the marks and running totals an alley monitor shows, and a total waits for its
strike or spare bonus exactly as the real sheet does. It is drawn at a readable minimum size, because a
3.2-metre sign seen from the approach is only about 90 pixels wide.
04Observability and performance engineering
Frame budgets, measured in the real loop. A performance tour drives one world through the actual frame
loop to its most expensive places (a metropolis, a launch in progress, a jet low over mountains) and reports
percentiles, the subsystem that dominates slow frames and CPU time by function. A fixed benchmark is saved per
machine so hardware can be compared.
Spike reports. Any frame over 250 ms captures what it was doing and its stack. Frame-time logs attribute
each spike to the game, the OS, the display or a travel load.
Crash reports you can replay. A crash writes a report with the recorded input and a snapshot of the
save. A replay tool re-runs it headless and exits 0 when it reproduces, which turns a player's bug report into a
test case.
Profile to rank, instrument to size. The profiler finds the candidates; targeted timers give the numbers
that go in the log. In practice the cost is almost always redundant work, not a poor algorithm: the next sections
show a per-building search every two seconds costing a 667 ms frame.
Resource governance on shared hardware. The development GPU also serves a desktop and local AI inference.
A machine-wide lock allows one rendering run at a time, helper processes die with their parent, and every run
reports the driver faults logged during it. It's an SRE habit (know your blast radius) applied to a workstation.
05Delivery model: from request to done
The workflow is deliberately lightweight, but it has the traceability of a larger process:
Requirements in the stakeholder's words. Each request becomes a row in an open-items file, quoted
verbatim, with its status, the checks that prove it and how to try it. A row closes when the feature is confirmed
in the product, not when tests pass: definition of done as acceptance, not as a build status.
A branch per change, trunk kept releasable. Merges require the targeted checks, the documentation check
and the code-map check to pass. The full suite runs in parallel workers on demand.
An engineering log as the decision record. Every change ends with a dated entry holding the evidence (the
measurements, what was tried, what was rejected), and a one-line lesson, tagged by area, in a lessons file that's
read before touching that area again.
Every job leaves a tool. A question that was hard to answer once becomes a tool that answers it always:
41 so far, all on one result protocol with JSON output, saved baselines and comparisons, and one local web page
that runs them.
06Incident reviews
Blameless, short, and focused on the detection gap: why the existing checks didn't catch it, and what now does.
A draw-path-only crash that every check missed
Impact: the game would have crashed on the first frame of play. It was caught by pre-release screenshots
before reaching a player. Root cause: new code in the draw path called the settings lookup with a default
value; the lookup takes only a name. Detection gap: all checks for the feature exercised its logic directly,
and no check draws a frame during play, so that line never ran under test.
Remediation: a static check parses the source and fails on any settings lookup with the wrong
arity, and a practice of capturing a GPU screenshot after any change to the draw path.
GPU driver resets during test runs
Impact: repeated graphics driver resets on the development machine, taking down every program using the
GPU. Root cause: the Windows event log showed 1,334 driver errors in one day, all during test runs: up to
seven concurrent runs, their workers each creating and destroying 30 to 50 OpenGL contexts, beside an inference
server holding 7 GB of video memory. The driver's context-switching engine failed. Detection gap: the dying
workers looked like crashes in whatever check they happened to be running, so each failure first read as a
different bug.
Remediation: one GL context per process, at most four rendering workers, a machine-wide lock,
workers bound to their parent's lifetime, and a driver-fault count printed at the end of every run.
Collision tunnelling in the bowling physics
Impact: a perfect pocket roll knocked down two pins. Root cause: at 8 m/s the ball advanced 40 cm
per tick against a pin 13 cm wide, so most pins were stepped over. Detection gap: none. The first
end-to-end check caught it before merge, which is the point of writing that check first.
Remediation: adaptive substeps capping any fast body's movement at 4 cm (up to 24 per tick).
A pocket roll now takes nine pins, and just off the head pin is a strike.
A market with no volatility
Impact: simulated share prices barely moved. Root cause: prices summed four draws from CRC32 of the
time; CRC32 is linear, so the draws were correlated and their sum nearly constant. Detection gap: the check
asserted determinism (the same prices for everyone), not the distribution.
Remediation: a seeded PRNG per block of 256 ticks, which is deterministic across clients and
swings 0.3–2.4× over a long run. The check now covers the range as well.
Content that existed in code and never in a world
Impact: two new building types never generated. Root cause: the planetarium lost every weighted draw
because of a bias in the hash on the lot grid, and the garden centre never fit a town lot. Detection gap:
unit checks built each building directly; nothing measured the population.
Remediation: adjusted weights and footprints, and a census tool that plans 500+ settlements
and reports types that never appear, and why. A type that never generates is treated as a defect.
Eighteen screens nobody had looked at
Impact: screen games shipped with placeholder visuals: the home run derby was two white rectangles and a dot.
Root cause: the games were verified by bots and checks, which don't judge appearance. Detection gap:
looking required the GPU, which is rationed.
Remediation: a software renderer for the games' view description, which draws any game's screen
without a GPU. The first redraw cost 475 shapes a frame; lines and fields became polygons (64 shapes), and a gated
budget now caps every game screen at 160.
A learning agent that couldn't express the policy
Impact: after blackjack moved to real casino rules (six-deck shoe, 3:2 blackjack, doubling, dealer peek),
the learned player lost 10 chips a round, and the balancing gate failed. Root cause: a linear scorer over raw
features can't represent a conjunction such as “12–16 and the dealer shows 7+”. Detection
gap: none. The gate did its job.
Remediation: the domain concepts as features. The agent now plays level with textbook basic
strategy, which over 300 rounds ends about 3 chips down: the house edge, as in the real game.
A font without the glyphs
Impact: card suits would have rendered as empty boxes. Root cause: the UI font atlas covers ASCII
only; the default font's metrics for ♥ ♦ ♣ ♠ were one identical box. Detection gap:
caught by measuring before building.
Remediation: suits composed from the view's own primitives (discs, rectangles, one polygon
for the diamond, after a GPU screenshot showed stacked bars reading as a plus sign).
07The rendering engine
The engine's problem statement: draw an endless, editable voxel world at a few milliseconds a frame when the host
language pays heavily per call. The answer follows the direction the industry has taken since the mid-2010s:
compute on workers, GPU-driven submission, and as few API calls per frame as possible.
08The stack, and why Python
VoxelBlox is Python 3.13 with pygame for the window and input, moderngl for OpenGL
(3.3 core, 4.3 where the card has it), and numba compiling the heavy loops over numpy arrays.
There are no asset files: every texture, model, sound and piece of music is generated in code at load.
Python is slow per call, so the whole design follows one rule: Python touches single blocks; numba touches
many. Generation, lighting, meshing, raycasts over many rays and bulk queries are compiled kernels over arrays.
Python orchestrates, handles the player's own edit, and issues as few GL calls as possible. A browser port was
researched and declined; Windows + Python + numba was a deliberate choice.
09The world's data
The world is a grid of columns, 16×16 blocks wide and 384 tall (y from −128 to 255: under y 0
lie 128 blocks of deep rock, magma and mantle). A column is keyed by integer cells,
(x >> 4, z >> 4), never by Python object identity (Python reuses ids, and a sister project once
collapsed 8 missiles into 2 that way).
A cell is two bytes and a byte: a uint16 block id and a uint8 state (facing,
water level, upside-down). Anything that stores or sends blocks carries both: the world, the journal, saves, the
network, blueprints.
Light is a byte per cell: sky and block light, flooded by a numba kernel.
Floating origin: every vertex is column-relative, and the camera offset is subtracted in float64 on the
CPU. The view matrix is rotation only, so float32 never sees a large coordinate. A check renders the same scene at
the origin and at 30 million blocks out: 93 of 57,600 pixels differ.
Endless by construction: generation is a pure hash of int64 coordinates, so any column regenerates
identically, in any order.
The journal is the save: per column, only the cells that differ from generation. What you build stays built
across unload, reload and save.
At render distance 32 the raw columns would be 1,114 MB. A ColumnStore keeps columns beyond the near
ring zlib-packed (blocks about 30×, meshes about 3×) and unpacks them on read: 181 MB.
10Generation and streaming off the main thread
The terrain kernel, the light flood and the mesher are all compiled nogil, so they run truly in
parallel on a thread pool (cores − 2, at most 8) while the main thread gathers inputs and integrates results. Output
is byte-identical to a single thread. Loading at render distance 16 went from 1.97 s to 0.78 s.
Where the work runs. The main thread never generates: the compiled kernels do it on the workers.
Two lessons shaped the scheduling:
Ground where the player is comes first. Flying fast at render distance 18, generation once held every
worker slot until the whole ring (~1,300 columns) was generated, and the ground within 4 columns fell to 0% drawn.
Now never-drawn ground near the player gets half the workers as soon as it's ready: 100% drawn at 40 m/s, and the
whole view in 2–3 s when stopped.
The deep is meshed only round the player. Digging straight down once lowered the mesh floor for the
whole view and re-meshed about 1,000 columns on the main thread: a one-second freeze per step. Now the deep is meshed
in a zone 4 columns round the player that follows a tunnel, and a dig re-meshes only what that edit dirtied:
148 ms → 8 ms a dig.
11Meshing: 12 bytes a vertex
The mesher takes a column padded by one cell on every side (258 × 18 × 18), so faces on the border cull
against the neighbours and ambient occlusion and smooth light see across the seam. It emits 4 vertices a visible face,
each six int16:
// 6 x int16 = 12 bytes a vertex
x, y, z column-local corner, in 1/16 block (shaped blocks need the sixteenths)
layer texture layer in a sampler2DArray (every texture generated at load)
info face:3 | ao:2 | sky:4 | flow:4 | glow:1 | fluid-surface:1
block light 0-15, flooded by the light kernel
Per-vertex AO from three neighbours; a quad whose AO is anisotropic is emitted rotated by one vertex so its
diagonal follows the light (otherwise it shows as seams).
Fluids are drawn at their level: each top corner is the mean surface height of the fluid cells sharing it,
so neighbours meet exactly and runs slope downhill. The flow direction (8 ways, or falling) rides in 4 spare bits,
and the shader scrolls runs, pours falls and churns lava.
Shaped blocks (stairs, doors, furniture, wedges for 45° slopes) are box models turned by the cell's
state, lit with the same corner light as a full cube so a hillside of slopes shades like the blocks round it.
One shared index buffer (0,1,2, 0,2,3 per quad) serves every mesh.
A slow reference mesher in plain Python is the oracle: a check requires the numba mesher to match it face
for face.
Not built yetGreedy merging. The shader's UVs are already world-space and tile with
fract(), so merged quads would drop in unchanged, but so far vertex count hasn't been what limits a frame.
12The terrain in one draw call
In Python, every GL call costs real CPU time, so the aim is to keep terrain to a handful of calls. All column meshes
live in one arena buffer, sub-allocated first-fit in spans of 64 quads and doubled on the GPU when full. Each
frame, numpy culls the columns against the view (six planes against every column's box at once), sorts them
near-first, and writes one indirect command per visible column:
# one row per visible column: (count, instances, first index, base vertex, base instance)
cmd = np.zeros((n, 5), dtype=np.uint32)
cmd[:, 0] = (quads - wet) * 6 # the dry quads' indices
cmd[:, 1] = 1
cmd[:, 3] = base # where this column's vertices start in the arena
cmd[:, 4] = np.arange(n) # baseInstance -> this column's camera-relative offset
vao.render_indirect(commands, moderngl.TRIANGLES, count=len(cmd))
Each column's camera-relative offset is a per-draw attribute read through baseInstance, so the
floating origin survives without one uniform per column. The whole terrain is oneglMultiDrawElementsIndirect; water is a second, sorted far-first for blending, from the wet quads kept at
the end of each column's mesh.
Column meshes sub-allocated in one buffer; culled columns simply get no command.
The per-column path (GL 3.3: one buffer and one draw per column) stays as the fallback, and a check requires the
two to render identical pixels (0 of 57,600 differ). Paired smoke runs: terrain draw 2.01 / 2.17 ms per column
→ 1.78 / 1.77 ms multi-draw, and the p99 frame 11.0 / 8.8 ms → 5.7 / 5.9 ms.
Upload only what changed. A sister project once re-uploaded a 112 MB terrain every frame while streaming and
fell from 40–50 FPS to 1–4. Here a new mesh is written into its span with a byte offset, and nothing else
moves.
13The far land, and the holes a jet outruns
Beyond the meshed columns, a far land of coarse height tiles (a region each) runs out to the horizon, built on
the workers from the same generation rule the columns follow. Towns, cities and megastructures appear on it: each
tile samples the heights of its structures' own voxel programs, coloured by their top blocks.
One call for all tiles. The tiles live in an arena too, with a shared index buffer (fine grid, coarse
grid, box pattern) and one render_indirect. At 4 km: 228 calls → 1, CPU 0.5 → 0.2 ms,
pixel-identical to the per-tile path.
The holes pass. Flying the jet at 112 m/s at render distance 18, the ground was only drawn 98–114 m
ahead (p5; p50 165–178 m): the far land's shader discards everything inside the render distance because "the columns are drawn
there", and where they weren't yet, there was sky. The fix draws the far tiles again inside the ring, depth-tested
with the columns' own projection, keeping only fragments over a column that isn't on the GPU yet. That's an R8
texture with one texel a column (u_cover), rebuilt only when a mesh comes or goes. Holes went from
4.8% of the view to 0.01%, and the ground is drawn to the full 1,800 m ahead (p50).
The far land fills exactly where the near columns aren't ready, and nowhere else. A check requires every
pixel on a drawn column to be identical with the fill on and off.
14Occlusion: letting the depth buffer do it
Eniko Fox's excellent software-rendered
occlusion culling in Block Game rasterises nearby opaque blocks into a small CPU depth buffer and skips hidden
chunks, cutting a lot of chunks (50–60% above ground, up to 95% in caves). It's a great fit for an engine that pays
per chunk it draws.
VoxelBlox made the opposite call, for two reasons:
The draw calls are already gone. All visible terrain is one indirect call, so skipping a column saves only
GPU vertex work, not CPU time per chunk. The CPU cost of drawing doesn't grow with the number of columns in view.
A CPU occlusion test where the GPU writes depth was a bug as well as a cost. In the sister project,
VoxelStrike, machines blinked as they crested hills: the CPU's answer lagged or disagreed with what the GPU
drew. The rule since: where the GPU writes depth, the CPU never tests occlusion.
Where the article's idea still pointsTwo things would fit the rule if a GPU measurement ever
shows terrain vertices are the bound: cave culling by connectivity (flood each section's air at mesh time, record
which faces see each other, and walk only open air from the camera: conservative, on the workers, with no depth and so
no blinking), and GPU occlusion (a compute pass testing boxes against a depth pyramid of the GPU's own depth,
writing the indirect command list directly). Neither is built: measure first.
15Frame pacing and the CPU side
The first windowed profile had p50 16.6 ms but p99 35–57 ms, while the logged work summed to ~6 ms. Every spike
was in flip: the driver's windowed vsync stalled 2–4 refreshes. With vsync off, p99 was 5.6 ms. So the game
paces itself: sleep, then spin the last 1.5 ms, at the monitor's rate, and the result was 15.15 ms intervals with
0.00 jitter over 1,125 frames.
On the CPU, profiling ranked things nobody would guess:
The HUD: 800k tuple appends plus np.array over them was 2.6 ms of a 4.2 ms draw. A flat
array('f') and cached label layouts cut draw from 4.2 to 1.1 ms.
Posing creatures one by one in numpy: batched into one pass (pose_many), identical rows,
1.2 → 0.5 ms. Parked vehicles are posed once per column and uploaded together.
Resting creatures skip physics until the world changes (41% of creature ticks).
No full-screen CPU surfaces: UI and effects are GL quads and shader uniforms. A CPU HUD upload once cost
1.1 ms GPU at 1080p and 4.4 ms at 4K.
Net result on the reference smoke world (render distance 8): frame p50 3.88 → 2.69 ms, draw 2.43 →
1.33 ms.
All figures measured on the development machine and
recorded in the project's engineering log with the run that produced them.
17Patterns from adjacent projects
Several of the rules here were paid for first in sister projects. Each one generalises well beyond games.
Never let two systems answer the same question. In VoxelStrike, a sister voxel game, vehicles flickered
on hill crests because a CPU visibility test and the GPU's depth buffer disagreed. The rule since: where the GPU
writes depth, the CPU doesn't test occlusion. It's the same principle as having one source of truth in any
distributed system.
Identity is data, not memory addresses. Keying objects by Python's id() collapsed eight
missiles into two when ids were reused. Everything is keyed by coordinates or stable names.
Move deltas, not state. Re-uploading a 112 MB terrain every frame took a project from 40–50 FPS to
1–4. Upload only what changed, into its own slot: the same logic as incremental builds or change data
capture.
Test tooling is production software on the machine it runs on. In a dictation project, a test harness
launched a headless browser with a fresh profile each time. Its password manager attempted a blank-password
Windows sign-in, and the account locked eighteen times in three days before the cause was found. Harnesses now run
from one launcher with a prepared profile, and the event log is the first place to look.
Shared hardware needs governance. The GPU resets and the account lockouts have one root cause: tools that
assumed they had the machine to themselves. Serialise, bound lifetimes, report side effects.
18Should you copy this?
Honestly: copy the techniques and the habits; think hard before copying the language. Here's each piece
graded for someone starting their own voxel game.
Piece
Verdict
Why
One buffer + multi-draw indirect
KEEP
The standard modern answer to draw-call cost, in any language. The single biggest win here.
Kernels off the main thread
KEEP
Generation, light and meshing as pure functions over arrays, run on workers, nearest ground first. It's language-independent and it removes most hitches.
Pure-hash generation, journal save
KEEP
Any chunk regenerates identically in any order, and a save stores only what the player changed. That makes streaming, saves and multiplayer simple.
Floating origin
KEEP
Cheap, and it removes a whole class of far-from-origin jitter bugs for good.
Your own frame pacer
KEEP
Windowed vsync stalled 2–4 refreshes here. Measure your flip before you blame your code.
The far land + holes pass
KEEP
A cheap horizon, and the fill trick means fast travel never shows sky through the ground.
Witnessed checks, perf tours
KEEP
The habit that paid most: every fix was measured through the player's path and every check shown to fail.
Python + numba as the engine
ONLY IF
It works because the heavy loops are compiled and the draw calls are gone. But everything left in Python (creatures, UI, feature passes) costs many times what it would in C#, Rust or C++, first-run compiles are slow, and shipping is heavier. Choose it if you love Python or want to learn the ideas. For a commercial game, Godot, Unity, Bevy or a C++ engine will get you there with less effort.
Columns as the draw and cull unit
IMPROVE
A 16×384 column is simple, but it's coarse. 16³ sections cull caves and empty sky separately and enable connectivity culling.
12-byte vertices, no greedy merge
IMPROVE
Fine so far, but greedy meshing or vertex pulling with packed 4–8-byte faces would cut memory and vertex work several times.
CPU occlusion culling
SKIP HERE
It's a big win when you pay per chunk. With one indirect call you don't, and a CPU answer that disagrees with the GPU's depth makes things blink. Prefer connectivity culling or GPU-side tests.
19What's missing, and what I'd do differently
Sections and cave culling first. Split columns into 16³ sections, record at mesh time which faces of
each section can see each other through air, and walk only open air from the camera. It's conservative, it runs on the
workers, and it catches the underground case, where Block Game's article saw up to 95% of chunks hidden.
Greedy meshing or vertex pulling. The shader already tiles world-space UVs, so merged quads drop in. Vertex
pulling (face data in a storage buffer, the vertex shader expanding quads) would shrink the arena further.
A real sky-light flood. Sky light is exposure from the height map today; a flood would light overhangs and
cave mouths properly.
GPU occlusion, if measured to matter. A compute pass testing boxes against a depth pyramid of last frame's
depth, writing the indirect commands itself. The command list is already on the GPU, so it would slot in.
Measure on more than one machine. Every number here comes from one development PC (a big desktop GPU,
played over Remote Desktop). The ratios transfer; the absolute milliseconds won't.
20What's next: faster, then finer
Two research tracks, each item with a measurement to take first and a condition that stops it. Nothing here ships
until it beats the current path on more than one machine and matches its pixels.
Faster: first inside Python, then out of it
Measure better first. Build a GPU lab tool (per-pass GPU time, quads, overdraw, arena use for every
screenshot scene), get a second, weaker baseline machine, and gate frame budgets per system so a regression fails a
check instead of being noticed by feel.
Sections and cave culling, then greedy meshing and packed vertices, then
GPU-driven culling (a compute pass writing the indirect list), each only if the GPU is shown to be the bound.
The known CPU cost: the per-column feature passes still on the main thread while streaming, and the
per-object Python work (creatures, people, vehicle tools) that grows with the scene.
Leaving Python, piece by piece. Replace the numba kernels one at a time with native ones (Rust or C++,
called from Python), each behind a setting beside its twin and required to produce byte-identical output: the mesher
first, since its reference oracle already exists. Only when the kernels are native, profile again and decide whether
the renderer and streaming core should follow. A full port to an engine would be the last resort, because it throws
away the tools and checks that made this work.
A second graphics API (wgpu: Vulkan, Metal, DX12) behind the same renderer interface, if macOS or GPU
culling ever needs it.
Finer: less blockiness, built on what already works
Heightfield shapes. The 45° wedges are already a height function within a cell; generalise it to dome
caps, quarter-rounds, rounded corners and arches, one primitive for all of them. Add side slopes, posts and pillars.
Micro-blocks. A cell split into 4³ or 8³ sub-voxels for carved detail: rounded edges, domes,
railings, sculpture. Only detailed cells pay, and far away they draw as a plain block. This grows what a cell is
from (id, state) to (id, state, detail), so every system that stores or sends blocks has to carry it, with
round-trip checks.
Smooth terrain, the grid way first: slopes and corners placed automatically on generated hills, as a
per-region trait so some land stays rugged. Truly organic ground (surface nets) is a research spike for a separate
world type, not a change to this one.
Rounder vehicles and objects. Start with shading that fakes a bevel on every model part, which is a
shader-only change, then add rounded primitives (cylinders, rounded boxes, tapers, domes), then prototype a vehicle
modelled as a fine voxel grid (meshed once, with AO, and instanced), comparing screenshots and the parked-city frame
cost side by side.
21Principles
The instrument is wrong more often than the engine. A CPU meter read the game at 0% (a truncated process
handle); a stall blamed on the code was the driver's vsync. Check the tool before believing its number.
The cost is almost always redundant work, not a bad algorithm. Parked vehicles re-posed on every column
streamed in, the deep re-meshed for the whole view, a whole terrain re-uploaded every frame. Look for work done twice
before rewriting anything.
A green test suite proves nothing unless a check asks the observable question. Read back pixels and
positions, not internal flags.
A check never shown to fail isn't a check. Revert the line it protects and watch it go red. Never derive the
bar from the setting it checks.
Ship a new render path off until an on-device frame rate says it's fine, and keep the old path as a
pixel-identical oracle.
Python pays per call. Batch everything that repeats: one indirect draw, one numpy pass for all creatures, a
flat array('f') for the HUD instead of 800,000 tuples.
Never key anything by object identity, and never draw from a global random stream during play. Both
produced bugs that only showed up much later.
Where you put a player, ask for room and for ground. "Not inside a block" alone once put a driver on top of
the hill above the tunnel they'd just dug.
If a line only runs while drawing, only a picture tests it. Take a screenshot after touching the draw
path, and have a check read the code for the mistakes pictures find.
Look before you polish. A rough software drawing of a screen showed more in a minute than a day of
reading its code.
Physics needs small steps. Anything fast moves further in one tick than the thing it should hit is wide.
Split the step.
“It exists in the code” isn't “it exists in the world”. Count what actually gets
generated.
Give a learner the concepts, not just the numbers. A simple scorer can't discover “this
and that” on its own.
Code that only runs while drawing needs its own detection. Either render in the test or read the source
for the failure shape. Neither happens by default.
Test the population, not only the unit. A generator whose every part passes its checks can still never
produce a given output. Measure what comes out.
Budget what you can't see. A render cost cap in the gate stopped a 7× regression that no functional
check would have noticed.
Discrete simulation needs step control. Anything that moves further in a tick than the thing it should hit
is wide will tunnel. Substep adaptively.
Give models the domain concepts. Feature engineering beat more training for a simple learned player.