Skip to content

The shader VM

The shader VM is the part of the panel’s firmware that runs compiled shaders. It is an interpreter, but an unusual one: each instruction it reads is applied to 64 pixels before it reads the next. Reading and decoding an instruction is the slow part of any interpreter, and here that cost is shared by 64 pixels.

Tile registers. A register is not one number but a row of 64 floats, one per pixel of the tile being shaded. A program may use up to 32 of them, 8 KB in all. One set is shared by all the shaders of a script, since they run one after another.

Scalars. Three kinds of value are the same for every pixel, so they are stored once: the constants, the uniform values, and the frame registers. They sit in one array in that order.

An operand of an instruction is either a register (a different value for each pixel) or a scalar (one value for all of them).

When a script calls gfx.fill(shader):

  1. The frame budget is charged with the shader’s cost. If the frame has no budget left, the script stops with an error before any work is done.
  2. Uniforms are copied in from the script’s uniforms table, for every name the shader declared.
  3. The prologue runs once. Its instructions fill the frame registers. This is everything that depends only on uniforms and constants.
  4. The body runs for each tile. A tile is up to 64 pixels of one row. On a 64-pixel-wide panel, that is one row per tile and 64 tiles per frame.

Take the body from the earlier pages, on row 10:

r0 = gt r0, 31
r1 = div r1, 63
r2 = select r0, p2, 0
r0 = select r0, 0, p2
ret r2, r1, r0

First the VM fills in the inputs: register 0 gets the x coordinates and register 1 the row number.

r0: 0 1 2 ... 31 32 33 ... 63
r1: 10 10 10 ... 10 10 10 ... 10

Then it goes through the instructions. For each one it looks at the opcode once, and then runs a simple loop over the 64 values.

r0 = gt r0, 31 compares every x with the constant 31. The register now holds the mask:

r0: 0 0 0 ... 0 1 1 ... 1

r1 = div r1, 63 divides all 64 values by 63:

r1: .159 .159 .159 ... .159 .159 .159 ... .159

r2 = select r0, p2, 0 takes the frame register p2 (the glow, say 0.8) where the mask is 1, and 0 where it is 0. r0 = select r0, 0, p2 does the opposite, and overwrites the mask, which is no longer needed:

r2: 0 0 0 ... 0 .8 .8 ... .8 red
r0: .8 .8 .8 ... .8 0 0 ... 0 blue

ret r2, r1, r0 names the three registers that hold the colour. The VM converts 64 colours and writes them straight into the row of the canvas.

Five instructions were decoded and 64 pixels were shaded. Run interpreted, the same row would have been 64 calls into Lua.

A register can be the result of an instruction and one of its inputs at the same time, as in r1 = div r1, 63. That is safe because each pixel’s result depends only on that same pixel’s inputs.

A processor running one pixel would jump over the code it does not need. The VM cannot: the 64 pixels of a tile take different paths. So a condition is data, not control flow.

  • A comparison writes a mask register: 1 where it holds, 0 where it does not.
  • Both sides of an if are computed for all 64 pixels.
  • select then picks, for each pixel, the value from the side that pixel took.

Computing both sides is wasted work for some pixels, and is still far cheaper than leaving the 64-at-a-time loop. Computing a side a pixel did not take is harmless, because instructions only do arithmetic: nothing can fail, and nothing outside the registers is changed.

When a branch is expensive (it contains a sin, a pow, noise), the compiler puts a jmpnone in front of it. The instruction checks the mask: if it is 0 for all the pixels in the tile, it skips the block.

This pays off when a condition is the same across whole rows or large areas, such as “below the horizon” or “inside this circle”. It is purely an optimization. Running the block with an all-zero mask would change nothing that is read afterwards, so skipping it cannot change the picture.

Jumps only go forward. There is no way to express a loop, so a program always ends.

The three results are floats, normally between 0 and 1. For each pixel the VM:

  1. Clamps each channel to the range 0 to 1. A NaN becomes 0.
  2. Scales it to the canvas’s precision: 32 levels of red and blue, 64 of green.
  3. Dithers. Where the value falls between two levels, a small fixed 4 × 4 pattern decides which pixels round up and which round down, in proportion to how far between the levels it is. Across a few pixels the eye sees the in-between shade, so a slow gradient does not break into bands. The pattern depends only on the pixel’s position, so a still picture does not flicker.

There is no gamma step here. The LED driver applies its own brightness curve to everything drawn on the panel.

An interpreted shader’s colours go through exactly the same code.

A script cannot make the panel hang by uploading a heavy shader:

  • The program has no loops, so its running time is its length times the number of tiles.
  • Every opcode has a cost, and the cost of the whole body is added up when the program is loaded.
  • Each gfx.fill charges that cost to the frame. A frame may spend 1,000 units per pixel, about 0.45 seconds on a 64 × 64 panel in the worst case. One more fill past that is refused and the script stops.

Lua code is bounded differently, by counting instructions as they execute. That is not needed here because the cost is known before the first pixel.

The same shader produces the same pixels on every panel and in the reference implementation used for testing. Three things make that true:

  • sin, cos and tan are the panel’s own polynomial versions, not the C library’s, which round differently between systems.
  • The noise functions use integer hashing.
  • The firmware is built so that the processor does not merge a multiply and an add into one differently-rounded operation.

The firmware’s benchmark carries a checksum of a reference tile for this reason: a panel whose VM did not reproduce it bit for bit would fail the run. The slow library functions (pow, exp, log, asin, acos, atan) are the exception. They are the same on every panel, but can differ in the last digit from a desktop computer.

On the panel’s ESP32-S3 processor with a 64 × 64 display:

Shader Interpreted Compiled
Plasma: 18 arithmetic instructions, 6 sin/cos 28.7 µs per pixel 6.0 µs per pixel
Clouds: one four-octave fbm 24 µs 15.5 µs

From those two: an arithmetic instruction costs roughly a tenth of a microsecond per pixel and a sine about six times that. Four octaves of noise cost about 15 µs however they are called, which is why the clouds gain less. Copying the finished canvas to the LEDs takes a further 8 ms per frame for any kind of script.

The plasma at 6 µs per pixel is 25 ms per frame: about 30 frames per second, against 8 when interpreted.

The VM is small because it does little. It has no stack, no function calls, no memory to read or write besides its registers, and no loops. The compiler removes all of those before the program reaches the panel: helpers are copied into their callers, loops are written out, vectors are split into numbers. What arrives is straight-line arithmetic, which is the one thing that can be done 64 pixels at a time.