The shader VM
The shader VM is the part of the panel’s firmware that runs compiled shaders. It is an interpreter, but an unusual one: each instruction it reads is applied to 64 pixels before it reads the next. Reading and decoding an instruction is the slow part of any interpreter, and here that cost is shared by 64 pixels.
What it holds
Section titled “What it holds”Tile registers. A register is not one number but a row of 64 floats, one per pixel of the tile being shaded. A program may use up to 32 of them, 8 KB in all. One set is shared by all the shaders of a script, since they run one after another.
Scalars. Three kinds of value are the same for every pixel, so they are stored once: the constants, the uniform values, and the frame registers. They sit in one array in that order.
An operand of an instruction is either a register (a different value for each pixel) or a scalar (one value for all of them).
One fill
Section titled “One fill”When a script calls gfx.fill(shader):
- The frame budget is charged with the shader’s cost. If the frame has no budget left, the script stops with an error before any work is done.
- Uniforms are copied in from the script’s
uniformstable, for every name the shader declared. - The prologue runs once. Its instructions fill the frame registers. This is everything that depends only on uniforms and constants.
- The body runs for each tile. A tile is up to 64 pixels of one row. On a 64-pixel-wide panel, that is one row per tile and 64 tiles per frame.
One tile
Section titled “One tile”Take the body from the earlier pages, on row 10:
r0 = gt r0, 31r1 = div r1, 63r2 = select r0, p2, 0r0 = select r0, 0, p2ret r2, r1, r0First the VM fills in the inputs: register 0 gets the x coordinates and register 1 the row number.
r0: 0 1 2 ... 31 32 33 ... 63r1: 10 10 10 ... 10 10 10 ... 10Then it goes through the instructions. For each one it looks at the opcode once, and then runs a simple loop over the 64 values.
r0 = gt r0, 31 compares every x with the constant 31. The register now holds the mask:
r0: 0 0 0 ... 0 1 1 ... 1r1 = div r1, 63 divides all 64 values by 63:
r1: .159 .159 .159 ... .159 .159 .159 ... .159r2 = select r0, p2, 0 takes the frame register p2 (the glow, say 0.8) where the mask is 1, and 0 where it is 0. r0 = select r0, 0, p2 does the opposite, and overwrites the mask, which is no longer needed:
r2: 0 0 0 ... 0 .8 .8 ... .8 redr0: .8 .8 .8 ... .8 0 0 ... 0 blueret r2, r1, r0 names the three registers that hold the colour. The VM converts 64 colours and writes them straight into the row of the canvas.
Five instructions were decoded and 64 pixels were shaded. Run interpreted, the same row would have been 64 calls into Lua.
A register can be the result of an instruction and one of its inputs at the same time, as in r1 = div r1, 63. That is safe because each pixel’s result depends only on that same pixel’s inputs.
Conditions without jumps
Section titled “Conditions without jumps”A processor running one pixel would jump over the code it does not need. The VM cannot: the 64 pixels of a tile take different paths. So a condition is data, not control flow.
- A comparison writes a mask register: 1 where it holds, 0 where it does not.
- Both sides of an
ifare computed for all 64 pixels. selectthen picks, for each pixel, the value from the side that pixel took.
Computing both sides is wasted work for some pixels, and is still far cheaper than leaving the 64-at-a-time loop. Computing a side a pixel did not take is harmless, because instructions only do arithmetic: nothing can fail, and nothing outside the registers is changed.
The one jump
Section titled “The one jump”When a branch is expensive (it contains a sin, a pow, noise), the compiler puts a jmpnone in front of it. The instruction checks the mask: if it is 0 for all the pixels in the tile, it skips the block.
This pays off when a condition is the same across whole rows or large areas, such as “below the horizon” or “inside this circle”. It is purely an optimization. Running the block with an all-zero mask would change nothing that is read afterwards, so skipping it cannot change the picture.
Jumps only go forward. There is no way to express a loop, so a program always ends.
From numbers to colour
Section titled “From numbers to colour”The three results are floats, normally between 0 and 1. For each pixel the VM:
- Clamps each channel to the range 0 to 1. A NaN becomes 0.
- Scales it to the canvas’s precision: 32 levels of red and blue, 64 of green.
- Dithers. Where the value falls between two levels, a small fixed 4 × 4 pattern decides which pixels round up and which round down, in proportion to how far between the levels it is. Across a few pixels the eye sees the in-between shade, so a slow gradient does not break into bands. The pattern depends only on the pixel’s position, so a still picture does not flicker.
There is no gamma step here. The LED driver applies its own brightness curve to everything drawn on the panel.
An interpreted shader’s colours go through exactly the same code.
Why the time is bounded
Section titled “Why the time is bounded”A script cannot make the panel hang by uploading a heavy shader:
- The program has no loops, so its running time is its length times the number of tiles.
- Every opcode has a cost, and the cost of the whole body is added up when the program is loaded.
- Each
gfx.fillcharges that cost to the frame. A frame may spend 1,000 units per pixel, about 0.45 seconds on a 64 × 64 panel in the worst case. One more fill past that is refused and the script stops.
Lua code is bounded differently, by counting instructions as they execute. That is not needed here because the cost is known before the first pixel.
The same result everywhere
Section titled “The same result everywhere”The same shader produces the same pixels on every panel and in the reference implementation used for testing. Three things make that true:
sin,cosandtanare the panel’s own polynomial versions, not the C library’s, which round differently between systems.- The noise functions use integer hashing.
- The firmware is built so that the processor does not merge a multiply and an add into one differently-rounded operation.
The firmware’s benchmark carries a checksum of a reference tile for this reason: a panel whose VM did not reproduce it bit for bit would fail the run. The slow library functions (pow, exp, log, asin, acos, atan) are the exception. They are the same on every panel, but can differ in the last digit from a desktop computer.
Measured
Section titled “Measured”On the panel’s ESP32-S3 processor with a 64 × 64 display:
| Shader | Interpreted | Compiled |
|---|---|---|
Plasma: 18 arithmetic instructions, 6 sin/cos |
28.7 µs per pixel | 6.0 µs per pixel |
Clouds: one four-octave fbm |
24 µs | 15.5 µs |
From those two: an arithmetic instruction costs roughly a tenth of a microsecond per pixel and a sine about six times that. Four octaves of noise cost about 15 µs however they are called, which is why the clouds gain less. Copying the finished canvas to the LEDs takes a further 8 ms per frame for any kind of script.
The plasma at 6 µs per pixel is 25 ms per frame: about 30 frames per second, against 8 when interpreted.
What it leaves out
Section titled “What it leaves out”The VM is small because it does little. It has no stack, no function calls, no memory to read or write besides its registers, and no loops. The compiler removes all of those before the program reaches the panel: helpers are copied into their callers, loops are written out, vectors are split into numbers. What arrives is straight-line arithmetic, which is the one thing that can be done 64 pixels at a time.