Skip to content

From source to pixels

This page and the two after it describe how compiled shaders work underneath. You do not need any of it to write shaders. It is here for the curious, and for anyone who wants to know exactly what the panel will do with their code.

The panel’s processor is a 240 MHz microcontroller. A 64 × 64 frame has 4,096 pixels, and at 30 frames per second there are about 25 milliseconds to shade them: roughly 6 microseconds, or 1,500 processor cycles, per pixel.

Calling a Lua function for a pixel costs about 5 microseconds before the function has done anything, and each Lua instruction inside it about 1 more. Vector math in Lua creates a small object per operation. So an interpreted shader spends its whole budget on overhead: the plasma takes 28.7 µs per pixel.

The compiled path removes the overhead instead of speeding up Lua:

  • No call per pixel. The shader becomes a short list of instructions, and each instruction processes 64 pixels of a row in one tight loop.
  • No vectors at run time. A vec3 becomes three plain numbers during compilation.
  • No objects, no garbage collection, no interpreter. The program is fixed arithmetic.

The compiler runs in your browser (or on a developer’s machine), never on the panel. The panel only receives the result.

Every stage below shows this shader:

-- @uniform time float
function fragment(x, y, u)
local glow = 0.5 + 0.5 * sin(u.time)
if x > 31 then
return glow, y / 63, 0
end
return 0, y / 63, glow
end

The left half of the panel fades blue, the right half fades red, and both get greener towards the bottom.

The compiler reads the file into a tree: this function contains a local, then an if, then a return, and so on. Anything outside the supported subset is refused here, with a message that names the construct and the line. A clear “tables are not allowed in shaders” at this point is more useful than a confusing error later.

Functions whose names start with an underscore are skipped without being read, which is why _init, _update and _draw can contain anything.

Lua does not declare types, so the compiler infers them. x and y are numbers. A vec3(...) is a vec3. u.time is whatever its -- @uniform line says. Each operation follows one rule: two values of the same type combine, and a number combines with any vector. Everything else is an error, such as adding a vec2 to a vec3.

Helper functions are checked once per call, with that call’s argument types, so one helper can serve numbers and vectors.

Vectors disappear. A vec3 becomes three separate values, a vec2 * number becomes two multiplies, length becomes multiplies, adds and a square root. Helper calls are replaced by their bodies and loops are written out once per pass.

What is left is a list of instructions that each do one thing to plain numbers. r0 and r1 hold x and y; every other r is a value the compiler introduced:

r2 = sin u.time
r3 = mul 0.5, r2
r4 = add 0.5, r3
r5 = gt r0, 31
if r5 {
r6 = div r1, 63
ret r4, r6, 0
}
r7 = div r1, 63
ret 0, r7, r4

Sixty-four pixels run together, and some will take the if while others will not. There can be no jump, so the if goes away: both sides are computed for every pixel and a select keeps the right result for each one.

r2 = sin u.time
r3 = mul 0.5, r2
r4 = add 0.5, r3
r5 = gt r0, 31
r6 = div r1, 63
r7 = div r1, 63
r8 = select r5, r4, 0
r9 = select r5, r6, r7
r10 = select r5, 0, r4
ret r8, r9, r10

r5 is the mask: 1 for the pixels where x > 31, 0 for the others. select r5, a, b gives a where the mask is 1 and b where it is 0. The two return statements have become three selects, one per colour channel, and the program now has a single ret at its end.

The same idea covers everything else that depends on a condition:

  • An assignment inside an if becomes select mask, new value, old value.
  • Nested conditions combine their masks with and.
  • Code after an early return runs under the mask of the pixels that have not returned.

When a branch is expensive, the compiler also wraps it in a jmpnone, which lets the panel skip the whole block for a group of pixels in which nobody takes it. That is only a shortcut. The result is the same whether the block runs or not.

Five steps. The first four repeat until the program stops shrinking, because each can uncover work for the others; then registers are assigned once.

  1. Give every assignment its own register, so that a register always means one value.
  2. Simplify. Compute what can be computed now, replace copies by their source, and merge identical computations. In the example, r6 and r7 are the same division, so one goes away.
  3. Hoist. An instruction whose inputs are only constants and uniforms gives the same answer for every pixel. It is moved to a short list that runs once per frame. Here that is the whole of glow.
  4. Remove dead code: instructions whose result nothing reads.
  5. Assign registers. A value only needs a register from where it is made to where it is last used, so registers are reused. The example needs three.
frame:
p0 = sin u.time
p1 = mul 0.5, p0
p2 = add 0.5, p1
pixel:
r0 = gt r0, 31
r1 = div r1, 63
r2 = select r0, p2, 0
r0 = select r0, 0, p2
ret r2, r1, r0

Three instructions run once per frame, and four instructions plus the ret run for the pixels. The sin is no longer per pixel at all.

Every simplification is exact. It produces the same 32-bit float the original instruction would have, which is why x + 0 is left alone (it is not the same as x when x is negative zero) and why the compiled shader matches the interpreted one pixel for pixel.

The instruction list is written out in a compact binary form, 102 bytes for this shader, and attached to the script as hex digits. The next page takes those bytes apart.

Next: The bytecode.