Skip to content

Add shape primitives and sprite flipping - #18

Open
TomHoenderdos wants to merge 8 commits into
atomvm:mainfrom
TomHoenderdos:shape-primitives
Open

TomHoenderdos wants to merge 8 commits into
atomvm:mainfrom
TomHoenderdos:shape-primitives

Conversation

@TomHoenderdos

@TomHoenderdos TomHoenderdos commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds vector shape primitives and sprite flipping to the display list, plus a rendering benchmark and opt-in LCD profiling.

{rounded_rect, X, Y, W, H, Radius, Color}
{line, X1, Y1, X2, Y2, Thickness, Color}
{circle, CX, CY, R, Color}
{ellipse, CX, CY, RX, RY, Color}
{arc, CX, CY, R, Thickness, StartDeg, EndDeg, Color}   % clockwise from StartDeg to EndDeg
{polygon, [{X, Y}, ...], Color}                        % even-odd, up to 256 points
{scaled_cropped_image, ..., [flip_x, flip_y], Img}     % or {flip_x, true}

Exact pixel rules, limits and validation are documented in docs/primitives.md.

Display list validation (#19) and the one-pass walk (#21) are merged; the SDL damage fix is in #20 and the per-row item walk in #22.

Commits

  1. Add flip_x and flip_y to scaled_cropped_image
  2. Add rounded_rect primitive: also adds the shape module (shape.c, integer-only geometry, exact scanline runs cached per row), the renderer hooks and tests/shapes.
  3. Add circle primitive
  4. Add ellipse primitive
  5. Add line primitive
  6. Add arc primitive
  7. Add polygon primitive
  8. Add rendering benchmark and LCD profiling: host benchmark and an opt-in -DATOMGL_PROFILE=1 log line (parse / draw / max line / spi wait / frame).

Each commit builds and passes tests/shapes (UBSan) and tests/items (ASan+UBSan) on its own.

Tested on hardware

ESP32-S3 badge, ST7789 320×240 RGB565 at 80 MHz SPI, ESP-IDF 5.5.2, running a racing game with 71–104 items per frame (road/kerb polygons, full-width rects, a scaled sprite, an arc, HUD text) with -DATOMGL_PROFILE=1, on the previous version of this branch. The draw time numbers from that run are in #22.

Other testing

  • tests/shapes, tests/items (ASan+UBSan), tests/bench on the host, for every commit.
  • This touches the same lines of dcs_lcd_display_driver.c and display_items.h as Walk only items that cover the current row #22; whichever is merged second will be rebased.
  • On the previous version of this branch: randomized differential checks (valid rect/text/image scenes render byte-identical to main) and upstream CI reproduced locally in Docker (AtomVM release-0.6, IDF 5.4.3 and 5.5.2, esp32 and esp32s3, with and without ATOMGL_PROFILE).
  • Not tested: the SDL plugin as a whole (it doesn't build against release-0.6 on main either, ufont_manager_register arity).

Compatibility

Existing display lists render the same; the previously ignored Opts of scaled_cropped_image now accepts flips, other entries are still ignored. BaseDisplayItem stays 56 bytes on ESP32; each shape item allocates its geometry separately.

🤖 Generated with Claude Code

@TomHoenderdos
TomHoenderdos marked this pull request as draft September 29, 2026 11:34
@TomHoenderdos
TomHoenderdos marked this pull request as ready for review September 29, 2026 19:19
@TomHoenderdos

Copy link
Copy Markdown
Contributor Author

@bettio, tested this on the racer game and it works amazing :) image

@bettio bettio left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's start sorting this PR, would you mind extracting "display_items: Validate display list items" commit (6d59cb0) and the related tests: b0bfe1e and make a new PR just with those 2 commits?

TomHoenderdos and others added 7 commits October 2, 2026 10:43
Sprites often need a mirrored version, such as a character facing the
other way. Without flips, applications have to keep a second, mirrored
copy of the sprite sheet in memory.

The Opts element of scaled_cropped_image, which was ignored, now
accepts flip_x and flip_y, as bare atoms or as {flip_x, true} and
{flip_y, true}. Other entries are still ignored, and so is an Opts
that is not a list, so existing display lists draw as before.

A flip mirrors the drawn pixels in place rather than the whole source
image: with flip_x, the pixel at offset C from the item's left edge
shows what the unflipped item shows at offset Width - 1 - C. This
also holds when Width is not a multiple of the scale factor, or is
cut short by the end of the source image.

The source pixel lookup is shared by all renderers through helpers in
display_items.h. The renderers work out the source row, the flip and
the clamping to the image once per run of pixels, and step through
the row in either direction. The SDL display compares the flips when
it looks for changed items.

tests/bench/test_flip checks that every pixel of flipped and unflipped
items lies inside the source image and that the per-run walk matches
the per-pixel lookup. tests/items now also renders every flip
combination and checks the parsing of Opts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Drawing anything but rectangles, text and images meant building an
image of it in Erlang. Add a filled rectangle with rounded corners:

  {rounded_rect, X, Y, Width, Height, Radius, Color}

The geometry lives in a new shape module, shape.c, independent of
AtomVM and of any display, that the following shape primitives build
on. A shape answers two questions: whether a pixel is inside it, and
how long the run of pixels from a given one that are all inside or all
outside is. Renderers draw a row of pixels at a time, so shape_run()
lets them fill or skip a whole span with one call instead of testing
every pixel. Convex shapes cache the inside span of the last row.

Pixels are sampled at their centers. The corners are quarter circles
centered on the pixels Radius pixels in from each corner and follow
the midpoint rule. The radius is clamped to (min(Width, Height) - 1)
div 2, which keeps a straight edge of at least one pixel on every
side. Every value must be within +-32767, which keeps all intermediate
math within 64 bits; anything else makes the item invalid and logged.

A shape paints only the pixels inside it: the rest of its bounding box
shows whatever is below it in the display list. The DCS LCD, mono,
e-paper and SDL renderers ask the shape for the run of pixels from the
current one, fill an inside run with the rect renderer and skip an
outside run. Under an item that leaves a run unpainted, the items
below may draw up to the end of that run rather than a single pixel,
and each shape item remembers its last outside run on a row, so that
the items above it don't make it compute the same run again pixel by
pixel. The SDL display compares shapes when it looks for changed
items.

tests/shapes checks the pixel rules, checks that shape_run() agrees
with shape_contains() in and around the bounding box, and covers the
value limits. It has no dependencies and can also run under UBSan:

  cmake -S tests/shapes -B build/shapes -DSHAPES_UBSAN=ON
  cmake --build build/shapes && build/shapes/test_shapes

tests/items covers the parsing, validation and allocation failures of
the item. sdl_display/test_display.erl gets a shapes/0 demo scene, and
its requests now use the {'$call', {Pid, Ref}, Request} format, the
only one AtomVM's port_parse_gen_message() accepts: the old
{Pid, Ref, Request} tuples were rejected and nothing was drawn.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Add a filled circle centered on a pixel:

  {circle, CX, CY, R, Color}

The shape module draws it as an ellipse with equal radii, so that the
ellipse primitive can share it. A pixel at offset {DX, DY} from the
center is inside when DX * DX + DY * DY < R * R + R, the midpoint rule
that rounded_rect corners use as well, so a circle of radius R is
2 * R + 1 pixels across and a 2 * R + 1 square rounded_rect with radius
R draws the same pixels. Each row is a single run, found by a binary
search on either side of the center and cached like the rounded_rect
rows.

The center and the radius must be within +-32767 and the radius must
be at least 1, or the item is invalid.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Add a filled ellipse centered on a pixel:

  {ellipse, CX, CY, RX, RY, Color}

It uses the shape module's ellipse, which applies the midpoint rule of
circles to each axis: a pixel at offset {DX, DY} from the center is
inside when DX^2 / (RX^2 + RX) + DY^2 / (RY^2 + RY) < 1, so the ellipse
is 2 * RX + 1 pixels wide and 2 * RY + 1 pixels high, and an ellipse
with equal radii is the circle of that radius.

The center and the radii must be within +-32767 and both radii must be
at least 1, or the item is invalid.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Add a straight line of any angle and thickness:

  {line, X1, Y1, X2, Y2, Thickness, Color}

Endpoints are pixel centers and both are covered. A line is Thickness
pixels thick measured along its minor axis, so a line that is wider
than it is tall covers Thickness pixels in every column from X1 to X2,
and a 1 pixel line has one pixel per column (or row) and no gaps. With
an even thickness the extra pixel goes above (or left of) the ideal
line. From thickness 3 on, both ends get a round cap that never sticks
out above or below the line, and a line with equal endpoints is a disc
on that point.

A line is convex, so each row is a single run. The shape works out
the first and last candidate pixel of the row from the line equation,
then finds the exact edges with a galloping search around them.

All values must be within +-32767 and the thickness at least 1, or the
item is invalid.

tests/shapes compares random lines with a reference definition of the
pixel rule and checks the thickness of each column or row, that 1 pixel
lines have no gaps and that swapping the endpoints draws the same
pixels.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Add a ring segment, which becomes a pie slice when the thickness
reaches the radius:

  {arc, CX, CY, R, Thickness, StartDeg, EndDeg, Color}

Angles are integer degrees, 0 at 3 o'clock, and the arc sweeps
(EndDeg - StartDeg) mod 360 degrees clockwise. Distinct angles that are
equal mod 360 draw a full ring and equal angles make the item invalid.
The angles can be any 64-bit integer and are reduced mod 360 when
parsed. Every other value must be within +-32767,
and the radius and thickness must be at least 1.

A pixel is drawn when its center is inside the ring, by the midpoint
rule for both the outer radius R and the inner radius R - Thickness,
and either its center lies within the sweep, both edges included, or
one of the two edge rays passes through the pixel. The rays keep
narrow arcs visible: a 1 degree arc still draws the ring pixels its
edges cross. Edge directions come from a Q14 sine table, so all math
stays in integers.

An arc row can have several inside runs. The shape finds the columns
where the row can change, at the outer and inner circles, the edge
half-planes and the edge rays, and caches the resulting runs per row.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Add a filled polygon:

  {polygon, [{X, Y}, ...], Color}

Unlike the other shapes, polygon vertices lie on pixel corners: the
pixel at {X, Y} spans X..X+1 and Y..Y+1 and is drawn when its center is
inside, using the even-odd rule. The fill is half-open, so
[{0, 0}, {10, 0}, {10, 10}, {0, 10}] fills exactly 10 by 10 pixels, and
polygons that share an edge neither overlap nor leave a gap.

A polygon needs 3 to 256 points, each within +-32767, or the item is
invalid. The limits keep every edge step within 64 bit math and let
the edge indexes fit 16 bits.

The shape keeps its edges sorted by top row. For each row it updates
the active edges incrementally when moving down one row, as renderers
do, and caches the sorted columns where the fill toggles, so a run is
a lookup of the next toggle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
Add tools to measure what the display list costs to draw, on the host
and on a device.

tests/bench/bench times dcs_lcd_draw_x(), the DCS LCD scanline
renderer, over a set of scenes on a 240x240 screen: single shapes,
grids of small shapes, overlapping circles, lines and arcs crossing
the same rows, a 256-point polygon, a text UI and full-screen sprites
with and without flips. Scenes with shapes are timed again with each
shape replaced by a rect of its bounding box, which shows what the
shapes cost over flat rects. The README describes the scenes, how
to read the results and an estimate of each scene's slowest line on
an ESP32-S3, the number that matters there since each line is drawn
while the previous one is sent over SPI. How long a line may take
depends on the panel: width * bits per pixel / SPI clock, 96 us for
240 px at 40 MHz but 64 us for 320 px at 80 MHz.

Building the firmware with -DATOMGL_PROFILE=1 makes the DCS LCD
driver log one line per update to stderr: item count, parse time,
draw time, slowest line, time spent waiting for SPI and frame time.
Draw and SPI wait are kept apart so a slow frame can be pinned on the
renderer or on the bus. Profiling is off by default and adds nothing
to normal builds; the esp_timer library is linked only when it is on.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Tom Hoenderdos <tomhoenderdos@gmail.com>
@bettio

bettio commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

So we are getting close and closer.

Here some action points:

  1. Rebase on main. The only conflicts are two hunks that add the flip and shape cases to cmp_display_item(), which sdl: Fix plugin build and remove broken damage tracking #23 removed; drop them. Item comparison is moving to a shared display_diff module where shape_equal() and the flip fields will be compared.

  2. Fold the per-item outside-run memo into shape_run(). The shape already caches its row, so the memo only saves a call and a switch, while it costs three fields in ShapeItemData, two helpers in display_items.h and the same remember/lookup dance in four renderers. With the gap check inside shape.c, the item holds just the pointer and each renderer's shape case shrinks to a few lines.

  3. Transparent runs for every drawer. draw_x now carries a transparent_run down to the items below, but only shapes feed it; text, image and scaled image still return 0 and the items below get one pixel per call, so every background pixel of a label or a sprite costs a full draw_x round. Have those drawers report their transparent run the same way, and unify the contract on the return value (positive: pixels drawn, negative: transparent for that many pixels) so the shape drawer's out-parameter can go and draw_x has one rule. Measured on the host on top of the per-row lists of Walk only items that cover the current row #22, output identical on 2000 random scenes: a 25-label text UI 111 → 63 us/frame, ufont-style glyph images 106 → 56, text with no background 159 → 55, racer-like scene 53 → 40, sprites with holes about 1.3x (that run also changed the scaled stepping, which does not apply to your helper, so expect less there). Your call whether it is a last commit here or a follow-up stacked on this PR; it supersedes what I wrote in Walk only items that cover the current row #22 about doing it on our side.

  4. display_items.h is becoming an implementation file. The flip/scaled row helpers are about 110 lines of inline math with a negative divisor for mirroring and an INT_MAX sentinel mode in ..._span(), and display_items_scaled_cropped_pixel() implements the same mapping a second time. Comments or a plainer representation, please. If you prefer, the flips commit could also become its own small PR so this one is shapes only.

  5. Polygon parsing. The point count is known before the points are read, so they can be parsed straight into the shape allocation instead of a temporary array handed to shape_new_polygon_owned(), an API that exists only to free its input; that is one allocation less per polygon per update. Also consider a cap below 256 points: at 38 bytes per vertex a maximal polygon is close to 10 KB per item per frame, and nothing realistic needs it.

  6. Diagnostics. Every shape failure logs "wrong arity, bad argument or out of memory", where the other primitives name the cause. Checking the arity first and saying so is trivial and keeps the error lines useful.

  7. Arc segment arrays. seg_start[13] and seg_inside[12] are written at seg_count + 1 and seg_count after the loop, which is in bounds only because a row never has more than 6 inside/outside transitions (I checked by fuzzing and by exhaustive search over radii, thicknesses and angles). One entry of slack, or an assert, would make that explicit.

  8. Nits. init_shape_item() has an unused ctx (new -Wextra warning on the host); new_polygon() checks malloc without IS_NULL_PTR/UNLIKELY; shape_contains() could take const; the rewritten scaled loops in sdl_display/display.c, mono_draw.c and epaper_draw.c kept (*pixels >> 24) next to the READ_32_UNALIGNED value, UBSan flags it with unaligned image data, use img_pixel; clang-format on the new test files.

  9. Not blocking, worth a measurement. On the host, with the shape case present in draw_x, the existing primitives render 20-50% slower per draw_x call (removing only that case restores main's speed; neither jump tables, stack protector, inlining nor alignment explain it). On the device that would be roughly 1% of the draw time, so a before/after ATOMGL_PROFILE on a shape-free scene would settle it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants