Skip to content

Show every Ruby's speed side by side on the results site - #252

Draft
JuanVqz wants to merge 6 commits into
mainfrom
feature/results-site-view-2
Draft

JuanVqz wants to merge 6 commits into
mainfrom
feature/results-site-view-2

Conversation

@JuanVqz

@JuanVqz JuanVqz commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Adds an "Across Rubies" view to the results site, and the site now opens on it: for every benchmark, how fast each idiom is on each Ruby, from 2.1 to head, JRuby and TruffleRuby, with and without the JIT.
  • Every build of a benchmark runs on the same machine, one after the other, so i/s can be compared across Rubies. Today each Ruby runs as its own CI job on its own machine (6 CPU models in one run), which only allows comparing idioms inside one Ruby.
  • A new weekly workflow does those runs in 6 shards and publishes the whole site, both views, from them. benchmarks.yml keeps checking every benchmark on every Ruby and building a preview for PRs, but no longer deploys.

Changes

Step 1: Run every build on one machine, with the reference in between

  • script/run_cross_ruby.rb runs every 6th benchmark file (one shard) on all 26 builds, in a random order seeded by the run id.
  • The builds come from the rake matrix in benchmarks.yml, so a Ruby added to CI is measured too. The newest released MRI (Ruby 4.0 today, 5.0 once it is in the matrix) is the reference: it runs first, after every 3 builds, and last, to show how steady the machine was.
  • A Ruby's interpreter and JIT builds run back to back, so each image is fetched once and removed after its last pass.
  • The collector records the shard, the pass, whether it is the reference and when each report finished; empty for the regular CI jobs.

Step 2: Run the builds across Rubies every week

  • .github/workflows/benchmarks-cross-ruby.yml: Sundays 03:00 UTC and by hand, 6 shards.

Step 3: Build both views from the run across Rubies

  • --cross-ruby DIR builds both views. cross-ruby.json has every build's i/s and error as measured, relative to the reference's median, and how much the reference varied over its passes (standard deviation over mean). Over 5% a benchmark is marked noisy, and when the typical benchmark is, the page warns.
  • "Does it hold?" comes from the same run (the reference's middle pass for the reference build), so every build of a benchmark now shares one machine.
  • Without --cross-ruby, the builder still makes the "Does it hold?" preview from the regular CI jobs, as today.

Step 4: Show Across Rubies first on the results site

  • Per benchmark: a chart with the Rubies on the x axis, one line per report (the claimed one green), a JIT switch with the interpreter dashed underneath, i/s or relative to the reference, a table with every build, and how steady the machine was.
  • Two Rubies count as same-ish when their error bars overlap or their gap is within how much the reference varied: runs minutes apart vary more than benchmark-ips' error bars say.
  • ?view=rubies and ?view=advice pick the view; the build choice is shared by both views.

Step 5: Publish the site from the run across Rubies

  • The weekly workflow builds the whole site and deploys it from main; it is the only workflow that deploys. A merged benchmark reaches the site with the next weekly run, or sooner by running the workflow by hand.

Step 6: Explain how to run the builds across Rubies locally

  • CONTRIBUTING: run a few files on a few builds and build the site from them; which builds to leave out on an Apple silicon Mac.

Test plan

  • The workflow ran once on GitHub runners from this branch (run 37380780390, through a temporary commit now removed): all 6 shards passed in 152 to 196 minutes (timeout 330), setup 11 to 15 minutes per shard, 87 GB of free disk at the start.
  • Compared with the regular CI run on the same commit (37380786678):
    • The typical jump in i/s between neighbouring MRI versions drops from 30.2% to 7.0%.
    • The win ratios agree (median difference 3.1% over 1946 benchmark x build pairs).
    • The reference varied 2.3% on the typical benchmark, with no drift over the run, and 7 of 78 benchmarks went over 5%.
  • Correcting each build by the reference's speed at the time it ran was tried and dropped: on that run it made the lines jumpier (7.6% against 7.0%), and in two local pilots it made two runs agree less.
  • The site built locally from that run's artifacts: a headless-Chrome harness covering both views (every Ruby x build of "Does it hold?", every build x scale of Across Rubies, desktop and phone width) ran 2318 checks, 0 problems. The PR-style preview (regular jobs only) ran 407 checks, 0 problems.
  • Unit tests: entries, steadiness, "Does it hold?" from the same run, inputs (19); the existing verdict tests (14); adding ruby_5.0 to the CI matrix makes it the reference and the Ruby the site opens on (4).
  • actionlint: 0 errors.

Notes

  • After merge, the live site keeps today's "Does it hold?" until the first weekly run. Running "Benchmarks across Rubies" by hand from the Actions tab publishes the new site right away.
  • Two findings from the first run, for their own PRs: format-vs-round-and-to-s has a different winner on EPYC 7763 and EPYC 9V45 machines, and in loop-vs-while-true YJIT never compiles the methods (each call runs 100 million iterations), so YJIT shows the same speed as the interpreter.
  • This PR changes compose.yaml and docker/, so CI runs every benchmark on all 26 jobs.

The results site can only compare idioms inside one Ruby today: each Ruby
runs as its own CI job on its own machine, so i/s across Rubies is mostly
the machine. script/run_cross_ruby.rb runs every build of a shard's files
on one machine, one after the other, so their i/s can be compared.

The builds come from the rake matrix in benchmarks.yml, so a Ruby added to
CI is measured here too. The newest released MRI is the reference: it runs
first, again after every 3 builds, and last. It is the same code every
time, so its passes show how steady the machine was during the run.

A Ruby's interpreter and JIT builds run back to back, so each image is
fetched once and removed after its last pass (with --fresh-images, in CI),
which keeps the runner's disk free.

The collector records the shard, the pass, whether it is the reference and
when each report finished; those fields are empty for the regular CI jobs.
A new workflow runs script/run_cross_ruby.rb in 6 shards, each on its own
runner: every 6th benchmark file, all 26 builds. A full run is about 3
hours per shard, under the 6 hour job limit, so it runs once a week
(Sunday 03:00 UTC) and by hand. Each shard uploads its results.
The site builder reads a cross-Ruby run (--cross-ruby DIR) and builds both
views from it. cross-ruby.json has, per benchmark, every build's i/s and
error as measured and relative to the reference's median, and how much the
reference varied over its passes (standard deviation over mean). A
benchmark over 5% is marked noisy; when the typical benchmark is, the page
warns that small differences between Rubies may be the machine.

"Does it hold?" comes from the same run: every build's result, and the
reference's middle pass. Each verdict still compares reports inside one
run, and now every build of a benchmark shares one machine, recorded per
benchmark. Without --cross-ruby the builder still makes the "Does it hold?"
preview from the regular CI jobs.

On the first run on GitHub runners the reference varied 2.3% on the typical
benchmark, with no drift over the run (last third within 0.1% of the
first). Correcting each build by the reference's speed at the time made
the lines jumpier (7.6% against 7.0% between MRI versions), as it did in
two local pilots, so the numbers are shown as measured.
The page gets a second view, Across Rubies, and opens on it; "Does it
hold?" opens first only when the page has no cross-Ruby data (a PR
preview). Per benchmark: a chart with the Rubies on the x axis (MRI as a
line, JRuby and TruffleRuby as points), one line per report with the
claimed one in green, a JIT switch with the interpreter dashed under it,
i/s or relative to the reference, a table of every build, and how steady
the machine was.

Two Rubies count as same-ish when their error bars overlap or their gap is
within how much the reference varied on that machine: runs minutes apart
vary more than benchmark-ips' error bars say. "Does it hold?" says when
every build of a benchmark ran on one machine, so its columns compare.

Both views keep their choices in the URL (view, and build, which they
share) and leave each other's alone.
The weekly workflow builds the whole site from its shards and deploys it
from main; it is now the only workflow that deploys. benchmarks.yml keeps
checking every benchmark on every Ruby and builds the "Does it hold?"
preview for PRs, but no longer publishes: one source for the site means no
second view to fetch, merge or keep in step.

A merged benchmark reaches the site with the next weekly run, or sooner by
running the workflow by hand.
@JuanVqz
JuanVqz force-pushed the feature/results-site-view-2 branch from 73b7656 to fa5a769 Compare October 6, 2026 05:28

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant