Conversation
The results site can only compare idioms inside one Ruby today: each Ruby runs as its own CI job on its own machine, so i/s across Rubies is mostly the machine. script/run_cross_ruby.rb runs every build of a shard's files on one machine, one after the other, so their i/s can be compared. The builds come from the rake matrix in benchmarks.yml, so a Ruby added to CI is measured here too. The newest released MRI is the reference: it runs first, again after every 3 builds, and last. It is the same code every time, so its passes show how steady the machine was during the run. A Ruby's interpreter and JIT builds run back to back, so each image is fetched once and removed after its last pass (with --fresh-images, in CI), which keeps the runner's disk free. The collector records the shard, the pass, whether it is the reference and when each report finished; those fields are empty for the regular CI jobs.
A new workflow runs script/run_cross_ruby.rb in 6 shards, each on its own runner: every 6th benchmark file, all 26 builds. A full run is about 3 hours per shard, under the 6 hour job limit, so it runs once a week (Sunday 03:00 UTC) and by hand. Each shard uploads its results.
The site builder reads a cross-Ruby run (--cross-ruby DIR) and builds both views from it. cross-ruby.json has, per benchmark, every build's i/s and error as measured and relative to the reference's median, and how much the reference varied over its passes (standard deviation over mean). A benchmark over 5% is marked noisy; when the typical benchmark is, the page warns that small differences between Rubies may be the machine. "Does it hold?" comes from the same run: every build's result, and the reference's middle pass. Each verdict still compares reports inside one run, and now every build of a benchmark shares one machine, recorded per benchmark. Without --cross-ruby the builder still makes the "Does it hold?" preview from the regular CI jobs. On the first run on GitHub runners the reference varied 2.3% on the typical benchmark, with no drift over the run (last third within 0.1% of the first). Correcting each build by the reference's speed at the time made the lines jumpier (7.6% against 7.0% between MRI versions), as it did in two local pilots, so the numbers are shown as measured.
The page gets a second view, Across Rubies, and opens on it; "Does it hold?" opens first only when the page has no cross-Ruby data (a PR preview). Per benchmark: a chart with the Rubies on the x axis (MRI as a line, JRuby and TruffleRuby as points), one line per report with the claimed one in green, a JIT switch with the interpreter dashed under it, i/s or relative to the reference, a table of every build, and how steady the machine was. Two Rubies count as same-ish when their error bars overlap or their gap is within how much the reference varied on that machine: runs minutes apart vary more than benchmark-ips' error bars say. "Does it hold?" says when every build of a benchmark ran on one machine, so its columns compare. Both views keep their choices in the URL (view, and build, which they share) and leave each other's alone.
The weekly workflow builds the whole site from its shards and deploys it from main; it is now the only workflow that deploys. benchmarks.yml keeps checking every benchmark on every Ruby and builds the "Does it hold?" preview for PRs, but no longer publishes: one source for the site means no second view to fetch, merge or keep in step. A merged benchmark reaches the site with the next weekly run, or sooner by running the workflow by hand.
JuanVqz
force-pushed
the
feature/results-site-view-2
branch
from
October 6, 2026 05:28
73b7656 to
fa5a769
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
benchmarks.ymlkeeps checking every benchmark on every Ruby and building a preview for PRs, but no longer deploys.Changes
Step 1: Run every build on one machine, with the reference in between
script/run_cross_ruby.rbruns every 6th benchmark file (one shard) on all 26 builds, in a random order seeded by the run id.benchmarks.yml, so a Ruby added to CI is measured too. The newest released MRI (Ruby 4.0 today, 5.0 once it is in the matrix) is the reference: it runs first, after every 3 builds, and last, to show how steady the machine was.Step 2: Run the builds across Rubies every week
.github/workflows/benchmarks-cross-ruby.yml: Sundays 03:00 UTC and by hand, 6 shards.Step 3: Build both views from the run across Rubies
--cross-ruby DIRbuilds both views.cross-ruby.jsonhas every build's i/s and error as measured, relative to the reference's median, and how much the reference varied over its passes (standard deviation over mean). Over 5% a benchmark is marked noisy, and when the typical benchmark is, the page warns.--cross-ruby, the builder still makes the "Does it hold?" preview from the regular CI jobs, as today.Step 4: Show Across Rubies first on the results site
?view=rubiesand?view=advicepick the view; the build choice is shared by both views.Step 5: Publish the site from the run across Rubies
main; it is the only workflow that deploys. A merged benchmark reaches the site with the next weekly run, or sooner by running the workflow by hand.Step 6: Explain how to run the builds across Rubies locally
Test plan
ruby_5.0to the CI matrix makes it the reference and the Ruby the site opens on (4).Notes
format-vs-round-and-to-shas a different winner on EPYC 7763 and EPYC 9V45 machines, and inloop-vs-while-trueYJIT never compiles the methods (each call runs 100 million iterations), so YJIT shows the same speed as the interpreter.compose.yamlanddocker/, so CI runs every benchmark on all 26 jobs.