C++ search - #210
Draft
ms609 wants to merge 1335 commits into
Draft
C++ search#210ms609 wants to merge 1335 commits into
ms609 wants to merge 1335 commits into
Conversation
ms609
marked this pull request as draft
March 25, 2026 14:21
ms609
added a commit
that referenced
this pull request
Mar 28, 2026
ms609
added a commit
that referenced
this pull request
Mar 28, 2026
ms609
added a commit
that referenced
this pull request
May 18, 2026
In R CMD check, R runs as a non-interactive subprocess with captured stdout. R_FlushConsole() calls fflush() on that pipe; when the buffer fills the call blocks indefinitely, causing the 6 h GHA timeout seen on every ubuntu runner for PR #210. Gate the \r-overwrite progress line and the flush behind R_Interactive (FALSE in batch/check contexts). Interactive sessions are unchanged. At verbosity >= 2 in batch mode, emit plain \n-terminated lines so diagnostic logs still carry progress detail without the flush risk. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
ms609
added a commit
that referenced
this pull request
May 18, 2026
In R CMD check, R runs as a non-interactive subprocess with captured stdout. R_FlushConsole() calls fflush() on that pipe; when the buffer fills the call blocks indefinitely, causing the 6 h GHA timeout seen on every ubuntu runner for PR #210. Gate the \r-overwrite progress line and the flush behind R_Interactive (FALSE in batch/check contexts). Interactive sessions are unchanged. At verbosity >= 2 in batch mode, emit plain \n-terminated lines so diagnostic logs still carry progress detail without the flush risk. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This was referenced May 19, 2026
ms609
added a commit
that referenced
this pull request
Jul 3, 2026
…as invalid Withdrawn same day, before any fix was written, so nothing was built on it. The claim (pushed in 7d4d255) was that the MPT-set symptom survives a canonicalised report path, evidenced by "2 distinct scores (183, 182) among 32 MPTs at a common tip-1 rooting". The measurement rooted each tree at its OWN tip.label[1], which after Renumber() is a DIFFERENT TAXON for different trees -- so it was a different rooting per tree, not a common one. Re-measured with the rooting taxon named explicitly (dev/red-team/heavy-tests/t385-diagnose-rooting.R): all 32 returned trees are already rooted at the kernel's tip 0 (32/32, root degree 2), and at a genuinely common rooting they all score 183 -- 32 distinct topologies, 0 disagreement. The pool IS self-consistent, so pool re-filtering is NOT required and n_topologies/collapse semantics do not need to move. That materially simplifies T-385. What survives is the P1 itself: reported 178 is not the score of the tree at the rooting it is returned at (183). And the objective is genuinely rooting-sensitive -- the same topology spans 178-183 over 8 rootings, inside the Sum nSec = 12 bound, with 178 attained at 2 of them. Fourth environment/measurement artefact this session (after %in%-on-Splits dispatch, ARM64 tip ordering, and covr timing). Same root cause each time: asserting on something incidental -- here "tip.label[1]" as a stand-in for a fixed taxon -- instead of naming the thing I meant. See [[loadall-is-not-rcmdcheck]] lesson 3.
Pre-fix reproduction against cpp-search tip a8fbba8, written before any fix so the failure is on record independently of the change that addresses it. The finding was recorded today but the tip has moved since (T-391 landed), and [[redteam-verify-against-current-tip]] says to re-verify rather than assume. All three symptoms reproduce on a 36-tip / 6-block / nSec=2 synthetic dataset (seed 1, search seed 11, maxReplicates 4): reported attr(res, "score") 178 TreeLength(res[[i]], ...) as returned 183 for all 32 trees -> gap 5 same topology, 8 rootings 178..183, spread 5 (bound Sum nSec = 12) 32 MPTs at a common tip-1 rooting 183 and 182 -> 2 distinct Note the reported 178 IS attained, by 2 of the 8 sampled rootings: the engine is not computing a wrong number, it is recording a score at one rooting and returning a topology rooted elsewhere (ts_collapse_pool's tip-0 canonicalisation being the trigger site named in T-374). The script reports rather than asserts, and exits 1 while the gap is open, so it doubles as the acceptance check for the fix.
Implements the ALREADY-DECIDED Option 3 of dev/plans/2026-07-29-t374b-xform-rooting-policy.md, whose three commits on cpp-search (7a18a4b et al.) turned out to be docs only -- no code implemented "make MaximizeParsimony's reported score and TreeLength() agree on one rooting". ## The defect XFORM's step matrix is asymmetric (gain = nSec + 1 against loss = 1), so a tree's length depends on where it is rooted -- unlike the symmetric criteria. MaximizeParsimony reported `result$best_score`, recorded mid-search at whatever rooting the replicate held, while ts_collapse_pool hands every tree back re-rooted on tip 0. So the reported number was not the length of the tree returned, and re-rooting a returned tree changed it again. Reproduced on tip a8fbba8 (36 tips / 6 blocks / nSec 2), pre-fix: reported 178 | TreeLength(returned) 183 on all 32 trees | gap 5 same topology over 8 rootings: 178..183, spread 5 (bound Sum nSec = 12) The reported number was not WRONG -- 178 is attained by 2 of those 8 rootings -- it was a value at a rooting the user never receives. ## The fix Canonicalise at both boundaries, on the dataset's first taxon (the tip-0 rooting ts_collapse_pool already imposes): - TreeLength(): root before scoring, in BOTH the single-tree and multiPhylo methods. The multiPhylo path previously rooted only trees that arrived unrooted, so an already-rooted tree kept its own rooting. Rooted by NAME with tips re-aligned afterwards, never by an R reroot on the edge matrix (na-validation-alignment-gotcha). - MaximizeParsimony(): rescore the returned pool through that same path and report the result -- |pool| evaluations. Post-fix the acceptance script is clean: gap 0, spread 0 across 8 rootings. Scores may therefore differ from previous versions and will not decrease; the value is now the length of the tree in hand, an upper bound on the rooting-free minimum exceeding it by at most Sum nSec (attained by 87-98% of rootings). Deliberately NOT done: min-over-rootings reporting. It is what the plan calls the better variant, but it reports a quantity the search never compared and costs (2n-3)x on the Sankoff term, so it stays Option 4. HSJ reporting is also untouched -- there rooting-invariance is REQUIRED by the method (two-state DP, T-374's open half), so canonicalising would convert a wrong objective into a stably-wrong one. Gated on XFORM alone, with that reason in the code. ## Scope boundary, and a warning where it bites Pool membership is still decided on search-time scores taken at differing rootings, so the returned trees need not share the canonical length. That is T-374's open residue, out of scope here -- but no longer silent: MaximizeParsimony now warns and reports the smallest. The warning immediately found a live instance in the existing suite (all-hierarchy data spans 7 to 9), which independently corroborates the earlier session's "4 of 6 trees do not share a score" observation. ## Tests Two regression tests, both verified to FAIL against the pre-fix tip built in a throwaway worktree (spread 2 not 0; reported 46 against min 48) and pass after: - The deterministic one carries its own anti-vacuity guard: it first ASSERTS via ts_sankoff_test -- untouched by this fix -- that the raw kernel really is rooting-sensitive on its 8-tip input, so it cannot pass by scoring rooting-insensitive data. - The end-to-end one asserts the actual contract (reported == min of the returned trees' canonical lengths), not the stronger "every tree shares it", which holds for its data but is not what the code promises. Also fixes the three false root-invariance comments the plan lists (ts_tbr.cpp:123 and :3152, ts_rcpp.cpp's collapse reroot, MaximizeParsimony.R's collapse note), and documents the behaviour in ?MaximizeParsimony, ?RecodeHierarchy, NEWS.md and vignettes/search-algorithm.Rmd. Suite green: xform, tree_length, hsj, t330, recode-hierarchy, resample-hierarchy, t306, prune-reinsert -- 0 failures.
… rename Not part of T-385 -- surfaced by running check_man() for it. 419168d changed the roxygen example from "one notch less / effort = -1" to "one notch more / effort = 1L" but man/MaximizeParsimony.Rd was never regenerated, so the shipped example contradicted its own source. Regenerated from the roxygen block, which is authoritative.
Coordination-only. The reporting half of T-374 is implemented on claude/t385-xform-report-agreement (stacked on PR #277): TreeLength() canonicalises the rooting in both its single-tree and multiPhylo methods, and MaximizeParsimony() rescores the returned pool through that path before reporting. Acceptance script: gap 5 -> 0, rooting spread 5 -> 0. Worth the row update on its own: the new "returned trees do not share a length" warning fired immediately on an EXISTING test -- all-hierarchy xform data returns a pool spanning 7 to 9. That independently corroborates T-374's original "4 of 6 trees do not share a score", which my own 36-tip attempt to reproduce had measured wrongly (retracted in 9296e87). The residue is real, reachable, and now visible at runtime instead of silent.
The acceptance script carried its own copy of MakeXformData() identical to the one the rooting diagnostic sources, so the two could silently drift. Sourced from the shared file instead. Acceptance script re-verified: exit 0, gap 0, spread 0 (checked WITHOUT a pipe -- the earlier 'EXIT: 0' was tail's status).
Three follow-ups, no behaviour change. - test-ts-xform.R's all-hierarchy test now expect_warning()s the "do not share a length" warning instead of emitting it into an otherwise-green suite. That matrix provably triggers T-374's residue (pool spans 7 to 9), so the warning is an expectation, not noise -- and if it ever stops firing, the residue has been fixed or masked and this test says so. Assign INSIDE expect_warning(): it returns the condition, not the expression's value, so `res <- expect_warning(...)` silently bound the warning object and broke the two assertions below it. - Both diagnostic scripts source() a path relative to the package root, so they only work when invoked from there. They now stop with that message instead of failing obscurely -- the same relative-path trap that made five tree_length tests "fail" earlier today when test_dir() changed the working directory out from under a relative lib.loc. - Recorded the real reason min-over-rootings reporting is declined. "Reports a quantity the search never compared" is the weaker half; the deciding objection is cost -- naively (2n-3)x on the Sankoff term, so a 4000-tip pool of 100 trees needs ~800k boundary evaluations. Affordably it needs an all-rootings up-down DP, which is Option 4's own piece of work.
All 8 CI jobs green on run 30662279563. The one ubuntu-24.04 (4.1) failure was transient -- it died in 'Set up R dependencies' with the Bioconductor 3.14 repos unreachable, so 'Check package' was skipped and no code was ever tested; green on re-run of that job alone.
T-385: report the x-transformation score of the tree actually returned
…f 30) Array 18126933, 150/150 cells at pin 644b5b1, every stopping rule left as shipped (maxSeconds = 0, maxReplicates = 96, targetHits = max(10, ntax/5)). Raw cells in dev/profiling/na-certify-stop-hamilton.csv. THE HEADLINE, and it corrects documentation written two commits ago: tripling `targetHits` changed floor attainment on NOT ONE of 30 matrices, while costing 2.58x the wall (26 of 30 slower, 22 by >10%, p = 6e-05). Only 33 of 150 cells were identical in score AND replicate count, so the runs genuinely went longer -- they just never found anything better. The mechanism is the disjoint- population argument measured at corpus scale rather than on one cell: on hard matrices the replicate cap binds before the hit target is reached (hitCapBound rises 22% -> 47%, i.e. the extra demand only pushes more runs into the cap), and on easy matrices the optimum is already in hand. The extra work lands exactly where it cannot help. The `effort` ladder doubles `targetHits` at rungs 5+, so this is evidence against a choice made in b2dfc3a6/2d871b41. It is kept -- without it a notch would be INERT on every dataset that stops early, and under implied weights it deepens the ratchet, which an equal-weights panel cannot see -- but the docs no longer imply it buys reach. Rungs 5+ now read as buying confidence and distinct trees on easy data, and reach on hard data. Gating still costs reach under the real stopping rules, so the default still does not flip: B_gate is 0 better / 4 worse (p = 0.125, uniformly negative), losses concentrated on the hard tail (Zanol2014 -0.6, Wortley2006 -0.4, Aria2015 -0.2, Zhu2013 -0.2). C_final is the arm worth pursuing: statistically indistinguishable from baseline (-0.013, 1 better / 2 worse, p = 1.000) at a THIRD of the wall, and its certifications are the productive ones -- 4349 of 11000 sweeps found a real improver (40%) against A_default's 9555 of 44889 (21%). Certifying the tree a replicate actually reports is four times cheaper and twice as likely to pay. It still regresses on Zanol2014 (0.6 -> 0.0), so it is a candidate, not a change to make now. Settling it wants the hard tail at more seeds, not another corpus-wide sweep: 24 of the 30 matrices sit at 1.0 in every arm and can only dilute the signal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…epos
R reads EITHER ./.Rprofile OR ~/.Rprofile, never both (?Startup). GHA steps
run with cwd = the checkout, so the root .Rprofile added on cpp-search (for
the Rcpp CRLF patch) silently suppressed the user profile -- which is exactly
where r-lib/actions/setup-r writes
options(repos = c(RSPM = <Posit PM>, CRAN = ...), Ncpus, HTTPUserAgent)
(see setup-r/src/installer.ts:659-757). Consequence on every cpp-search job:
Posit Package Manager never entered getOption("repos"), the User-Agent that
makes it serve binaries was never set, and pak fell back to its default CRAN
mirror -- resolving 100% of dependencies from source. ~10 min of compiling per
leg, and it also explains why the earlier Ncpus and RSPM/env experiments
measured nothing: those values are written into the same shadowed file.
Evidence (cache-immune -- pak prints this before installing anything):
main arm64 leg: repo_status() row 1 = RSPM <Posit PM>; 110 aarch64
binary resolution targets, 1 source
cpp-search leg: repo_status() row 1 = CRAN https://cran.rstudio.com,
Posit PM absent entirely; 0 binary targets, 101 source
Fix: chain to the user profile before installing the Rcpp hook, with a guard
for the cwd == home case. inst/Parsimony/.Rprofile is deliberately left alone
-- it ships in the tarball and has no bearing on dependency setup.
cache-version bumped to 2 on every setup-r-dependencies step: the v1 caches
were populated while the shadowing was in effect, so every package in them
was built from source. Retire them once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two cuts to the preliminaries on the fast legs, now that dependencies actually
resolve as binaries again.
1. Pin Posit PM to a dated snapshot (2026-07-30, the date proven to serve
aarch64/noble binaries) on `sense-check` and on agent-check's ubuntu leg.
`setup-r-dependencies` keys its cache on the RESOLVED dependency versions,
so `latest` misses the cache whenever any of ~110 dependencies gets a new
CRAN release -- most days -- and re-downloads the lot. A fixed snapshot
keeps the cache valid for weeks, so dependency setup becomes a restore.
`core`'s ubuntu legs deliberately keep `latest`: catching breakage from
bleeding-edge dependencies is the R-devel recipe's job, not the fast
"did this commit build?" signal's.
2. Drop the texlive apt step from `sense-check` (~45 s/run). It was never
needed: `check-r-package@v2` defaults to args = c("--no-manual",
"--as-cran") and build_args = "--no-manual", and all eight vignettes are
`rmarkdown::html_vignette`, i.e. pandoc-only. `core` and `full` still
install texlive, so a real LaTeX dependency would still be caught.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…(T-375, T-376) HSJ's primary_present test and fitch_label_char() both compared tip_labels' token (allLevels/contrast-row) indices directly against state (levels) indices -- absent_state, inapp_state, and bit positions in state_sets are all in state space, but tip_labels is in token space. The two spaces coincide often enough that earlier fixes and their test batteries (which permuted levels, not allLevels) stayed green, but on real reductively-coded data (e.g. Vinther2008.nex) absent/inapplicable primaries could read as present while "?" read as absent, and the score was provably not a function of the data under contrast-row permutation alone. DataSet now stores token_states (the token -> state-set bitmask already computed transiently in build_dataset()) and n_levels, so the HSJ kernel can translate a token label into its state set before testing set membership against absent_state/inapp_state, and before building Fitch state_sets for secondary characters (fixing "?" scoring as a concrete, conflicting state instead of a wildcard). Also corrects the stale "token index" docstrings on .HSJAbsentState() and its caller, which described a levels index.
…s ~70 s/run Profiling the (now-binary) dependency step showed the download of 108 binaries takes 6 s and installing them ~10 s. The time is elsewhere: pak install + resolve + metadata DB ~17 s download 108 binaries 6 s pak system-requirements apt pass ~75 s <-- ppa:xtradeb/apps + Chromium install 108 binaries ~10 s MaxMin (GitHub source): Packaging 37 s + Building 29 s = 67 s shinytest2 -> chromote declares `chromium` as a system requirement, so on Linux pak runs `add-apt-repository -y ppa:xtradeb/apps`, `apt-get update`, then installs Chromium and its GTK/X11 tree. That happens on EVERY run: the sysreqs pass is not covered by the package cache (confirmed on a warm-cache run whose dependency step still took 83 s with 112/117 packages kept). No Linux leg needs it. shinytest2 appears only under inst/Parsimony/tests, which `.Rbuildignore` excludes, so it is not even in the tarball R CMD check examines; the app suite runs in the dedicated `shiny` job on Windows, which is left untouched. MaxMin's 67 s is deliberately NOT addressed here: it is a genuine trade (tests/testthat/test-WideSample.R would skip on the affected leg) and is the maintainer's call. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…evels (T-375, T-376) make_hsj_dat()'s construction keeps allLevels in the same relative order as levels, so none of the existing absolute-value tests could have caught the token/state index confusion -- confirmed by running this file's original tests unmodified against a pre-fix build (all still pass, unchanged). Adds: (1) contrast-row permutation invariance, lifted from dev/red-team/heavy-tests/hsj-token-permutation.R, for zero and one secondaries; (2) a MatrixToPhyDat()-emergent token/state misalignment (levels = "- 0 1", allLevels = "1 0 -", the same shape as the shipped Vinther2008.nex dataset) with a hand-derived absolute score, plus the paper's alpha=0-equals-Fitch identity on the same data. Verified both new tests fail against a throwaway pre-fix build of cpp-search and pass against the fix.
Unrelated to the HSJ fix: R/MaximizeParsimony.R's @examples block was updated by an earlier commit (18397db, "regenerate a MaximizeParsimony example left stale by the effort rename") but man/MaximizeParsimony.Rd wasn't regenerated to match. Caught by running devtools::check_man() before dispatch, as this branch's fix required no doc changes of its own.
…374 fix Verified while validating the T-375/T-376 fix: check [3] now runs live (the alpha term is no longer inert) and passes on this small, Fig.1-derived matrix -- but a fresh sweep over larger/messier matrices shows T-374's rooting-dependence still reproduces readily post-fix, exactly as expected since T-374 needs the paper's two-state DP, which this fix does not touch. Add a comment so a future session doesn't read an all-green run here as T-374 being closed.
…after) Answers whether a container with R (or R + all deps) preinstalled beats setup-r + setup-r-dependencies. Bar to beat: ~40-70 s for R alone, ~3 min for R + dependencies. Measures raw pull time/size for extrapolation, plus the runner's own "Initialize containers" step, which is the cost that actually lands on a container-based leg. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…branch) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…hypothesis RETRACTION. Panel 2's headline -- "targetHits escalation is pure cost, 0 of 30 matrices improved at 2.6x the wall" -- is true and useless. 24 of the 30 matrices are saturated at 1.0 in every arm, so "no improvement" there is the expected result, not a finding; and floor attainment is binary, so it cannot see an arm getting closer without arriving. I flagged the dilution in the same document and then led with the diluted number anyway. Re-analysed on score, restricted to the six matrices with headroom, and counting what the arm actually bought: extra replicates from targetHits x3, all 30 matrices : 4409 (4.2 CPU-hours) ... on the six matrices with headroom : 243 ... on Zanol2014 / Zhu2013 / Geisler2001 : 0 On the three hardest the 96-replicate cap bound BOTH arms on every seed (21 of 30 headroom cells), so A_hits3 did byte-identical work to A_default -- identical scores, identical replicate counts. 94% of the extra work landed on matrices that were already solved. Only 9 cells anywhere gained a replicate. So the correct claim is not "raising targetHits does not help" but "targetHits cannot act once maxReplicates binds", which is a stronger statement about mechanism and a much weaker one about effort. The hypothesis -- low effort misses the optimum, high effort takes longer and recovers a better score -- was never tested, because the arm could not act where it mattered. PANEL 3 tests it. Hard matrices only (the six with headroom), scored on SCORE not binary attainment, arms = `effort` 0/+1/+2 crossed with certification. Arms are `effort` itself, not a hand-rolled budget, because a first draft that raised only maxReplicates was INERT on Aria2015: targetHits stopped every run at 39 replicates, far below even the 96 cap, so all three budgets returned identical scores. That is the exact mirror of panel 2's mistake. Only moving both knobs escalates every dataset -- which is what the ladder does -- so testing the shipped argument is both more honest and more informative. Two things the smoke test exposed, recorded rather than silently patched: * Panel 2 pinned `strategy = "default"` for every matrix, so on 65-119-tip matrices it did NOT measure what a user gets (`auto` selects `thorough`). * Notch 3->4 is DEAD on hits-bound datasets: rungs 1-4 never touch targetHits and rung 4 raises only the cap, so `effort = +1` and `+2` gave literally the same run on Aria2015. Rung 4 is an auto pick calibrated for >=120 tips, so this is not restructured on no evidence -- panel 3 will quantify it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
arm64, ubuntu-24.04-arm runner: rocker/r-ver:latest pull 10s, 865MB unpacked rocker/tidyverse:latest pull 25s, 3073MB unpacked container: rocker/r-ver -> "Initialize containers" = 12s (pull+create+start) => ~10s fixed + ~7ms/MB; image ships R 4.6.1 + 31 packages Conclusion: an R-ONLY image is roughly a wash -- setup-r spends only ~10s on R itself, its other ~52s being the hard-coded devscripts/ghostscript apt install. An image with R + all ~110 dependencies baked in would land near 1.3-1.6GB, predicting ~15-20s of container init against the 170-250s the two setup steps cost today. That is the variant worth building, if we build one. Incidental: rocker sets repos in Rprofile.site, to p3m.dev/cran/__linux__/noble/latest -- the rolling "latest", NOT a dated snapshot. Rprofile.site is read unconditionally, so unlike the ~/.Rprofile route it is immune to the project-.Rprofile shadowing this branch fixes. Whether it also sets HTTPUserAgent is unverified and would need checking before relying on it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
as warning permissible but not required
…port-calculation-857d5f
For `attr(tree, order) = NULL` support
Suppresses NOTE.
…857d5f Bremer (decay) support via converse-constraint + pool
Replace the R-based codemeta workflow, which pushed to the default branch, with ms609/actions/update-codemeta: a PR-only job on ubuntu-slim that regenerates codemeta.json in seconds (stdlib Python, no R) and pushes any update to the PR branch. codemeta.json is regenerated here with the same script so the first PR starts from a current file. Also runs update-csl and the paths-filter job on ubuntu-slim. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Manual testing underway; shiny app in particular has some usability issues.