Benchmark and paper package for validating operational self-model and agency constructs in controlled artificial neural architectures.
-
Updated
Jul 6, 2026 - Python
Benchmark and paper package for validating operational self-model and agency constructs in controlled artificial neural architectures.
Construct-validity audit of the standard blood–brain barrier (BBB) peptide benchmark: an identity-controlled re-evaluation + shared-source provenance/overlap map, with an open, CPU-reproducible evaluation harness. Do these predictors measure penetration, or their benchmarks?
Reproducibility source, frozen configurations, manifests, and provenance for the SROP research programme.
Supplementary materials for the following publication: Davydenko, A., & Goodwin, P. (2021). Assessing point forecast bias across multiple time series: Measures and visual tools. International Journal of Statistics and Probability, 10(5), 46-69. https://doi.org/10.5539/ijsp.v10n5p46
Inspect eval: do LLMs writing case notes separate observation from interpretation, and can they evade a lexical validator? Deterministic grader, pre-registered rubric.
Five falsification studies of an instrument's own limits — every figure tied to an executed measurement with confidence intervals, or marked BLOCKED. Bounded evidence lattice, critic calibration, domain transfer, verification scaling curve, Goodhart-guarded self-improvement.
Repository for "Testable or Not? A Pre-Registered Validity Protocol for Architecture Comparisons." Includes pre-registrations, per-seed data, analysis, and paper source.
Construct validity of sycophancy interventions in Llama-3.1-8B — reading a concept is not controlling it
Local-first construct and operational-definition canvas for research-methods planning.
R and Python replication of a psychometric study validating brief (18-item) versions of the MCQ and DLQ for assessing delay discounting of gains and losses (Wan et al., 2025, The Psychological Record).
Maritime Intent Probe is a Phase 1 research programme on construct validity in neural probing. It introduces BC1 and uses a preregistered maritime routing counterexample to establish the identifiability requirements that motivate a crossed-design Phase 2 validation.
Reproducible benchmark of security-oracle construct validity on 140 real CVE fixes.
Ecological study on administrative diabetes indicators and the care cascade using NDB Open Data, Japan (335 secondary medical areas, FY2023-2024)
R script analyzing the Swahili RCADS-25 among Kenyan adolescents, assessing internal consistency, construct validity, convergent and divergent validity, and measurement invariance to evaluate the psychometric propertiesscale.
Code, per-item results and figures for a study dissociating prompt quality from response compliance in automated prompt-engineering assessment. Four-agent MATLAB evaluator on a locally hosted Qwen 2.5-7B judge, over 498 prompts from IFEval, LMSYS-Chat-1M and WildChat.
Measurement validity, construct validity, and unsupported claims derived from AI agent telemetry.
Data, code and verification gates for "The Granularity Gap" (arXiv:2606.05183): a continuous-scale audit of sycophancy across 8 Gemini variants and 8,830 responses, with 10,792 per-vote judge logs.
Guided R workflow for scale development and construct validation: item screening, factor retention, CFA, reliability, bifactor and higher-order models, measurement invariance, scoring, and theory-specified nomological networks. Every method is placed in its literature and cited.
To associate your repository with the construct-validity topic, visit your repo's landing page and select "manage topics."