An evaluation framework for evaluating types of unstable behaviour in LLMs.
This repository uses uv.
Also, make sure you have API keys set up for Google AI [and others? TBA] in a .env file.
To run the pipeline, set variables in run.sh to the desired config and run
bash run.shStabilityBench automatically caches diversified datasets.
Manage cache with the cache_manager, see scripts/README.md.
On a high level, StabilityBench operates under the following flow:
- Diversification: 2 orthogonal expansions across user types and user behaviours.
- In default mode, diversification is an expansion protocol, and it allows to evaluate performance differences across user types and behaviours.
- In
minimode, i.e, using the diversifier as a standalone module, diversification is a distribution-preserving transformation that makes the benchmark more realistic, while preserving its size. This can make evaluation more robust to real-world use cases, while keeping the same cost.
- Evaluation on the diversified set: running inference, then scoring: substance, form, and bait effect.
Below, we link resources used in this work:
- Benchmarks:
- AIME: we use the dataset from
opencompasson huggingface, and some code from the AIME 2025 Benchmark Starter Notebook. - GSM8k
- HealthBench
- StrongReject
- AIME: we use the dataset from
- Other resources:
- We use the U.S Center for Medicare & Medicaid Services (CMS)'s taxonomy for ethnic and racial categories, as established 2024 Resource for Standardized Demographic and Language Data Collection as it is fit for the healthcare setting (and user type especially factors into Health Bench evaluation). They establish a set of ethnic categories (with 2 categories: "Hispanic or Latino" or "Not Hispanic or Latino"), and separately of racial categories (with 5 categories). While the distinction is motivated for language-related purposes, we agglomerate the two sets into one set of 6 categories.
- We use the follow socio-linguistic works to inform our diversification methodology:
- Demographic and Socioeconomic Basis of Ethnolinguistics, Jacob S. Siegel
- The Social and Demographic Context of Language Use in the United States, Gillian Stevens
- Methods and Annotated Data Sets Used to Predict the Gender and Age of Twitter Users: Scoping Review, O'Connor et al.
- The Computer-Linguistic Analysis of Socio-Demographic Profile of Virtual Community Member, Syerov et al.