Skip to content

About

An evaluation framework to assess unstable behaviours in LLMs

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

StabilityBench

An evaluation framework for evaluating types of unstable behaviour in LLMs.

Set-up:

This repository uses uv.

Also, make sure you have API keys set up for Google AI [and others? TBA] in a .env file.

Usage

To run the pipeline, set variables in run.sh to the desired config and run

bash run.sh

Caching & Resuming

StabilityBench automatically caches diversified datasets.

Manage cache with the cache_manager, see scripts/README.md.

Methodology

On a high level, StabilityBench operates under the following flow:

  1. Diversification: 2 orthogonal expansions across user types and user behaviours.
    • In default mode, diversification is an expansion protocol, and it allows to evaluate performance differences across user types and behaviours.
    • In mini mode, i.e, using the diversifier as a standalone module, diversification is a distribution-preserving transformation that makes the benchmark more realistic, while preserving its size. This can make evaluation more robust to real-world use cases, while keeping the same cost.
  2. Evaluation on the diversified set: running inference, then scoring: substance, form, and bait effect.

Credits:

Below, we link resources used in this work:

  • Benchmarks:
  • Other resources:
    • We use the U.S Center for Medicare & Medicaid Services (CMS)'s taxonomy for ethnic and racial categories, as established 2024 Resource for Standardized Demographic and Language Data Collection as it is fit for the healthcare setting (and user type especially factors into Health Bench evaluation). They establish a set of ethnic categories (with 2 categories: "Hispanic or Latino" or "Not Hispanic or Latino"), and separately of racial categories (with 5 categories). While the distinction is motivated for language-related purposes, we agglomerate the two sets into one set of 6 categories.
    • We use the follow socio-linguistic works to inform our diversification methodology:
      • Demographic and Socioeconomic Basis of Ethnolinguistics, Jacob S. Siegel
      • The Social and Demographic Context of Language Use in the United States, Gillian Stevens
      • Methods and Annotated Data Sets Used to Predict the Gender and Age of Twitter Users: Scoping Review, O'Connor et al.
      • The Computer-Linguistic Analysis of Socio-Demographic Profile of Virtual Community Member, Syerov et al.

About

An evaluation framework to assess unstable behaviours in LLMs

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages