Skip to content

About

A common set of design research analysis tools

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

design-research-analysis

CI Coverage Examples Passing API in Examples Docs PyPI Version Python Versions

design-research-analysis is the analysis and interpretation layer in the CMU Design Research Collective design-research ecosystem. It turns recurring design-study records and exported experiment artifacts into reproducible findings.

It provides typed, reusable workflows for sequence, language, embedding-map, and statistical analysis over recurring event logs.

Quality Signals

  • Coverage reports total line coverage for the default deterministic test suite; CI requires at least 95%.
  • Examples Passing reports checked-in example scripts that execute successfully in the examples workflow.
  • API in Examples reports curated top-level __all__ exports referenced by runnable examples. N/N means every supported top-level export appears in at least one example, and CI requires 100%.

Run make coverage, make examples-test, and make examples-coverage to reproduce these checks locally.

Overview

This package centers on reproducible analysis workflows with a curated top-level API:

  • Unified-table coercion, validation, and mapper-based derived columns
  • Dataset profiling, schema checks, and codebook generation
  • Sequence modeling (Markov chains, discrete HMM, Gaussian HMM)
  • Language analysis (semantic convergence trajectories, topic modeling, sentiment scoring)
  • Embedding maps (PCA, t-SNE, UMAP, PaCMAP, TriMap) with clustering, comparison, and trajectory-plotting helpers
  • Statistical wrappers (group comparisons, OLS regression, mixed-effects models, nonparametrics, and power)
  • Portable analysis-result records and evidence-linked paper contributions
  • Runtime provenance capture for reproducibility manifests
  • Deterministic, integrity-checked paper-draft bundles with narrow data selection
  • Top-level artifact handoff helpers for experiment exports
  • A thin CLI for deterministic pipeline runs

Quickstart

Requires Python 3.12+. Maintainer workflows target Python 3.12 (.python-version). For a VS Code path that starts from PyPI and then shows the repository example workflow, see VS Code Start.

Install from PyPI:

python -m pip install --upgrade pip
python -m pip install design-research-analysis

The base install supports unified-table validation and Markov analysis:

import design_research_analysis as dran

rows = [
    {"timestamp": "2026-01-01T00:00:00Z", "session_id": "s1", "event_type": "ideate"},
    {"timestamp": "2026-01-01T00:00:05Z", "session_id": "s1", "event_type": "refine"},
]
report = dran.validate_unified_table(rows)
model = dran.fit_markov_chain_from_table(rows)
print(report.is_valid, model.states)

Common install profiles:

python -m pip install "design-research-analysis[seq]"
python -m pip install "design-research-analysis[lang,embeddings]"
python -m pip install "design-research-analysis[maps,embeddings]"  # text-driven maps
python -m pip install "design-research-analysis[maps]"             # numeric features
python -m pip install "design-research-analysis[stats,data]"
python -m pip install "design-research-analysis[all]"

Unified-table coercion, validation, and derived-column helpers ship in the base install, so there is no separate table extra.

For contributor workflows:

python -m venv .venv
source .venv/bin/activate
make dev
make test

Run a compact end-to-end example:

PYTHONPATH=src python examples/basic_usage.py

For dependency profiles and release-check guidance, see Dependencies and Extras.

CLI

The package installs a design-research-analysis CLI:

design-research-analysis validate-table --input data/events.csv --summary-json artifacts/validate.json
design-research-analysis run-sequence --input data/events.csv --summary-json artifacts/sequence.json --mode markov
design-research-analysis run-language --input data/events.csv --summary-json artifacts/language.json --trajectory-csv artifacts/language_trajectory.csv
design-research-analysis run-embedding-maps --input data/events.csv --summary-json artifacts/embedding_maps.json --map-csv artifacts/embedding_maps.csv
design-research-analysis run-stats --input data/events.csv --summary-json artifacts/stats.json --mode regression --x-columns x1,x2 --y-column y

The Python API can start from files too at the main ingestion points. For example, coerce_unified_table("data/events.csv") uses the base install; profile_dataframe("data/events.csv") requires the data extra.

Examples

Start with examples/README.md for runnable scripts across all analysis families.

Docs

See the published documentation for quickstart, workflow guidance, schema details, CLI reference, and API docs. Use the design-research umbrella documentation for canonical whole-stack orientation and cross-package examples.

Build docs locally with:

make docs-check
make docs-build

Public API

The supported public surface is whatever is exported from design_research_analysis.__all__.

Selected primary entry points include:

  • Package metadata: __version__
  • Artifact handoff helpers: validate_experiment_events, build_condition_metric_table_from_artifacts, build_event_table_from_artifacts
  • Table contracts: UnifiedTableConfig, UnifiedTableValidationReport, coerce_unified_table, derive_columns, validate_unified_table
  • Sequence: fit_markov_chain_from_table, fit_discrete_hmm_from_table, fit_text_gaussian_hmm_from_table, decode_hmm, plotting helpers, and result types
  • Language: compute_language_convergence, compute_semantic_distance_trajectory, fit_topic_model, score_sentiment
  • Embedding maps: embed_records, build_embedding_map, cluster_embedding_map, compare_embedding_maps, plot_embedding_map, plot_embedding_map_grid
  • Statistics: compare_groups, fit_regression, fit_mixed_effects, permutation_test, bootstrap_ci, power helpers
  • Dataset + runtime: profile_dataframe, validate_dataframe, generate_codebook, capture_run_context, attach_provenance, write_run_manifest
  • Paper support: build_analysis_result, write_analysis_result, load_analysis_result, collect_analysis_paper_contributions
  • Verified bundle: create_research_bundle, verify_research_bundle, BUNDLE_SCHEMA_VERSION

Contributing

Contribution workflow and validation gates are documented in CONTRIBUTING.md.

About

A common set of design research analysis tools

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

Generated from cmudrc/python-template